Portrait of Abhishek Bhandari

Abhishek Bhandari

Ph.D. Student in Computer Science & Engineering · IIT Jodhpur

I am working on OCR and ASR error correction for Indic languages. My work also explores multimodal document understanding and cross-lingual reasoning in vision-language models.

I submitted my doctoral thesis, Context-Assisted Post-Text-Recognition Error Correction for Indic Languages, at IIT Jodhpur in July 2026.

OCR & ASR Correction Low-Resource Indic Languages Vision-Language Models

Selected Publications (view all )

Schematic of the pipeline: an OCR sentence is matched by CharBM25 to similar examples, which prompt an LLM to produce corrected text.

Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

Abhishek Bhandari, Gaurav Harit

arXiv preprint arXiv:2609.21595 · 2026

The first systematic evaluation of 3B-32B LLMs for training-free post-OCR correction in Hindi and Marathi. A proposed character n-gram BM25 retrieval (CharBM25) for in-context examples outperforms random selection by up to 4 points of absolute WER.

A Framework and Dataset for Contextual Post-OCR Correction

A Framework and Dataset for Contextual Post-OCR Correction

Abhishek Bhandari, Gaurav Harit

ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) · 2026

A correction framework uses the preceding sentence as context to resolve OCR errors in Hindi, Marathi, and Gujarati. The accompanying dataset supports evaluation with and without sentence context.

Figure 1 from the paper: the proposed multi-view architecture with character-wise gated fusion.

Post-ASR Correction for Low-Resource Rajasthani Language

Abhishek Bhandari, Gaurav Harit

ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) · 2026

A character-level correction model combines transcripts from Whisper and MMS through gated fusion, using complementary recognition cues to correct errors in low-resource Rajasthani speech transcription.

Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

Somraj Gautam, AS Penamakuri, Abhishek Bhandari, Gaurav Harit

Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025) · 2025

MMCRICBENCH-3K tests numerical and cross-lingual reasoning on English and Hindi cricket scorecard images. It exposes difficulties in interpreting table structure and transferring visual understanding across scripts.

All publications

News

Education

  • Indian Institute of Technology Jodhpur
    Indian Institute of Technology Jodhpur
    Computer Science & Engineering
    Ph.D. (thesis submitted July 2026)
    Since 2020
  • University of Hyderabad
    University of Hyderabad
    M.Tech. in Information Technology
    2017 - 2019
  • Government Engineering College Ajmer
    Government Engineering College Ajmer
    B.Tech. in Information Technology
    2012 - 2016