Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction
Abhishek Bhandari, Gaurav Harit
arXiv preprint arXiv:2609.21595 · 2026
The first systematic evaluation of 3B-32B LLMs for training-free post-OCR correction in Hindi and Marathi. A proposed character n-gram BM25 retrieval (CharBM25) for in-context examples outperforms random selection by up to 4 points of absolute WER.