Can Morphological Analyzers Improve the Quality of Optical Character Recognition?
DOI:
https://doi.org/10.7557/5.3467Abstract
Optical Character Recognition (OCR) can substantially improve the usability of digitized documents. Language modeling using word lists is known to improve OCR quality for English. For morphologically rich languages, however, even large word lists do not reach high coverage on unseen text. Morphological analyzers offer a more sophisticated approach, which is useful in many language processing applications. is paper investigates language modeling in the open-source OCR engine Tesseract using morphological analyzers. We present experiments on two Uralic languages Finnish and Erzya. According to our experiments, word lists may still be superior to morphological analyzers in OCR even for languages with rich morphology. Our error analysis indicates that morphological analyzers can cause a large amount of real word OCR errors.Metrics
Metrics Loading ...
Downloads
Published
2015-06-17
How to Cite
Silfverberg, M., & Rueter, J. (2015). Can Morphological Analyzers Improve the Quality of Optical Character Recognition?. Septentrio Conference Series, (2), 45–56. https://doi.org/10.7557/5.3467
Issue
Section
Articles