Automatic creation of bilingual dictionaries for Finno-Ugric languages

  • Eszter Simon Research Institute for Linguistics Hungarian Academy of Sciences
  • Ivett Zs. Benyeda Research Institute for Linguistics Hungarian Academy of Sciences
  • Péter Koczka Research Institute for Linguistics Hungarian Academy of Sciences
  • Zsófia Ludányi Research Institute for Linguistics Hungarian Academy of Sciences

Abstract

We introduce an ongoing project whose objective is to provide linguistically based support for several small Finno-Ugric digital communities in generating online content. To achieve our goals, we collect parallel, comparable and monolingual text material for the following Finno-Ugric (FU) languages: Komi-Zyrian and Permyak, Udmurt, Meadow and Hill Mari and Northern Sami, as well as for major languages that are of interest to the FU community: English, Russian, Finnish and Hungarian. Our goal is to generate proto-dictionaries for the mentioned language pairs and deploy the enriched lexical material on the web in the framework of the collaborative dictionary project Wiktionary. In addition, we will make all of the project’s products (corpora, models, dictionaries) freely available supporting further research.
Published
2015-06-17