The Finno-Ugric Languages and The Internet Project

Heidi Jauhiainen; Tommi Jauhiainen; Krister Lindén

doi:10.7557/5.3471

Authors

Heidi Jauhiainen University of Helsinki Department of Modern Languages
Tommi Jauhiainen University of Helsinki Department of Modern Languages
Krister Lindén University of Helsinki Department of Modern Languages

DOI:

https://doi.org/10.7557/5.3471

Abstract

This paper describes a Kone Foundation funded project called "The Finno-Ugric Languages and The Internet" together with some of the achieved results. The main activity of the project is to crawl the internet and gather texts written in small Uralic languages. The sentences and words of the found texts will be assembled into a freely available corpus. Crawling is done using the open source crawler Heritrix, which is developed by the Internet Archive. Heritrix crawls through the pages and passes the found texts to a language identifier.

We are using a state of the art language identifier, which has been further developed within the project and has been evaluated using 285 languages. We describe the language identification evaluation results concerning the 34 Uralic languages known by the language identifier. We also describe the initial observations and results from the first five large crawls which were done in the national internet domains of Finland, Sweden, Norway, Russia, and Estonia.

The Finno-Ugric Languages and The Internet Project

Authors

DOI:

Abstract

Downloads

Published

Issue

Section

License

How to Cite

Language

Information

Make a Submission

Latest publications

Keywords