Shillong: Pnar is spoken by hundreds of thousands in Meghalaya’s Jaintia Hills, but it is effectively invisible to digital tools. A new study shows the language lacks the basic data needed to train modern machines. Researchers recently finished the first machine-translation study connecting English and Pnar. They built a 10,234-sentence parallel corpus using text from the newspaper Wyrta. Data is scarce. The team used 9,563 pairs to train systems and 371 for testing.
Performance remains low. The best system hit a BLEU score of 14.97 for Pnar-to-English translation. English-to-Pnar translation scored only 11.16. Researchers claim these results provide the first quantitative benchmark for the pair. The team noted specific technical hurdles. They faced missing vocabulary, complex word ordering, and frequent Khasi code-mixing in the source material. Structural differences between Pnar’s subject-object-verb pattern and English’s subject-verb-object structure added to the load.
The study highlights a grim reality for indigenous tongues. AI systems now dominate search, education, and public services. Languages without massive digital footprints risk total exclusion. Researchers struggled most with simply finding enough material to train their model. They now look toward moving from statistical systems to neural and multilingual AI models. As the team observed, "For Pnar, the first step may already have been taken—not by a large technology company, but by researchers who had to assemble thousands of sentences simply to give a computer enough language data to begin learning."

Comments