We weave language models the way Assam weaves silk: thread by thread, from scratch.
Navdyut AI trains foundational and fine-tuned models on curated, scanned Assamese sources (books, newspapers, journals) instead of the open internet.
Read the evidence: corpus method Eri release record
One research core. Two ways we deliver it.
The same Assamese stack feeds both tracks. Only access and constraints change.
Eri
Foundation weights trained from scratch on our curated corpus. Released so low-resource pretraining can be studied and reproduced.
Hugging Face ↗ DeployMuga and beyond
Language, documents, and voice: fine-tuned and deployed inside your boundary for enterprise cost control and government sovereignty.
How we deploy →Your model. Your machines. Your data never leaves.
Most Assamese-capable AI today is a wrapper around someone else's API: priced per token, hosted outside the state, and impossible to audit. Our models ship for self-hosted deployment instead.
How we deploy →Language, documents, and voice.
200 million tokens, scanned from the printed record.
We scan the printed record of Assamese that the open web never indexed well. Cleaner sourcing shows up directly in output quality.