Welcome to MaLA-LM 🌍
MaLA-LM (Massive Language Adaptation of Large Language Models) focuses on adapting large language models to support hundreds of languages, including many underrepresented ones. Our models are multilingual, scalable, and optimized for diverse linguistic tasks.
Featured 🗣️ Check out our multilingual LLM collections, featuring models trained to handle 500+ languages, ideal for global, multilingual applications. Dive into the HuggingFace collections: EMMA-500, MaLA corpus and MaLA-500.
News
Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation? is accepted to Findings of EMNLP 2026 🎉 📄arXiv:2606.03782
MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models (aka EMMA-500) is accepted to COLM 2026 🎉 📄MaLA
We release the preprint of Model-Based Quality Assessment for Massively Multilingual Parallel Data 📄arXiv:2606.00285
Data-Centric Continual Pre-training for 500+ Languages: A New Bilingual Translation Corpus and Multilingual Models (aka EMMA-500 Gen 2) is accepted to Findings of ACL 2026 🎉 📄ACL Anthology
Test-Time Scaling of Reasoning Models for Machine Translation is accepted to EACL 2026 🎉 📄ACL Anthology
We launch the FineOPUS project for refining parallel texts in many languages 🌐FineOPUS
We release EMMA-500 Llama 3/3.1 models and MaLA bilingual corpus in 2,500+ language pairs 🌐EMMA-500 Gen2
We release MaLA OPUS bilingual corpus (2410), aka, parallel corpus, in 16,000+ language pairs 🤗MaLA-LM/mala-opus-dedup-2410
We release a series of CPT models that study the data mixing in continual pre-training 🤗MixCPT
We release the preview of GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models GlotEval
We release PolyWrite, a synthetic benchmark for open-ended generation 🤗MaLA-LM/PolyWrite
We release the EMMA-500 Llama 2 model and MaLA monolingual corpus in 939 languages 🌐EMMA-500
We release Lucky 52 models that study the number of languages for instruction fine-tuning 🤗Lucky 52
We release the MaLA-500 v2 model 🤗MaLA-LM/mala-500-10b-v2
We release the MaLA-500 v1 model based on Llama 2 and LoRA 🤗MaLA-LM/mala-500-10b-v1
Research directions
Open monolingual and parallel corpora covering hundreds of languages and thousands of language pairs.
Extending open LLMs to 500+ languages and studying how data mixes shape cross-lingual transfer.
How many languages, and which ones, make multilingual instruction tuning work.
Consistent, non-English-centric benchmarking of LLMs across dozens to hundreds of languages.
Our works
A massively multilingual translation corpus and pipeline, refining the OPUS collection of parallel texts in many languages.
The MaLA bilingual translation corpus in 2,500+ language pairs and the EMMA-500 Llama 3/3.1 models.
The EMMA-500 Llama 2 model and the MaLA monolingual corpus in 939 languages.
Vocabulary extension and continued pretraining on Llama 2 with Glot500-c to cover 534 languages.
36 CPT configurations comparing monolingual, bilingual and code-augmented data across 30+ languages.
A lightweight framework for seven tasks spanning dozens to hundreds of languages.
The Lucky 52 models, studying the number of languages for instruction fine-tuning.
Comparing monolingual and multilingual instruction tuning under a fixed budget.
Our corpora, models and code are open. Join the community to ask questions, share results or contribute data for the languages you care about.