MaLA-LM · Massive Language Adaptation

Large language models that speak hundreds of languages

We adapt large language models to hundreds of languages, including many underrepresented ones, through open corpora, continual pretraining, reasoning, instruction tuning and massively multilingual evaluation.

939 languages in the MaLA monolingual corpus
2,500+ language pairs in the MaLA bilingual translation corpus
16,000+ language pairs in the deduplicated MaLA OPUS corpus
671B tokens of continual pretraining for EMMA-500 Llama 3

Welcome to MaLA-LM 🌍

Closing the gap between high- and low-resource languages.

MaLA-LM (Massive Language Adaptation of Large Language Models) focuses on adapting large language models to support hundreds of languages, including many underrepresented ones. Our models are multilingual, scalable, and optimized for diverse linguistic tasks.

Featured 🗣️ Check out our multilingual LLM collections, featuring models trained to handle 500+ languages, ideal for global, multilingual applications. Dive into the HuggingFace collections: EMMA-500, MaLA corpus and MaLA-500.

News

Latest Updates

  1. 2026.08

    Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation? is accepted to Findings of EMNLP 2026 🎉 📄arXiv:2606.03782

  2. 2026.07

    MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models (aka EMMA-500) is accepted to COLM 2026 🎉 📄MaLA

  3. 2026.06

    We release the preprint of Model-Based Quality Assessment for Massively Multilingual Parallel Data 📄arXiv:2606.00285

  4. 2026.04

    Data-Centric Continual Pre-training for 500+ Languages: A New Bilingual Translation Corpus and Multilingual Models (aka EMMA-500 Gen 2) is accepted to Findings of ACL 2026 🎉 📄ACL Anthology

  5. 2026.01

    Test-Time Scaling of Reasoning Models for Machine Translation is accepted to EACL 2026 🎉 📄ACL Anthology

  6. 2025.11

    We launch the FineOPUS project for refining parallel texts in many languages 🌐FineOPUS

  7. 2025.06

    We release EMMA-500 Llama 3/3.1 models and MaLA bilingual corpus in 2,500+ language pairs 🌐EMMA-500 Gen2

  8. 2025.05

    We release MaLA OPUS bilingual corpus (2410), aka, parallel corpus, in 16,000+ language pairs 🤗MaLA-LM/mala-opus-dedup-2410

  9. 2025.04

    We release a series of CPT models that study the data mixing in continual pre-training 🤗MixCPT

  10. 2025.04

    We release the preview of GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models GlotEval

  11. 2024.09

    We release PolyWrite, a synthetic benchmark for open-ended generation 🤗MaLA-LM/PolyWrite

  12. 2024.09

    We release the EMMA-500 Llama 2 model and MaLA monolingual corpus in 939 languages 🌐EMMA-500

  13. 2024.04

    We release Lucky 52 models that study the number of languages for instruction fine-tuning 🤗Lucky 52

  14. 2024.03

    We release the MaLA-500 v2 model 🤗MaLA-LM/mala-500-10b-v2

  15. 2024.01

    We release the MaLA-500 v1 model based on Llama 2 and LoRA 🤗MaLA-LM/mala-500-10b-v1

Research directions

From raw text to evaluated models, in every language.

📚 01

Data Construction

Open monolingual and parallel corpora covering hundreds of languages and thousands of language pairs.

📜 02

Continual Pretraining

Extending open LLMs to 500+ languages and studying how data mixes shape cross-lingual transfer.

🔮 03

Instruction Fine-tuning

How many languages, and which ones, make multilingual instruction tuning work.

🛠️ 04

Evaluation

Consistent, non-English-centric benchmarking of LLMs across dozens to hundreds of languages.

Our works

Projects & papers

New project · 2025

FineOPUS

A massively multilingual translation corpus and pipeline, refining the OPUS collection of parallel texts in many languages.

Read more
Continual pretraining

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

The MaLA bilingual translation corpus in 2,500+ language pairs and the EMMA-500 Llama 3/3.1 models.

Read paper
Continual pretraining

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models

The EMMA-500 Llama 2 model and the MaLA monolingual corpus in 939 languages.

Read paper
Continual pretraining

MaLA-500: Massive Language Adaptation of Large Language Models

Vocabulary extension and continued pretraining on Llama 2 with Glot500-c to cover 534 languages.

Read paper
Continual pretraining

Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

36 CPT configurations comparing monolingual, bilingual and code-augmented data across 30+ languages.

Read paper
Evaluation

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

A lightweight framework for seven tasks spanning dozens to hundreds of languages.

Read paper
Instruction fine-tuning

How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM

The Lucky 52 models, studying the number of languages for instruction fine-tuning.

Read paper
Instruction fine-tuning

Monolingual or Multilingual Instruction Tuning: Which Makes a Better Alpaca

Comparing monolingual and multilingual instruction tuning under a fixed budget.

Read paper

Bring your language.

Our corpora, models and code are open. Join the community to ask questions, share results or contribute data for the languages you care about.