
Google's 1,000 Languages Initiative: what it has delivered
January 06, 2023There are more than 7,000 languages in the world, yet most software, and most AI, works well in only a small fraction of them. Google and its competitors want their products to reach the widest possible audience, and in November 2022 Google made that goal concrete with the 1,000 Languages Initiative: a commitment to build AI models that support the 1,000 most spoken languages. At the time, Google Translate handled 133.
The plan announced in 2022
Google presented the initiative at an AI event on 2 November 2022 and described it as a project that would take many years. The idea was to train a single model on many languages at once, so that languages with plenty of data help those with very little. As a first step, Google showed a speech model trained on more than 400 languages, and it said it would fund the collection of audio recordings and written texts in low-resource languages. It expected the work to feed into products from Google Translate to YouTube captions.
What Google has delivered since
In March 2023 Google published the details of that speech work, the Universal Speech Model (USM): a family of speech models with 2 billion parameters, trained on 12 million hours of speech and 28 billion sentences of text spanning more than 300 languages. It was built for use on YouTube, for example for automatic captions, and can recognise speech in under-resourced languages such as Amharic, Cebuano, Assamese and Azerbaijani.
Google Translate followed. In June 2024 it added 110 languages at once, its largest single expansion, with help from the PaLM 2 large language model. The new languages, from Cantonese and Afar to Manx and Tok Pisin, have more than 614 million speakers, about 8% of the world's population, and about a quarter of the new languages come from Africa. By September 2026 Google said Translate covered more than 250 languages and that its products overall work in more than 300.
Open models carry the work further. Gemma 3, released in March 2025 as part of Google's open Gemma family, was pretrained on more than 140 languages and supports over 35 out of the box. TranslateGemma, a set of lightweight open translation models trained across 55 languages, runs on the device, without a connection to the cloud. For speech, Google says Gemini 3.5 Live Translate handles real-time spoken translation in 70 languages.
The 1,000-language target itself has not been reached yet. To get closer, Google now relies on community data projects such as WAXAL, an open speech dataset for 27 Sub-Saharan African languages, and Project Vaani, which has collected more than 30,000 hours of speech in 109 Indian languages.
Meta is working on the same problem
Google is not alone. In July 2022 Meta released NLLB-200 (No Language Left Behind), an open model that translates between 200 languages. In May 2023 its Massively Multilingual Speech project published speech-to-text and text-to-speech models for more than 1,100 languages and language identification for more than 4,000.
What it means for app localisation
For product teams, these models make it realistic to ship an app in languages that once required a specialist agency or were skipped altogether. Machine translation is now a sensible first draft for interface strings, help centres and support replies. The Google Cloud Translation API, the Gemini API and open models such as Gemma can be built into a localisation pipeline, and open models are an option when text has to stay on the device or on your own servers.
Quality is still uneven. Output for low-resource languages is usually less reliable than for widely used ones, and idioms, plurals, grammatical gender and levels of formality remain common sources of error. Human review is still essential for legal text, payments, health information and marketing copy, and translations should be checked in context on real screens, where truncated strings and right-to-left layouts cause problems of their own. Our article on localisation in mobile app testing covers that side of the work.