Image source: Public Domain
LILT launched AURORA, the industry's first Multilingual AI Leaderboard that measures frontier models on non-English agentic, multimodal, and socio-cultural tasks through scientifically rigorous evaluation.
Enterprises are deploying agents worldwide, yet nearly every published measure of frontier progress relies on English-centric or translated benchmarks. With agents doing customer-facing work, performance in every language now carries commercial weight.
"The industry is choosing models on a scoreboard that stops at English. That gap used to cost you an awkward translation. Now that agents are taking real actions in the real world, it costs you a wrong decision, and you find out from your users." — Spence Green, CEO and co-founder, LILT
Measuring Multilingual Performance:
AURORA addresses the multilingual gap by testing models against LILT's multilingual benchmark suite featuring tasks designed and verified by native-language domain experts. These tasks represent real-world enterprise applications, such as software development, customer support, and complex workflows, each grounded in language, region and culture.
Specifically, AURORA provides visibility across LILT's multilingual benchmarks, including:
AURORA is developed and managed by LILT's Applied AI practice, a PhD-led research team with 10+ years' experience in multilingual AI.
Their latest analysis shows model quality can differ substantially between languages. For example, in coding, GPT 5.5 performs best in Spanish, Claude Opus 5.5 wins in Japanese, while Muse Spark 1.3 leads in Serbian.