
Multilingual AI has a longitude problem. For centuries, sailors could fix their latitude from the stars, but not their longitude, until a Yorkshire clockmaker built a timepiece precise enough to survive the ocean. Marzieh Fadaee, Head of Cohere Labs, has described her own research in almost identical terms: she is "exploring the longitude problem of AI" (X), the puzzle of giving language models a genuine sense of place when nearly everything about how they are built, measured and rewarded still points due English.
The scale of that problem is easy to underestimate. English is the majority language on only 49.5% of the web's identifiable content (W3Techs), yet it remains the default lens most frontier labs use to judge whether a model is "intelligent." There are roughly 7,159 living languages in active use worldwide (Ethnologue), and even the most ambitious multilingual models typically cover no more than about 100 of them (LMU Munich research). Fadaee, who took over as Head of Cohere Labs in September 2025 after a decade researching machine translation (BetaKit), joined an interview with WSAI’s Editor, Fawn Hudgens, to explain why closing that gap is not simply a matter of scaling up, and what leaders building AI products right now should be doing differently.
Why English can't be AI's proxy for intelligence
Fadaee's starting point is a habit she says almost every lab at the frontier shares, often without noticing it. "English and the English language can be a universal proxy for intelligence," is the implicit assumption, she says, and "the value that English data and the English language brings is undeniable." The problem is what gets lost when that becomes the only measuring stick. "Languages are not just wrappers around intelligence," she argues. "What is essentially captured in different languages around the world, in my opinion, is all of this history of human knowledge and intelligence in different contexts." Build for English alone and you are not just missing some vocabulary, you are missing entire parts of the distribution of human knowledge the model was supposed to learn from.
How leaderboards hide AI's long tail of language failures
Fadaee's second warning is aimed squarely at how the industry keeps score. Averaging performance into a single leaderboard number, she says, is useful for comparison but dangerous for multilingual systems specifically, because "in this sort of simplification and averaging of averaging... what you lose is the long tail of performance." Her clearest example is Cohere Labs' own Tiny Aya family, a compact, 3.35-billion-parameter open-weight model line covering more than 70 languages, released in February 2026 to run offline on everyday devices (TechCrunch; arXiv). Chasing a single benchmark score would have pushed the team to over-sample the handful of languages evals reward best, undermining the actual goal of broad coverage. Instead, the team deliberately reported the messier, more nuanced results. "People appreciated that sometimes this simplicity isn't, shouldn't be, the end goal," she says. Researchers at universities, startups and big tech companies told her it "was a breath of fresh air" to see that bird's-eye view rather than a tidy, digestible headline number.
How to build AI for an underserved language on a tight budget
Asked how a small team without a frontier-lab budget should approach building for one underserved language ethically, Fadaee's answer is refreshingly unglamorous:
- Start with evaluation, not training. Build a small but high-quality, representative eval set for the target language and the specific tasks real users need, ideally with native speakers verifying it in the loop, since this becomes what she calls "the North Star of developing the model."
- Use machine-translated data to bootstrap. It has limitations and "there will be a point that then you want to go beyond translation," but for many languages it remains "a really good starting point."
- Lean on existing open checkpoints rather than training from scratch. Scale unlocks real capability, but "that doesn't have to be the starting point."
- Treat scale as the second question, not the first. Once the recipe works at a smaller size, then decide how to transfer it upward.
She is just as direct about what to avoid. The first trap is trusting your eval too much: "you should revisit what are you really building towards," she says, "rather than making the assumption that this is one fixed goal that we should reach." The second is uncritical reliance on synthetic data. Nearly every model pipeline uses it now in some form, but a recipe that worked six months ago "might not work right now," and teams that never revisit how they generate it or train reward models on it often only spot the resulting problems much later than they should.

Why open science is critical as AI labs close their doors
Cohere Labs has built its reputation on open research, and Fadaee ties that directly to the multilingual mission. "If your goal is to build global AI, you cannot really do that with only one cultural, geographical perspective," she says. Compute is inevitably centralised, but "the expertise and the knowledge that we are trying to capture is very distributed," which is precisely what open science is designed to pool. She is candid that the tide has been moving the other way. Openness has been in retreat "more and more the last couple of years," she says, even though it "brings its own accountability and transparency and safety" that addresses many of the concerns closed labs cite for staying shut.
Beyond translation: what cultural intelligence means for AI
Fadaee's most forward-looking point is that fluent output is no longer the bar. Even in well-supported languages, models can stay "at this surface level" rather than grasping "the nuances of that particular culture and language," where the right answer often depends on context-specific norms, values and social expectations. Her test for progress is simple and human: when a native speaker uses the technology in their own language, "how much do they feel like it's a little bit weird and unnatural versus how much they feel like it captures all of these specific nuances that comes from my language?" Sometimes the gap is a missing fact about a culture's history or literature; sometimes it's a verb that is semantically correct but socially off. Both are hard to measure, and both matter.
The biggest misconception about building AI for a global audience
Pressed on what people get most wrong, Fadaee doesn't hesitate: "Maybe that scaling solves everything." Bigger models will keep mattering, and reasoning capability will keep expanding beyond maths and code into softer judgement calls about context and appropriateness. But representation gaps, whether by language, culture or behaviour, need deliberate design choices at every stage, not just more parameters. "Bringing diversity into every step of the way, whether it's about your eval, your data, just like the annotators that you use," she says, "I think that's also quite undervalued."
For any team building AI products for a genuinely global audience, that's the practical takeaway to sit with: know which languages your benchmark is quietly ignoring, keep interrogating the eval you built six months ago, and remember that "good in English" was never the finish line. Fadaee takes this argument to the World Summit AI stage in October.
InspiredMinds! Community Hub
Inside the InspiredMinds! Community Hub, you’ll find deeper insights, expert perspectives, and practical discussions from the people building and deploying AI across Europe and beyond.
Join our regular LinkedIn newsletter - Global AI Dispatch!
World Summit AI global Summit series
5 – 9 October 2026
Amsterdam, Netherlands
World Summit AI Qatar
World Summit AI Canada
World Summit AI USA
.png?width=259&name=WSAI%20Amsterdam%20Orange%20no%20dates%202000x300%20(1).png)
.png?width=263&name=IM_Mothership_assets_LOGO_MINT%20(2).png)


