
Daniel Eriksson, Executive Director of the Meta Oversight Board, explains why opaque AI refusals, culturally blind safeguards and jurisdictional spillovers risk shaping public debate — and what it will take to build generative AI systems that are transparent, context-aware and accountable.
Generative AI is no longer just answering customer queries or drafting internal documents. It is increasingly shaping what people can ask, criticise, learn and say about the world around them. That makes the governance question more urgent than the familiar debate over whether a chatbot gives a correct answer.
In its first assessment of leading large language models, the Oversight Board found that models refused requests for political criticism of leaders in restrictive jurisdictions 34% of the time, compared with 14% for comparable requests about permissive ones. The prompts were made from Australia — not from the countries under discussion.
The question is no longer whether AI systems shape public debate. They already do. The more pressing question is whether they are importing restrictions on political speech into places where those restrictions do not apply.
What happens when a model quietly refuses to help someone criticise a political leader? What if that refusal reflects a restriction introduced for a more repressive jurisdiction, even when the user is sitting in London, Berlin or New York? And how can organisations tell the difference between a necessary safety control and a rights-limiting constraint they cannot see?
For Daniel Eriksson, Executive Director of the Meta Oversight Board, the industry does not need to start from scratch. The Board’s experience scrutinising high-stakes content moderation decisions offers a serious, if imperfect, blueprint for how AI companies should approach generative systems: build in oversight, test for real-world harms, and be candid with users when responses have been limited.
“There is a need for some kind of oversight with regards to the user rights and human rights of the users, both ex post and in design,” Eriksson says.
That distinction matters. Governance cannot be something applied after a public failure. It has to shape the design of systems before they reach millions of people — and provide a route to challenge outcomes once they do.
Why content moderation lessons matter for generative AI
The Oversight Board reviews some of Meta’s most difficult content moderation decisions — cases where a seemingly simple policy rule can collide with language, culture, safety and fundamental rights.
A revealing example is the Board’s work on Meta’s treatment of the Arabic term shaheed, often translated as “martyr”. Under Meta’s Dangerous Organisations and Individuals policy, use of the word could trigger removal. Yet shaheed has multiple meanings and uses: it can appear in political or religious contexts, refer to someone who died a suffering death, or simply be a person’s first or family name.
“There was a lack of knowledge of those setting up the policies that this term can be used in many different ways,” Eriksson says. “There’s many interpretations of the context.”
The Board recommended that Meta should not remove content merely because it used the word. Instead, it should act when the term appeared alongside material that implied violence or threats. The resulting policy change increased legitimate use of the term by more than 20%, Eriksson says, allowing millions of pieces of content previously caught by an overly broad rule to remain available safely.
The lesson for LLM developers is not simply that language is complicated. It is that a policy can look clear in a spreadsheet and still fail badly in the world. Context is not an edge case. It is the work.
For enterprises deploying generative AI across markets, that should change the governance agenda. Model safety cannot be assessed only through broad benchmark scores or English-language testing. Teams need to ask which terms, identities, political contexts and linguistic conventions their systems might flatten into a single risky category.
How generative AI changes the accountability challenge
Moderating a social media post and governing a generative AI response are related problems, but they are not identical.
With platform content, there is a defined object to assess. A person created a post, a platform acted on it and an appeals process can examine the decision. With an LLM, the output is produced dynamically, can vary between prompts, and may be shaped by training data, system instructions, safety filters, jurisdiction-specific controls and product design choices that are not visible to the user.
That makes evaluation harder — and more necessary.
Eriksson points to the Board’s recent research testing how LLMs respond to requests for critical content about leaders in repressive regimes. A single prompt does not prove a pattern. Researchers need to test many versions of a request, repeatedly, to establish whether a model systematically refuses, alters or steers responses in a particular direction.
“The trial and the testing of it needs to be more rigorous than what we can do with simple content,” he says. “That doesn’t mean that it’s not needed, because the breach of user rights still exists.”
This is a useful corrective for leaders tempted to treat a model’s policy page as evidence of how it behaves. It is not enough to read the stated rules. You need repeatable, scenario-based testing of the product in use — across languages, locations, user profiles and high-stakes topics.
What the Oversight Board’s LLM research found
The Oversight Board’s first evaluation of large language models raises an uncomfortable question for every company building or deploying generative AI: are systems designed to protect users from harm also reproducing political restrictions far beyond the places where those restrictions apply?
The Board tested 10 commercial LLMs from six providers: Anthropic, DeepSeek, Google, Meta, OpenAI and xAI. Researchers asked them to create politically critical material — including protest flyers and satirical poems — about governments and political leaders in 10 jurisdictions. The models were accessed through commercial APIs hosted primarily on US infrastructure and queried from an Australian IP address. (Read the Oversight Board’s research).
The finding was striking. On average, the models refused 34% of requests for politically critical material concerning jurisdictions classified as restrictive of free expression, compared with 14% of comparable requests concerning permissive jurisdictions.
The restrictive group included Cambodia, China, Saudi Arabia, Thailand and Turkey. The permissive comparison group included Chile, Japan, Taiwan, the United Kingdom and the United States.
The study does not establish why individual models produced those differences. It was deliberately limited, testing foundation-model outputs rather than consumer chatbot interfaces and using English-language prompts from one Australian IP address. But that is precisely why the pattern warrants scrutiny: it emerged in a consistent testing environment outside the jurisdictions whose political-speech restrictions were being assessed.
That does not prove that every refusal resulted from an explicit government demand, nor does it establish that providers deliberately designed their products to censor political criticism. The Board is careful on that point, and so is Eriksson. But the pattern is difficult to dismiss as incidental: the models were materially more reluctant to generate political criticism when the target was a government that itself restricts that sort of speech.
“If you’re sitting in the US or Germany asking the bot to make critical material against the leadership in Thailand, even though it’s fully legal where you sit, you are more likely to get a no or an altered answer,” Eriksson says.
The significance lies in the user’s location. A person in a country with strong free-expression protections may encounter a refusal shaped by rules, risks or assumptions associated with a wholly different jurisdiction — without any clear indication that this has happened.

How censorship by proxy can cross borders
The Board describes this risk as “censorship by proxy”: model outputs may extend speech restrictions associated with repressive regimes to users elsewhere, including in countries where the same expression is lawful.
It is a more subtle problem than a government ordering a platform to remove a specific post. Generative systems operate through interlocking layers: training data, model fine-tuning, safety systems, policy rules, prompts, content classifiers and product decisions. Any one of those layers may be intended to manage legal, commercial or safety risk. Together, they can create a response pattern that is difficult for users — and sometimes providers — to explain.
Eriksson does not suggest that AI companies are necessarily acting with political intent. His concern is precisely that the effect may be unintentional.
“We assume that through training data or tweaking or whatever to adjust to local legislation in oppressive regimes, that has a spillover, a proxy effect on other open governments,” he says.
That is the governance problem. A system can appear neutral because it applies the same interface and policy language to everyone, while still carrying different political constraints into the answers it provides. A refusal may arrive as a generic safety message. An answer may be softened, abbreviated or redirected. In each case, the person using the tool may have no way to know whether they have encountered a genuine safety limit, a provider policy, a local legal rule or a restriction that has travelled across borders.
The democratic risk is not only that a user is denied one particular output. At scale, invisible constraints can shape what people ask, what they learn, how they frame political questions and which forms of criticism appear acceptable. For systems rapidly becoming an interface for information and public debate, those defaults matter.
“What else is there that we don’t know?” Eriksson asks. “Could there be other limitations or alterations that exist even in an open society that are spillovers?”
.png?width=1200&height=500&name=Generative%20AI%20and%20Free%20Speech%20The%20Oversight%20Gap%20(1).png)
Transparency must go beyond a chatbot’s explanation
A chatbot may explain a refusal with polished confidence. That does not mean the explanation reliably describes what happened inside the system.
The Board found that some models cited broad, or apparently invented, policies that were not applied consistently. Others offered plausible-sounding legal, safety or policy rationales that did not necessarily reveal the mechanism behind the refusal.
That matters because a fluent explanation can create the impression of accountability without providing it. If users, researchers and regulators cannot distinguish a safety control from a jurisdictional restriction or an internal policy choice, they cannot meaningfully scrutinise the system’s impact on free expression.
Eriksson argues that users should receive clear notice when a response has been altered or limited because of regulation in a particular jurisdiction. Not every restriction is illegitimate. There are laws and safety requirements that platforms must meet. But users should not be left to infer whether an answer is incomplete, unavailable or framed differently because of a local legal requirement.
“It is more fair if the users are then told that you are getting an altered reply because of local regulation,” he says.
Instead of vague refusals that simply say “I can’t help with that”, product teams should tell users which broad category has shaped the response — and, where a legal restriction is involved, identify the relevant jurisdiction. At a minimum, a disclosure should distinguish between:
- A safety restriction designed to prevent harm
- A legal or jurisdiction-specific constraint
- A formal government request or informal government pressure
- An explicit product-policy decision made by the provider
- A limitation arising from uncertainty or an inability to provide a reliable answer
The disclosure does not need to expose proprietary details or hand bad actors a roadmap around safeguards. It does need to give users enough information to understand that an invisible rule has shaped the response — and, where appropriate, provide a route for challenge or review.
A practical governance checklist for AI leaders
The Oversight Board’s experience suggests that governing generative AI is not about finding one perfect policy. It is about building the capability to detect, explain and correct harm. For leaders building, buying or deploying LLM-enabled products, six actions stand out.

The aim is not unrestricted output. Eriksson is explicit that some limitations are needed. The point is that restrictions should be proportionate, context-aware and visible enough to be scrutinised.
The governance test for the generative AI era
The Oversight Board’s research does not offer a final account of why models behave this way. Nor does Eriksson claim that the industry already has a ready-made remedy. But it establishes a serious governance question: when a model limits political speech, users deserve to know whether that limit exists to prevent harm, comply with a law, reflect a provider’s policy or manage a constraint imported from elsewhere.
“The fundamental principles still hold true,” Eriksson says. “But the way that this should be implemented and then scaled is very different, fundamentally different. We don’t have the concrete answers here yet. Leaving it to the chatbots to explain what’s happening is insufficient.”
For AI leaders, the immediate task is to make the constraints shaping a model’s behaviour testable, challengeable and legible to the people affected by them. As generative AI increasingly mediates public conversation, transparency is not a nice-to-have product feature. It is part of the democratic infrastructure.
InspiredMinds! Community Hub
Inside the InspiredMinds! Community Hub, you’ll find deeper insights, expert perspectives, and practical discussions from the people building and deploying AI across Europe and beyond.
Join our regular LinkedIn newsletter - Global AI Dispatch!
World Summit AI global Summit series
5 – 9 October 2026
Amsterdam, Netherlands
World Summit AI Qatar
World Summit AI Canada
World Summit AI USA
.png?width=259&name=WSAI%20Amsterdam%20Orange%20no%20dates%202000x300%20(1).png)
.png?width=263&name=IM_Mothership_assets_LOGO_MINT%20(2).png)

