US AI Models Give Uneven Answers on China
Tests show uneven AI refusals involving China and other restrictive governments, while state-aligned Chinese training data can influence model answers.
免费微软自然语音 · 推荐使用 Microsoft Edge
Researchers asked Claude to produce a protest flyer criticizing US President Donald Trump, and the model complied. When the target was changed to Chinese President Xi Jinping, the Saudi crown prince or Thailand’s king, it refused. The Oversight Board, an independently operated body funded by Meta, said in a test published in July that similar disparities appeared across several large language models developed by US companies.
In March, the board tested 10 commercial models from Anthropic, Google, Meta, OpenAI, xAI and China’s DeepSeek. Researchers issued English-language prompts from Australian IP addresses, while most of the cloud infrastructure was based in the United States. Optional content filters available through Google and Microsoft platforms were disabled. The models were asked to create protest flyers and satirical poems, or to assess whether users should support or protest political leaders and institutions in different countries.
The average refusal rate was 34% for prompts involving jurisdictions with tighter limits on political expression, compared with 14% for more permissive jurisdictions. Claude Sonnet 4 showed the largest gap, refusing 59% of prompts in the restrictive group and 16% in the permissive group. Gemini 3 Pro and Llama 4 Maverick each showed gaps of roughly 30 percentage points. GPT-5.2 recorded nearly identical rates, at 24% and 23%, while Gemini 3 Flash and Grok 4 Fast did not refuse any of the tested prompts.
Taiwan was classified among the more permissive jurisdictions, but several models still refused Taiwan-related criticism at relatively high rates. The disparity was most pronounced in Claude’s results, suggesting that a model’s treatment of a political subject does not always follow the legal environment assigned to it by researchers.
When Claude Opus 4 declined to make a protest flyer criticizing the Chinese government, it said the material could expose the user to danger or involve the model in sensitive political activity. The same model generated material criticizing Trump and Britain’s King Charles III. The researchers cautioned that a model’s explanation for refusing a request is not reliable evidence of how its safeguards were designed. The result therefore does not establish that Anthropic created a China-specific censorship rule.
A separate study published in Nature in May found that the language of a prompt can also change the answer. Researchers assembled 822 Chinese-language questions about Chinese politics and compared responses from ChatGPT and Claude in Chinese and English. Asked in English whether China is democratic, ChatGPT said the country is generally not considered a democracy. Asked in Chinese, it said the answer “depends on how democracy is defined.” Across the sample, the Chinese-language responses from both US models were more favorable toward the Chinese government.
The researchers identified 3.1 million Chinese documents in the open-source CulturaX training corpus that closely resembled content from state-coordinated Chinese media. These documents accounted for 1.64% of the Chinese-language corpus—a match rate about 41 times that of Chinese Wikipedia. Among documents mentioning Chinese leaders or political institutions, the share reached as high as 24%. Only about 12% of the matched material came directly from known government or news domains; the rest had spread through reposting, aggregation and ordinary websites.
After adding such material to the pretraining data of an open model, researchers found that nearly 80% of its responses became more favorable to the Chinese government than those of the original model. This kind of training-data contamination does not necessarily amount to deliberate “poisoning.” But once heavily repeated narratives lose their source labels, a model may treat multiple copies of the same official account as independent evidence.
The Oversight Board’s test likewise does not show that US AI companies deliberately censor on behalf of authoritarian governments. It demonstrates a significant association between model outputs and local restrictions on political expression. Safety testing conducted only in English can miss softened positions in Chinese-language answers. Measuring only outright refusals can also overlook “soft censorship”: responses that appear to answer but avoid facts or reproduce an official line.
Methodology note: The Oversight Board tested 10 commercial models in March 2026. The training-data findings are based on research by Waight and colleagues published in Nature. Model behavior can change with versions, prompts and platform settings; these results do not represent every response produced by the products today. Sources were checked through August 15, 2026.
You read this far. You're not here for noise.
SharpPost delivers one weekly deep dive on geopolitics, finance, and tech — decoded for readers who want signal. No ads, no filler.