Chinese Military Researchers Used US AI Outputs
Chinese defence-linked researchers used GPT-3.5 and Claude outputs to train local models for code summaries, content monitoring and drone imagery.
首次合成约需 20 秒,之后再访即点即听
Papers published by a People's Liberation Army unit and defence-linked universities show that Chinese researchers used outputs from GPT-3.5 and Claude to train smaller domestic models. Reuters, after reviewing more than 80 papers and patents, said the work covered code summarisation, social-media monitoring and drone imagery. Public records do not show that the resulting systems entered military service.
The common technique in these studies is known as knowledge distillation. Researchers prompt a more capable model, collect its answers and use them to train a smaller system for a narrower task. Once training is complete, the smaller model no longer needs a continuous connection to the original provider, making it suitable for closed networks or edge hardware such as drones.
One paper, “Code summarization based on large model knowledge distillation,” lists PLA Unit 96941 as the affiliation of every author. The researchers asked GPT-3.5 to write one-sentence descriptions of Java methods, then used 2.15 million code-and-summary pairs to train a model called Jam. Its 350-million-parameter version could run on a single 16GB consumer GPU. When 15 programmers compared the outputs, 52% preferred GPT-3.5 and 46% preferred Jam.
The experiment used funcom-java-long, a public dataset of Java methods and human-written summaries. The paper does not show that sensitive military source code was uploaded to OpenAI, nor does it say that Jam was connected to an operational PLA system. The test asked whether code-summarisation ability could be transferred from GPT-3.5 to a smaller model so that new code could later be processed locally.
Reuters, citing public records and research by the Washington-based Jamestown Foundation, linked Unit 96941 to intelligence and cyber-operations research. The paper itself gives the authors' names, unit number, methodology and results, but does not describe the unit's duties.
Another paper reviewed by Reuters came from North University of China. The university says it is administered by the Shanxi provincial government and jointly supported by the Ministry of Industry and Information Technology and the State Administration of Science, Technology and Industry for National Defense. Reuters reported that its researchers used Anthropic's Claude 3 Haiku to generate synthetic data for a model designed to classify, monitor and moderate social-media text. Ruibao has not obtained the full paper, and the available material does not show that the system was used in live surveillance.
Reuters also reviewed material from researchers at the PLA's National University of Defense Technology. It said they used distillation to shrink image-processing models for hardware such as drones, where they could assist navigation and target recognition without a network connection. Ruibao has not obtained the full paper or patent cited in the report. Public records do not establish that the system was fielded or used in combat.
Reuters found no evidence that OpenAI or Anthropic supplied model weights to the Chinese military or entered a dedicated partnership with it. The training material described in the documents came from model outputs. A smaller model can imitate one narrow skill, such as code summarisation or text classification, after seeing enough examples; it does not thereby acquire the full capabilities of a general-purpose model.
US technology controls on China have largely focused on advanced chips, manufacturing equipment and high-performance computing services. Text and images returned by commercial models through websites or application programming interfaces are harder to track individually. Once those outputs have been assembled into a training set, a local model can operate without reconnecting to the original provider, leaving that company with little ability to police its use through account controls or online safeguards.
Knowledge distillation is widely used across the AI industry. The dispute turns on how a model is accessed, the scale of extraction, service terms and the eventual use of the resulting system. Anthropic this year accused several Chinese AI companies of using fraudulent accounts to make large numbers of Claude queries and gather training data. Responding on July 27 to a separate dispute involving commercial models, China's Commerce Ministry said US accusations lacked a factual and legal basis and criticised Washington for threatening sanctions over distillation. That response did not address the military-linked papers reviewed by Reuters.
The military unit and universities named in the report had not issued public responses by publication time. The papers also omit a complete account of their training data, downstream users and deployment dates. The public record shows that outputs from US commercial models entered Chinese defence-related research. Whether the prototypes became formal military systems remains unverified.
Source note: This report draws on Reuters' July 31 review of more than 80 Chinese papers and patents, the Unit 96941 code-summarisation paper and its published results, North University of China's institutional profile, and Reuters' reporting on the US-China dispute over model distillation. Ruibao verified that the Unit 96941 experiment used a public Java dataset; it does not support the claim that GPT-3.5 processed sensitive military source code. Details and evidentiary limits for the North University of China and National University of Defense Technology projects follow the Reuters investigation.
You read this far. You're not here for noise.
SharpPost delivers one weekly deep dive on geopolitics, finance, and tech — decoded for readers who want signal. No ads, no filler.