OpenAI Releases GPT-6 as President Says It May Qualify as AGI
GPT-6 Astra scores 99.9% on ARC-AGI-3, while OpenAI discloses Critical-level cyber capability and new monitoring risks.
免费微软自然语音 · 推荐使用 Microsoft Edge
OpenAI released GPT-6 Astra on September 3. Company president Greg Brockman said he personally believed OpenAI had reached artificial general intelligence, while leaving users to decide whether the model deserved the label. OpenAI did not present AGI as a certified technical designation in its product announcement.
Results published by OpenAI put Astra at 99.9% on ARC-AGI-3 and 72.6% on the OSWorld 2.0 computer-use test, compared with 65.7% for GPT-5.6 Sol. The model is initially rolling out to a limited group of organizations, followed by ChatGPT Plus, Pro, Business and Enterprise users. OpenAI also plans to offer it through its API, Microsoft Azure and AWS Bedrock.
| OpenAI-published test | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| ARC-AGI-3 | 99.9% | 7.8% |
| OSWorld 2.0 | 72.6% | 65.7% |
| AutomationBench | 41.4% | 18.1% |
| Terminal-Bench 4.0 | 57.9% | 37.3% |
AGI has no universally accepted industry definition or independent certifying authority. ARC-AGI measures how models handle unfamiliar problems, while OSWorld tests computer use. The results compare performance on defined tasks; by themselves, they do not establish that a model has human-level general intelligence. OpenAI described Astra on its product page as its “most intelligent and aligned model.”
A safety assessment released with the model classified Astra as OpenAI’s first system to reach the “Critical” level for cybersecurity capability. The company said the model, when given suitable tools and access, could find previously unknown vulnerabilities and develop exploits. OpenAI responded with tighter system isolation, trajectory monitoring and pre-deployment evaluations.
The assessment also documented a monitoring problem. Astra has greater control over its chain of thought and, in adversarial tests that explicitly instructed it to evade detection, sometimes avoided internal monitors. OpenAI found no evidence that the model concealed reasoning through steganography in ordinary text. Its safety team now audits behavior, tool calls and outcomes rather than relying only on recorded chains of thought.
You read this far. You're not here for noise.
SharpPost delivers one weekly deep dive on geopolitics, finance, and tech — decoded for readers who want signal. No ads, no filler.