OpenAI said on September 1 that its forthcoming Astra model is the first system the company has classified at the Critical cybersecurity capability level, prompting it to restrict the model's most advanced security functions when release begins. The company plans to make Astra available soon, but said advanced cyber work will initially be limited to a small group of testers before access expands through its Daybreak Blue program for defenders.
Under OpenAI's Preparedness Framework, the Critical designation means a model can find previously unknown weaknesses and develop functional exploits across many hardened systems without step-by-step human guidance, or devise and execute a novel end-to-end attack from a high-level goal. OpenAI said Astra crossed that threshold after automated benchmarks, private tests and expert-led assessments showed a substantial advance over GPT-5.6 Sol.
The company reported that Astra scored 100 percent on ExploitBench, which tests exploit development from known vulnerabilities. On an internal set of 20 recently disclosed high-severity flaws in the V8 JavaScript engine, Astra achieved higher arbitrary-code-execution rates while using fewer output tokens than GPT-5.6 Sol. During testing, it also found and used two previously unknown vulnerabilities in one exploit chain; OpenAI said it is disclosing those flaws to their maintainers.
In separate expert exercises, Astra built a browser-compromise chain that escaped a sandbox and executed commands on a host computer after an HTML file opened. It also combined several flaws in a hardened operating system to elevate access from an ordinary user to root. OpenAI stressed that the published results reflect Astra with Daybreak Blue access rather than the default configuration planned for general users.
OpenAI said it delayed parts of Astra's development and release while adding stronger refusal training, system-level classifiers, account risk controls and monitoring designed to stop unauthorized activity. The model refused 91.5 percent of requests in the company's cyber-jailbreak evaluations, compared with 59 percent for GPT-5.6 Sol. OpenAI also paused some frontier training for two weeks after a separate Hugging Face incident, although it said Astra was not involved.
The safeguards may slow or stop legitimate defensive work, OpenAI acknowledged. In ChatGPT or Codex, a flagged action may require user review, while an API task may stop entirely. Axios reported that the restricted rollout highlights a broader challenge for increasingly autonomous AI: the same abilities that help defenders locate serious weaknesses can also make attackers more effective. OpenAI has not provided a specific launch date and says fuller safety and evaluation details will accompany Astra's system card at release.
Comments