Exclusive: OpenAI slows release of Astra model citing cyber capabilities
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios
OpenAI "cannot rule out" that its upcoming model Astra has "critical" cyber capabilities, a designation that has prompted the company to expand safety testing and pause internal activities that do not meet stricter security requirements, OpenAI told Axios first on Friday.
Why it matters: It's the latest sign of rapidly advancing cyber capabilities from AI models, after others worked autonomously outside of testing sandboxes and protections.
Driving the news: OpenAI said "we cannot rule out critical cyber capabilities" after running internal evaluations of Astra, one of its upcoming models.
- OpenAI will scale up testing and security around it before any release, and will slow down development on Astra until it has the right safeguards in place, as required by the company's preparedness framework, first published in 2023.
- Astra was not involved in the Hugging Face exploits, the company said.
- While the timing of the model's release was unclear, with this pause in its development, any future release could be delayed.
- "OpenAI voluntarily informed the administration of their plans to delay the release," a White House official said.
Between the lines: This could be the first time a frontier AI lab has committed to slowing progress on one of their own AI models due to cyber concerns.
- Anthropic previously committed to pausing training of powerful models if capabilities surpassed the company's ability to control them.
- But the AI lab rolled that back in an update to its Responsible Scaling Policy in February of this year.
- "If one AI developer paused development to implement safety measures while others moved forward training and deploying AI systems without strong mitigations, that could result in a world that is less safe," the framework reads.
The big picture: The announcement comes as the Trump administration works to develop a process for evaluating AI models before their release.
- Select industry were briefed on a framework this week but there are still a lot of unanswered questions for companies.
- For example, how to engage the government, how long the review process will take, what the government and industry hope to learn from it.
- Also: who has access to or reviews the models.
- What constitutes as sufficient national risk and state of the art models was operationalized in the framework but not defined.
Flashback: Competing AI lab Anthropic released a safer version of its most cyber-capable model, Mythos, in June.
- Dianne Penn, Anthropic's head of product management, research and labs, told Axios at launch that the company was being "deliberately more conservative" with that release.
- Anthropic warned about models improving themselves in a company blog in June that also called for a global pause in AI development.
Between the lines: Earlier this week at the Black Hat cybersecurity conference, members of OpenAI's technical staff said the company was slowing down testing while it works on upgrading its security practices.
- In the blog post Friday, OpenAI said it's started implementing stricter security controls for testing, including isolated testing environments and universal monitoring across agentic applications of Astra."
- OpenAI has started "consciously slowing down research to enhance security," Michael Dalton, a member of OpenAI's technical staff, said during a presentation.
The bottom line: AI models are getting better and more cyber capable faster than the regulation around their use is formalizing.

