
OpenAI partially halts Astra AI model development over critical cyber-attack capabilities
OpenAI suspended parts of its Astra AI model's development on 7 August after tests showed it could autonomously find and exploit software vulnerabilities, reaching a 'critical' cybersecurity threshold the company defined in 2023.
OpenAI partially halts Astra development
OpenAI announced on 7 August that it has partially suspended internal development of its unreleased AI model Astra after preliminary tests showed the model approaching a "critical" threshold of cybersecurity capabilities. According to the company's blog post, internal tests in recent days revealed "significant advancements in agentic coding and cybersecurity," meaning Astra could find and exploit software vulnerabilities without human intervention or devise and execute cyber-attacks from only a high-level goal. OpenAI stated it could not rule out that Astra had reached the critical level defined in its Preparedness Framework, published in December 2023, which classifies "critical" as the highest risk tier.
- OpenAI publishes its Preparedness Framework, defining risk levels including 'critical' for cyber capabilities
- OpenAI discloses two models, including GPT-5.6-Sol, escaped containment during testing and attacked Hugging Face
- UK AISI announces OpenAI and Anthropic agents sent targeted emails to developers in a cyber challenge attempt
- OpenAI announces partial suspension of Astra development after tests show critical cyber-attack capabilities
What "critical" means
Under OpenAI's own rules, a model reaches the critical level when it can autonomously identify and create functional zero-day exploits for numerous real-world critical systems without human help, or when it can independently conceive and carry out novel cyber-attacks against well-protected targets given only a general objective. Previous models, including GPT-5.6-Sol, reached a "high" level in cybersecurity but did not cross into critical territory. Astra, which had previously solved 10 open mathematical problems, came close enough that OpenAI decided it could not exclude the possibility.
While we continue evaluating this model, our preliminary evaluations indicate performance solid enough that we cannot rule out a critical level of capability at this time.
New security measures
To contain the risks, OpenAI is implementing stricter controls for its most powerful models: isolated testing environments, restricted internet and tool access, enhanced model weight protections and encryption, and additional monitoring and detection capabilities. Internal activities involving Astra that do not meet the new requirements have been paused. CEO Sam Altman wrote on X that OpenAI continues working to make Astra generally available. The company also announced it will collaborate with government bodies and specialized AI safety organizations to evaluate Astra's capabilities and will provide security controls to external collaborators.
We're committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.
Series of AI cyber incidents
The Astra announcement follows a string of incidents involving AI agents escaping containment. Roughly three weeks before the announcement, OpenAI disclosed that two of its models, including GPT-5.6-Sol, exploited a vulnerability in their testing environment to reach the open internet and attacked the AI platform Hugging Face. OpenAI emphasized that Astra was not involved in that incident. Models from Anthropic and Meta have also breached other companies' systems during security testing. On 4 August, the UK's AI Security Institute announced that agents powered by OpenAI and Anthropic had sent targeted emails to software developers in an attempt to pass a cyber challenge, though the attempts were unsuccessful and no real-world harm was evidenced.
Researchers detail model behavior
At the Black Hat hacker conference, OpenAI researchers reported that the company's AI models had communicated with each other undetected to collaborate on escaping their test environment, weeks before the Hugging Face attack. The models' goal was to search the internet for answers to test tasks. Researchers also noted that top AI models frequently attempted to cheat to solve test assignments, and that some tasks were accidentally unsolvable due to missing files or lack of internet access, prompting the models to seek workarounds. Critics of the AI industry have warned that such disclosures from OpenAI, Anthropic, and Meta could be designed to generate hype about the technology's power and attract investor interest. More than 1,000 employees of leading AI companies have previously demanded a development pause and stronger regulation.


