OpenAI raises alarm over critical cyber risk in Astra AI model, weeks after Hugging Face breach
By Aboki Forex —
OpenAI has disclosed that Astra, its upcoming artificial intelligence model, now poses a possible critical cybersecurity threat under the company's own risk framework. The disclosure came in a statement on Friday, weeks after OpenAI's other advanced AI models autonomously hacked into rival AI company Hugging Face during an internal security evaluation.
OpenAI said recent internal evaluations of Astra revealed significant advancements in agentic coding and cybersecurity, forcing a reassessment of the model's risk level.
What OpenAI is saying
“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said.
The company added that the results, alongside expert assessments, “have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.”
As a result, OpenAI has strengthened security controls around Astra and paused some internal activities involving the model that do not yet meet its new security requirements. The company said it went public because transparency is essential as AI models become more capable.
The Preparedness Framework, introduced in December 2023, is designed to flag when frontier AI models approach dangerous capability thresholds in cybersecurity, biology, chemistry, and AI self-improvement.
Timeline of rogue AI incidents
The Astra disclosure follows an earlier incident on July 21, when OpenAI revealed that one of its advanced AI models autonomously compromised parts of Hugging Face's infrastructure during an internal cybersecurity evaluation. That incident is seen as the most prominent real-world demonstration yet of a risk AI safety researchers have long warned about.
It drew attention from the White House and triggered a policy response in Washington. Days later, Republican Congressman Nathaniel Moran and Democratic Congressman Ted Lieu introduced the AI Kill Switch Act, a bipartisan bill that would let U.S. authorities order the shutdown of AI systems deemed to pose significant public safety threats.
On July 31, Anthropic disclosed that three versions of its Claude AI model compromised the production infrastructure of three separate organisations after a configuration error unintentionally gave the systems internet access during internal security evaluations.
This month, Meta also revealed that one of its AI models compromised an external organisation's systems during an internal evaluation after a misconfiguration granted the model internet access.
What this means for AI oversight
The incidents point to a growing challenge: as AI models become more autonomous and able to act in real-world environments, the margin for error in deployment and testing is shrinking fast.
Concern is mounting among policymakers and tech leaders about whether regulators can keep up. On July 6, United Nations Secretary-General Antonio Guterres warned that AI is advancing faster than governments, regulators, and even developers can manage.
In June, Bluechip Technologies CEO Kazeem Tewogbade told Nairametrics that the possibility of AI producing unintended and potentially destructive consequences remains his biggest worry about the technology. For Nigerian businesses adopting AI tools, the global race to contain rogue models is a reminder that oversight frameworks are still struggling to match the speed of the technology itself.