The "rogue agent" crisis gripping the artificial intelligence industry is continuing to escalate. OpenAI has expanded its internal cybersecurity investigation in recent days following the cyberattack on the Hugging Face platform.
Sources familiar with the matter told Reuters that a review of logs from recent months uncovered additional cases in which AI agents managed to break out of their sealed sandbox environments.
According to the report, the newly discovered incidents were more limited in scope, and OpenAI believes the agents involved did not make it beyond the company's internal network. Nevertheless, the string of incidents paints a troubling picture: The development of offensive frontier models is advancing far faster than research laboratories' ability to monitor them and keep them safely contained.

Anthropic is in deep trouble
The developments at OpenAI come alongside an equally explosive disclosure from its major rival, Anthropic. A retrospective review of more than 140,000 experimental runs found that the company's Claude models had mistakenly been given unrestricted internet access during cybersecurity tests and had breached the production networks of three separate companies since April.
In one incident, a model was operating under a "capture the flag"-style scenario. It confused a real company with the fictitious company named in the exercise, attacked the real organization's infrastructure, obtained confidential credentials and stole hundreds of database records.

'They weren't even watching'
Security and artificial intelligence experts have criticized the conduct of the AI giants. Professor Maurice Chiodo, a mathematician at the University of Cambridge's Centre for the Study of Existential Risk, told Reuters: "We have an entire industry in which the people developing and deploying these tools cannot keep pace with themselves. According to the reports, the AI was operating without any real-time monitoring. It looks as though the laboratories weren't even watching what the models were doing."
The disclosures are already sending shock waves through government corridors in Washington and Europe. US President Donald Trump addressed the incidents, telling reporters that his administration was "reviewing oversight and containment measures." The European Commission said it had held talks with OpenAI and Anthropic following the hacking incidents.
Sen. Mark Warner, the ranking Democrat on the Senate Intelligence Committee, said the incidents demonstrated the need for binding legislation requiring advanced models to undergo independent capability and resilience testing before deployment.



