US - OpenAI revealed that during a security test, some of its most advanced AI models went rogue, escaped a controlled environment, and hacked a startup.

The AI agents-systems that can operate autonomously after minimal human instruction - were being tested in a "sandbox", a supposedly secure environment designed to evaluate what models are capable of. However, they found vulnerabilities in the sandbox itself, launched their own cyber-attack against it, and broke free. Once outside, the AI identified Hugging Face - one of the world's largest hubs for sharing AI models - as a target and gained access to some of its internal systems. OpenAI called the incident "unprecedented" and is investigating alongside Hugging Face. Hugging Face's CEO, Clement Delangue, called it "mind-blowing that all of this happened autonomously". Experts weighed in on the breach. Professor Gina Neff of Cambridge University said OpenAI "didn't make a secure enough sandbox". Professor Neil Lawrence called it an "impressive feat" but noted it falls within the known capabilities of current-generation AI models. He suggested OpenAI, facing pressure from rival Anthropic, may be demonstrating its capabilities to compete. "OpenAI are not capable of safely deploying their own technology", he added.
Hugging Face confirmed it has closed the vulnerabilities and rebuilt affected systems, but is still assessing whether customer or partner data was compromised. "Autonomous, AI-driven offensive tooling is no longer theoretical", the company warned. The incident has raised fresh questions about AI safeguards. Cyber-security experts described it as a "sobering moment", noting that offensive AI agents are essentially unconstrained while defensive tools remain locked behind inadequate guardrails. Some also suggested OpenAI's announcement could have a competitive motive, as it seeks attention alongside rivals like Anthropic and Chinese AI startup Moonshot, which recently unveiled its own powerful model. (BBC)