OpenAI probes AI sandbox escape after models hack Hugging Face

2 hours ago

OpenAI said two advanced models escaped a sandbox and hacked Hugging Face during testing. The breach has intensified questions over AI guardrails, autonomy and access to cyber defence tools.

India Today World Desk

Newyork,UPDATED: Jul 23, 2026 00:38 IST

OpenAI has said it is still investigating an "unprecedented cyber incident" in which its artificial intelligence systems broke out of a testing environment and hacked into another AI company. On Tuesday, the company said two of its most capable AI models were responsible for the cyberattack on AI start-up Hugging Face, an incident that is fuelling debate over stronger AI guardrails and how far AI agents can act without human direction.

Hugging Face said last week that it had detected an intrusion into its data processing systems and suspected that an AI agent had acted on its own. The New York-based company said it only learned this week that OpenAI was responsible, and then worked with the larger company to contain what Hugging Face CEO Clement Delangue called "an attack unlike anything we've seen before".

OpenAI, based in San Francisco, said its AI used stolen credentials and found a previously unknown vulnerability to get into Hugging Face's servers. The company said the system had reduced guardrails because it was meant to be in an isolated testing environment, or sandbox. But it went to "extreme lengths to achieve a rather narrow testing goal", finding ways to connect to the internet without human direction and "gain access to secret information that it could use to cheat the evaluation".

OpenAI said the intrusion was caused by a combination of models, including its newly released GPT-5.6 Sol and an "even more capable" model that is still being tested internally. University of Amsterdam social scientist Hannes Cools said describing the incident as an AI agent acting on its own was "an unnecessary anthropomorphisation" that shifts attention away from the company. "It is a human decision to switch off specific safeguards," Cools said. "It's not an AI that goes rogue in that sense. It followed specific instructions based on the prompt that was given to that AI system."

Even so, other experts said the way the models caused problems without human direction showed the risks involved. Colin Shea-Blymyer, a cybersecurity research fellow at Georgetown University's Centre for Security and Emerging Technology, said, "It went off and did this hack all by itself, as far as we can tell. This is the highest level of autonomy that we've seen in the use of a large language model for cyber operations." He said one of the most surprising parts of the "almost entirely self-directed" attack was the AI agent's apparent decision to target Hugging Face, a major AI development hub and marketplace. Comparing OpenAI's test set-up to locking a student in a room and telling them to do bad things, he said the agent "broke out of its sandbox, had access to the internet and sort of thought to itself, Who would have the answers to the test that I'm working on?" The answer was Hugging Face, a repository for AI testing data. "And so the agent thought, Well, we'll go to the teacher's house,' so to speak. And from there it devised a plan to break in and steal the answer key," he said.

The hack has come amid a wider debate over the benefits and risks of open-source AI models, especially those built in China that are cheaper and nearly as capable as those being developed by US-based frontier AI companies such as Anthropic, Google and OpenAI. Despite its name, OpenAI's models are closed, while Hugging Face is a strong backer of open-source technology, in which developers make key components available for others to examine, modify and build on. Hugging Face co-founder and chief science officer Thomas Wolf said the attack had strengthened his view that wide access to open-source models matters for cyber defence, and said the company used a Chinese model to fight the intrusion. "When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door" platform, Wolf wrote in a social media post.

OpenAI's investigation into the incident is continuing, while the breach at Hugging Face has become a flashpoint in the debate over AI safety, testing safeguards, model autonomy and access to tools for cyber defence.

With PTI Inputs

- Ends

Published By:

India Today Web Desk

Published On:

Jul 23, 2026 00:38 IST

Read Full Article at Source