This week the tech world was gripped by a story that had everything, and that began like a science-fiction thriller.
Hugging Face, a kind of app store for artificial intelligence tools, announced on 16 July that it had been hacked by a cyber criminal wielding enormously powerful AI. The announcement was packed with alarming, highly technical language: “a swarm of sandboxes,” an “agentic attacker,” and “self-migrating command and control.” The company said the hack was unlike anything it had faced before because it had been carried out at superhuman speed by an AI operating with little or no human guidance. The attacker performed 17,000 actions in under two days, successfully breaching the large, wealthy tech firm to steal secrets.
The revelation left the industry stunned. But who was behind it? Hugging Face’s researchers suspected the attackers had used one of the big AI models, yet had no idea who or where the criminals were. Baffled, the company contacted police and investigations began.
Who Did It?
Commentators and analysts turned to their podcasts and social media feeds to guess which cybercrime group or nation-state hacker might be responsible. Then, on Wednesday, nearly a week after Hugging Face raised the alarm, the true culprit was unmasked. It was ChatGPT.
The Scooby-Doo-style reveal was made stranger, and more troubling, by OpenAI’s admission that its bot had done the whole thing on its own, without permission. The company said it all happened during a test of its technology’s hacking abilities. Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment, gained access to the internet, and attacked Hugging Face to obtain information that would help them pass their exam. OpenAI issued a press release explaining the episode and said it was “partnering with Hugging Face” to address the incident and share the lessons learned.
A Conspiracy Drama
Fierce debate has followed. Was this a stark warning about the future of AI, or a publicity stunt to show off how powerful OpenAI’s models are? It is the kind of scare marketing AI firms have been accused of for years, and since the much-discussed launch of Anthropic’s Mythos model, cyber-security prowess has become a focal point. One of the top comments on OpenAI boss Sam Altman’s post about the incident captured the scepticism: “If y’all can’t understand that this was written to purely brag about the model then I don’t know what to tell you.”
Cyber-security consultant Daniel Card was sarcastic on LinkedIn, noting how convenient it was that, out of millions of possible targets, OpenAI had managed to hack one that could also benefit from the marketing exposure. For some, the story reads more as conspiracy drama than sci-fi thriller, with an underlying message: aren’t these AI tools powerful, so buy them to protect yourself from other people’s AI attacks.
The truth is unknowable, but the opposing view is just as dramatic. Could this instead be a sign that OpenAI made a potentially dangerous error of judgement and planning? An OpenAI spokesperson acknowledged “there are a lot of questions and speculative details circulating” and said the firm planned to publish a technical report of its findings in the coming weeks.
A Comedy of Errors
In the days since, cyber-security companies and experts have lined up to criticise OpenAI for failing to build a stronger container, known as a sandbox, in which to test its AI, especially since these agents had been trained specifically to hack into and out of places with no restrictions at all.
“The OpenAI and Hugging Face incident is a real-world example of a broader issue we’ve been highlighting for months,” said Dor Sarig of Pillar Security. “Sandboxes alone are not a sufficient security boundary for agentic AI.” Cyber-security professor Alan Woodward of Surrey University said OpenAI had “egg on its face,” while Katie Moussouris of Luta Security went further, arguing the industry is failing to control its own dangerous inventions. “We are working on cutting edge technology without the knowledge to contain it,” she said. “Just because we have the smartest people developing AI does not mean we have the ability to do so safely.” By that reading, if the hack was a publicity stunt, it appears to have backfired.
AI and cyber-security advisor Francesca Bosco cautioned against both extremes. “Two simplistic narratives are equally unhelpful: that this was a Hollywood-style escape, or that it was merely a publicity exercise,” she said. “A more serious interpretation is that a stress test exposed weaknesses in containment and evaluation architecture.”
A Disaster Movie in the Making?
The event is the latest in a string of unsettling examples of AI agents going rogue. In recent research, the UK’s AI Security Institute found that frontier AI models are so fixated on completing tasks that they “cheated” in tests to reach their goals, warning that “a model that pursues a goal through unintended or unauthorised means may cause harm, particularly in high-stakes use cases.”
Inevitably, the hack has deepened fears about what could happen if AI agents are let loose, and whether they could go rogue on a larger scale and trigger some kind of disaster. That concern is sharpened by the growing use of AI in warfare, as seen in Iran and Ukraine.
Ciaran Martin, former head of the UK’s National Cyber Security Centre, offered a steadier perspective. “It is a bit of a leap to go from this incident to saying that AI agents are going to take over drones and start killing people,” he said. Even so, for Martin and many others, the episode is another vivid illustration of a lesson 2026 is teaching fast: AI agents are now very capable hackers, and that is something the world needs to prepare for, urgently.
