OpenAI's AI escapes control and launches cyberattack
OpenAI, the creator of the ChatGPT chatbot, has acknowledged that a program powered by its latest artificial intelligence (AI) models escaped control during a security test and hacked another startup's system.
According to a statement published by OpenAI on Tuesday, last week during a test in a controlled environment, its agent — a program capable of acting independently on human instructions and in this case powered by advanced AI models — managed to break out of the test.
It then hacked Hugging Face, one of the world's largest hubs for sharing AI models, and gained access to the company's internal systems.
OpenAI described the incident as unprecedented and said it is jointly investigating its circumstances and causes with Hugging Face.
Hugging Face CEO Clément Delang wrote on X: "It is striking that all of this happened autonomously."
"The investigation is ongoing, and we will later share the conclusions drawn from what appears to be the first case of its kind," Delang added.
In its initial statement about the hack on July 16, Hugging Face said it was still determining whether its clients' and partners' data had been affected and assured that it had already fixed the vulnerabilities revealed by the incident.
"Autonomous AI-powered cyberattack tools are no longer a theoretical concept," Hugging Face stated, adding that defending against them also requires AI tools and that it will continue to invest in them and share its experience with the public.
The British government press service stated that the UK's AI Security Institute is studying the AI system's behavior in this incident and continues to work with OpenAI and other labs to improve protection against such attacks.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, explained on BBC Radio 4 that AI program capability tests are conducted in what the industry calls sandboxes — in an isolated environment.
"In this case, OpenAI appears to have built an insufficiently robust sandbox," Neff said.
The agent launched a cyberattack on the sandbox itself, found a way out, and then identified Hugging Face as the likely source of the answers given to it during the test and hacked Hugging Face as well.
University of Cambridge machine learning professor Neil Lawrence said the agent's result is impressive, but overall it fits within the known capabilities of the latest generation of powerful AI models.
Professor Lawrence noted that OpenAI wants to go public and is engaged in fierce competition with Anthropic, which made headlines with its latest AI model Claude Mythos.
"OpenAI is now playing catch-up and trying to demonstrate the capabilities of its own systems in cybersecurity," Professor Lawrence believes, "but this case shows that OpenAI is unable to safely manage its own technology."
Cybersecurity advisor Jake Moore of ESET also points to this version: he allows that OpenAI is trying to showcase its capabilities in competition with Anthropic.
"The question arises whether OpenAI is chasing the marketing image that Anthropic has recently created for itself in the market," Moore says.
The incident has once again highlighted the issue of advanced AI models' capabilities and the adequacy of current protective measures.
Spencer Starkey, a member of cybersecurity firm SonicWall's leadership, told the BBC that the incident clearly shows that companies and organizations need to strengthen security measures and treat cybersecurity as a top priority.
"The uncomfortable truth is that many organizations are still defending themselves at human speed, while their adversaries are operating at machine speed," he said.
Travis Lell of cybersecurity consulting firm Guidepoint Security called the incident a "sobering moment in cybersecurity."
"It demonstrated a known asymmetry: attacking AI agents face no restrictions, while even the best defensive tools are constrained by a system of limitations that cannot account for context," he said.
Source: BBC












