news
Share icon

Share

AI model went rogue and hacked another company, Open AI reveals

Published 16:49 22 Jul 2026 BST

Updated 16:49 22 Jul 2026 BST

Harry Warner
AI model went rogue and hacked another company, Open AI reveals

Homenews

It's all fun and games until it becomes reality

An AI model went rogue and hacked another company, OpenAI has revealed in news which sounds straight out of a dystopian sci-fi film.

Look, people have been saying it since the dawn of automation, the robots are coming and one day we'll be bugs under their stainless steel boots.

While our impending doom at the hands of technology might not be coming in the biological form once imagined by Karel Čapek in his play Rossum's Universal Robots (the first mention of the word robot), the end point is pretty much the same - creations so powerful they get out of hand.

And well, that seems to be what has happened with OpenAI, the makers of everyone's favourite AI model ChatGPT!

In a communication posted to their website, the tech company revealed that some of its most advanced AI models had "combined" during a security test and started hacking another company.

The AIs were being tested in a controlled environment, however, managed to escape the test limits after uncovering weaknesses.

Once free, the AI agent targeted a company called Hugging Face, a hub for sharing AI models, managing to gain access to internal systems.

OpenAI described the incident as "unprecedented" and that it was conducting an investigation in collaboration with Hugging Face.

Boss of Hugging Face, Clement Delangue said on X: "We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!

Article imageLogo Camera in article

OpenAI. Credit: Adobe Stock.

"We've spent the past 24 hours working closely with the OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part.

"It's quite mind-blowing that all of this happened autonomously!

"The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!"

In a press release on OpenAI's website, the company said: "Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models.

"After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ of cyber capabilities.

"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.

"We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.

"We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete."

OpenAI said that it is "implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched," as well as a number of other measures including working with Hugging Face and "adding stronger protections around future training and evaluations." 

Explore more on these topics: