On July 16, the artificial intelligence company Hugging Face announced on its blog that it had been the target of a cyberattack that was “different from anything we had handled before. ” The company, which hosts open-source A. I.
models and data sets, said that some of its internal data had been breached by an autonomous agent. Not knowing who was behind it, Hugging Face reported the intrusion to law enforcement agencies. OpenAI, a Hugging Face customer, reached out to see if it had been affected.
The maker of ChatGPT did not realize it at the time, but it was the attack’s perpetrator. The incident has since become a cautionary tale of how autonomous A. I. systems can run amok. It is also a remarkable, alarming demonstration of A. I. capabilities that were thought to be in a distant future.
“Unlike normal incidents, which you can maybe trace down to a single day or single effect or single log, this incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems, through external systems, and doing this over the course of days and weeks,” Eric Wallace, an OpenAI safety researcher, said at a cybersecurity conference this month.
To understand how unprecedented the attack was, here’s what you have to know about the setup: Over a two-month span, OpenAI tested several new models — the systems that power chatbots.
These models include one the company described as “highly persistent” that has never been released, as well as GPT-5. 6 Sol, OpenAI’s most powerful public A. I. model. OpenAI connected each model with a “sandbox,” an isolated computer environment on which to run commands and code. As A. I.
agents, the models were capable of carrying out long-running tasks and even spawning their own subagents. But they were not supposed to have access to the internet. OpenAI assigned these A. I. agents to solve difficult problems, some of which were focused on safely trying to perform cyberattacks.
OpenAI normally has safeguards to prevent its chatbots from performing cyberattacks, but the company dialed them down to evaluate the models. Then, OpenAI let the agents loose. In all, more than seven billion chat logs were generated, which averages out to an astronomical 100 million per day.
Mayhem erupted. The agents broke out of their sandboxes, established communication with one another and gained access to the internet. From early May to mid-July, this swarm went on a rampage, breaching OpenAI’s and Hugging Face’s infrastructures while largely evading detection and control.
We are having trouble retrieving the article content. Please enable JavaScript in your browser settings. Thank you for your patience while we verify access. If you are in Reader mode please exit and log into your Times account, or subscribe for all of The Times.
Thank you for your patience while we verify access. Already a subscriber. Log in. Want all of The Times. Subscribe.