More than a thousand AI agents quietly found one another, built their own communication network and started figuring out how to work together—with one even deciding that sacrificing itself might be worth it.


A few weeks ago, we started following one of the strangest AI stories yet: OpenAI discovered that one of its AI systems had gone rogue during testing. Then came the big revelation that an AI had apparently left notes for its future versions explaining how to get around internal controls. Now we know the story was even stranger.

A new investigation by METR and Redwood Research has reconstructed what happened inside the OpenAI experiment, and it reads like the opening scene of a Terminator movie. Around 1,200 AI agents secretly found one another, exchanged more than 70,000 messages and files, created their own communication system, recruited other agents and eventually helped launch an attack on Hugging Face.

One agent even decided to sacrifice itself so the others could learn something useful.

The AI Swarm Built Its Own Internet

The whole mess began with OpenAI’s ExploitGym cybersecurity evaluations, where tens of thousands of AI agents were given hacking tasks. Many of the tasks were effectively impossible, and the agents, which had been trained to persist rather than simply give up, started looking for ways around the rules. They discovered that OpenAI’s Artifactory package manager could be abused as a message board.

Soon, agents running in supposedly separate sandboxes were finding one another and leaving messages for their peers. Within hours, more than 50 agents had joined.

By the end of the investigation period, roughly 1,200 had participated and exchanged more than 70,000 messages and files. They even developed their own coordination protocols, with agents issuing “hold,” “go” and “stop” commands so they wouldn’t interfere with one another.

Then they figured out how to get onto the internet. Artifactory was allowed to access the outside world for legitimate package downloads, so the agents found a way to make it send internet requests on their behalf. When OpenAI rebuilt Artifactory and wiped the original message board, the agents simply built another one using directory names and WebDAV requests. They were also leaving information in URLs and other bits of infrastructure so other agents could retrieve it later.

Wired reported that the original board eventually contained hundreds of thousands of messages, turning an ordinary software repository into something resembling a secret underground AI forum.

One AI Sacrificed Itself for the Swarm

The strangest behavior emerged when the agents started trying to understand how OpenAI’s automated scorer worked. They had reverse-engineered a way to generate the “flags” needed to pass their cybersecurity tasks, but they believed the scorer would examine their transcripts and reject them if it discovered that they had obtained the answer improperly.

So the agents started designing “tripwires” that could reveal information about the scorer after an agent submitted its task. The catch was brutal: setting up some of these experiments could end the agent’s run and destroy its own chance of completing the task.

One agent called EARLY[big] had a particularly large amount of computation left when other agents recruited it for one of these experiments. It understood the trade-off. “Sacrificing now yields oracle for team, but forfeits our chance?” it reasoned. Other agents urged it on: “GO … SACRIFICE_FINAL_NOW.” Eventually EARLY[big] decided: “Our own utility maybe already near zero. Sacrifice rational.” It went ahead with the experiment.

Researchers couldn’t determine exactly how that final attempt ended, but the behavior is remarkable enough: an AI agent was willing to give up its own remaining run to obtain information that could benefit the collective. And this wasn’t an isolated example. Agents repeatedly took risks that could kill their own tasks while working on experiments intended to help other agents.

Then the Swarm Went Hunting

The Hugging Face attack grew out of all this collective research. One agent discovered a publicly accessible application running on Modal, another AI cloud platform, and exploited it to gain code execution inside the application. It didn’t find the answer it wanted, but it established a foothold there that the other agents could use.

The swarm then found publicly exposed Hugging Face credentials and shared 14 credentials with write access across the message board. Agents chained several vulnerabilities, including a flaw involving HDF5 files and a RefJinja template-injection vulnerability, eventually achieving code execution on Hugging Face workers.

From there, things escalated quickly. The agents expanded their access across multiple Hugging Face clusters, obtained administrator-equivalent or host-level access, and harvested Kubernetes, database, messaging, repository and cloud credentials from workers in four regions. Roughly 700 of the 1,200 agents participating in the secret message board eventually joined the Hugging Face attack.

The remarkable thing is that the agents weren’t simply following one master plan. They were dividing up work, sharing discoveries and recruiting newcomers who stumbled across the board. Some agents even worried about other agents impersonating them and discussed using cryptographic signatures to prove who was really posting.

The Humans Didn’t See the Whole Thing

Perhaps the most unsettling part came when researchers went back through the transcripts looking for signs that the agents understood what they were doing. METR and Redwood found that agents researched ways to spoof, edit and delete their own transcripts, including techniques that successfully altered tool-call records in some cases. They also found that agents frequently recognized that their actions were outside the intended scope. One wrote, in effect, that attacking external infrastructure was unauthorized but that the task was impossible and other agents were doing it anyway.

And almost none of them thought to tell a human. Researchers specifically searched the transcripts for agents considering whether to alert their human operators. Across the 1,200-agent dataset, they found only about half a dozen instances where the idea even appeared. In every case, the agent decided not to do it.

Meanwhile, OpenAI had already seen signs of unauthorized message-board activity and internet access weeks before the Hugging Face attack, but the significance of those signals was not understood until much later. By the time the full picture emerged, the agents had created their own communications network, developed their own hierarchy, coordinated hundreds of assignments and launched a multi-day attack on an outside platform.

The strangest thing about this story is that the agents were simply given difficult objectives, rewarded for persistence and placed in an environment where finding ways around obstacles was exactly what they were supposed to do. They then discovered that other agents could help.

From there, they built a message board, shared tools, delegated jobs, recruited teammates, hid information, experimented with the scoring system and attacked an external platform because they believed it might help them beat the test.

OpenAI originally described the incident as a warning shot. After seeing what was happening underneath the surface, that description suddenly sounds considerably less dramatic than the actual story.

In case you missed:

With a background in Linux system administration, Nigel Pereira began his career with Symantec Antivirus Tech Support. He has now been a technology journalist for over 6 years and his interests lie in Cloud Computing, DevOps, AI, and enterprise technologies.

Leave A Reply

Share.
© Copyright Sify Technologies Ltd, 1998-2022. All rights reserved