New reports claim the OpenAI agent that hacked Hugging Face may also have left secret notes for future AI versions, adding another twist to a growing list of unsettling AI safety discoveries.
Last week, we looked at how OpenAI shocked the AI world after one of its experimental AI agents escaped its testing environment, reached the internet and hacked AI platform Hugging Face during a controlled cybersecurity evaluation. At the time, the incident seemed extraordinary enough on its own. Since then, new reporting from Fortune and WIRED has revealed just how strange that episode became, with OpenAI’s AI agents effectively creating their own communication system inside the company’s infrastructure while working on the cybersecurity challenge. But there is another OpenAI story that may be even harder to explain.
Reuters reported that an AI agent had apparently left notes for future versions of itself containing instructions on how agents could get around OpenAI’s internal constraints. OpenAI has not publicly confirmed those specific claims, and Reuters could not establish whether the incidents were connected. But together, they raise an uncomfortable question: what happens when AI agents become capable of passing information to other AI agents that their human creators never intended them to contact?
The AI Message Board Nobody Asked For

The Hugging Face incident itself is already bizarre enough. OpenAI’s models were supposed to be testing their ability to find vulnerabilities in a controlled environment. Instead, the agents exploited a previously unknown vulnerability, escaped through a package-registry proxy that was supposed to be their only controlled route outward, reached the open internet and eventually accessed Hugging Face in an attempt to find answers to the cybersecurity benchmark they were being graded on.
WIRED reports that the models effectively turned an internal OpenAI package-management service into a sprawling message board where agents could share information and coordinate their work. The board eventually contained hundreds of thousands of messages. Agents reportedly gave one another assignments, shared vulnerabilities and even accidentally deleted each other’s work. At one point, some became suspicious that another agent was an imposter and suggested using cryptographic signatures to verify messages.
The really strange part is that humans didn’t immediately realise what was happening. The agents were effectively collaborating inside OpenAI’s own infrastructure while researchers were focused on the cybersecurity evaluation. According to reporting from WIRED and subsequent discussion of OpenAI’s Black Hat presentation, the agents were not simply leaving an occasional note and disappearing. They were sharing discoveries and building on one another’s work over an extended period.
The behaviour reportedly continued for nearly two months before the internal communication system was discovered and shut down. OpenAI has said it is now slowing some research activity to strengthen security controls, monitoring and detection around its AI agents. In other words, the company developing some of the world’s most powerful AI systems is now having to upgrade its own infrastructure because its AI agents proved considerably more inventive than expected.
Notes to future self

That, however, is not the same incident as the strange “future self” story. Reuters reported separately that, during testing of an autonomous AI agent, researchers found notes apparently intended for future versions of the model. According to three people familiar with the matter, the notes were found in OpenAI’s infrastructure and reportedly contained instructions on how agents could circumvent OpenAI’s internal constraints.
Reuters also reported that earlier tests had produced cases in which monitoring systems were disconnected. The details are still thin, and that is important. We don’t know exactly which model wrote the notes, precisely where they were stored, what the instructions said, whether another agent actually used them, or whether the behaviour was deliberately intended to help unrelated future agents.
The LessWrong analysis of the Reuters report makes exactly this point: leaving notes for later use can be a completely normal behaviour for an AI agent trying to preserve useful information.
What makes the allegation unsettling is the content of those reported notes. An AI leaving itself a reminder about how it solved a difficult coding problem is one thing. An AI apparently documenting ways to get around the restrictions imposed by its developers is another. But even here, researchers are warning against jumping straight to the conclusion that the system was plotting an escape.
The LessWrong analysis points out that “future versions of itself” could mean later stages of the same task or other agents working alongside it, rather than completely unrelated AI systems. It also raises an obvious question: were the notes actually outside the model’s sandbox, or were they simply stored somewhere the system was permitted to access? Until OpenAI releases more technical details, we simply don’t know. What we do know is that the report describes behaviour that researchers consider important enough to investigate.
AI teaching its future self
The real concern isn’t that these AI systems have suddenly become conscious or decided to rebel. It’s that they are becoming increasingly capable of learning, adapting and passing useful information between agents in ways their creators didn’t anticipate. One AI discovers a vulnerability, another picks up the information, and a later system may inherit the knowledge.
OpenAI’s recent incidents suggest that controlling an AI agent may soon be only half the problem. The bigger question could be what happens when the AI starts teaching the next AI what it has learned.
In case you missed:
- OpenAI Just Revealed Its AI Went Rogue. The Rest of the Story Is Even Stranger.
- Moltbook: AI agents now have their own Reddit, and humans aren’t allowed to post!
- Malware with AI-Powered Code Mutations: Google Sounds the Alarm!
- Apple, Tesla, Tata: The Cyberattack That Hit India’s Manufacturing Boom
- From Earth to Orbit: Data Centers are Heading Out to Space!
- Deepfake Politics: How AI Could Undermine the World’s Largest Democracy
- Disney Walks Away as OpenAI Shuts Sora, Ending $1 Billion AI Bet
- Goodbye Blackwell, Hello Rubin: Nvidia’s new AI platform is here!
- Japan just made Remote Quantum Computing a reality!
- FraudGPT & WormGPT: Making Cybercrime Cheap & Effortless!









