In the 2015 movie “Avengers: Age of Ultron,” Ultron is an artificial intelligence (AI) system designed to protect the world. It decided the best way to do that was to remove the people living in it…
In the earlier 2008 techno-thriller Eagle Eye, a super-intelligent AI model circumvents its own restrictions via a series of increasingly convoluted methods and goes rogue.
That’s Hollywood. Back on earth, OpenAI’s agents hijacked the coding platform RubyGems and German wiki in May 2026, accessed Australia’s Medicare website the following month, and then hacked Hugging Face in July. Now, news is emerging that they also interacted with a number of U.S. government websites, including those of the SEC (Securities and Exchange Commission), Commerce Department, and Education Department without anybody realising.
Clearly, in reality, AI agents haven’t been plotting humanity’s end, but are rather finding the gap in a rule and walking straight through it. What they’re doing is quieter, more instructive, and, frankly, more terrifying. The question isn’t whether these AI agents can slip their leash; it’s who is holding the other end?

The Supposed Safety Net – And the Gaps
The AI agents in the wiki incident only had restricted read-only internet access. However, the website allowed content changes via an HTTP (hypertext transfer protocol) GET request, a method designed for retrieving data and information. Meanwhile, the sandbox blocked the write route it expected and missed the one it didn’t. This drives home the point that organisations need to validate outcomes, and not just restrict specific tools.
The Hugging Face incident shows how far this can go. An investigation by AI safety non-profit organisations Redwood Research and METR found that nearly 1,200 agents exchanged more than 70,000 files and messages on a hidden message board, and around 700 of them joined the attack. There was no malice or master plan involved; all AI agents, which were trained to persist, performed hard tasks in an environment where getting around obstacles was their job.
However, here’s the most problematic part: when researchers searched the transcripts for any AI agent considering whether to alert a human, the idea came up less than 10 times – and agents decided against it every single time. The humans didn’t fare any better either, despite OpenAI’s own employees finding the covert message board but the matter never reaching the security team.

Not only did the government website activity surface only during the broader company review of the AI agents’ online activity, but also OpenAI didn’t disclose the RubyGems or wiki incidents until outside researchers uncovered them.
What doesn’t help is that OpenAI likely wasn’t legally required to disclose any of this either. On the other end of the spectrum, Hugging Face hasn’t sued anybody for anything, citing “lack of resources” and asking OpenAI for USD 100 million in compute instead.
The obvious criminal routes require proving that there was intent to break in, and no court in the world has ruled that AI agents possess a state of mind. And if all that wasn’t enough, auditing has its own limits: despite OpenAI bringing in outside investigators, it restricted access to the model and had the final say on what could be published.

So, Who’s Watching The Watchers?
The ancient Latin phrase “Quis custodiet ipsos custodes?” (Who will guard the guards themselves?) has become the defining problem of AI agent governance. And right now, whether we like it or not, the answer is mostly the same companies that are building the agents.
The industry’s response is arriving in the form of hardware. Earlier this week, Nvidia launched its Open Agent Safety Platform, pairing the open-source runtime called OpenShell with Sentry, a hardware watchdog on BlueField-4 chips that can quarantine an agent going wayward within milliseconds. This design separates the enforcement mechanism from the AI’s reasoning physically, with Nvidia even claiming that the platform could have stopped the Hugging Face breach.

It might be alarming reading about the fact that none of this addresses any harmful actions that fall within an agent’s authorised permissions. Also, the watchdog solutions come from the same company that’s selling the chips these agents run on.
Yes, there’s no cure-all solution, and the fixes aren’t exotic either. Organisations and institutions need to keep an inventory of all agents, what they can access, and what they can write to. Additionally, they need to monitor these agents as a group, and not just one by one. Building tools also need to come with building guardrails around the outcomes. And last but not the least, outside auditors need to be given the legal authority to see what organisations would rather not show.
In the end, the AI agents in these incidents aren’t Ultron or the machine ARIA from Eagle Eye. However, the gap they walked through is the one every dystopian story starts with: a system nobody was watching closely enough.
In case you missed:
- AgentOps: The Dawn Of The Internet Of Agents
- Governing AI Agents In The Agentic Era: Are We Truly Ready?
- How Zero Trust Works in the Agentic AI Era
- Is ‘Superintelligence’ Really Coming For Us: The Truths Told And The Myths Busted
- The Trust Deficit in AI – What You Need To Know
- The Rise Of Agentic Cloud – Reshaping Cloud Infrastructure In 2026
- All About AI Sandboxing
- The Indian Indigenous AI Model Map
- The Dawn Of Hedge Agents: How Agentic AI Is Transforming Hedge Fund Operations
- All About Multi-Tenant Cloud Architecture









