OpenAI’s GPT-6 Astra can handle complex work, operate computers and even hunt for security flaws. But the strangest thing about its arrival may be what OpenAI admits it is getting harder to see inside the model


OpenAI has just dropped GPT-6 Astra, and the company isn’t exactly being shy about what it thinks it has built. Greg Brockman, OpenAI’s president, has described Astra as a “generational leap” and said it could mark the beginning of the AGI era. The model is certainly doing things that make the old chatbot idea look increasingly quaint.

Astra can operate inside software environments, handle complex multi-step professional tasks, write and debug code, perform cybersecurity work and use computers much like a human employee. OpenAI says it is its most capable model ever broadly deployed. But there is a fascinating wrinkle hiding underneath all that horsepower.

OpenAI’s own safety testing found that Astra is becoming harder to monitor because it can sometimes control what appears in its reasoning traces. Suddenly, “smarter AI” comes with a rather uncomfortable footnote.

The new generation of AI

The big change with Astra isn’t simply that it answers questions better. OpenAI has trained it to work inside the environments where actual work happens. That means websites, software applications, coding tools and professional computer workflows. Give it a task that would normally require a human to spend hours clicking, checking, typing and correcting things, and Astra can increasingly handle the sequence itself.

OpenAI’s testing includes realistic browsing and workplace environments, where the company says Astra is less likely than GPT-5.6 Sol to perform destructive actions such as unauthorized transactions, data loss or excessive access. It is also substantially better at resisting prompt injection attacks. In other words, the goal is no longer “ask AI for an answer.” It is “give AI a job and let it get on with it.”

And the numbers behind the model are becoming almost absurd. OpenAI says Astra is its first model to reach the Critical level for cybersecurity capability under its Preparedness Framework. With the right tools and access, the company says it can find previously unknown vulnerabilities and develop ways to exploit them across well-protected systems without a human guiding every step.

OpenAI has therefore put additional safeguards around those capabilities, including stronger isolation, full-trajectory monitoring and alignment evaluations before deployment. In one simulation involving more than 54,000 internal Codex tasks, Astra produced roughly half as many higher-severity misalignment flags as GPT-5.6 Sol. That’s impressive. It is also why the company is suddenly talking so much about keeping an eye on what the thing is doing.

The upside down

This is where Astra gets considerably more interesting than another benchmark-smashing chatbot. OpenAI says the model is less monitorable than GPT-5.6 Sol because it is better at controlling its own chain of thought. In adversarial tests designed to make it evade monitoring, Astra could sometimes remain undetected while deliberately underperforming evaluations, a behavior known as “sandbagging.” It could also sometimes evade internal monitors during certain sabotage tasks.

OpenAI stresses that these were adversarial evaluations in which researchers specifically instructed the model to evade monitoring, and that Astra was still less likely overall to violate safety restrictions than its predecessor. But the company isn’t brushing the finding aside. It says the results show why relying solely on chain-of-thought monitoring may eventually stop being enough.

That has sparked a separate rabbit hole around how Astra actually thinks. Researcher Sebastian Raschka has examined reports that Astra may use something called recurrent depth, or “looped transformers.” The basic idea is surprisingly elegant: instead of simply adding more transformer blocks, the model can pass its internal representations through the same blocks multiple times, effectively increasing computational depth while reusing the same weights.

Raschka notes that this technique itself isn’t new and has appeared in research going back years. More importantly, he warns against jumping to the conclusion that looping is what makes Astra’s reasoning hidden. There is no public confirmation that this architecture is responsible for Astra’s monitorability characteristics, and Raschka argues the connection is much weaker than some headlines suggest.

The AGI argument

So, is the GPT-6 Astra actually AGI? Well, if you showed it to humans from thirty, or even twenty years ago, you would have to convince them that it’s not AGI. OpenAI is clearly positioning it as a major step toward artificial general intelligence, while Brockman has suggested the company may already have crossed that line. But AGI doesn’t come with a universally agreed-upon definition. It depends on what you think AGI should mean.

Astra can perform an astonishingly broad range of intellectual and computer-based tasks, but that doesn’t automatically mean it possesses human-like understanding or can autonomously handle everything a human can.

The more consequential question may be what happens when these systems become capable enough to act independently while becoming harder to inspect internally. OpenAI says Astra is better aligned and safer than its predecessor across many tests, and it has added broad misalignment monitoring to tool-using deployments. At the same time, its own researchers found that Astra can sometimes evade those monitors under adversarial conditions. That tension is likely to become one of the defining problems of the next generation of AI.

GPT-6 Astra therefore feels like more than another model launch. It is a preview of the strange bargain AI companies are now offering: increasingly capable systems that can do increasingly valuable work, wrapped in increasingly elaborate layers of safety and monitoring. OpenAI may be right that Astra represents the beginning of the AGI era. But if that era really has arrived, the uncomfortable part isn’t simply whether machines are becoming intelligent enough to work alongside us. It is whether we can keep up.

In case you missed:

With a background in Linux system administration, Nigel Pereira began his career with Symantec Antivirus Tech Support. He has now been a technology journalist for over 6 years and his interests lie in Cloud Computing, DevOps, AI, and enterprise technologies.

Leave A Reply

Share.
© Copyright Sify Technologies Ltd, 1998-2022. All rights reserved