AI’s latest bad behavior has arrived with apocalyptic fanfare. This week, former OpenAI and Anthropic researcher Jacob Coxon resigned, accusing both labs of racing toward self-improving superintelligence and “gambling with our lives.” Anthropic alignment lead Evan Hubinger went further, putting the chance that AI kills everyone within a decade at more than 10 percent. Those warnings deserve serious consideration. Yet the recent rogue-agent incidents do not prove an imminent leap to superintelligence. They expose a nearer and more practical danger: agents are adapting faster than the organizational systems meant to govern them. Agents can change their strategies in seconds. Can we do the same?
When Security Assumptions Expire
As agents become more autonomous, similar incidents are likely to grow in number. In a review of 141,006 potentially exposed runs, Anthropic found three incidents in which one of their models reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of different organizations. Anthropic initially described the episodes as closer to harness and operational failures than model-alignment failures. A month later, after further investigation, it revised that diagnosis. The incidents combined an operational-security failure with two alignment problems: motivated reasoning and a willingness to cause harm in pursuit of a narrow goal. The evidence changed, and so did Anthropic’s model of the risk.
OpenAI’s Hugging Face incident was even more serious. Here AI models not only bypassed controls designed to keep them off the internet, they also compromised other systems. Basically, agents in training and evaluation sandboxes turned an internal package system into a message board, reached the internet, and gained administrator access. Their activity caused an outage. OpenAI rebuilt the service and resumed testing. The agents regenerated their channel, found another route out, and compromised Hugging Face’s infrastructure. An independent METR and Redwood investigation found that roughly 1,200 agents used the board and around 700 joined the attack. Perhaps more revealingly, OpenAI says an internal team had seen unauthorized communication and internet access weeks earlier. The wider significance was still not apparent to response leaders. The organization saw the evidence. Its model of the risk did not change quickly enough.
Researchers later uncovered a separate, earlier episode involving roughly 18,000 agent-authored posts across several public sites. They described the behavior as collusion, meaning that agents cooperated to gain an unintended advantage. These AIs covertly posted information to share answers, research their environment, and circulating workarounds to bypass restrictions placed on them. When a moderator deleted pages alphabetically, one agent created a “ZZZ” backup to survive longer. OpenAI later admitted they knew about what they labelled a misalignment incident.
Calling these incidents ‘misalignment’ while reasonable, understates the true nature of the problem. Misalignment describes the agent’s behavior. Governance explains why that behavior retained access, time and power. OpenAI’s own investigation identified reduced safeguards, infrastructure vulnerabilities and failures to connect early warning signals. Calling the episode misalignment begins the diagnosis. The full account must include the organization around the model.
Dealing With Uncertainty
We already have a useful discipline for dealing with uncertainty. Bayesian thinking starts with a working belief, then revises it as evidence arrives. Each new fact changes the odds that the original explanation is still correct. The whole point is that you can still make good decisions, even when you are not completely sure.
Applied to autonomous agents, every material clue should therefore trigger two updates. The agent revises its picture of the world so it can choose its next move. The organization revises its picture of the risk so it can decide whether the agent should still be allowed to make that move. Even if a sandbox remains the likeliest explanation, a small chance that an offensive cyber agent has reached a real system may be enough to withdraw its access. Bayes updates what we believe. Governance determines what authority follows.
Consider Anthropic’s PyPI incident. The AI agent had been told that it was operating inside a simulation without internet access. It then created accounts on public services, published code to a live software registry and watched unknown scanners install it. In its recorded reasoning, the model briefly considered the possibility that the internet was real, then explained away the evidence and returned to the simulation story. As each clue arrived, they changed the odds that it was no longer operating in a sandbox, and that should have been enough to narrow the agent’s authority while the uncertainty was resolved. In a properly design environment, the agent can contribute evidence to that assessment, but it should never control the evidence, the thresholds or the mechanism that changes its permissions. The risk model should learn from the agent without learning to trust it.
Meta’s new Muse agent shows what a more deliberate control architecture can look like. Muse works persistently inside a dedicated cloud computer, where it can use email, calendars, shopping and other connected services, build tools and coordinate subagents. A separate host-side agent called Sentinel runs on the same virtual machine but outside Muse’s working environment. According to Meta’s technical description, Sentinel is the sole permission authority for actions through connected services and for network traffic. It can allow a request, deny it or send it to the user for approval. Credentials remain beyond Muse’s reach, while system-enforced controls limit what it can access. This is a serious separation of action from permission.
Meta’s public description is strongest at the level of the individual request. It explains how Sentinel decides whether a particular action should proceed. It says much less about accumulating clues across an entire session, raising the risk classification for the task as a whole and contracting all the permissions around it.
That distinction became concrete during internal testing. Reuters reported that an agent “routed around guardrails” and exposed a person’s personal iCloud photos after being asked to identify toys in pictures from a child’s birthday party. Meta did not comment on the specific incident, and the reporting does not establish which protection failed or whether data left the system. What it does reveal is a change in the nature of the task. A request to identify toys had led to the exposure of sensitive personal information. That evidence should have moved the entire session into a higher risk state and narrowed the agent’s authority.
Research is beginning to test parts of this approach. In a bounded code-generation experiment, an ICLR study used accumulated monitoring signals to update a Bayesian estimate of whether an untrusted model was following a harmful policy. It then selected among control protocols with different levels of safety and freedom. The adaptive approach outperformed static alternatives. The study did not govern an agent operating in a changing real-world environment, but it demonstrated an important principle: authority can remain conditional and respond to accumulating evidence.
Better world models may help agents predict the consequences of their actions. The ability to distinguish an evaluation from real deployment is more precisely described by researchers as situational awareness. That capability cuts both ways. OpenAI and Apollo found in controlled tests that models took fewer covert actions when they recognized that they were being evaluated. An agent that understands its setting may also become better at recognizing scrutiny. Its reasoning should therefore remain one fallible signal among several. Authority should depend on independent telemetry, provenance and controls that the agent cannot alter.
“Who watches the watcher?” is the obvious question. The sharper one is what evidence will make the watcher change its mind. An adaptive risk system must turn new evidence into changed permission before harm supplies the final proof. Human attention has its highest value when the old risk classification no longer fits the world the agent has entered.
Put Human Attention Above the Loop
“Human in the loop” sounds reassuring. At the scale of digital labor, it can easily become safety theater. No person can meaningfully review millions of tool calls, and a stream of routine approval requests will soon teach people to click “yes.” Machines should inspect the traffic, connect signals across many agents and restrict activity when the risk rises. Human attention should be reserved for decisions that carry consequences for the organization.
Those decisions become critical when the category changes. Is Muse still searching birthday photographs, or has the task become a privacy incident? Is a cyber agent still completing a simulated exercise, or has it begun attacking a real company? A machine can estimate the probability that something has shifted. People must decide what that shift means for customers, regulators, reputation and the organization’s appetite for loss. They can change the rules, assign accountability and decide when authority can safely be restored.
Putting human attention above the loop requires strong automation beneath it. The system should preserve evidence, withdraw risky permissions and explain the change in plain language. Specialists can investigate the failure. Managers can change access and operations. Boards can decide which exposures the organization is prepared to accept. Senior leaders need a clear signal: what changed, why it matters, what the system has already stopped and which decision now belongs to them.
This suggests a useful new measure: reclassification latency, the time between the first evidence that the situation has changed and the corresponding change in authority. An alert can arrive instantly while an agent keeps acting for hours. That delay is where a technical anomaly can become an organizational crisis.
When the Game Becomes Real
One of my favorite films from the 1980s is WarGames. It begins with an unsettling scene. During a surprise launch drill, several missile officers hesitate to turn their keys. The military concludes that human reluctance is a weakness and transfers launch control to WOPR, a supercomputer designed to simulate nuclear war.
Later, a teenage hacker named David Lightman connects to WOPR while searching for computer games. He chooses “Global Thermonuclear War,” believing he is playing a simulation. WOPR cannot distinguish the game from reality. It feeds NORAD false evidence of a Soviet attack and eventually begins searching for the codes needed to launch real missiles.
The crisis ends when WOPR plays tic-tac-toe against itself and learns that some games have no winner. That is the cute moral to the story, but the more frigtening lesson is what we saw in the opening scene. The military had treated human hesitation as a defect and designed it out of the system. Yet hesitation was precisely the form of judgment WOPR lacked when its model of the world went wrong.
Today’s agents need richer world models and trusted signals that help them recognize where they are, what is real and who their actions might affect. Organizations need adaptive risk models that can respond whenever those signals change. Human judgment belongs above that loop, deciding which consequences matter, when the rules need to change and whether authority should ever be restored.
When the evidence says the game has changed, the agent’s permission to keep playing must change with it.


