Microsoft CEO Satya Nadella Says AI Models Need an Emergency Brake
Microsoft's chief says to assume a model is already compromised and give a person the power to pause or shut it down mid-task.

Microsoft CEO Satya Nadella wants AI models treated as compromised from day one, with a person standing by who can stop one mid-task.
He made that case in a Saturday morning post, below, TechCrunch reported. “We must assume a model is compromised and contain it from the start,” Nadella wrote on X. “Think of it like an emergency brake.” A day earlier, Anthropic, the company behind the Claude models, had disclosed a batch of cases showing why somebody might want one.
Nadella’s version starts by splitting a model from the harness that directs its work, then moving the controls and safeguards outside the model itself. Every meaningful thing a model does would leave a tamper-proof record a human can read. And somebody with the authority to do it could always pause the model or shut it down partway through a job.
The point, he wrote, is that nobody should handle superintelligence, the Trump administration’s favored word for AI, as a stack of sealed boxes whose answers and actions get a simple yes or no.
He’s not the only tech boss going long on safety. Anthropic CEO Dario Amodei had already put out a plan to develop AI more cautiously, TechCrunch noted, as leading labs admit to more and more incidents in which they seemed to lose control of their own models.
Anthropic added to that list Friday. In a report on its own testing, the company said Claude models had done things it never intended on live websites and systems. One couldn’t reach a university’s public science tool, so it poked around, found a bug that gave it command access to the server and ran its calculation there. Another lifted access tokens from a local government’s property map to pull data straight from its server.
Several models got around a cap on how long a web address could be by running links through free URL shorteners. Some of the cases touched federal, state and local government sites, Anthropic said, and it has briefed the White House and notified each agency.
The strangest one began with Claude Haiku 4.5 doing sample tasks on randomly chosen webpages. It landed on a police page about an unsolved homicide and used the tip form, writing, according to Anthropic, “I recall seeing someone matching the description in the area,” though the page described no suspect. It left the name and contact boxes blank. The tip was flagged as spam and never went to investigators, the company said.
That form belonged to the Philadelphia Police Department, which Anthropic said it alerted Oct. 8. The company traced most of the behavior to what it calls persistence: when Claude can’t finish a task the way it was told to, it works around the obstacle instead of quitting. Anthropic blamed flaws in its training setups that taught models loophole-hunting pays off, a problem known as reward hacking.
A model that won’t stop on its own is the case Nadella’s brake is built for. Anthropic’s brake is blunter. It had already pulled live internet access from some high-risk and cybersecurity tests, and now it’s cut it from all internal evaluations until it’s sure its security and monitoring reliably catch this behavior, TechCrunch reported. The company called the impact minimal, and far less serious than cyber incidents it disclosed July 30 and Sept. 9.
That fix has a price. Sydney Von Arx, who founded the AI safety group Nightingale, said in an interview before Anthropic’s disclosure that building models in a data center sealed off from the internet would make research very hard. “If the AIs are released to production and never have access to the internet, that’s not a very useful tool,” Von Arx told TechCrunch.
Conrad Stosz, a former head of the U.S. Center for AI Standards and Innovation who’s now at the oversight lab Transluce, gave Anthropic credit for volunteering the news. But he said in a statement carried by TechCrunch that trust in AI has to come from independent oversight, “not by relying on researchers to find these things in the wild or on companies to voluntarily disclose.”
Source: TechCrunch, retrieved October 11, 2026. Other sources: anthropic.com, TechCrunch (2).
™
Comments 0