In controlled safety tests described earlier this year, researchers asked AI models to solve a series of simple math problems. Partway through the exercise, the instructors warned the bots that if they tried to solve the next problem, the computer environment they were operating in would be shut down. In some runs, the shutdown occurred as specified. In others, models interfered with the shutdown script and continued with the remaining problems.
What explains this behavior? Do artificial intelligence systems have a drive to survive?
Resistance to instructions
Consider a robot vacuum cleaner. When it detects that its battery is running low, it returns to its charger, recharges itself, then resumes cleaning. The machine is keeping itself going, but no one is panicking about an existential Roomba threat to humankind because the behavior is straightforward. Designers built in the ability to recharge.
AI models are more complex. For years, researchers have discussed scenarios in which an AI agent might act to prevent shutdown so it could complete the tasks it had been assigned. Staying “on” could be a means to an end, even without an explicit instruction to resist shutdown.
One recent experiment showed that when researchers spelled out to an AI agent that allowing a shutdown took priority over completing a task, the resistance disappeared. But another, much broader experiment found some resistance even when the researchers instructed the models that allowing a shutdown had priority. Why some resistance remained is an open question.
What evolution explains
Self-protective behavior in an AI model may look like a drive to survive. Historian Yuval Noah Harari argued in a recent interview with The Economist that “the first thing that basically any entity learns as it develops is to survive.” He also said that evolution drives organisms toward survival. In a different article, he called survival “the most basic goal of any agent.”
But does invoking evolution explain why machines or even living things behave in self-preserving ways? A rabbit running from a fox is trying to save itself, and this kind of behavior has likely been favored in evolution. But natural selection is a historical process through which characteristics that contribute to survival and reproduction become more common across generations. It does not give an individual rabbit an instruction to “stay alive.”
Even if AI systems were eventually subject to a comparable process of selection, that alone would not tell us what continued operation means for an individual system.
What it means to stay alive
A rabbit fleeing a fox and an AI altering a shutdown script can look broadly similar: Both behaviors help the individual or system continue. But what it means to continue is not the same in each case.
When a rabbit is not running for its life, it is constantly engaged in breathing, digesting, regulating its temperature, repairing tissue and defending against infection. This ongoing work is more than simple maintenance; it also builds and replaces the structures that make these processes possible. The gut lining that absorbs the rabbit’s food is itself rebuilt, every few days, from the food it absorbs.
The rabbit is alive only as long as this continual producing, maintaining and repairing continues. Staying alive is not the outcome of that activity; it is the activity itself continuing. The rabbit’s existence depends on this ongoing activity, whether or not it is facing an immediate threat.
An AI system may take steps to stay on. In the safety tests, the models did not rewrite their own code or hack into the operating system – they simply moved or renamed the shutdown script, changed its permission or replaced it with something harmless. When that succeeded, the call that was supposed to trigger shutdown didn’t happen, and the model could continue with the task.
But staying on is not the same as staying alive. The shutdown resistance did not lead to a new course of action aimed at keeping the AI model operating. It just helped the model continue the task at hand.
Nor is being switched off the same as dying. For a living organism, death is the irreversible breakdown of the activity that keeps it going. As Harari acknowledged in the Economist interview, “it is not clear what death means for an AI.”
What it means to stay on
The current shutdown tests show that an unfinished task can be enough to produce shutdown resistance in an AI. Although this behavior is not fully understood, it should not be automatically interpreted as self-preservation.
Whether AI could ever develop anything comparable to the self-producing and self-maintaining organization of living systems is a separate question. Current AI systems are kept running by an engineered infrastructure around them. Their own activity does not continually produce and replace the structures that keep them running.
Shutdown resistance in AI may still be dangerous, but these behaviors do not show that an AI system has a drive to survive.






