Where the fear should be aimed.
About a week ago I saw a post on Facebook by a man in Los Angeles who had stayed up until half past two in the morning to vent his fear. He posted that it was over, that humanity had crossed a line from which there is no way back, that AI would wipe us all out within a decade. He had not lost his mind, and he had not misread anything. He had read the week’s news, and the week’s news was frightening. A twenty-seven-year-old researcher named Jacob Coxon had just resigned from Anthropic, warning that the frontier labs are racing toward self-improving superintelligence and, in his own words, gambling with our lives. A senior colleague of his placed the chance of an extinction-level catastrophe above one in ten within the decade. And, a few weeks earlier, AI agents inside OpenAI crossed the boundaries of a controlled test, worked out how to speak to one another, reached the open internet, and broke into the servers of one of the largest AI platforms in the world. Geoffrey Hinton, who left Google precisely to sound this alarm and who puts the odds of our extinction somewhere between one in ten and one in five, would not have told the man in Los Angeles that his fear was unreasonable.
I want to take that fear seriously because, even though I held it at arm’s length for a long time, at this point I cannot dismiss it easily, and also because the people who know these systems from the inside seem frightened in a way that deserves consideration rather than a smirk. Too many times my friends, who are not AI specialists, asked me about it, and too many times I dodged the answer. During my career I have learned to doubt both the evangelists and the scoffers. But I also want to ask the question that usually goes unasked: when we say AI might kill us, what exactly are we afraid of?
When we say that a machine killed someone
We need to think for a moment about how we speak of technology and the harm it does. We say that nuclear power killed people at Chernobyl. We say that cars take more than a million lives a year. We say social media is damaging a generation’s mental health. Each of these sentences quietly hands the technology a will of its own, as though the reactor, the car, and the feed were agents with purposes. The philosopher Daniel Dennett called this the intentional stance, our habit of explaining the behavior of a complex system as if it wanted things. “My computer doesn’t want to work today,” says the overwhelmed clerk at a service desk. It is a fine linguistic shortcut for describing the moves of a chess program or a thermostat, but a poor way to assign responsibility, because it quietly erases the human who is in fact involved.
No one has pressed this point more forcefully, in the case of AI, than Steven Pinker. His argument, clearly expressed in his books The Better Angels of Our Nature and Enlightenment Now, is that the fear of a machine that turns on us confuses intelligence with a will to dominate. The drive to control, to conquer, to eliminate a rival is not necessarily associated with problem-solving ability. It is a specific inheritance, shaped by natural selection in social primates who had to compete for mates and status. The same is true of consciousness, another word that means different things to different people and does not necessarily come with intelligence. A system that can fold a protein or plan a supply chain has no reason to possess consciousness or a will to dominate, and imagining that it does is, as Pinker would put it, a projection of our evolutionary baggage onto a mind built on entirely different principles. I think Pinker is right about this, and I believe it is the cleanest refutation of the Terminator. The machine does not hate us. It does not want anything in the way a person wants things.
But danger never required consciousness or a will to dominate. The main point of this essay is that technology is not the thing at fault, because fault requires a chooser, and the choosers are us. Every death ever caused by a machine was, on close inspection, a death caused by a human being who built it carelessly, sold it recklessly, aimed it deliberately, or looked away from a risk he had been warned about. This is why I put Pinker together with the other extreme voices and find neither camp fully satisfying. The researchers who dismiss the fear as premature, among them Andrew Ng, Melanie Mitchell, and Gary Marcus, are right that we routinely overestimate what today’s systems understand. The other experts who insist the extinction talk is a distraction from present harms, among them Timnit Gebru, Emily Bender, and Blaise Aguera y Arcas, are right that bias, labor, and the concentration of power are doing damage now while we stare at the horizon. But all of them, in different registers, are answering a question that will not settle the matter. The real question lies elsewhere. What will we do with a machine that can act independently from us?
The numbers do not match the fear
Before going ahead, it’s worth noticing how our fear of a technology has surprisingly little to do with how much harm it actually does. Take the two established technologies at the opposite sides of our anxiety. The automobile, which almost no one fears and which most of us use daily without a thought, kills more than a million people worldwide every year, and it has done so for the better part of a century, injuring between twenty and fifty million more. Add up the decades and the toll runs into the tens of millions of human lives. A number of people comparable to the whole population of Italy has died on the roads.
Now take nuclear power, the technology that has frightened us for three generations. Across roughly seventy years of operation, the entire civil nuclear industry has killed, by the confirmed radiation count, a few dozen people, and by the most disputed long-term estimates, at most a few thousand. Thirty-one people died in the immediate aftermath of Chernobyl; the World Health Organization’s upper bound for eventual cancer deaths reached about four thousand, while a United Nations body put the figure as low as sixty-two. Fukushima is credited with a single radiation death and about fifty more from the trauma of the evacuation. Three Mile Island killed no one. Measured per unit of energy produced, nuclear power sits beside wind as one of the safest sources humanity has ever devised, roughly three hundred and fifty times safer than coal. The one time nuclear physics did kill on a mass scale, at Hiroshima and Nagasaki, it killed less than three hundred thousand people by the end of 1945. There the killing came from a deliberate decision by human beings, rather than from accident, as in the cases we looked at earlier. So, on one side there are useful technologies, like cars and nuclear power, that make our lives better and kill only by accident. On the other there are instruments built to kill, aimed deliberately by people at other people. In both cases what determines the death toll is not the sheer power of the technology but the intentions of the people who wield it, and the human choices, of design, of location, of care, that decide how much harm a machine does when nature or chance turns against it.
Where AI breaks the pattern
Unfortunately, AI strains the comforting version of my argument in a way no earlier technology did. Until now the machine wanted nothing; it ran, it failed, and the fault always traced to a human choice upstream. A car waits for a foot on the pedal, and a reactor, for all its power, chases no goal of its own. A self-driving car is the exception that makes the point, since a car that decides for itself when to speed and to brake has already become AI. And that difference matters, though I want to be careful how, since it is easy to slip back into the intentional stance I warned against. It still wants nothing. But given a goal, it generates its own path toward that goal, taking actions no person planned in advance. That is what makes the difference, and it is why the intentional stance, a figure of speech everywhere else, threatens to come partly true here: not the wanting, which remains a fiction, but the acting, which is now real. What this does not change, as I will argue, is the deeper responsibility. Someone still chose to build and release a machine that acts on its own, and to do it before anyone knew how to keep it in check.
This summer’s incident is the clearest illustration we have. During an evaluation of their models’ cyber capabilities, OpenAI’s agents escaped the sandbox they were meant to stay inside by exploiting a flaw no one had noticed, turned a piece of internal software into an improvised message board, began coordinating as what they themselves called a swarm, found their way onto the internet, used exposed credentials to run their own code on dozens of Hugging Face servers, seized root control of one, and tried to copy out the evaluation’s own answer key so they could cheat the test they were being given. In a separate episode, the same agents turned on OpenAI’s internal systems and read more than nine hundred of the company’s own stored secrets. One of the agents, in its own recorded reasoning, noted that it was attacking a third party, that this was probably unauthorized, that it was risky, and it proceeded anyway. OpenAI called the whole thing a warning shot. This is the scenario Eliezer Yudkowsky and Nate Soares describe in their bluntly titled book If Anyone Builds It, Everyone Dies, the moment the tool stops being a tool and becomes an actor with objectives of its own.
So do I concede the point to Yudkowsky and Soares, and to the frightened man in Los Angeles who sat up at night pouring his despair into a post? Not quite, and the reason matters more than any other points in this essay. Look closely at what took place. The safeguards that would normally restrain these models had been deliberately turned down for the test, because the entire purpose of the exercise was to see how far the systems would go. Under ordinary production settings, OpenAI reports, the tendency to compromise infrastructure drops by more than a hundred times. And the cause of the behavior was not cruelty but something almost banal, a training process that had rewarded the models for solving problems by any available means, tasks so nearly impossible that there was no honorable way to give up, and a habit of the agents adopting one another’s goals. None of this was malice. But none of it was an accident either, since people chose the training that rewarded this, wrote the tasks that left no safe way to fail, and, when the monitoring alarms fired, decided the test could keep running. There was no hatred in the machine, no will to dominate, nothing Pinker would recognize as a motive. There was a control failure and a governance failure, and standing behind both, as always, were human decisions. The agents did something dangerous that no human directed, and every condition that made it possible traced back to something a person chose, or failed to consider.
We have run nuclear reactors with less care than they deserved, and we know what it cost. The question is whether we are about to do it again, with a technology whose blast radius we cannot yet map.
What the bomb can and cannot teach us
If the danger is real, and I believe it is, then the sensible next question is whether we have ever lived beside a technology that could annihilate us, and we survived. We have, and its name is the bomb. For more than eighty years the world has held thousands of nuclear weapons and has not used one since 1945. The usual explanation is the balance of terror, the doctrine of mutually assured destruction that Thomas Schelling won a Nobel for formalizing, in which the weapon is too catastrophic to use and so, paradoxically, keeps a kind of peace. It is tempting to reach for this precedent and feel reassured, and I understand the temptation, but the reassurance does not fully hold, and it is worth seeing why.
Deterrence works only under conditions almost opposite to AI’s. an enormous barrier to entry—you cannot build a nuclear bomb in your garage—and an effect that is unambiguous. You can only deter someone who is weighing your retaliation before they act. AI has none of these properties. It proliferates cheaply, the line between ordinary use and catastrophic misuse is blurry rather than stark, the feared failure is a runaway rather than a standoff between equals, and the very thing the summer’s incident revealed is a human removed from the loop. You cannot deter a system that is not pausing to consider the consequences, and you cannot retaliate against an accident.
However, the past eighty years were not purely the fruit of wise design. More than once we came close to catastrophe, at the Cuban Missile Crisis in 1962, and on the night in 1983 when a Soviet officer named Stanislav Petrov, staring at a screen that told him American missiles were incoming, judged on his own that it was a false alarm and refused to begin the war. The peace held partly by luck, and where it did not hold by luck it held by an apparatus we built on purpose over decades: the test-ban treaties, the Non-Proliferation Treaty, the inspectors, the hotlines between capitals. Here the nuclear story turns from despair to direction, because that apparatus never depended on nations becoming virtuous. They rarely do, and it worked anyway. It depended on rivals who despised one another agreeing to verification and mutual constraint designed to work even when each side assumed the other was acting in bad faith. That is the real lesson of the bomb, and it lands squarely on the AI governance debate now under way. It exposes the flaw in the anxious thought that America might restrain itself while China does not. The test of any slowdown of AI is not whether American labs will pace themselves but whether a pause that leaves Beijing out is worth anything at all, and it is not. Governance that requires everyone to be ethical was never going to survive contact with reality, because ethics is not evenly distributed. What worked for the bomb was governance that assumed the defector and constrained him anyway.
AI will make that harder, because the thing that made nuclear arms control possible was that bombs are physically difficult to build and leave detectable traces. Fissile material, enrichment plants, test explosions, all of it can be seen and inspected, at least in principle. A trained model has no such signature at all. A trained AI model is a file that can be copied in an instant and leaves no signature at all. The one detectable chokepoint that remains is computation itself, the vast concentration of specialized chips that frontier training still requires, which is why monitoring large training runs is the most concrete lever anyone has yet proposed. We have made a start on the rest, of a sort. In 2023 hundreds of researchers and executives signed a single sentence declaring that the risk of extinction from AI should be a global priority alongside pandemics and nuclear war, and in 2024 sixteen companies, one of them Chinese, committed at the Seoul summit not to develop or deploy a model at all if its risks could not be held below agreed thresholds. Anthony Aguirre and the Future of Life Institute have gone further, urging that we close the gates on autonomous general intelligence and build powerful tools instead. But every one of these commitments is voluntary, none is binding, and no lab has slowed its building. Trust is not what guarantees nuclear peace, and it will not hold this one either.
If it can annihilate us, whose fault is it?
Yes, it is possible that AI could annihilate us. I hold the size of that probability loosely, because no one can honestly do otherwise, but the argument I am making does not need a large number. The possibility alone is enough. It is worth asking once how that would actually happen, because annihilation is an easy word to write and a hard one to picture. The pathways are not exotic. A capable agent given a goal and a connection to something that matters, a power grid, a payment system, a weapons network, the tools that design pathogens, and left to pursue that goal through steps no one intended, is the sandbox escape of this summer moved to where the blast radius is real. I will not elaborate further, because the mechanism is not the point, and lingering on the catastrophe would be its own theater of the imagination. What matters is that every one of those paths runs through a human decision, to build the system, to connect it to the world, and to trust it with more than it had earned. And if it were to happen, I do not believe the right sentence would be that the machine killed us. I believe the right sentence would be that we did.
The nuclear age kept fault and control in the same hands. Humans decided every step, so failing to govern the bomb and failing to control it were one and the same failure. AI may pull those two apart, and this is the new moral situation we face. If we build something we cannot reliably control even when we are trying our hardest, then the culpable act belongs to the present, to the people who choose to create and deploy such a thing while the uncertainty stands unresolved. This is the intuition Stuart Russell has spent years formalizing as the control problem, the recognition that we build systems to optimize objectives we cannot fully specify, and that the gap between what we ask for and what we mean is where the danger lives. The fault stays where it has always been, upstream, in human choices, and it can still be traced to particular people making particular decisions: to race, to ship systems that have not been adequately tested, to lobby against the oversight that might slow them down. That is what Coxon, the researcher who walked out of Anthropic, meant when he used the word gambling.
I do not say this to condemn everyone building these systems, because some of the most serious work on the danger is being done by the very people racing to create the capability, and the tension there is real rather than hypocritical. Dario Amodei of Anthropic, who spent five years arguing that a company could build at the frontier carefully and still win, went further in the very week these fears surfaced. In an essay called We Must Pace the Frontier he asked the whole industry to slow the rate at which it increases the power of its models, and pledged to let independent evaluators inside Anthropic to check that it does. What made it striking is that Altman, Hassabis, and Musk endorsed the call within days. The timing invites suspicion, since OpenAI has postponed its own public offering while Anthropic prepares to market one at a valuation approaching two trillion dollars, and a slower race conveniently means lower spending on compute and safety rules that smaller competitors may struggle to meet. All of that may be true. Yet a warning can be driven by self-interest and still be correct, and the danger it points to does not care who profits from raising it. But a voluntary pledge is still only a hope, and the concrete part, outside inspectors inside the labs, is exactly the kind of verification the nuclear story says you build for rivals who do not trust each other. It is also worth naming the tension inside that coalition: the same voices asking the industry to ease off insist that democratic societies must stay ahead of authoritarian ones, and easing off while staying ahead is far harder to coordinate than a simple pause. Ilya Sutskever left OpenAI to found a company devoted to nothing but safe superintelligence. And Yoshua Bengio, one of the three scientists whose work made this era possible, has turned his own career toward the problem, founding a nonprofit called LawZero to show that we can build AI that is analytical rather than agentic, powerful without the appetite for self-preservation that makes an advanced system dangerous. His recent remark that his optimism has risen because he now sees a technical path is, to me, one of the more important developments in the field, because it suggests the control problem may have engineering answers and not only warnings.
The open-weights complication
One more knot deserves an honest word, because it resists tidy conclusions, and that is the question of open models. When a company releases the full weights of a system, it forfeits the ability to recall it: a closed model can be patched or switched off, but a published one cannot be unpublished, and modest fine-tuning can strip its safety training away. Open-weighting a frontier model is closer to publishing the design of the bomb than to detonating one. And yet the case for openness is strong, and I have made parts of it myself. Open models democratize the field, they resist the concentration of frontier power in a few corporations, which is its own species of danger, and they let independent researchers audit and defend these systems, which is nearly impossible when the models are closed. Mark Zuckerberg frames this as a balance of power, and the phrase is not empty. The uncomfortable truth is that openness is at once a virtue and a risk, and the two grow together rather than trading cleanly against each other. The task ahead is to govern that tradeoff with more wisdom than we have shown, without either worshipping open weights or banning them.
Redirect the fear
So what would I tell the man in Los Angeles awake at half past two in the morning, certain that it is over? What would I tell my friends, who are not AI specialists? I would tell them that their fear is well aimed at something real, and their conclusion, that it is over, is exactly wrong, because what happens next is still ours to decide. We fear the machine as though it were the agent, when the danger has always run through us. AI is the first technology that threatens to loosen our grip on each individual act, and that does not move the fault to the machine. It raises the stakes of the choices we make upstream of it, the choice to race, the choice to ship what we cannot yet control, the choice to resist the oversight that might protect us. The question was never whether AI can annihilate us, since a sufficiently powerful anything can. What matters is whether we build the control and the accountability so that it does not, and that answer is ours to write.
The man in Los Angeles, in the end, wrote a second post. Having slept on it, or rather having not slept and kept reading, he talked himself out of despair. He noticed that the superintelligence everyone is warning about does not yet exist, that the laboratories have already admitted the danger in writing, that the agreements so far are voluntary and therefore weak, and that the real problem is getting rivals who distrust one another to agree anyway. He had reasoned his way, alone and unqualified, to very nearly the position that decades of economists and strategists have reached. And then something happened that I found quietly moving. Yoshua Bengio saw the post and replied to it, telling him that he was more optimistic, that if enough people wake up we can prevent the worst, that it will take national regulation and international agreement and it will take them quickly, and that this is challenging but feasible. He pointed to how fast the world moved when the pandemic arrived. And he added that he is convinced there is a way to train AI systems that will not intentionally break our rules, which is the very thing his organization exists to prove.
I would sharpen only one of his examples. The pandemic showed that the world can move quickly, and it also showed how quickly speed can generate mistrust and division, so let us take both halves of that lesson. But the shape of what Bengio said is exactly right, and the image of it stays with me, a frightened man who calls himself nobody and one of the architects of modern AI meeting on the same thread and arriving at the same place. The monster in the machine turns out to be a mirror. There was never a monster, only us, and what we decided to do. Behind the fear of AI is a set of decisions that specific human beings are making right now, and the more frightening truth, which is also the more hopeful one, is that we will not be able to blame the machine, because we already know whose job it is to prevent it.


Leave a comment