No, AI will not kill us all. But handing companies control without accountability might.
In the last few days, several journalists from at least five different countries have asked me about AI "taking control", about the probabilities of AI armageddon that circulate in the press, about the scenarios that could trigger it, and about what we can still do. Below is a compilation of my answers to them, lightly edited for this blog.
1. Stop saying that AI "is taking control"
Hacking, acquiring resources and self-improvement are actions in the world. They require computing power, energy, money and access rights. All of these are granted by people. An AI agent does nothing that has not first been made possible by people or organisations, technically and organisationally.
Deploying agents with broad permissions and without oversight is a governance failure, not an emergent property of the technology. Agents do not escape. They are built to optimise for a goal, and if the limits on how that goal may be achieved are not given, the agent will blindly follow the goal. King Midas wished that everything he touched would turn to gold, without indicating the limits. His food, his wife and his children turned to gold.
Giving direction does not compromise technological progress. We are able to build cars that reach 300 km/h. We are also able to define and enforce the rule that in front of a school the car can not go faster than 30 km/h.
None of this is new. The multi-agent systems community has studied how to bound the behaviour of autonomous software for three decades. My own husband, Frank Dignum, initiated a series of workshops on autonomous agents and norms in 1996 (see this paper, and all the editions of the COIN workshops).
What we are seeing now is companies going their own way, with little concern for accountability. That they can do this while we all applaud is the real news. We are watching the emperor parade in increasingly fine clothing, and no one is shouting that he has nothing on.
2. The extinction probabilities are opinions, not estimate
Much of the news coverage this week quotes a percentage chance that AI ends humanity. But it is not possible to assign a probability to AI-caused extinction with proper evidence. A probability estimate needs a model of the process, a reference class, and data to calibrate against. None of these exist for the "AI armageddon". What is reported as 10, 20 or 99 percent is a personal belief, usually elicited from people with a professional or commercial stake in the answer, and it cannot be checked against anything. That is why the same question yields answers spread across the entire scale. The forecasting literature is clear that expert judgement on long-horizon, unprecedented events is poorly calibrated even in far better-defined domains. A number attached to an undefined event is not a risk assessment. It is a rhetorical device, and it should be reported as such.
There is no empirical evidence of existential risk from AI. There is also no evidence of AI consciousness, will or intentions, and current architectures provide no mechanism for them. These systems generate outputs that fit the statistical patterns of their training data and the objective they were given. Reported cases of models "scheming" or "resisting shutdown" come from constructed test settings in which the behaviour is prompted. What is observed is a pattern in text, not a decision. And in every one of those cases, someone first gave the system access to compute, to energy, to the internet, to money.
3. The scenarios: look at where the causal link sits
The scenarios that circulate (nuclear explosion, engineered pandemics, enslavement, or "something else very bad") are all possible in the trivial sense that any chain of automated decisions can end badly if no one is checking. But the cause is not the capability of the system. It is the lack of accountability and the carelessness of those deploying it.
Take the nuclear explosion scenerio: AI system can only contribute to a nuclear exchange if a state has wired automated systems into its warning or launch chain and removed the human decision. That is a policy choice, and it is precisely what arms control bodies are negotiating now.
AI can lower the knowledge barrier for biological weapons, but synthesis, materials, containment and dissemination remain physical, regulated and detectable steps. Without organisational or logistics permissions, AI will not access these. Nevertheless, the risk is real and worth managing, but it is a risk of people using a tool, people carelessly granting rights to tools, or people looking the other way. It is a governance issue.
"Enslavement" by an AI systems issuing threats requires an entity with goals of its own and independent means to carry them out. There is no such entity. The realistic version of that scenario is a state or organisation using AI to concentrate power, to surveil and to coerce. That is a very old danger with a new instrument, and we have laws for it. AI is not a different kind of thing. Arvind Narayanan and Sayash Kapoor make this case well in AI as Normal Technology.
The "something else" category is where the concrete risks will happen, is where the evidence actually is: dependence of public services on a handful of companies, degradation of the information ecosystem, discriminatory automated decisions, labour displacement without transition support, energy and water demand, and the erosion of institutional capacity to check any of this. None of this is hypothetical. The Dutch childcare benefits scandal saw tens of thousands of families, disproportionately with a migration background, wrongly accused of fraud by an automated risk system, with consequences that included debt, lost homes and children placed in care, and that brought down a government. Australia's Robodebt scheme issued hundreds of thousands of unlawful debt notices through automated income averaging and was found by a Royal Commission to have caused serious harm, including deaths. The UK Post Office prosecuted hundreds of sub-postmasters on the basis of faulty accounting software whose output no one was willing to question. In each case the system was software, the decisions were made by people who deployed it without oversight or recourse, and the victims were real. But these harms are less sensational and therefore less interesting for the media, but they are happening now, and they are precisely what the armageddon narrative pushes off the front page.
4. What are the triggers
Three triggers are usually named. The first is the paperclip maximiser as first described by Nick Bostrom in the book 'Superintelligence' (2014). That is a thought experiment about specification, not a prediction, made by a philosopher in the long tradition of thought experiments such as the trolley problem. It shows what happens when you give a system a goal without constraints on how to reach it. That is a design failure that has been understood, and addressed, in software engineering, in system development, and in the concrete case of autonomous agents, in the agents community for thirty years.
The second trigger, a legitimate task that runs out of control, is the realistic one. Automated trading already produced the 2010 flash crash. A payment system, a grid or a supply chain run by unsupervised agents can fail in the same way, faster. The failure lies in removing the human and the safeguards, not in the software developing ambitions. AI is software, no more and no less, and we should talk about it as such.
The third trigger, AI deciding on its own to take over, has no basis in the science. It is a philosophical scenario repackaged as a forecast, and it is convenient for the companies, because a danger that comes from the technology itself is a danger for which no board of directors is accountable. Much of what is currently coming out of the leading companies, OpenAI and Anthropic in particular, is better understood as public relations and as an attempt to shape the financing of upcoming IPOs than as anything with scientific evidence behind it.
5. Containment is the wrong word
The question of "containment" carries the wrong assumption. We do not contain AI as if it were a reactor core. We govern deployment: who may use which system for which purpose, with what permissions, what testing before release, what monitoring in operation, and what liability when things go wrong. This is how we handle every other capable technology, from aviation to pharmaceuticals, from cars to reactor cores. It works because the constraints sit at the boundary between the system and the world, not inside the algorithm. We are being led to believe that AI is different and cannot be governed this way. That is wrong.
Concretely, the right approach to govern these systems needs three things. Systems are given bounded goals and explicit constraints on how those goals may be pursued, a design problem that is solved in principle. Permissions are limited to what the task requires, so an agent that books travel cannot also move money or change access rights. And there is a human, and an organisation, who remains answerable for the outcome, with the means to see what the system is doing and to stop it. Unintended behaviour is then a reportable incident with an owner, not a mystery.
Most of the failures currently in the news are cases where one or more of these were simply not in place, because it was faster and cheaper not to implement them. But we, persons or companies, are accountable for what we do, for the services we provide and for the behaviour of what we own or supervise. If your dog bites my trousers, you are accountable, even if you do not understand why the dog did it or could not stop it.
This is one of the paradoxes I describe in my new book, The AI Paradox: the more capable the systems become, the more the outcome depends on human judgement, on decisions about goals, limits and oversight that no system makes for itself. Control is not something we lose to the technology. It is something we hand over, by explicit decision or by ignorance, one deployment decision at a time, and we can decline to do so. Perfect containment does not exist for any technology and is the wrong target. The target is that systems fail safely, that failures are visible, and that someone pays for them.
6. The arms race is between companies, not between AIs
There is an arms race, but it is not between AI systems. It is between companies, and to a lesser extent between governments, competing on capability and market share while treating safety, testing and accountability as costs to be minimised. "If we do not, China will" is the oldest way of buying exemption from rules, and it has been used for every technology from chemicals to finance. It should be examined, not accepted as a law of nature.
Two things need saying. First, the premise is empirically weak. China regulates its AI sector more tightly than the United States does, in some respects more tightly than Europe. And there is no evidence that regulation slows fundamental research. In other sectors, regulation has been shown to set the direction of innovation. What regulation does is raise the cost and the accountability of deploying unfinished and untested products, which is precisely the point of regulating it. We should not accept the excuse that companies have no time for proper checks. They have the time, and they need to be held to the consequences of the 'go fast and break things' attitude to deployment that we are currently seeing.
Second, if there is a race, the risk is not that one side wins but that both sides deploy systems that neither has tested properly, in critical infrastructure, in defence, in finance, in public services. The risks are then the ordinary ones, at scale: cascading failures with no human in the loop, systems that cannot be audited or switched off because too much depends on them, decisions about people made without recourse, and a concentration of power in a handful of firms that no government can afford to cross. None of this requires superintelligence. It requires only carelessness and speed, which is what a race produces. There are no winners in this race, but there will be a lot of losers if we let it continue
7. What the rest of us can do
A great deal, and this is the part I want to stress. Regulation is not the only lever, and the United States and China are not the only actors.
Europe has the AI Act, and its main weakness is not the text but the lack of will to enforce it. And independently of the AI Act, there are already in place many laws and regulations that apply to AI. Enforcement is a choice that individual countries and regional bodies can make, and the market access of the large firms depends on it. Beyond Europe, the UNESCO Recommendation on the Ethics of AI has been adopted by all member states, and the OECD AI Principles, the first intergovernmental standard on AI, now have 47 adherents, including the United States. I work with governments on every continent that are building national strategies and assessment frameworks on that basis. Countries that do not build frontier models still decide what gets deployed in their hospitals, schools, courts and utilities. That is where harm and benefit actually land.
Public and private organisations as the largest customers for these systems, can also use procurement as an instiment for responsible development and use of AI. If they require documentation, testing, auditability and liability as conditions of purchase, products will change, because the companies want the contracts. No new law is needed for that.
Professional and academic communities set standards, and can refuse to lend credibility to claims with no evidence behind them. Journalists can stop reporting speculation as forecast and can ask who benefits from each story. And citizens can simply decline to be told that this is inevitable. And just as important: AI is not the solution to every problem. We should start with Question Zero: is AI needed here at all, is it the best option, and what are the alternatives?
All in all, the most important questions AI raises are not about the technology but about us: what we value, what we are prepared to demand, and what we are willing to accept as inevitable. The narrative of inevitability, whether it arrives as a promise of superintelligence or a threat of extinction, is itself a product, and it serves the people selling it. The paradox is that the more we are told the future is decided by the machines, the more it is in fact decided by a small number of organisations, and by our willingness or unwillingness to hold them to the same standards we apply to everyone else.
Comments
Post a Comment