"This is fear marketing."
Every major risk industry in history — tobacco, lead, asbestos, pesticides, fossil fuels — has spent fortunes persuading the public that its products were safe. Why would the AI industry want to convince its customers, investors, and governments that its products could cause catastrophes?
This is all the less convincing given that serious incidents are typically minimized or concealed by the companies involved. In the OpenAI–Hugging Face incident, for example, OpenAI initially presented a far less severe account than subsequent investigations revealed.
Above all, these concerns do not come only from company executives. They are shared by many employees of those companies, several of whom have left their positions to speak out about these risks — well before the widely covered case of Coxon, there were Jan Leike, Daniel Kokotajlo, Miles Brundage, and Mrinank Sharma, and before them Geoffrey Hinton (Nobel laureate), who left Google in order to speak freely. These concerns are also shared by a substantial portion of the scientific community and by some of the world's leading specialists in the field (see below).
"Predictions about these risks have no scientific basis."
There is obviously no dataset from which one could experimentally calculate the probability of a catastrophe caused by a technology still under development. That does not mean the risk is imaginary, nor that the scientists studying it regard it as such.
The survey by Katja Grace and her co-authors is the largest ever conducted among AI researchers. It polled 2,778 researchers who had published at the field's most prestigious scientific conferences. On the question specifically addressing the risk that humanity might lose control of future advanced systems — leading to extinction or an equally severe and permanent loss of power — 661 researchers (randomly selected from the initial sample) responded. Their mean estimate was 19% and the median was 10%. While these figures do not constitute a measure of the actual probability of a catastrophe, they do show that such scenarios are taken very seriously by the majority of specialists in the field. (AI Impacts)
The subject now has a substantial scientific literature. The International AI Safety Report is one of the leading international syntheses of the state of knowledge, disagreements, and uncertainties around advanced AI risks. Its first edition brought together around a hundred international experts under the direction of Yoshua Bengio, with a committee representing 30 countries as well as the UN, the OECD, and the European Union. The report explicitly examines risks related to loss of control, cyberattacks, and biological weapons. (International AI Safety Report)
The exact probability of these scenarios is open to debate, but describing them as "unscientific" no longer reflects the state of the research. For more technically inclined readers, we refer to the literature on the Goodhart effect (for example, the 2019 paper by Manheim and Garrabrant), instrumental convergence (notably Alex Turner's dissertation and his accompanying commentary), and the characteristic problems with evaluating current models (for example, van de Weij et al. 2024). For an introduction to the field, we suggest the AI Safety Atlas, a textbook written by CeSIA.
"This is meant to win over investors ahead of an IPO."
When calls to slow down frontier models intensified in mid-September, markets reacted in precisely the opposite direction from what this theory would require. Between 4 and 14 September, Nvidia's share price fell 8.4%. On 14 September, Nvidia dropped 3.36% and the PHLX Semiconductor Index fell 5.9%. Reuters explicitly linked the decline in chipmakers' stocks to AI lab executives' calls to slow AI development. (Reuters)
More broadly, this hypothesis is incompatible with what investors actually do when told that the sector may have to slow down. Convincing markets that the industry might need to decelerate, absorb more regulatory constraints, and consume less compute is a peculiar way to maximize a valuation ahead of a stock market listing.
OpenAI has in fact just postponed its IPO, with Sam Altman explaining that going public now would be ill-advised given the safety issues still to be resolved. (Reuters)
"The OpenAI–Hugging Face incident was staged by OpenAI."
This hypothesis fits the facts particularly poorly.
OpenAI did not foreground the full severity of the incident. A crucial part of what we know today comes instead from an independent investigation conducted by METR and Redwood Research.
The independent investigation itself operated under significant constraints imposed by OpenAI's requirements and restrictions. Three researchers worked on-site for only six days. They had no direct access to the relevant infrastructure and had to request additional data from OpenAI as needed. They could not query the primary model involved. Their investigation focused mainly on a very limited time window; earlier incidents during training, the subsequent compromise of OpenAI's infrastructure, OpenAI's internal investigation process, and remediation measures were all out of scope. OpenAI also requested changes to the structure, emphasis, clarity, and tone of the published report. (Metr)
This does not look much like a publicity operation orchestrated to dramatize the incident. If there is a clear bias in the initial communication, it runs unmistakably in the other direction.
"This is to prevent competitors from catching up."
Regulatory capture is real. But it provides a very poor explanation for a demand to slow down the companies that are furthest ahead.
Slowing frontier models hits those at the frontier first: OpenAI, Anthropic, Google DeepMind, xAI, and the other players capable of training the most expensive and most advanced models. Meanwhile, companies with less powerful models can continue to progress, exploit existing techniques, and close the gap.
In other words, "slowing the frontier" mechanically reduces the advantage of those already there. If the goal were simply to lock in market position, far more effective strategies exist — such as restricting access to compute.
Some have argued that China's negative reaction to these announcements proves they were actually designed to prevent Chinese companies from catching up with their American counterparts. That is not what happened. The Chinese reaction targeted the very explicit passages in which Dario Amodei argued for maintaining the United States' technological lead over China, tightening export restrictions on advanced chips, and combating the "distillation" of American models by Chinese laboratories. These geopolitical positions can be contested, but they are distinct from the argument that the development of frontier models should slow down to allow safety techniques time to advance.
"Talking about existential risks distracts us from immediate harms."
There is no contradiction between AI exposing us to immediate harms and, simultaneously, to future catastrophic risks. Claiming we must choose between the two is like arguing that the flammability of fossil fuels exempts us from worrying about their contribution to climate change, or that the risks of radioactivity make it unnecessary to think about nuclear weapons. A 2025 study published in PNAS found that exposing the public to AI's catastrophic risks increased concern about those risks without reducing concern about AI's immediate harms. Discussing catastrophic risks tends to reinforce oversight of these systems, which can also have positive effects in addressing near-term harms.
"Recent incidents are primarily the result of human error."
Aviation, nuclear, chemical, and space disasters almost always involve a chain of human decisions. No one concludes from this that aircraft safety, nuclear plant safety, or rocket safety is a non-issue. The role of safety engineering in any high-stakes industry is to design systems that remain safe despite the foreseeable errors of their designers and users.
The argument becomes even less convincing as systems acquire greater autonomy. In the OpenAI–Hugging Face incident, the agents were supposed to be isolated from one another. They independently discovered a means of communicating, coordinated in the hundreds, took actions outside the intended scope, compromised third-party systems, and attempted to deceive the evaluation of their own behavior. Reducing this to "a human misconfigured the test" dismisses precisely the phenomenon that needs to be studied.
Let us take the argument to its logical conclusion. If tomorrow a model were to circumvent its safeguards and provide a terrorist group with the knowledge needed to design a biological weapon, would we say the problem was simply that a human asked the wrong question, or that the system's designers had made an unfortunate mistake?
The relevant question is whether extremely powerful systems remain safe when humans make errors, misuse them, or give them imperfect objectives. This matters all the more because current legal frameworks were designed for products and software far less autonomous than these, and because companies largely escape legal liability when their systems fail.
