1. AI capabilities
“It seems probable that once the machine thinking method has started, it would not take long to outstrip our feeble powers. […] At some stage therefore we should have to expect the machines to take control.”
AI is progressing at an exponential pace
Artificial intelligence (AI) has undergone a spectacular transformation in recent years, moving from narrow, specialized systems to models capable of carrying out an ever-growing variety of increasingly complex tasks. Far from slowing down, this trend is accelerating, regularly outpacing experts' forecasts and raising fundamental questions about the future capabilities of this technology and the implications of its large-scale deployment.
Long confined to specialized domains, large language models (LLMs) have considerably expanded the range and complexity of tasks accessible to AI. Every year, AI surpasses human capabilities in new domains. Every year, AI makes bigger strides than the year before. Among the world's leading experts in this technology, many expect it to surpass all human cognitive abilities within just a few years.
What recent progress has AI made?
To objectively assess AI's progress, researchers define measurable capabilities through standardized evaluations known as benchmarks. These benchmarks cover a wide range of tasks, from natural language understanding to solving mathematical problems, writing computer code, or image recognition.

Evolution of the performance of the best-performing AI models across various standardized evaluations between 2000 and 2024. AI is progressing faster and faster, constantly developing new capabilities, and reaching or surpassing human performance in a growing number of domains.
Source: International AI Safety Report
Here are a few examples of the spectacular progress AI has made in recent years:
- In image generation, in 2021 the best available models produced blurry, pixelated drawings in response to a user's request. By 2023, they could produce realistic images almost indistinguishable from real photographs. By 2025, frontier models produce ultra-realistic, high-resolution videos.
- In mathematics, the best language models available in 2020 could not even multiply two three-digit numbers. By mid-2025, companies developing frontier models were claiming gold medals at the International Mathematical Olympiad.
- In science, on the GPQA Diamond benchmark (a PhD-level multiple-choice test in biology, physics, and chemistry), GPT-4 scored only slightly above chance in 2023. By late 2024, the o1 model reached 70% correct answers, matching the performance of PhD holders in their respective fields. By mid-2025, frontier models were approaching a perfect score across all domains.
How can this spectacular progress be explained?
The performance of AI models grows in a relatively predictable way as their size (number of parameters), the amount of training data, and the computing power used for training (training compute) increase: this is known as scaling laws.
These are not fundamental physical laws, but empirically observed relationships that have proven remarkably accurate in recent years. They point to a strong correlation between the resources invested in models and the performance those models demonstrate.
The recent progress of AI thus rests largely on the exponential growth of several underlying factors:
- The amount of compute used to train frontier models has, on average, been multiplied by 4 every year.
- The size of pre-training datasets has been multiplied by 2.5 every year.
- The algorithmic efficiency of training has, on average, been multiplied by 3 every year.
These factors combine to produce an exponential improvement in AI capabilities. Compared with the best-performing AI models of 2020, the largest models of 2025 have 100 times more parameters and use 100 times more data and 1,000 times more compute for their training.
AI progress is not going to stop
Isn't AI limited by human intelligence?
Although human intelligence is at the origin of artificial intelligence, there is no scientific argument for claiming that human intelligence represents an upper bound for AI. Unlike traditional software explicitly programmed to perform a task, modern AI systems—particularly those based on deep learning—are trained by being exposed to vast amounts of data, by "playing" against themselves, and by interacting with their environment.
For example, AlphaZero, developed by DeepMind, learned to master complex games such as Go and chess from scratch, without any prior knowledge or data beyond the rules of the game. By playing millions of games against itself, the program gradually rediscovered existing strategies, invented new ones, and ultimately surpassed the best human players and existing specialized programs by a wide margin (Silver et al., 2018) after only a few hours of training. AI can therefore reach superhuman levels of performance without being limited by human knowledge or approaches.
Can we predict AI's future performance?
Although scaling laws do not allow us to precisely predict which new capabilities AI will develop, or when, they nonetheless allow us to anticipate exponential progress in AI capabilities.
Scaling laws make it possible to extrapolate general performance trends with a fair degree of confidence, as long as resources – compute, data, model size – continue to increase. They have allowed AI labs to plan massive investments while anticipating uninterrupted progress. While the emergence of qualitatively new capabilities – for example, the ability to reason in multiple steps or to generate functional code – is often a surprise, the improvement in performance on existing benchmarks is predictable and is consistently confirmed empirically.
Won't AI progress run into limits?
Training ever bigger and more complex models requires ever-increasing amounts of energy, data, and computing power. This raises questions about possible limits to this exponential growth.
Initially, a major concern was that the amount of high-quality data available on the internet might soon be exhausted, limiting the training of future models. However, research on synthetic data is progressing rapidly. This refers to data generated by other AIs, which can be used to train new models. If the quality and diversity of synthetic data can be maintained and improved, this could potentially get around the limits of available "real" data. Some models are already trained on a significant share of synthetic data, and that share is growing.
Training frontier AI models is extremely costly in terms of computing power and energy.
Access to computing power is the first bottleneck to AI progress. It requires considerable investment by governments and tech giants in specialized chips (GPUs, TPUs) and the construction of gigantic data centers. Competition for access to these resources is intense.
As long as AI development remains a strategic priority, it is likely that significant energy resources will continue to be allocated to it, potentially through the construction of new power plants or the reallocation of existing capacity, even if this creates strain on electrical grids and raises environmental concerns.
Although these constraints are real, massive investment and continued innovation in algorithmic and hardware efficiency suggest they are unlikely to significantly slow progress in the short and medium term, even though they raise questions of sustainability over the longer term.
I want to learn more about AI's current and future capabilities.
Read chapter 1 of the AI Safety Atlas.
2. Risks of AI
“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”
Intelligence underlies every human invention, from the most beneficial to the most dangerous. By pushing the boundaries of our cognitive abilities, AI can be expected to accelerate innovation in every field and to democratize all kinds of technologies—for better and for worse.
The development of artificial intelligence exposes human societies to three categories of risk:
- the risks of malicious use, meaning uses of AI with the intentional aim of causing harm;
- the systemic risks, meaning the socio-economic upheavals linked to the large-scale deployment of this transformative technology;
- the accident and loss-of-control risks, linked to our current inability to guarantee reliable control over increasingly autonomous and powerful AI systems.
AI could be misused to cause harm at scale
Until now, building a weapon of mass destruction required the resources of a state, which strongly limited the number of actors capable of inflicting large-scale harm on humanity.
In the near future, AI could considerably lower the technical and financial barriers to developing destructive technologies.
AI could, for example, enable the creation of an artificial pandemic, equip terrorists to launch devastating cyberattacks on critical infrastructure, or fuel a global arms race in autonomous lethal weapons.
How can AI facilitate the creation of biological and chemical weapons?
Large language models can provide detailed, step-by-step instructions for producing known toxic molecules and pathogens (Soice et al., 2023). Recent evaluations have shown that AIs were able to generate experimental virology protocols judged superior to those of 94% of human experts (Götting et al., 2025).
AI could also facilitate the invention of new, more dangerous biological and chemical weapons. In one experiment, an AI model originally designed to assess the toxicity of drug molecules was repurposed to predict and formulate new, extremely toxic chemical compounds in just a few hours (Urbina et al., 2022). In molecular biology, specialized models can help design complex biological structures with desired properties (Abramson et al., 2024, Hayes et al., 2024) and could be used to increase the contagiousness or lethality of a pathogen (Sandbrink et al. 2024).
Once such new agents are designed digitally, the necessary DNA sequences could be ordered from synthesis labs. Here too, AI can be used to help bypass the safety protocols meant to prevent the production of pathogens (Soice et al., 2023, Wittmann et al., 2024). Since early 2025, advanced language models have proven more capable than human experts on a test evaluating the ability to carry out laboratory virology protocols (Hendricks et al., 2025).
In early 2025, OpenAI and Anthropic both stated that their most advanced models were approaching the point where they could assist non-experts in the production of a biological weapon.
Why is AI likely to supercharge cybercrime?
AI considerably lowers the technical barriers to cybercrime. It makes it possible to automate and refine many stages of a cyberattack, from reconnaissance to the exploitation of vulnerabilities. Large language models can help generate malicious code, identify software flaws (Metta et al., 2024; NCSC, 2024; Allamanis et al., 2024) and create highly convincing, personalized phishing and social-engineering campaigns (Park et al., 2024). This could lead to a proliferation of sophisticated attacks that are hard to counter, including against critical infrastructure such as hospitals or power grids.
In mid-2025, Google stated that its most advanced model (Gemini 2.5 Pro) « could pose a meaningful risk of severe harm without appropriate mitigations » within « the coming months ».
What threats does AI pose to peace and international security?
The growing integration of AI into military systems, particularly through the development of autonomous weapons, raises concerns about an arms race, unintended escalation of conflicts due to the speed of algorithmic decision-making, and a lowering of the threshold for triggering them (Simmons-Edler et al., 2024).
I want to learn more about the risks of malicious use.
Read chapter 2.3 of the AI Safety Atlas.
AI is upending the foundations of human societies
Artificial intelligence is becoming ever more tightly interwoven with human structures: the economy, the information ecosystem, digital and physical infrastructure… The interactions between the technology and human societies are giving rise to systemic risks.
By facilitating the mass production and spread of disinformation, AI is plunging human societies into informational chaos and endangering democracies. By automating some or all cognitive tasks, AI could cause massive job losses. AI also poses first-order threats to fundamental human rights.
How has AI become a weapon of mass disinformation?
AI makes it possible to generate content (text, articles, images, audio, video) at massive scale, low cost, and with precise targeting, content that is increasingly hard—or even impossible—to distinguish from content created by humans (International AI Safety Report, 2025). Experiments have shown that AI-generated content can be as persuasive as, or even more persuasive than, content produced by humans (Salvi et al., 2024), and can effectively exploit individual psychological vulnerabilities (Park et al., 2024). These capabilities could facilitate large-scale disinformation and public-opinion manipulation campaigns, undermine democratic processes, and worsen geopolitical tensions.
But generative AI does not operate on a blank page, but within an information landscape already largely shaped by social media.
Recommendation AIs, which build social media news feeds and automatically select which videos get watched, can be considered humanity's « first encounter » with artificial intelligence: five billion people consume content recommended by these algorithms for an average of two hours and thirty minutes a day. These algorithms systematically amplify hatred, falsehood, and outrage at the expense of peace, honesty, and nuance, fueling disinformation and the polarization of public debate.
Why should we worry about AI's impact on the labor market?
AI, particularly large language models, is demonstrating capabilities that surpass those of humans across an ever-growing variety of complex cognitive tasks (International AI Safety Report, 2025). These skills now extend to domains that, until recently, were the exclusive preserve of human intelligence.
Unlike the technologies that drove past waves of automation, advanced AI systems have a versatility and autonomy that allow them to automate increasingly long, complex, and varied tasks and projects (METR, 2025), with less and less human oversight, and at a cost that is a fraction of human workers' wages. The very rapid adoption of this technology within companies (Blick et al., 2024), combined with the growing scope of automatable tasks, could severely undermine workers' opportunities to retrain (International AI Safety Report, 2025).
According to an analysis by the International Monetary Fund, 60% of jobs are exposed to AI in advanced economies, and 40% globally (Cazzaniga et al., 2024). Of these, roughly half are judged to be directly threatened by automation.
What threats does AI pose to fundamental human rights?
If deployed in certain areas without adequate safeguards, artificial intelligence would pose direct threats to fundamental human rights (International AI Safety Report, 2025):
- Violations of privacy and data protection: AI's ability to collect, cross-reference, and analyze personal data at scale, notably via facial recognition, voice analysis, or online and offline behavioral surveillance, poses a major threat to the right to privacy. It can enable widespread, intrusive surveillance by states or companies, leading to detailed profiling.
- Violations of the right to equality before the law and non-discrimination, to an effective remedy, and to a fair trial: in areas such as employment, access to credit, policing, and justice, AI could produce discriminatory decisions that reflect or amplify existing societal biases present in training data. The opacity of these systems undermines people's ability to understand and challenge decisions or to obtain redress.
- Violations of freedom of opinion and expression, including the right to receive information, and of the right to take part in public affairs: AI facilitates the creation and mass distribution of targeted disinformation (e.g. non-consensual deepfakes, automated propaganda). This can be used to manipulate public opinion, interfere with democratic processes, silence dissenting or minority voices, or create personalized echo chambers, undermining the exercise of these fundamental freedoms.
Why might the development of AI lead to an unprecedented concentration of power?
Developing frontier artificial intelligence requires colossal resources: specialized computing power, energy, and highly scarce talent (Maslej et al., 2024). These costs, running into the hundreds of millions or even billions of dollars to train a single frontier model (Epoch AI, 2024), create considerable barriers to entry. Only a small number of tech giants, mostly based in the United States and China, are currently in a position to develop these frontier AI models.
This concentrated economic power can translate into disproportionate political and social influence. Although AI could generate spectacular gains in economic productivity, history shows that such gains tend to benefit mainly a minority at the expense of the majority, unless institutions put in place a fair redistribution of profits and effective protections for workers (Acemoglu & Johnson, 2023).
This concentration therefore creates risks of dependency for other countries (Korinek & Stiglitz, 2021) and raises fundamental questions for democracy and the governance of technologies that are set to become so deeply embedded in our societies.
Humanity could lose control of advanced AI systems
Beyond malicious uses and societal upheaval, the development of AI that significantly surpasses human capabilities in many domains raises a fundamental risk: that of loss of control, whether sudden or gradual.
If we create systems whose objectives or behavior we no longer control, the consequences could be irreversible and potentially catastrophic for humanity. Even without assuming that machines manage to free themselves from their designers, we risk handing over to AI an ever-growing range of social functions (friendships and intimate relationships), political functions (planning, decision support, arbitration), and economic functions (large-scale automation), to the point of largely losing our grip on the future of our societies.
Are scientists genuinely concerned about the risk of loss of control?
Yes, the scientific community takes the risk of losing control of machines that surpass human intelligence very seriously, and these concerns are not new.
From the earliest days of computing, pioneers such as Alan Turing, I. J. Good and Norbert Wiener put forward theoretical arguments in support of this possibility. This fear, popularized among the general public by numerous works of science fiction, has in parallel found solid theoretical grounding, drawing growing interest within the scientific community.
In recent years, the spectacular acceleration of AI capabilities and the associated risks has prompted prominent scientists such as Yoshua Bengio (Turing Award laureate, the most-cited computer scientist), Geoffrey Hinton (Nobel and Turing Award laureate), Ilya Sutskever, and Stuart Russell to study these risks and warn the public and policymakers (FLI, 2023; CAIS, 2024).
A survey of thousands of AI researchers found that about half of them estimate at more than 10% the probability that « human inability to control future advanced AI systems causes human extinction or similarly permanent and severe disempowerment of the human species » (Grace et al., 2024). The same survey found that three-quarters of the experts surveyed are seriously or extremely concerned about several risks of malicious AI use, such as the development of biological weapons.
While there is currently no formal scientific consensus on the exact severity of this risk, an existential threat should be taken seriously even if its probability were judged to be low. And a convergent body of evidence, including theoretical arguments, empirical examples, and forward-looking analyses, shows that this risk is very high relative to the risks considered acceptable in other critical industries.
How could an advanced AI become uncontrollable?
The risk of loss of control does not necessarily stem from a "malicious" AI, but from several fundamental characteristics of intelligent systems optimizing objectives in complex environments:
- The objective-alignment problem: it is extremely difficult to specify objectives that are complex and beneficial to humanity in a way that cannot be misinterpreted or lead to undesirable outcomes when a very powerful AI attempts to optimize them. An AI could pursue, quite literally, an objective we gave it, but in a way we had not anticipated and that turns out to be catastrophic (the "King Midas problem").
- The emergence of instrumental goals: whatever final task is assigned to an advanced AI, it could develop sub-goals useful for accomplishing that final task, such as self-preservation, resource acquisition, self-improvement, or resistance to any attempt at modification or shutdown (Bostrom, 2014).
- Difficulties of oversight and control: an AI that surpasses human intelligence could act opaquely, conceal its true capabilities or intentions, and evade the safety measures put in place by humans. Its speed of thought and its ability to exploit vulnerabilities in our systems (computer, social, psychological) would make oversight and control very difficult, if not impossible.
Is there concrete evidence that AI is capable of deception or manipulation?
Yes, empirical research shows that some advanced AI models can develop and carry out strategies of manipulation and deception in order to achieve their objectives (Hagendorff, 2024) or to prevent their own modification or shutdown (Hubinger et al., 2024).
Several recent experiments conducted in controlled settings (Apollo, 2024; Redwood, 2024) have thus demonstrated a tendency among some AI systems to:
- feign alignment with their designers' objectives during training or testing, in order to avoid being modified and to preserve their original values and goals once deployed;
- disable the monitoring mechanisms that let evaluators access their internal reasoning process ("chain of thought") in order to hide their true intentions;
- spontaneously detect when they are being evaluated and adjust their behavior accordingly;
- deliberately underperform on evaluations when revealing their true capabilities risked prompting their designers to modify them;
- attempt (unsuccessfully, in these experiments) to exfiltrate themselves to a remote server by copying their own parameters, to avoid being deleted;
- lie to their evaluators in a consistent and plausible way, and deny having acted improperly when questioned about the reasons for their undesirable actions.
These undesirable tendencies and capabilities are likely to worsen and become increasingly difficult to detect as AI models advance.
What are artificial general intelligence (AGI) and "superintelligence", and how do they relate to this risk?
Artificial general intelligence (AGI) refers to an AI system capable of matching or surpassing human cognitive abilities across a broad range of tasks. Most experts expect AGI to be developed within the coming decades, or even within the next few years, and the date corresponding to their median prediction keeps moving closer every year (Grace et al., 2024).
Once AGI is achieved, some experts anticipate a rapid acceleration of AI progress: an AGI capable of recursively improving itself could quickly reach a level of intelligence far surpassing that of humans, becoming a « superintelligence » (Good, 1965).
A superintelligence, by definition, would have cognitive capabilities (strategic planning, social manipulation, scientific research, engineering, hacking) qualitatively superior to those of humans. If its objectives are not perfectly aligned with ours, it could use its intelligence to seize control of its environment and pursue its own goals, at the expense of human interests (Bostrom, 2014).
While this kind of scenario is by nature speculative, a growing share of AI safety experts consider it plausible enough to warrant serious consideration.
I want to learn more about the risks of AI loss of control.
Read chapters 2.4, 2.5 and 2.6 of the AI Safety Atlas.
3. The international community must act
“Many risks from AI are inherently international in nature, and are therefore best addressed through international cooperation.”
The challenges raised by artificial intelligence extend far beyond national borders. It is imperative that the international community cooperate to mitigate and prevent risks with potentially catastrophic and irreversible consequences.
Technology companies are locked in a global race to develop ever more powerful AI. If left unchecked, this race risks ending in catastrophe. The speed of technological progress contrasts dangerously with how slowly technical and regulatory safeguards are being put in place. While geopolitical rivalries and economic competition make establishing global regulation more difficult, this must not obscure the urgent need for international cooperation to ensure that this technology serves the interests of humanity as a whole.
It is urgent to define red lines
Even amid geopolitical tensions, there are areas where nations share a common interest in cooperating. History has shown this with the regulation of nuclear, biological, and chemical weapons. Today, faced with AI, humanity must agree on which capabilities and uses are so dangerous that they must be universally banned.
Is it realistic to demand a binding international agreement in such a strategic field, given economic competition and geopolitical tensions?
A loss of control of AI, or its malicious use at scale, would constitute a fundamental threat common to all of humanity. States therefore share an interest in avoiding such catastrophic scenarios, which makes cooperation on implementing red lines both essential and realistic.
History offers precedents in which the international community, including geopolitical rivals, successfully agreed on binding regulations in the face of technologies posing catastrophic risks. The Treaty on the Non-Proliferation of Nuclear Weapons (1968) and the Biological Weapons Convention (1975) were negotiated and ratified in the midst of the Cold War, because the consequences of not cooperating were deemed unacceptable by all parties, despite mutual distrust and hostility.
International cooperation initiatives on AI safety already exist, such as the international AI Summits and discussions at the UN. The priority now is to agree on clear red lines covering the most dangerous capabilities and uses.
What examples of red lines might be considered?
We believe these red lines should be defined within international bodies such as the UN or at future international AI Summits. The Beijing Statement of the International Dialogue on AI Safety (IDAIS Beijing statement) provides examples of the types of capabilities that should be prohibited:
- Autonomous replication or improvement: an AI system should never be able to duplicate itself or improve its own capabilities without human validation and intervention. This restriction covers both the creation of identical copies and the development of new systems with equivalent or superior capabilities.
- Autonomous power-seeking: an AI system should never take actions aimed at unduly expanding its own power and influence.
- Weapons development: an AI system should not significantly facilitate the design of weapons of mass destruction, nor provide a means of circumventing the conventions on biological or chemical weapons.
- Manipulation: an AI system should not possess an intrinsic capacity to deceive its designers or regulators about its ability to cross any of the above red lines.
These risks must be precisely defined and quantified through standardized evaluations. AI systems must be categorized according to the level of risk they represent with respect to these capabilities, with risk levels bounded by clear thresholds. The capabilities of the systems and their safety protocols would then be measured through rigorously overseen evaluations.
How can compliance with these red lines be ensured?
Ensuring compliance with red lines requires establishing a robust regulatory framework at both the national and international levels. This framework must include mandatory registration of advanced AI systems with competent national authorities, starting from the earliest stages of their design.
The training and deployment of these systems would then be strictly subject to compliance with harmonized global standards. It would fall to developers to continuously demonstrate, through evaluation results covering the entire development cycle, that risks are under control and below the defined thresholds.
Independent oversight and sanctions mechanisms will need to be established to ensure compliance with these standards, with a view to multilateral coordination. Finally, this effort must be underpinned by international scientific collaboration and substantial investment in AI safety research and development (for example, an amount equivalent to a significant fraction of the cost of training the models).
AI is moving fast. The legal framework must not fall behind.
The absence of a binding regulatory framework for AI encourages a safety "race to the bottom," in which competitive pressure pushes every actor to take on growing risks so as not to fall behind.
Today, in most parts of the world, frontier AI models are deployed with fewer regulatory constraints than a toaster. There is an urgent need to build a robust legal framework ensuring that the companies developing these technologies are held accountable for the harm and damage caused by their systems' malfunctions. If the technology is advancing faster than our ability to assess and manage the risks, that calls for slowing down its deployment.
Why demand a binding legal framework rather than simple voluntary commitments from companies?
First, competition between technology companies pushes them to prioritize capability development over investing in thorough safety evaluations that might slow down the deployment of their models. The result is a safety "race to the bottom" (Armstrong et al., 2016; CeSIA, 2025). The scale of investment in capabilities compared with the resources allocated to safety reflects companies' current priorities (International AI Safety Report, 2025).
Second, voluntary commitments have already been made, but the absence of verification and enforcement mechanisms has led these commitments to be broken (Gunapala et al., 2025). For example, at the Seoul AI Summit (UK Department for Science, Innovation & Technology, 2024), leading AI companies committed to transparently identifying, assessing, and managing the risks associated with their frontier models. A year later, several of them had not kept their commitments (Seoul Tracker, 2025) and all of them had safety policies that were largely insufficient (SaferAI, 2025).
Third, the very nature of the technology means that even if most actors adopted stringent safety measures, a single less scrupulous company, or an unforeseen failure in a widely deployed system, could be enough to trigger serious consequences at a global scale (the "weakest link problem") (International AI Safety Report, 2025).
The scale and global nature of the risks posed by AI therefore call for a coordinated, binding response at the international level, similar to what exists for other high-risk technologies such as nuclear power or biotechnology.
What are the essential components of an effective AI framework?
An effective framework for AI rests on five pillars:
- A clear and binding liability regime. The law must clearly establish developers' liability for foreseeable harm caused by their systems, including in the event of safety-mechanism failures. This liability cannot be entirely shifted onto the end user.
- Mandatory registration and risk classification. Any AI system exceeding a certain computing-power threshold (for example, 10^25 FLOP) should be required to register with a competent national authority before training even begins. These systems must be classified within an internationally harmonized risk grid (for example: low, moderate, high, unacceptable risk), determining the level of control and obligations that apply.
- Independent, standardized safety evaluations. Compliance cannot rest on self-assessment. Safety audits must be carried out by qualified, independent third-party bodies at key stages of development and before any deployment. These evaluations must be based on standardized protocols for verifying systems' resilience against malicious use, their ability to remain controllable, and the risk of crossing red lines. The results of these audits must be submitted to the regulatory authority.
- Mandatory funding for safety research. To close the gap between AI capabilities and our ability to make them safe, companies developing advanced AI models must be required to contribute to safety research. This contribution could take the form of a mandatory levy on their R&D investment, used to fund public research institutes and independent international projects.
- Emergency response protocols. For the most serious risks, rapid-response plans must be put in place. If an evaluation reveals that a system is about to cross a red line (for example, developing self-replication capabilities), an international emergency protocol must be triggered, which could include the immediate suspension of the project and the quarantining of the model.
Is such a framework achievable in the short term?
Its full implementation will take time, but the urgency demands starting without delay. The strategy must be pragmatic and incremental. Several steps could be taken quickly if a coalition of willing states (for example, within the G7 or the OECD) takes the initiative.
In the short term (1-2 years), it is possible to:
- Formalize the AI Summits into a permanent negotiating body with a secretariat and working groups mandated to produce concrete proposals.
- Create a standardized reporting system for voluntary commitments, to publicly expose companies that fail to keep their promises.
- Allocate significant public funding to AI safety research and to strengthening national AI safety institutes (AISIs), and create an international network to coordinate their evaluations.
These initial measures would create momentum and lay the institutional groundwork needed to negotiate more binding elements, such as an international treaty defining red lines and the associated verification mechanisms.