Politicians on AI
Our tracker analyzes 15,000 official statements by political figures on AI, revealing how key concerns evolve over time and vary across countries.
Built and maintained by Lennart Finke (ETH Zurich) and Markov Grey (CeSIA).
Updated 2026-09-29
About
This is an analysis of statements related to AI, its risks, and policy approaches, coming from officials at governments and intergovernmental organisations. The dataset consists of ~15,000 utterances on AI, drawn from ~200,000 documents put out by 19 governments and intergovernmental organizations. It is particularly useful for AI safety policy work. The complete dataset is public: you can download it as CSV, TSV, or JSON from the explorer page, and every export is archived as a permanently versioned public release.
Method
Our code is open-source, you can read about how we collected and processed the data and how this website renders the data.
Raw documents are collected from official government sources. See below for the full list. We select jurisdictions based on how much volume of AI discussion we expect, and how easy it is to collect data. We select sources per jurisdiction the same way, by volume of discussion, ease of scraping, and copyright. Note that some of the sources we use have significant latency, for example the Congress and House hearings, which are sometimes uploaded only months after the hearing is held. Note also that our sources are neither comprehensive nor necessarily representative. For example, when we say that officials in the United Kingdom talk much more about AI risk than lawmakers in the United States, the reason could be that we missed sources from the United States that cover AI risk. However, we aim to cover the most salient sources for each jurisdiction.
We split up the collected documents into utterances by single speakers. The utterances are filtered with a broad set of AI keywords in one of the 11 languages our sources are from. See here for a list of the keywords per language. Only utterances that contain one of the keywords are advanced to the next step.
The resulting utterances are annotated by three judge AIs (gpt-oss-120b, GLM 5.2, Deepseek Flash V4 0731) called via OpenRouter. See the prompt used. Humans do not check all labels. The judge AIs score each utterance for overarching topics, and annotates with labels: whether the utterance is substantive, whether the named speaker owns the statement, whether the speaker is a public official, whether the quote is direct or paraphrased, et cetera. The judge AI also annotates it with tags for risk and policy strategies. The tags for policy strategy are from the Existential Risk Observatory’s AI Governance and Regulatory Archive. The tags for AI risk are from MIT AI Risk Initiative’s Risk Repository.
The final dataset consists of utterances that passed the keyword filter, and where two of three judges agree on each of the following four conditions: The utterance has an AI score above a fixed threshold, it is substantive, it belongs to one speaker, and its associated speaker is a public official. We also filter out duplicate utterances.
For utterances included in the database, the judge AI also extracts a short quote and, if the utterance is in another language, translates the quote to English. We ensure that the verbatim quote appears in the utterance for all quotes, though abbreviations with brackets along the lines of “[...]” are allowed.
We annotate the sentiment of the utterance towards AI risk, into one of "optimistic", "neutral", "mixed", "concerned". This is done in a separate pass, so as to not bias the judges towards negative sentiment via the risk framing in the first pass. We use a majority vote of GLM 5.2, Deepseek V4 0731, and gpt-oss-120b. In rare cases where all three produce a different label, we annotate with "mixed". See the sentiment annotation prompt here.
A human spot-check on 100 English samples each was conducted to measure agreement between the LLM judge and a human annotator. Agreement was ~75% on inclusion in the database (Cohen's κ=0.52) and precision ~70% on the risk and governance strategy tags (Cohen's κ=0.56).
Data Sources
Quotes are fetched from official web sources of the governments and intergovernmental organisations, with a few fetched indirectly via a mirror as detailed below. Some sources are search APIs: For those, the generic AI-related keywords are used as search terms, analogous to the filtering step from sources that do not require search.
We give a summary of data sources by government or intergovernmental organisation:
- United States: Quotes are sourced from Congressional Records, Congress Committee Hearings via the search functionality of the GovInfo API (api.govinfo.gov), as well as from White House Briefings, Remarks, Fact Sheets via whitehouse.gov HTML pages and Presidential Documents via the Federal Register API (federalregister.gov).
- People's Republic of China: Quotes are sourced from the news and press-room listings of the State Council (gov.cn), the Ministry of Science and Technology, the Ministry of Industry and Information Technology, the Ministry of Foreign Affairs and the Cyberspace Administration of China, as well as from the front page of the People's Daily (people.com.cn). Editions of the People's Daily before December 2024 are no longer served by the publisher and are read from Internet Archive snapshots instead.
- European Union: Quotes are sourced from European Parliament Plenary Verbatim and from parliamentary questions together with the answers given to them, both via the Open Data API (data.europarl.europa.eu), with the older verbatim report PDFs as a fallback for sittings the API does not cover; from European Commission speeches, statements, press releases and Q&As via the presscorner JSON API (ec.europa.eu/commission/presscorner); and from Council of the EU and European Council press releases, which are read from Internet Archive snapshots of consilium.europa.eu.
- NATO: Quotes are sourced from Secretary General transcripts of speeches, press conferences and statements via nato.int HTML pages, enumerated from the site's sitemap.
- United Nations: Quotes are sourced from Security Council and General Assembly Verbatim Records, fetched as PDFs from the symbol-access API of the UN documents service (documents.un.org), and from the automatic transcripts of UN Web TV (transcripts.un.org), which cover the meetings, conferences and briefings that have no official record. The latter are machine transcriptions rather than records, and are marked as such throughout: their wording is not guaranteed verbatim and recognised speaker names are frequently wrong.
- United Kingdom: Quotes are sourced from Parliament Debates via the search functionality of the Hansard API (hansard-api.parliament.uk).
- Germany: Quotes are sourced from Bundestag Debates via the plenary protocol XML files of the Bundestag open-data service, discovered through its index endpoints (bundestag.de).
- France: Quotes are sourced from Assemblée nationale Debates and Sénat Debates via the verbatim XML archives the two chambers publish as ZIP files (data.assemblee-nationale.fr, data.senat.fr), as well as from Élysée speeches, statements and communiqués via elysee.fr HTML pages.
- Canada: Quotes are sourced from House of Commons Debates via the per-sitting Hansard XML of ourcommons.ca, using the English edition, in which French floor speech is translated.
- Japan: Quotes are sourced from National Diet proceedings via the search functionality of the Kokkai Kaigiroku API of the National Diet Library (kokkai.ndl.go.jp).
- South Africa: Quotes are sourced from National Assembly and NCOP Hansard, written ministerial questions and replies, and committee meetings, via the search functionality of the Parliamentary Monitoring Group API (api.pmg.org.za). Committee meetings do not have verbatim quotes available, they are paraphrases.
- Netherlands: Quotes are sourced from Tweede Kamer plenary and committee debates via the Gegevensmagazijn OData API (gegevensmagazijn.tweedekamer.nl), including the interruptions nested inside floor turns, and from Kamerstukken, Kamervragen and Eerste Kamer Handelingen via the search functionality of the KOOP repository (repository.overheid.nl).
- Russia: Quotes are sourced from Presidential Executive Office speeches, transcripts and readouts via the by-date event HTML archive of kremlin.ru, and from State Duma Plenary Stenograms via the database at transcript.duma.gov.ru.
- Singapore: Quotes are sourced from Parliament Hansard via the report backend of the parliamentary reports service (sprs.parl.gov.sg).
- Brazil: Quotes are sourced from Senado Federal floor speeches via the Dados Abertos API (legis.senado.leg.br), which lists every pronouncement of a period with the house's own summaries and subject tags; the verbatim text is then fetched for those whose metadata matches the keywords.
- Mexico: Quotes are sourced from Senado de la República plenary versiones estenográficas via senado.gob.mx.
- Australia: Quotes are sourced from House of Representatives and Senate Hansard via the transcript API behind the Hansard viewer of aph.gov.au.
- Taiwan: Quotes are sourced from Legislative Yuan gazette records via an unofficial mirror called LYAPI (v2.ly.govapi.tw).
- Switzerland: Quotes are sourced from the Amtliches Bulletin of the Nationalrat and Ständerat via the OData service of parlament.ch.