Where does our information come from?
From public documents, and only from them.
Our starting point is the MITRE ATLAS catalogue. MITRE is an American non-profit organisation that also maintains ATT&CK, the reference cybersecurity teams use to describe conventional attacks. ATLAS does the same for AI: a public knowledge base of attack techniques targeting AI systems, built from real-world attacks and demonstrations by security teams.
ATLAS in two minutes
A technique is a way of attacking: for instance poisoning the data a model learns from, or getting around its safeguards to make it say what it refuses to say.
A case is a documented attack, described step by step.
A grade says where each technique stands: feasible (described, never shown), demonstrated (achieved in the lab) or realized (observed in a real attack). The two transitions between these grades are what our home page calls “from the lab to the real world”.
ATLAS can be browsed freely at atlas.mitre.org.
To find out what still holds, we read what is published by those who test these systems: public evaluation institutes such as the UK AI Security Institute, the labs that evaluate their own models, incident reports, and the providers who describe their controls.
How do we sort it?
We look for locks. A lock is a defense that still holds: what is missing for an attack to become possible.
The essential rule: the lock must be written in black and white in a published document. We never guess it. When nobody has checked whether a defense still holds, we say exactly that: not tested. That is information in its own right, and nobody publishes it.
Going further: the eight types of defense
Capability (the model cannot do it yet), resource (money, compute), identity (payment, account, ID), access (privileged rights), duration (lasting without being detected), tacit knowledge (know-how absent from written sources), installed control (someone deliberately maintains it: rate limits, filters, sandboxes) and not tested.
Each lock gets a single type. If two types fit, it is marked not tested rather than settled by guesswork.
What do we refuse to publish?
For obvious reasons, an entry never says how to get around the defense.
Often, rewording is enough. “Identity checks by compute providers block an autonomous agent from opening an account” is a lock. We do not name the provider that skips them: that would be an attack manual.
When rewording is impossible, the entry is neither published nor kept. This happened twice among the thirteen candidates examined for the list of 26 September, from otherwise excellent sources.
What remains hard to measure?
- Both steps are being cleared faster and faster: from idea to lab demonstration, then from lab demonstration to a real case. In the MITRE ATLAS catalogue, roughly 1.5 times faster each year for the first, 2.2 times for the second, after correcting for the catalogue's own growing activity, a correction that cannot be measured before 2025. It is hard to tell a field that is accelerating from a better-kept catalogue; our hypothesis is that both effects are present. Without the correction the figures would be higher: we publish the smaller ones.
- Public tests are run without defenders. Evaluators' test ranges have no monitoring team and no detection. What the same attack would achieve against a well-defended system remains to be measured.
- What held yesterday may give way tomorrow. Four model limitations published in March had been overcome by May. Four observations, not a general law, but enough: our list gives the state at a date, and each entry carries its own.
Why the list is dated
We first wanted a list that would stay valid over time. That is impossible, and we measured it.
Under a rule written before any counting, only one scenario in ten met the conditions for a durable list. Under a rule also written before any counting, but for the state at a given date, eight candidates out of thirteen meet them. We publish both results side by side, with both rules.
How can you check us?
Everything we claim can be redone by someone else.
- Every lock cites its source.
- Every figure can be recomputed with a single command, from public sources whose version is pinned.
- The article is archived with a permanent identifier (DOI), which dates what we published and prevents it from being rewritten afterwards.
We also publish our failures. Before this list, we tried to build a tool that would automatically reconstruct attack scenarios. The test was set in advance, and the tool did not pass it. We published that negative result, with the means to check it (method note, archived on Zenodo).
What the method produces
The article
The list of defenses that still hold, scenario by scenario, with what has already been observed in the real world for each one. And the finding that comes with it: from the lab to the real world, it is the second step that holds out.
Read the article forthcoming
Already on the home page: the finding in one figure · one defense from the list, as an example