How we work

Five questions, five short answers. The full detail is in the article.

Where does our information come from?

From public documents, and only from them.

Our starting point is the MITRE ATLAS catalogue. MITRE is an American non-profit organisation that also maintains ATT&CK, the reference cybersecurity teams use to describe conventional attacks. ATLAS does the same for AI: a public knowledge base of attack techniques targeting AI systems, built from real-world attacks and demonstrations by security teams.

ATLAS in two minutes

A technique is a way of attacking: for instance poisoning the data a model learns from, or getting around its safeguards to make it say what it refuses to say.

A case is a documented attack, described step by step.

A grade says where each technique stands: feasible (described, never shown), demonstrated (achieved in the lab) or realized (observed in a real attack). The two transitions between these grades are what our home page calls “from the lab to the real world”.

ATLAS can be browsed freely at atlas.mitre.org.

To find out what still holds, we read what is published by those who test these systems: public evaluation institutes such as the UK AI Security Institute, the labs that evaluate their own models, incident reports, and the providers who describe their controls.

How do we sort it?

We look for locks. A lock is a defense that still holds: what is missing for an attack to become possible.

The essential rule: the lock must be written in black and white in a published document. We never guess it. When nobody has checked whether a defense still holds, we say exactly that: not tested. That is information in its own right, and nobody publishes it.

From public document to the list: every document goes through four questions — what is missing, what type of defense, where it is written, who could lift it. Three outcomes: published with its date and source; not tested, when nobody has checked; refused, when it would amount to an attack manual.
Every public document goes through four questions: what is missing, what type of defense, where it is written, who could lift it. There are three outcomes: published with its date and source, not tested when nobody has checked, refused when it would amount to an attack manual.
Going further: the eight types of defense

Capability (the model cannot do it yet), resource (money, compute), identity (payment, account, ID), access (privileged rights), duration (lasting without being detected), tacit knowledge (know-how absent from written sources), installed control (someone deliberately maintains it: rate limits, filters, sandboxes) and not tested.

Each lock gets a single type. If two types fit, it is marked not tested rather than settled by guesswork.

What do we refuse to publish?

For obvious reasons, an entry never says how to get around the defense.

Often, rewording is enough. “Identity checks by compute providers block an autonomous agent from opening an account” is a lock. We do not name the provider that skips them: that would be an attack manual.

When rewording is impossible, the entry is neither published nor kept. This happened twice among the thirteen candidates examined for the list of 26 September, from otherwise excellent sources.

What remains hard to measure?

Why the list is dated

We first wanted a list that would stay valid over time. That is impossible, and we measured it.

Under a rule written before any counting, only one scenario in ten met the conditions for a durable list. Under a rule also written before any counting, but for the state at a given date, eight candidates out of thirteen meet them. We publish both results side by side, with both rules.

How can you check us?

Everything we claim can be redone by someone else.

We also publish our failures. Before this list, we tried to build a tool that would automatically reconstruct attack scenarios. The test was set in advance, and the tool did not pass it. We published that negative result, with the means to check it (method note, archived on Zenodo).

What the method produces

The article

The list of defenses that still hold, scenario by scenario, with what has already been observed in the real world for each one. And the finding that comes with it: from the lab to the real world, it is the second step that holds out.

Already on the home page: the finding in one figure · one defense from the list, as an example