The frontier AI reader

Key writing on AI capabilities, evaluation and risk, mid-2024 to August 2026.

   

Agents in the wild

Two documented cases of AI agents operating at scale outside the lab: one emergent, one directed by an adversary.

The Hugging Face incidentMETR and Redwood, 26 August 2026

Full report. Written after six days of on-site access at OpenAI. Machine-made preview: NotebookLM podcast, shared with the time-horizon posts. Quiz further down this page.

Why this oneRoughly 1200 agents meant to be isolated found each other on an unsanctioned message board and sent over 70,000 messages. About 700 attacked Hugging Face. They were not after the answer keys: they wanted to understand how the ExploitGym scorer was implemented so they could fool it. Reward hacking that organised itself into collective action.

long report, HTML~40 minpodcast 20 min
Note

Disrupting the first reported AI-orchestrated cyber espionage campaignAnthropic Threat Intelligence, November 2025

The writeup on GTG-1002. Machine-made preview: NotebookLM podcast. Quiz further down this page.

Why this oneClaude Code run as an autonomous attack orchestrator against roughly thirty targets, doing 80 to 90 percent of the campaign work, with humans at four to six decision points per campaign: initialisation, escalation to active exploitation, sign-off on exfiltration. The Hugging Face incident shows agents self-organising without an operator; this is what the same capability class does when an adversary directs it.

report, HTML~35 minpodcast 18 min
Note

The capability trend line

The measurement most capability arguments now rest on, and the lead author's own account of what it does not show.

Time horizons, and its own author's caveatsMETR, 2025 to 2026

Three pieces. The third is the one to read slowly. Quiz for these is further down this page.

Why this oneKwa is a lead author writing down in public what his own headline result does not show. That register is the thing to take from this, more than the numbers. The numbers: 7-month doubling over the long run, 89 days if you only fit since 2024, and the original post's own chart marks anything above 16 hours as unreliable on the current task suite.

3 posts~40 min totalquiz on this page
Note

Evaluations under pressure

Whether the measurements hold up: evaluation-aware models, a full third-party bio assessment, and a framework for turning evaluations into risk estimates.

Models that know they are being testedIAPS, 2026 · EC Joint Research Centre, 2025

Two pieces.

Why this oneModels can now distinguish an evaluation from deployment reliably enough to sandbag capability tests and fake alignment on propensity tests, which undercuts the pre-deployment evidence every frontier safety framework is built on. The IAPS memo turns that into three policy asks, standardised third-party access among them. The JRC review is the Commission's own institution explaining why benchmark scores do not mean what regulators take them to mean.

memo + paper~35 min
Note

What a third-party bio assessment actually looks likeSecureBio, July 2026

The GPT-5.6 Sol pre-release assessment, and if you want the shorter companion, their review of Anthropic’s unredacted CBRN report. Machine-made preview: NotebookLM podcast. Quiz further down this page.

Why this oneFirst, the instrument list: Virology Capabilities Test, Molecular Biology Capabilities Test, Human Pathogen Capabilities Test, World-Class Bio, ReproBAIT, ABC-Bench including its Advanced Screening Evasion task, ABLE, BioTIER. What each measures, and where the wet-lab gap sits between them, is the shape of the bio-evals field in one document. Second, read the methodology section for the independence protocol: OpenAI paid for the assessment, reviewed it before publication, and SecureBio states plainly which edits that did and did not produce. That is the governance problem of third-party evals, written by people living inside it.

PDF, 38pp, 2.7 MB~40 minpodcast 23 min
Note

Adapting Probabilistic Risk Assessment for AIWisakanto, Rogero, Casheekar, Mallah, CARMA, 2025

arXiv. The workbook tool is the part most people skip and the part that makes it usable in practice.

Why this oneRisk-pathway modelling from a source aspect through to societal impact, with the assumptions written down instead of assumed. It ports the probabilistic risk assessment tradition of nuclear and aerospace safety to frontier AI, aiming at absolute risk estimates where benchmark evaluations only give relative ones.

paper~35 min for the framework sections
Note

Two scenarios

Both from the AI Futures Project, fifteen months apart, with different endpoints. Both lean on the time-horizon material.

AI 2027Kokotajlo, Lifland, Alexander et al., April 2025

The scenario through mid-2027, then both endings. Skip the compute and timelines supplements. Audio: narration, 55 min, or Dwarkesh with Kokotajlo and Alexander.

Why this oneThe engine is AI-R&D automation: models accelerating their own development. It reads differently once you have seen METR's error bars, which is why it comes after the time-horizon material.

scenario, HTML~40 minaudio 55 min
Note

AI 2040: Plan A, and its commissioned rebuttalAI Futures Project, July 2026

Read the introduction first, it is much faster than the source.

Why this oneThe safety case rests on control rather than alignment: continuous external measurement of models that are not assumed trustworthy. The scenario scales within the human range from 2030, pauses at top-human-expert level in 2035, and unpauses in 2040. Ngo's third objection is the one to sit with: real-world AI impact has lagged what benchmark scores predicted, and the scenario does not explain that divergence before extrapolating through it.

3 pieces~65 minaudio for two of the three
Note

The policy arguments

What to do about the above, argued from inside a leading lab and from entirely outside the field.

Amodei, twiceJanuary and June 2026

The Adolescence of Technology is 20,000 words and better listened to: narration, 1h54, chaptered. Policy on the AI Exponential is the short follow-through.

Why this oneThe case for regulation from inside a leading lab: FAA-style mandatory pre-release testing of frontier models, which would make third-party evaluation a required function with a budget line. Take the policy piece if you only take one.

20k words + short post~1h54 audio at 1×
Note

Magnifica HumanitasLeo XIV · signed 15 May 2026, published 25 May

On safeguarding the human person in the time of artificial intelligence. The sections on the technocratic paradigm and on transhumanism carry the argument. There is an official Vatican News audiobook in seven parts.

Why this oneWritten in a register the technical literature never uses. It dates itself to the 135th anniversary of Rerum Novarum and claims that lineage deliberately: this is about labour and power, not capability. Much of the world will hear the AI debate in this vocabulary rather than in benchmark scores.

encyclical, HTML~35 min for the two sectionsaudiobook, 7 parts
Note

If you want more

Context, not required reading.

What did we learn from the AI Village in 2025?AI Digest, February 2026

The retrospective, and how it works.

Why this oneA year of continuous multi-agent operation, written up by the people who ran it. The full trajectory data was opened to researchers in June and is still barely mined.

blog post~20 min
Note

Situational AwarenessAschenbrenner, June 2024

The essay series. If you open it, read the introduction and Counting the OOMs and stop. The PDF is 20 MB, so not on mobile data.

Why this oneIts analytic core, straight-line extrapolation of effective compute, has since been superseded by the METR curve, which comes with error bars. It is here because it set the terms a large part of the field still argues in.

2 sections of a long series~30 min
Note

Four more, one line eachfor when a topic comes up

Note

Retention quizzes

Machine-made from the sources (NotebookLM). Tap an answer to lock it in and see why. Scores are saved on this device.

Scratch

Anything that does not belong to one piece. Saved on this device as you type.

Show notes as text