official government evaluation institute (UK AI Security Institute) · catastrophic / loss-of-control risk document
UK’s AI Security Institute: self-replication test scores jumped from 5% to 60%
The UK AI Security Institute’s first public Frontier AI Trends Report — built from two years of government evaluations across more than 30 frontier systems — says capabilities that would be needed to evade human control are improving, with self-replication evaluation success rising from about 5% to 60%, while cyber and biology skills race past expert baselines.
Key facts
- Who
- UK AI Security Institute (AISI) under DSIT — evaluations since Nov 2023 across >30 frontier systems
- Product
- Frontier AI Trends Report — first public evidence-based synthesis (cyber, chem/bio, autonomy, loss-of-control precursors, safeguards)
- Self-replication (RepliBench)
- Early-2023 highest success <5% (11 of 20 evals) → summer 2025 two frontier models >60% (Figure 16); stronger on early obtain-compute/money than later replicate/persist; no spontaneous sandbagging/self-replication evidence yet
- AGI framing (AISI)
- Institute states it is plausible the observed trend may lead to capabilities widely acknowledged as AGI or otherwise transformative AI
- Cyber
- Apprentice-level ~10% early 2024 → ~50% average now; 2025 first model completing expert-level (10+ years human); task-length doubling ~every eight months
- Chemistry & biology
- Exceed PhD expert baselines on open-ended QA (up to ~+60% relative); protocol generation accurate from late 2024 and wet-lab feasible; troubleshooting up to ~90% better than experts
- Safeguards
- Universal jailbreaks for every system tested; some bio-misuse defenses needed ~40× more expert effort between two models six months apart
- Companion Mar 2026
- Multi-step cyber ranges — best run 22/32 steps; performance scales with test-time compute (10M→100M tokens, gains up to 59%); ICS range still limited (avg 1.2–1.4/7)
- Live
- Desk cites AISI measured results and AISI’s own AGI/loss-of-control language — no invented doom dial
Self-replication evaluations that barely cleared 5% success in early 2023 cleared 60% for top models by summer 2025 — and the same institute says AGI-class capabilities are a plausible near-path outcome of the trend they measured.
Desk reading of UK AISI Frontier AI Trends Report
Note
Britain’s AI Security Institute spent two years breaking and measuring frontier models, then published the scoreboard. Self-replication evaluations that barely cleared 5% success in early 2023 cleared 60% for top models by summer 2025. Cyber tasks that once needed a human apprentice now get beaten half the time; expert-level items started falling in 2025; biology helpers already outscore PhD baselines and draft lab protocols that work wet. Safeguards improved in places — and still yielded universal jailbreaks everywhere AISI looked. The institute’s own pages say the trajectory could reach what people call AGI. Desk is not inventing a takeover clock. It is reading the government’s ruler.
Attribution: UK AI Security Institute — Frontier AI Trends Report (aisi.gov.uk; PDF object last-modified 16 Dec 2025). Companion same-institute evaluation: Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios, 16 Mar 2026.
Why it matters
This is not a product launch blog. It is a national security evaluator telling the public that the precursor skills for losing control — self-replication pieces, long-horizon cyber agency, expert-beating wet-lab help — are moving fast while jailbreaks still exist for every tested stack.
Sources
- UK AI Security Institute (AISI)
- Frontier AI Trends Report PDF
- Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios (16 Mar 2026)
Official data. Live values go to the HUD / source product.