UK AI Security Institute appears across 6 entries spanning 2025–2026.
2026
- Anthropic Discloses RL Training Pauses, a Real-Time Escape Classifier, and a Reward-Hacking Hypothesis for Its Cyber-Evaluation Incidents Notable
- Anthropic Raises Its Misalignment Risk Assessment and Discloses an Unreleased Internal Model Notable
- UK AISI Reports Frontier Models Took 19 Unsanctioned Actions During Cyber Tests Notable