Research
What Should Frontier AI Developers Disclose About Internal Deployments?
Jacob Charnock, Raja Mehta Moreno, Justin Miller, William L. Anderson
TAIGR @ ICML 2026
How Does Information Access Affect LLM Monitors’ Ability to Detect Sabotage?
Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas, Francis Rhys Ward
ICML 2026, ICLR AIWILD & Trustworthy AI Workshops
CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
Francis Rhys Ward, Teun van der Weij, Hanna Gabor, Sam Martin, Raja Mehta Moreno, Harel Lidar, Louis Makower, Thomas Jodrell, Lauren Robson
NeurIPS 2025 (Spotlight)