What Should Frontier AI Developers Disclose About Internal Deployments?

Jacob Charnock, Raja Mehta Moreno, Justin Miller, William L. Anderson

TAIGR @ ICML 2026

arXiv

How Does Information Access Affect LLM Monitors’ Ability to Detect Sabotage?

Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas, Francis Rhys Ward

ICML 2026, ICLR AIWILD & Trustworthy AI Workshops

arXiv

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

Francis Rhys Ward, Teun van der Weij, Hanna Gabor, Sam Martin, Raja Mehta Moreno, Harel Lidar, Louis Makower, Thomas Jodrell, Lauren Robson

NeurIPS 2025 (Spotlight)

arXiv