Research
We are building a multi-step pathway to safe advanced AI.
At the heart of our breakthrough research is the Scientist AI, a novel approach conceived by Yoshua Bengio that represents a distinctive safety-centered path towards ASI.
The Scientist AI is inspired by an ideal scientist: a mind that has internalized the laws of nature and uses them to make predictions, but without predilection about how things unfold. It is a highly intelligent machine that uses probabilistic reasoning to understand the world, but with no hidden goals or preferences. Its predictions are transparent, auditable and verifiable.
As we build towards safe advanced AI, we expect the Scientist AI to accelerate scientific breakthroughs, provide guardrails and oversight for agentic AI systems while advancing our understanding of the risks posed by AI and how to avoid them.
Yoshua Bengio1,2,3, Oliver Richardson1,2,3, Tomáš Gavenčiak6,7, Michael Cohen4, Rory Svarc6, Damiano Fornasiere1,3, Gaël Gendron1, David Hyland8, Aton Kamanda1, Adam Oberman1,5, Francis Rhys Ward1, Anna Gavenčiak6, Jacob Livingston Slosser6,9, Vincent Mai1, Iulian Serban1, Joumana Ghosn1
1LawZero, 2Universit´e de Montréal, 3Mila, 4University of California, Berkeley, 5McGill University, 6Arb Research, 7Center for Theoretical Study, Charles University in Prague, 8University of Oxford, 9Sapien Institute
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. This paper presents a formal safety argument for the Scientist AI (SAI) Predictor, a model trained to approximate the Bayesian posterior conditioned on a fixed dataset of “epistemically contextualized” natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without
itself being an agent that selects outputs to steer those consequences. The separation rests on two design choices, one based on data representation and the other on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. Together with a posterior-seeking training objective and the many cross-context semantic constraints, this is intended to drive the Predictor toward calibrated, cautious predictions. The training process is constructed so that the downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails, and we call the resulting system that withholds high-harm predictions the Guardrailed Predictor. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries. We argue that such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.