Tom Zahavy
← Home

International Mathematical Olympiad · Bath, July 2024

AlphaProof

An AI system that taught itself to prove mathematical theorems in Lean, reaching silver-medal performance at the International Mathematical Olympiad through continuous reinforcement learning.

Illuminated proof steps threading through a field of dark mathematical statements.
28 / 42
points at IMO 2024
Silver
medal standard
First
AI system to reach medal level
P6
hardest problem of the year, solved by five contestants

An agent that self-taught itself Mathematics in Lean and achieved a silver-medal standard in the International Math Olympiad. Starting from a pre-trained LLM exhibiting proficiency in mathematics, AlphaProof embarked on a lifelong Reinforcement Learning journey: proving and disproving theorems, learning from them, and getting better and better.

During the IMO competition, AlphaProof first invented variations of the competition problems and performed a second (test-time) Reinforcement Learning phase, where it tested these variations while attempting to prove the main problems. Over time, its comprehension improved, allowing it to solve all the number theory and algebra questions, including a P6 problem that only five human competitors managed to solve.

“The fact that the program can come up with a non-obvious construction like this is very impressive, and well beyond what I thought was state of the art.”
Sir Timothy Gowers · Fields Medallist

Publication

Olympiad-level formal mathematical reasoning with reinforcement learning

Hubert, Mehta, Sartran, Horváth, Zahavy, et al. · Nature, 12 November 2025

From the blog

How we achieved an IMO medal, one year before any other AI system

My account of AlphaProof: the lifelong RL loop that taught it mathematics in Lean, the test-time RL phase that let it crack Problem 6, and what the result does and doesn’t mean.

Coverage

The July 2024 announcement was reported widely; the solutions were graded by Timothy Gowers and Joseph Myers under IMO rules.