Tom Zahavy
← Home

Creative Chess

What makes a chess idea creative? AlphaZero db, a league of superhuman players that think in different ways, and PuzzleGen, which composes original puzzles that grandmasters called beautiful.

Chess has been called the drosophila of AI, and engines have been superhuman at it for decades. That makes it the right place to ask a different question: not how to win, but what makes a chess idea creative — and whether a machine can have one.

Two projects approach it from opposite ends. AlphaZero db makes a single engine think in many different ways, and finds that a diverse team solves problems none of its members can. PuzzleGen turns generation itself into the task: composing original puzzles, judged by grandmasters on the same aesthetic criteria they apply to human compositions.

Diversity

AlphaZero db

A team of superhuman players that think differently

Artificial Intelligence (AI) systems have surpassed human intelligence in a variety of computational tasks. However, AI systems, like humans, make mistakes, have blind spots, hallucinate, and struggle to generalize to new situations. We explored whether AI systems can benefit from creative decision-making mechanisms when pushed to the limits of its computational rationality.

AlphaZero db is a league of AlphaZero agents, represented via a latent-conditioned architecture, and trained with quality-diversity techniques to generate a wider range of ideas. It then selects the most promising ones with sub-additive planning. AlphaZero db plays chess in diverse ways, solves more puzzles as a group and outperforms a more homogeneous team.

Four robot players around one chessboard, each with a different style.
2×
as many challenging puzzles solved as AlphaZero, including the Penrose positions
+50 Elo
over AlphaZero by picking the specialist player for each opening

Played from different openings, different players in the league specialise in different openings. Diversity bonuses emerge in teams of AI agents, just as they do in teams of people.

Aesthetics

PuzzleGen

Composing puzzles grandmasters call beautiful

A grid of chess puzzles generated by PuzzleGen.

Generative models are good at reproducing what they have seen, but genuinely creative, counter-intuitive output is much harder. Chess puzzles are an ideal test bed: strong players recognise beauty instantly, yet the elements that make a position surprising and elegant are hard to pin down.

PuzzleGen starts from a generative model trained on 4.4M Lichess puzzles, then applies reinforcement learning with rewards computed from chess-engine search statistics. Rather than optimising for difficulty, the rewards target the things that make a puzzle good: uniqueness, counter-intuitiveness, novelty, and realism. The result increases counter-intuitive puzzle generation roughly 10× — from 0.22% for the supervised model to 2.5% — beating both the training data’s own rate (2.1%) and the best Lichess-trained baseline (0.4%).

How it works

Each position the model samples is handed to a chess engine (Stockfish and AlphaZero), which scores it against four verification rewards. High-scoring puzzles are appended back into training, so the model keeps learning to produce more of them — a reinforcement-learning loop with machine feedback.

PuzzleGen pipeline: a generative model trained on the Lichess dataset samples positions that a chess engine scores against four verification rewards (counter-intuitiveness, uniqueness, novelty, realism); high-scoring puzzles are appended back into training and passed through aesthetic checks into a booklet.
The reinforcement-learning-with-machine-feedback loop (Figure 1 of the paper).

The rewards

The design of the reward is the heart of the method. Each is derived from engine search statistics rather than hand-authored heuristics:

Counter-intuitiveness — the search gap

Rewards positions whose winning move looks bad to a shallow search but is clearly best under deep search. That gap between first impression and truth is what makes a move surprising.

Uniqueness — best move ≫ second best

A real puzzle has one solution. Using Stockfish search statistics, we reward positions where the top move dominates every alternative, so the answer is unambiguous.

Novelty — w.r.t. the training data

The position should be genuinely new — not a near-duplicate of something already in the 4.4M-puzzle Lichess training set the model learned from.

Realism — a legal, natural position

The board must be legal and look like it could have arisen from a real game, rather than a contrived arrangement of pieces.

What the grandmasters said

Three world-renowned experts — all noted authors on chess aesthetics — reviewed the generated booklet and explained what made their favourites appealing.

“A valuable chess puzzle should be original and creative, with a surprising, counter-intuitive key move and a smart follow-up. The ideal puzzle is also aesthetically pleasing and offers a satisfying, flowing solution.”

Amatzia Avni · IM, chess compositions

“I favor natural positions resulting from reasonable play by both sides. Puzzles lose my interest if one side's pieces are clearly misplaced, or if a complex solution yields a minimal advantage.”

Matthew Sadler · Grandmaster

“AI is now capable of generating interesting chess positions, beyond just “mining” databases. The positions in this booklet represent a pioneering step in this human–AI partnership.”

Jonathan Levitt · Grandmaster

All three independently singled out one position as beautiful — its key move, a rook sacrifice, was described as “unorthodox” and “by no means a natural or obvious sacrifice.” It’s Puzzle 1 in the board below.

Try the puzzles

Nine of the generated positions — drag the pieces to play out your line. Open any one on Lichess for the full analysis board, or browse the whole set on chess.com.

Puzzle 1 of 9White to move
8
7
6
5
4
3
2
a1
b
c
d
e
f
g
h
Analyze on Lichess ↗

Watch

GothamChess
Hikaru Nakamura
Jen Shahade

Open source

Open models for puzzle generation

With Aalto University: controllable puzzles from diffusion models

A follow-up builds on PuzzleGen’s reward-driven approach and releases the first open-weight models for chess puzzle generation. It replaces left-to-right generation with a masked diffusion model, so puzzles can be conditioned on a tactical theme, a target rating, or a partially specified board.

Training the model to predict the best move alongside the position improves solution uniqueness by 11.6%, and reinforcement learning with Denoising Diffusion Policy Optimization raises the yield of unique, theme-matching puzzles by 89.1%. The code, weights and an interactive demo are public.

Papers

  1. T. Zahavy, V. Veeriah, S. Hou, K. Waugh, M. Lai, E. Leurent, N. Tomašev, et al. Diversifying AI: Towards Creative Chess with AlphaZero. 2023. arXiv:2308.09175.
  2. X. Feng, V. Veeriah, … T. Zahavy Generating Creative Chess Puzzles. NeurIPS 2025. arXiv:2510.23881.
  3. V. Veeriah, F. Barbero, … T. Zahavy Evaluating In Silico Creativity: An Expert Review of AI Chess Compositions. 2025. arXiv:2510.23772.
  4. A. Selkee, S. Rissanen, X. Feng, T. Zahavy, E. Malmi Conditional Generation of Creative Chess Puzzles with Diffusion Models. 2026. arXiv:2609.38577.