LLMs can't jump
A position paper examining a fundamental limitation of large language models: their difficulty with abductive reasoning, a capacity central to genuine scientific invention.

I explore the fundamental nature of scientific invention, highlighting the critical gap between the ability of humans to create computational systems from physical intuition and current artificial intelligence capabilities. While modern Generative AI excels at pattern recognition (induction) and logical proofs (deduction), we argue it fundamentally lacks the capacity for "abduction"—the intuitive leap required to generate novel explanatory hypotheses.
Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. Ultimately, we propose that integrating physically consistent, multimodal world models is the key to bridging this divide and unlocking true artificial scientific invention.
Position paper, January 2026. Presented at ICML 2026.
Reception
The paper was posted to PhilSci-Archive in January 2026, where it is among the ten most-downloaded papers. It was widely discussed on X, LinkedIn and Reddit, reached the top of Hacker News, and was covered in the science press, by essayists, and in a number of newsletters and explainers. Selected coverage:
Philip Ball, “The Einstein test: what happens when AI tries to rediscover relativity?”. Nature news feature, 9 September 2026.
On 'vintage' language models trained on pre-1911 data and whether they could rediscover general relativity. Presents the paper as the detailed case for the test, and its argument that such a breakthrough requires abduction rather than induction.
Ian Leslie, “Einstein, Churchill, and AI”. The Ruffian essay, 15 August 2026.
An extended discussion of the paper's central example — Einstein's 1907 thought experiment — and its claim that the insight was grounded in physical experience rather than in words or symbols.
Noah Smith, “The End of the Age of Heroes”. Noahpinion essay, 4 August 2026.
On AI systems solving open problems in mathematics; cites the paper for the possibility that novel conceptual leaps remain a general limitation of current models.
Responses
Yong Zheng-Xin, “Hot Take: LLM can ‘jump’”, 8 August 2026.
Argues that general relativity could also have been reached by a more deductive route, and that models able to spot gaps between fields may not need the abductive step.
HackerNoon, “LLMs Can’t Jump and They Shouldn’t Have To”, August 2026.
Accepts the premise and draws a division of labour from it: AI for synthesis and verification, people for the leap.
A note on scope
Some of the discussion framed the paper as an argument that LLMs cannot make scientific discoveries, or as an institutional view. It is neither. It is a personal position paper, and as a contributor to AlphaProof I have seen first-hand that LLM-based systems are already making real discoveries, and will continue to.
The paper is narrower than that: it asks what it would take for a system to make one specific kind of jump — Einstein’s formulation of the equivalence principle from thought experiments grounded in physical intuition — and argues that current recipes do not obviously supply it. Whether that capability is the most urgent thing to build next is a separate question; it may well be that scaling current systems is enough. I wrote up these clarifications in a short note after the paper began circulating.