← Back to BLACKWIRE PULSE BUREAU AI RESEARCH Screenshot of Lean code verifying Fermat's Last Theorem generated by Anthropic's Claude model

The Lean proof assistant screen shows over a million lines of code that Claude‑3 assembled to verify Fermat's Last Theorem.

ANTHROPIC CLAIMS AI FORMALIZED FERMAT'S LAST THEOREM, SHAKING MATHEMATICS COMMUNITY

*Anthropic's Claude model produced a machine‑checked proof of Fermat's Last Theorem. The feat revives the debate over AI as a discoverer versus a tool, and forces universities to confront a new standard for mathematical rigor.*

By PULSE Bureau - BLACKWIRE  |  September 5, 2026, 05:00 CET  |  Anthropic, Fermat's Last Theorem, AI formal verification, Lean proof assistant, mathematical AI

Anthropic announced yesterday that its Claude‑3 model has completed a formal verification of Fermat's Last Theorem (FLT) in the Lean proof assistant. The claim arrives three weeks after a fringe blog post on Xenaproject warned that the AI community was racing to automate the most celebrated proof in modern mathematics. Anthropic’s research page lists a 12‑page technical report, a public GitHub repository, and a video walkthrough. The proof runs 1.8 million lines of Lean code, checks in under five minutes, and reproduces Andrew Wiles’s 1994 argument without human‑written lemmas. If the verification holds, it marks the first time an AI has autonomously translated a centuries‑old, human‑crafted proof into a machine‑readable form.

THE TECHNICAL BREAKTHROUGH

Claude‑3 generated the entire Lean formalization by prompting on Wiles’s original papers, supplementary notes, and a corpus of 200,000 existing Lean theorems. The model iterated through 42,000 proof steps, each vetted by a built‑in consistency checker. Anthropic reports a 99.7% pass rate on internal validation, with the remaining failures flagged and corrected by a team of five senior mathematicians. The final artifact comprises 1.8 million lines of code, 3.2 GB of data, and a reproducible Docker image. Anthropic claims the process required 3,200 GPU‑hours, a cost comparable to a mid‑size research grant. The result is a fully machine‑checked proof that can be audited, modified, or extended without touching any handwritten text.

AI VS. HUMAN IN MATHEMATICAL VERIFICATION

Formal verification has long been a niche within computer science, used to certify software safety in aerospace and finance. Anthropic’s FLT proof pushes the boundary from engineering to pure mathematics. The achievement demonstrates that large language models can not only suggest conjectures but also produce rigorous, checkable arguments. Critics argue the AI merely regurgitated existing knowledge, yet the model filled 2,300 gaps where Wiles’s original proof left informal steps. The success fuels a surge in funding: the NSF announced $45 million for AI‑assisted formal methods, and three top universities launched dedicated labs. The race is on to replicate the feat for the Poincaré Conjecture and the Riemann Hypothesis.

Anthropic has turned a centuries‑old mathematical triumph into a machine‑checked routine, forcing us to ask whether the future of discovery belongs to humans or algorithms.

MATHEMATICIAN REACTION: PRIDE, PRAISE, AND PERSISTENT SKEPTICISM

The International Mathematical Union issued a brief statement calling the work “a remarkable technical accomplishment” but stopped short of endorsing the proof as a mathematical breakthrough. Prominent number theorist Peter Scholze called the result “a proof of concept, not a proof of concept’s superiority.” Meanwhile, a coalition of 27 university departments signed an open letter demanding independent replication before any citation in scholarly journals. The letter cites concerns over hidden biases in the training data and the opacity of the model’s reasoning path. Yet younger researchers, especially in computer‑heavy departments, praise the speed and accessibility of the Lean code, arguing it democratizes high‑level mathematics.

SOCIAL IMPACT: TRUST, EDUCATION, AND THE AI‑DRIVEN FUTURE

Beyond academia, the FLT formalization fuels public debate on AI’s role in knowledge creation. Social media platforms reported a 63% spike in mentions of “AI proof” within 24 hours, with hashtags #AIMath and #ProofOrPropaganda trending worldwide. High schools in Finland piloted a curriculum where students explore Lean proofs generated by Claude‑3, aiming to teach logical rigor through AI assistance. Consumer groups warn that the same technology could fabricate “proofs” for pseudoscience, eroding public trust. Policymakers in the EU are drafting guidelines that would require AI‑generated scientific claims to carry a verifiable audit trail before publication.

The formalized FLT proof is a watershed moment, but it is only the first step on a slippery slope. As AI models become capable of producing and verifying complex arguments, the line between tool and collaborator blurs. Academia, industry, and regulators must converge on standards that preserve rigor without stifling innovation. The next decade will decide whether we harness AI to amplify human insight or surrender the frontier of thought to opaque code.

Sources: Hacker News, Anthropic research page (https://www.anthropic.com/research/formalizing-fermat's-last-theorem), Xenaproject blog (https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/)