One Self-Improving Model, Eleven Number-One Titles: What That Takes
aimachinelearningllmopensource
TL;DR
One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions. The interesting part is not any single score, it is that a recursive self-improvement loop, bound to external verification, can push a whole spread of benchmarks to the top at once.
The eleven, at a glance
- Math: AIME 2026 100%, HMMT 2026 100%
- Science: GPQA Diamond 94.44%
- Knowledge and multimodal: MMLU-Pro 88.12%, MMMU-Pro 79.48%
- Law: LEXam 68.94%, LEXam-hard 45.72%
- Extraction and structure: ExtractBench 90.29%, IFStruct 98.95%
- Decisions: MDPBench 83.65%, S1MB Borda 89.58
Why a self-improving loop gets here
A model that only trains once is stuck with the data it saw. A self-improving (RSI) model runs a loop: it attempts problems, keeps the attempts that pass an external check, and trains on those. The key is that the reward is tied to verification outside the model, such as code execution, external evidence, or an answer key, not to the model grading itself. That is what keeps the loop from drifting into its own confident mistakes.
When the loop is honest, improvement on one reasoning skill tends to carry to others, which is why the records span unrelated fields instead of one narrow test.
The decision layer
For typed decisions, the family uses a zero-token method: it reads the problem in one forward pass and applies a calibrated probe to produce the decision, generating no text. That is deterministic and cheap, which is also what makes it a good fit for running decisions close to where the data is.
Open to use
- S1MB number one model: github.com/final-bench/s1mb
- ZTC method: github.com/final-bench/ztc
- On-device models: github.com/final-bench/pocket
FAQ
How many number-one records, and across what? Eleven at once, across math, science, law, structured output, and decisions.
What is a self-improving (RSI) model? A model that runs a loop: it solves problems, keeps the solutions that pass an external check, and trains on them, so it improves without new human labels.
Why tie the loop to external verification? If a model grades its own work, the loop can reinforce confident errors. External checks keep improvement real.
Is it open? Yes. The decision model and method are public under Apache-2.0.
Built by Ginigen. If this helps, a star on the repositories makes it easier for others to find.