— separate the strongest and weakest fleet scores.
The models behind
my work.
A focused view of the four AI systems in my daily stack, five independent evaluations reduced to one auditable fleet score.
One number. Clear signal.
Higher is better · M4Y0U Fleet Score v2 · dataset loading
Select any model to inspect its benchmark receipt.
— leads this five-evaluation basket.
Open weights account for two of the four tracked models.
Complex benchmarks,
made legible.
The displayed Fleet Score uses exactly five independently run evaluations: Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CritPt and AA-LCR.
The Value view answers a different question: how much measured capability a dollar buys. It divides the fleet score by the blended price per million tokens, using Artificial Analysis's 7:2:1 cache-hit, input and output ratio. The cache-hit rate is listed separately in every receipt, because cache-heavy agent workloads are billed mostly at that rate.
The basket was revised in September 2026. Artificial Analysis retired Terminal-Bench 2.1 and removed GPQA Diamond from its suite, so both stop being measured for any newly released model. Terminal-Bench 4.0 and CritPt take their places; the revision and the reason are recorded in the dataset.
Every raw result is a proportion. It is normalized to 0–100 by multiplying by 100, then the five normalized results receive equal 20% weight. Missing one benchmark means no aggregate is published.
The versioned snapshot lives in data/benchmarks.json. Updates are deliberate and reviewable; the dashboard never silently pulls new values.