Personal intelligence leaderboard

The models behind
my work.

A focused view of the four AI systems in my daily stack, five independent evaluations reduced to one auditable fleet score.

4models tracked
—average fleet score
Sep ’26latest release
Fleet score

One number. Clear signal.

Higher is better · M4Y0U Fleet Score v2 · dataset loading

Select any model to inspect its benchmark receipt.

01

— separate the strongest and weakest fleet scores.

02

— leads this five-evaluation basket.

03

Open weights account for two of the four tracked models.

Methodology

Complex benchmarks,
made legible.

The displayed Fleet Score uses exactly five independently run evaluations: Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CritPt and AA-LCR.

The Value view answers a different question: how much measured capability a dollar buys. It divides the fleet score by the blended price per million tokens, using Artificial Analysis's 7:2:1 cache-hit, input and output ratio. The cache-hit rate is listed separately in every receipt, because cache-heavy agent workloads are billed mostly at that rate.

The basket was revised in September 2026. Artificial Analysis retired Terminal-Bench 2.1 and removed GPQA Diamond from its suite, so both stop being measured for any newly released model. Terminal-Bench 4.0 and CritPt take their places; the revision and the reason are recorded in the dataset.

Every raw result is a proportion. It is normalized to 0–100 by multiplying by 100, then the five normalized results receive equal 20% weight. Missing one benchmark means no aggregate is published.

The versioned snapshot lives in data/benchmarks.json. Updates are deliberate and reviewable; the dashboard never silently pulls new values.

5 source results × 100 → equal-weight mean → 1 fleet score

Fleet
Score