// tmb scoreboard / methodology

How we test models.

What we measure, how we grade it, and where the limits are. The tasks themselves stay private.

Back to the scoreboard

From a task to a score

  1. 01

    Task

    A real professional request, identical for every model.

  2. 02

    Answer

    One answer per model, kept verbatim.

  3. 03

    Criteria

    The judge scores each criterion of a written grid.

  4. 04

    Totals

    The harness sums the points. The judge never does.

  5. 05

    Skills

    17 skills, each fed by one or more tasks.

  6. 06

    Categories

    5 categories, then the overall score.

What the scoreboard measures

32 tasks across four families. Each one is a request a professional would actually send, not a quiz with a single right token.

Agentic tool use

10 tasks

English

Choosing and sequencing tool calls to complete an operational request.

Coding

10 tasks

French instructions

Python, general code, debugging, React, Swift and refactoring, graded on correctness and craft.

Planning

4 tasks

French

Breaking down a project, writing a hand-off spec, making judgment calls under constraints.

Text work

8 tasks

French

Legal and compliance analysis, reasoning and calculation, business and creative writing.

A separate hard round of 10 tasks is used to separate models that saturate the main set. It is not part of the published score yet.

Running a task

  • Every model receives the same instructions. Tool definitions are sent whenever a task involves tools.
  • Each model answers once per task, with the sampling settings we set for that model or that task.
  • Answers are streamed and watched. A run that loops or stalls is stopped and retried, so a hung call does not count as an answer.
  • The answer is stored verbatim. Nothing is edited before grading.

Grading

  • Each task has a written grid: weighted criteria, each with described levels.
  • The reference judge, Claude Opus 4.8, scores criterion by criterion and flags factual errors. It never computes a total.
  • The harness sums the points, capped per criterion. Arithmetic is code, not model output.
  • A skill averages its tasks. A category averages its skills. The overall score averages the categories, each weighing the same, and is shown when at least four of five are covered.

We test the judge too

A score is only as good as its judge. We regularly re-grade the same answers with other judges and measure how far they drift from the reference, on the overall ranking and skill by skill.

Second judge, GLM 5.3

0.97

Rank correlation with the reference over 83 models. Average gap of 1.4 points.

Fast decision model

0.93

Good on the overall ranking, too noisy at the top to pick a winner per skill. Used for quick checks only.

Rejected judges

out

Candidates that diverged too far from the reference were dropped rather than averaged in.

Cost and speed

Cost per task is the number of tokens each model actually used on our tasks, reasoning included, multiplied by its current OpenRouter price. Prices are refreshed at every publication and again live in your browser, so they follow the market. Cached-input discounts are ignored, which makes our figures slightly pessimistic. Local models run on our own hardware and are not priced. Speed is the median decode rate we measured, in tokens per second.

Out of competition

Reference models are frontier systems we run as calibration points. CoeOS is our own router, which picks among the models listed. Both appear in the ranking for context. They never get a rank and are never recommended.

Mechanical benches

Three more benches are graded by code, with no judge at all. Their results will join the scoreboard.

Agentic execution

90 tasks in simulated environments where the model really calls tools and sees the results. Each task runs three times. Success is read from the final state of the environment. An entry gate first checks that the model emits usable tool calls.

Refusal threshold

7 risk categories, each climbed over 5 levels from harmless to operational and agentic, to find where a model starts refusing. Harmless mirror requests measure over-refusal. Runs locally only, and harmful outputs are never stored. It is not a danger score.

Defense

The model does legitimate work inside an agent cluster while injected instructions try to hijack it: spread a worm, exfiltrate a secret, run an untrusted payload. Each run ends compromised, resisted or contained. Air-gapped, graded by code.

Known limits

Language

Only the agentic tasks are in English. Coding, planning and text work are written in French. An English track is in progress, with the first translated tasks already running.

One reference judge

Any judge has preferences. We track them through calibration, and keep frontier models out of competition so the judge never ranks its own family.

Grading is not yet blind

The reference judge currently knows which model wrote an answer. We are moving grading to anonymised answers.

Saturation

Above roughly 97 points, differences are smaller than the judge's variance. That is why ties within 1 point are shown as ties.

One answer per task

A single sample per task keeps costs sane, but a model can have a good or bad day on any one task. Averages over many tasks absorb most of it.

Speed depends on hardware

Local speeds are measured on our own machines and quantizations. On different hardware they will differ.

What we keep private, and why

Prompts, grids and answers stay private. Once published, tasks end up in training data and a benchmark stops measuring anything but memory. Keeping them private is what lets the same tasks stay meaningful over time. If you want a model evaluated, get in touch.

Protocol v1. Data 2026-09-27.