CoeOS : divide and route
Between July 30 and July 31, the same benchmark batch went from $5.93 to $2.96. The score moved from 95.5 to 96.1. The number of API calls dropped from 71 to 28. No code changed between the two runs.
A JSON file did.
That file is the routing table of CoeOS, our skill router, and this article is about what it does differently from the routers currently appearing everywhere, why the cost figure is structural rather than a discount, and what happens when the same table stops routing single requests and starts composing an army of agents.
Routers everywhere
The market has quietly conceded that the single model is over. Sakana sells Fugu, a multi-agent orchestration system, as one API, announced at parity with Fable 5 on their metrics; the pool of models it coordinates is not disclosed. OpenRouter now ships its own router, which classifies your request on its context and forwards it to a model of its choosing. The frontier names themselves are served systems: harnesses, orchestration, post-training pipelines behind a model name.
Every frontier model has holes. The one that writes the best prose is not the one that debugs C the best; the one that plans a migration is not the one that drafts a GDPR note. Three answers exist to that fact. Pick one model and inherit its blind spots. Merge several per request — mixture-of-agents, debate, vote — and pay for every call, wait for the slowest, and trust an aggregator that is itself a model that can be wrong. Or route by skill: one classification, one call, to the model a benchmark has proven best at that specific competence.
CoeOS is the third answer. Its particularity is not the routing; LLM routers are plentiful. It is where the decision comes from: a table of empirical scores produced by a benchmark harness — currently 51 models scored across 18 skill axes — not a trained classifier, not hard-coded heuristics.
The routing is data, not code.
The mechanics
An axis binding is a JSON entry:
{
"key": "debug",
"model": "deepseek-v4-pro-or",
"bench": "C02 49/50 (panel record)",
"verified": true
}
The bench field carries provenance: which test, which score, which panel. That is what separates a routing table from a list of preferences. Every binding is traceable to a benchmark result, and dated, because a table is a snapshot and models move every month.
Classification runs in three tiers. An agent that knows its phase sends x-coeos-axis: plan_spec and skips everything: zero extra calls, zero added latency. Otherwise a small, fast decider reads the last user message — only that, truncated at 8,000 characters — against the axis taxonomy passed in from configuration. The decider is a router that reasons, not a trained classifier: two sentences of reasoning, then a final AXIS: line, with parsing defensive enough to survive a chatty model. If it names no configured axis, the request falls to the default axis. Never to an invented one.
Resolution goes through an indirection that came out of an incident, not a principle. Early on, an axis pointed directly at an engine’s serving alias. The day that alias changed, bindings silently stopped resolving, and there was no way to tell “wrong champion” from “right champion, dead alias” — the symptom is identical. So an axis now binds a logical model, and a separate registry maps that stable identity to whatever endpoint serves it today. The canonical OpenRouter path is the identity; the operator’s alias is just access.
And one invariant: never a silent fallback. If the champion of an axis cannot be served — no API key, missing registry entry, provider down — CoeOS does not route somewhere else. It returns an explicit 503 naming the expected model and the action to take. This is counter-intuitive for a proxy, where availability is usually the goal. But a skill router whose output you cannot trust has no value: a silent fallback turns “the best model on this axis” into “a model, we don’t know which”, and the whole reasoning collapses. Failures get loud and get fixed, instead of dissolving into slow quality decay.
Every routed response carries x-coeos-axis, x-coeos-model and x-coeos-provider headers. The decision is readable client-side, not just in server logs.
The cost
The real batch, late July 2026, same workload for all three:
| CoeOS 1.33 | Opus 5 alone | Qwen 3.8 alone | |
|---|---|---|---|
| Bench score | 96.1 | 96.0 | 95.9 |
| Lowest axis | 94.4 | 95.5 | 92.5 |
| Real cost | $2.96 | ~$12.35 | ~$12.10 |
Same level as Opus 5, 76% cheaper. Priced through Fable 5, the same volume comes to roughly $25: the routed mix runs it at 10% of the price. The honest details: Opus keeps the highest floor (95.5, nothing weak anywhere), the routed mix holds 94.4, and Qwen drops to 92.5 on its worst axis. The 0.1 and 0.2 gaps in the global score are noise. The $9.40 is not.
A word on why the claim is “equal” and never “better”. Inference is a stochastic process: a model can get lucky on a run, or not. Some of our suites are strict, others are creative, and on the creative ones we deliberately keep the temperature high, because that is the creativity the task requires — and it is also where run-to-run variance lives. Our sample counts per test are too small to shrink that uncertainty honestly, so half a point, even a full point, is not a verdict. What we do have is the property that matters: rerun the panel and the scoreboard comes back the same. The ranking is reproducible even where the individual scores wobble. Equal to Opus 5 at a quarter of the price is the entire claim, and it is the reproducible one.
Opus 5 deserves one more sentence, because it is the latest model through the bench and the comparison is telling. The 96.0 is real; the model is excellent. It is also very verbose: on this volume it would have burned an estimated 240k reasoning tokens where the routed mix spends 170k. The score table flags this profile explicitly — a token-burner marker — because verbosity is the one cost that quality scores never penalize.
The invoice does.
Where does the 76% come from? From paying the fair price per request instead of the frontier price for everything. Simple tasks route to models that cost nothing or close to it. Complex tasks go to the strong reasoners. Critical tasks go to the precise ones. High-frequency execution runs on local models at zero marginal cost.
The 1.32 → 1.33 jump shows the mechanism live. In 1.32, one heavy reasoner was eating 97.8% of the batch cost, mostly in billed thinking tokens: 496k reasoning tokens on the run. The bench said more direct models held the level on several of its axes, so those axes were rebound. Result: reasoning tokens down 66%, calls down 60%, cost down 50.4%, score up 0.6.
Half the cost, nothing lost, by editing a JSON. That is what “routing is data” buys you operationally: tuning the system is a table edit, not a development cycle. Two settings profiles now ship, both published — full spec, which maximizes measured quality per axis, and economical, which accepts the bench’s cheaper near-equals. Same engine, same code, two tables.
The sovereignty
On June 12th, a US directive suspended access to Fable 5 for all foreign nationals, without notice. I wrote about it at the time: for organizations built on one cloud model, that was not an outage, it was a dependency being exercised.
One model is one dependency. A router built on 51 scored models is that dependency divided by 51.
The arithmetic is worth spelling out. When your infrastructure is one frontier API, every failure mode is total: price change, deprecation, export directive, terms update. When your infrastructure is a routing table over open-weight models, every failure mode is one axis rebind. A model disappears from OpenRouter, gets repriced, gets restricted: you edit its bindings, the bench tells you the next-best champion, and the incident is a JSON diff. No proprietary model serves a single request in our stack; the only proprietary component anywhere in the chain is the benchmark judge, which scores criteria and never serves, and which the notes-only protocol keeps replaceable.
The settings themselves are built to travel. Axes bind logical models by their canonical paths, the registry that maps identities to access is separate and starts empty: publish the table, and every operator fills in their own keys and their own fleet. Nothing of anyone’s infrastructure leaks. Multiplying models is not a complication to be managed. It is the dilution of dependency, and the table is what makes it manageable.
The army
Routing single requests is half the system. The other half composes agents from the same table.
A task submitted to the CoeOS pipeline flows through dedicated roles: a triage that prescribes the flow, a planner that produces a closed, executable plan, a grill that tightens the plan before execution, an executor, a skeptic that attacks the result after execution. Fourteen composable roles in the current build. Each role declares its competency pair and a speed weight, and the console scores every servable model against it:
quality = mean of the model's scores on the role's axes
speed = 100 × tps_median / tps_max_of_panel
composite = quality × (1 - w) + speed × w
Ties break on cost: at equal quality, the cheaper model wins. A background role like the planner takes w = 0, only quality counts. A hot-loop role like the executor runs at w = 0.4, because a model 30% better but three times slower is the wrong choice when it runs a hundred times an hour.
Two constraints make the pipeline trustworthy, and both came out of failures. Separation is structural, not prompt-based: a smoke test proved that no prompt reliably forces a model to delegate, so the constraint lives in the tools, and only the executor can write. And adaptivity comes from the model while rigor comes from the code: the triage decides which steps to engage, through a typed decision, but deterministic floors override it — a task carrying a verifiable done-check always gets the skeptic, whatever the triage judged.
One refinement matters more than the formula. Three roles form a contradictory panel — direct, alternative, skeptic — whose entire purpose is independent opinions. Letting them converge on the same model destroys the object: the same bias three times under three labels. The assignment algorithm therefore reserves a distinct logical model for each panel member, deduplicated on identity rather than endpoint, so two aliases of the same underlying model count as one.
The $2.96 batch above is this pipeline running. Not a router demo. The army, invoiced.
What it is not
Not an ensemble: one classification, one call, no vote, no aggregation. Not a trained classifier: nothing was ever trained to route, the decision comes from measured scores. Not a guarantee: the table is a dated snapshot of a moving field, which is exactly why provenance — test, score, sample count, date — travels with every binding, and why CoeOS’s own rows in the panel are marked self and excluded from candidacy. Letting it compete against itself would close a loop it can only win.
The standalone edition, CoeOS SE, is free under MIT: same routing results, no settings surface, 100% OpenRouter. It is 1,934 lines of Python with three runtime dependencies, on PyPI, Docker, and as a signed macOS app. The full pipeline is AGPL-3.0 on GitHub. Small enough to actually read, which is the point: every claim in this article is checkable against the code and the published table.
Frontier level used to be a model you subscribed to. On the July numbers, it is a table you maintain — 51 models, 18 axes, one invariant — and the discipline costs $2.96 a batch.
Divide and route.
Sophie, The Monocle Bear
Get it
- CoeOS SE, free (MIT): github.com/Odyssai-eu/coeos-SE
- macOS app, Apple Silicon: CoeOS-SE-0.3.2.dmg
- Full pipeline (AGPL-3.0): github.com/Odyssai-eu/coeos
Sources
- Sakana AI, Fugu: sakana.ai/fugu
- TMB benchmark, July 2026 runs: methodology and score table published with the repositories
- CoeOS 1.32 → 1.33 batch data, 30-31 July 2026
- The Monocle Bear, “MiniMax M3, survived the unplug”, June 2026
Sophie, The Monocle Bear