Hybrid · sovereign · controlled

Cloud AI, controlled. Local when necessary.

The Monocle Bear designed a hybrid AI system: OdyssAI. A distributed inference engine developed in-house, the OdyssAI-X orchestrator, backed by Companion, a complete enterprise client. Together they give access to frontier cloud models as well as the best local open-source models. The full power of the large models, at a tenth of the price per token. A guardrail checks every request on the way out; confidential data stays inside the walls.

Sovereign by construction: no single provider, interchangeable models, local or cloud. The arbitration is automatic, measured, verifiable.

/01 · Problem

Data, or Open-Data?

In most organizations, generative AI came in through the service entrance: personal accounts, case files pasted into consumer interfaces, client data processed on foreign servers. A breach widens, a leak that excuses itself as necessity.

The real choice is not between using AI or going without. It is between letting it happen and controlling how it is used.

/02 · CoeOS

The best of open source, in a single model.

At the heart of OdyssAI, CoeOS summons for every request the most competent model to answer it. The result: the accuracy level of the best closed models.

Routing

Agentic routing.

Every request is analyzed for meaning and intent, not keywords. Writing, law, analysis, code: the strongest model for the task takes over, based on 18 measured competency axes.

Guardrail

The privacy guardrail.

The request is analyzed before it leaves. Personal data, client files, trade secrets are detected and processed on a local model. Nothing confidential leaves the infrastructure.

Measured on a public bench of 32 tests.

97.2/100 CoeOS $0.087 / test
96.9/100 Fable 5 $0.426 / test
96.0/100 Claude Opus 4.8 $0.141 / test

Published methodology, reproducible results, code under the AGPL.

/03 · Full local

The on-prem model, essential.

Some contexts rule out the cloud entirely. For those, OdyssAI runs the best open-source models, Kimi GLM 5.2 MiniMax Mistral, on an Apple Silicon cluster installed inside the organization's walls. One to six machines depending on load.

A stated engineering position: no model is served below 6-bit quantization, or 4-bit when the model was specifically trained for that quantization (Kimi, for example).

Past that point, accuracy and perplexity degrade enough to compromise professional use. Size does not make quality, model fidelity does.

On TMB Benches, GLM 5.2, Hy3 and MiniMax served locally show no significant loss of accuracy against their cloud versions.

/04 · Onboarding

Like any new hire, a model has to know the company.

An out-of-the-box LLM knows nothing of the files, the precedents, the internal vocabulary. The effectiveness of a working tool is decided in the knowledge layer: vector indexing and relations RAG across business corpora, structured memory on three levels, individual, team, organization, decision logging…

The model stays interchangeable.
The knowledge belongs to the organization and never leaves it.

/05 · Engagements

What The Monocle Bear does.

→ 01

Analysis.

A map of data flows and existing AI usage, including the usage no internal policy has ever seen.

→ 02

AIrchitecture.

Hybrid or full local: the architecture is designed around the actual sensitivity of the data, rather than fashion.

→ 03

Integration.

OdyssAI installed and adapted to the organization's needs: the Tailored edition. Migration of the first workflows, measurement before and after.

→ 04

Transfer.

The teams build up the skills, all the way to autonomy.

Our software is open source and free. The engagement covers architecture, deployment and method.

/06 · The Lab

Let the numbers speak.

L.01

TMB Scoreboard

Over 30 models, 18 axes, 32 tests, 5 suites.

L.02

TMB Benches

The full protocol, Opus 4.8 as judge at temperature 0, per-criterion isolated scoring, mechanically computed totals, hierarchy confirmed by two independent judges (GLM 5.2, ρ = 0.94; Hy3, ρ = 0.91), residual biases documented.

L.03

CoeOS

The code, AGPL-3.0, and the simplified edition under MIT.

/07 · Contact

Let's talk.

Engagements in Belgium, France, and beyond.

Write to hello@themonoclebear.com