← Blog

Qwen, Big Family

Qwen, Big Family

Qwen has just released the weights for version 3.8, open source. I took the opportunity to test it and compare it against the other versions in the family.

The 27B is the one going on the bench. Unfortunately, the 2.4 T is too large for our servers. It weighs 4.89 TB, and the cluster holds 1.5 TB of unified memory. It would have to come down below 3 bits per weight. And even then.

At that level of compression, this is no longer benchmarking, it is a séance. You question experts of whom only the ghost remains, and you record their answers as though they came from them.

The verdict is the same as for Kimi K3: the weights are public, the licence is open, and the object remains a datacenter model. Open source does not mean runnable at home.

But back to the 27B.

Two 27Bs, two collapses

The 3.8 comes as a 27B-VL, served at 8 bits, judged across the five benchmarks with no gaps in collection. On the TMB benchmarks it lands at 82.6 with thinking, 80.9 without. Fine in absolute terms, disappointing in context: its code falls to 68.4 without thinking, 77.8 with.

At strictly equal size, the non-multimodal 3.6-27B posts 91 on code, off a single test and therefore probably optimistic, and 97.4 on agentics. The cheapest hypothesis is that vision gets paid for. The 27B-VL spends part of its capacity on a modality the protocol, entirely textual, never calls on. For text work, the VL variant is a bad buy.

Except the 3.6-27B has a wall of its own, and it is worse. Reasoning: 49. It produces a schedule that violates the constraint it has just stated itself, classifies a processor where there is joint controllership, and breaks a case-sensitive sort. The 27B-VL holds 87 to 89 on the same benchmark.

Neither is a generalist. One is an agent that codes, the other an agent that reasons, and the reader looking for a 27B that is good at everything goes home empty-handed.

That is the first lesson of the family. At this size, agentics is a settled matter: the two 27Bs post 96 and 97, and the 4B itself puts up 90. Reasoning is not for sale at this size. You have to move up the range to get it.

The champion did not change its name, it changed its shape

The best local Qwen is still the 3.5-397B-A17B. What changed this summer is not its score, it is the number of machines it demands.

We used the campaign to build an in-house version: 6-bit quantization on the body of the model, heads kept in BF16. The recipe fits in one sentence. The 6 bits bring the weights down from 794 GB to 324 GB, and keeping the output layer at full precision protects the one thing quantization genuinely damages, the logit distribution. That is where rounding error turns into a wrong token, and a repeated wrong token turns into a generation that loops. A few gigabytes of precision in the right place buy stability on long formats.

Result: the model fits on a single 512 GB machine, without mobilizing the other nodes in the cluster. And it runs four times faster.

EngineRankingAgentCodePlanWritingTiebreak R2
qwen3.5-397B BF16 thinking91.998.887.393.088.589.8
qwen3.5-397B 6 bits heads BF16 non-think, 1 node89.696.884.794.082.896.4
qwen3.5-397B BF16 non-think88.297.685.086.583.695.4
Qwen3-Next-80B-A3B 8 bits87.098.277.889.582.485.4
Qwen3.8-27B-VL 8 bits thinking82.389.877.884.577.189.0
Qwen3.8-27B-VL 8 bits non-think80.496.468.481.075.787.0
qwen3.6-35B-A3B thinking75.589.258.580.573.876.4
Qwen3.5-122B 8H1674.484.276.186.051.188.8
qwen3.6-35B-A3B non-think69.179.256.375.565.451.0

Judged by Opus-4.8, TMB notes-only protocol. The ranking is the mean of the four first-level benchmarks. Reasoning is reported separately: it breaks ties between models already ranked, it does not rank them.

The quality cost of the operation: none. The 6 bits heads BF16 ranks 1.4 points above full precision in direct mode and takes one more point on the tiebreak, but those gaps are stochastic. That is sampling noise, not a gain. The finding worth keeping is that there is nothing to keep: holding the heads in BF16 preserves the model’s intelligence, all of it, and full precision has nothing left to offer.

This is not true of every model. On Qwen, the recipe is ideal. We reached the same conclusion on MiniMax M3, where BF16 adds nothing over 6 bits heads BF16.

ConfigurationTTFT P50Aggregate throughput20 requestsHardware
397B 6 bits heads BF160.77 s55.4 tk/s30.2 s1 node
397B BF1618.16 s14.4 tk/s116.7 s5 nodes
Gemma-4-31B 8 bits0.74 s54.1 tk/s25.2 s1 node

Separate stress runs, concurrency 4, 20 requests, 150 tokens, temperature 0.7.

Eighteen seconds before the first token. An aggregate throughput of 14.4 tokens per second under four concurrent requests, which is to say the throughput of a single request: the server was not batching. In an agent loop this model was unusable, and everyone concluded that a 397B was too large for local use.

The culprit was not the size. It was the multi-node.

The same model, compacted to 6 bits on a single machine, answers in 0.77 seconds and delivers 55 tokens per second. That is the exact latency profile of a 31B, with the quality of a 397B. The trade-off between quality and latency, which has shaped every local deployment decision for two years, has just disappeared for this particular model.

Four nodes freed, roughly 20,000 euros of hardware no longer tied up in this. That is not an optimization, it is a budget line changing sign.

The same lever, in reverse

The 122B in 8H16 quantization provides the counter-example, and it teaches more than an ordinary failure. Reasoning 88.8, plan 86.0, agent 84.2. Then writing: 51.1. Same model, same quantization, same run.

The detail that explains the gap sits in a single task, the long creative writing test with hard stylistic constraints, scored 27. A four-paragraph block repeated more than twenty times. The model has substance, its 88.8 in reasoning proves it. What quantization took from it is not competence, it is the ability to hold a long sequence without drifting.

So quantization does not damage uniformly. It spares short structured answers, and it breaks autoregression on long formats. A model judged only on brief tasks would have passed inspection with the flaw never showing.

Which leaves the matter of the right axis. The 122B is not an agent-tool model: when it answers that it cannot execute tools, that is not a defect, it is out of domain, and counting it as a failure amounts to scoring a novelist on typing speed. On its own ground, reasoning, planning, long-form generation, its only real handicap is the loop. At full precision, with the loop gone, it is a solid reasoning and long-writing model.

That is what makes the 397B recipe interesting beyond its own case. Protecting the output layer wins no points on a reasoning benchmark, where there was nothing to win. It prevents the one thing that actually breaks.

The thinking paradox

Three thinking and non-thinking pairs of the same model, measured under the same conditions, give a result I did not expect to be this clean.

The 35B-A3B gains 12.8 points when thinking is switched on, from 64.2 to 77.0. The 27B-VL gains 1.7. The 397B gains 1.8 on average, but loses 5.6 points on pure reasoning, from 95.4 to 89.8.

The more capable the model, the less explicit reasoning helps it, and past a certain level it hurts. The small model reasons poorly internally and needs to unroll in order not to lose the thread: the crutch is vital. The large model already reasons correctly on the inside, and thinking pushes it into over-deliberation. On R2, the 397B in thinking mode answers “it depends” where it used to decide.

The 27B-VL shows the same mechanic within a single model: thinking brings it 9.4 points on code, and costs it 6.6 on agentics. It analyzes instead of executing.

Which is routable. Thinking for the small models, direct for the large ones, and on the 397B a per-task call: direct for reasoning, thinking for planning and writing.

The 80B that beats the 397B on agentics

Qwen3-Next-80B-A3B at 8 bits posts 98.2 on agentics, ahead of the 397B BF16 non-thinking at 97.6. It holds 89.5 on planning, 82.4 on writing, and its creative writing test does not loop, which remains rare among heavily quantized Qwens, most of which drift into repetition on long generation.

At one fifth the size of the flagship, it captures the essentials. On tool driving, parameter count is visibly not the deciding factor.

Its reasoning score has just landed, and it is the number that changes its status: 85.4. The other A3Bs in the family collapse on this benchmark, 51 for the 35B-A3B, 49 for the 3.6-27B. The Next-80B is the only one that reasons properly, with three billion active parameters, which is to say exactly the compute budget of the 35B that fails.

So it is not active size that produces reasoning. It is the architecture, and what you put into it. A well-built MoE beats a heavier dense model on the ground where you would expect it least.

A complete generalist at 87.0, it becomes the number two local Qwen behind the 397B alone, conceding only code and top-tier reasoning.

What this ranking does not say

The three small models are judged on agentics only, and that is not a hole in the collection, it is their purpose. You do not ask a 4B to classify a data processing arrangement under the right GDPR article, you ask it to call the right tool with the right arguments. It does that at 90.2. A dedicated agent bench, considerably deeper than this one, is already running on its own: that will be the subject of a later article.

The 35B-A3B’s code score in thinking mode caps at 58.5, but four tests out of six could not be collected: the model loops during generation. That is a measurement floor, not a score.

Three agentic tests returned empty responses on some engines. That was a tool-call defect in the harness, not in the models, and the scores published here are post-correction.

A family, not a winner

The title promised a family tour, and that is what the campaign delivers. Not one model that wins, but a four-tier range, where each holds a role no other holds as well.

TierModelRole
The flagship397B-A17B, 6 bits heads BF16peak capability, at the latency of a 31B
The deep one122B, at full precisionreason and write long, not drive tools
The pilotQwen3-Next-80B-A3Bthe daily agent, five times smaller
The sprinterthe small ones, 4B and 30B-instvolume, instant answers, good enough

That is exactly the pyramid of a commercial product range, reassembled without meaning to out of open models. And it shifts the question. You are not looking for the best local model, you are routing the task to the right tier.

A stack that sends everything to the flagship pays for peak compute to rewrite emails. A stack that sends everything to the sprinter finds its limit at the first serious piece of reasoning. The ranking is not there to crown a winner, it is there to write the routing table.

The question that stays open

Qwen 3.8 exists. It just does not exist here, and compressing it to 3 bits to make it fit would amount to scoring a different model under its name.

The real game changer for local LLMs would be a Qwen 3.8 at 397B, sparse MoE, quantizable at 8 or 6 bits the way the 3.5 is. You would keep a reasonable hardware footprint and a usable processing speed, with performance close to frontier models.

Until then, a model you cannot run is worth nothing, whatever its licence.

Sophie, The Monocle Bear