Qwen 3.8 27B Can't Match Sonnet's Performance (On Health)

Updating Medmarks with recent 20B–40B parameter models

We're finishing our fall 2026 update to Medmarks-Verified, but before the full release we wanted to highlight some consumer/prosumer models in the 20-40 billion parameter range. The models you can run on a single 5090 or a well spec'd Mac.

What is Medmarks?

With 70 models evaluated across 89 configurations, Medmarks is the largest open-source benchmark suite for testing how well language models perform on a variety of medical tasks. For this update, we’re focusing on Medmarks-Verified which spans medical knowledge, expert-level clinical reasoning, medical calculations, and reasoning under uncertainty. Most tasks are multiple choice or similarly objective, and we combine the per-benchmark scores into a single win rate. The datasets, prompts, and grading code are all public, so the numbers can be reproduced on local open-weight models and proprietary APIs alike. Read the launch blog post or paper for more details.

What is a win rate, and how should I interpret it?

A win rate summarizes how a model performs relative to other models in the evaluation. Comparisons award a per question win 1 point, a tie half a point, and a loss 0 points. The headline score combines comparisons against the other models into a weighted mean across datasets. A score of 0.640 (64.0%) is a relative comparison score, not 64% accuracy on medical questions. It also depends on which models are included, so compare headline scores within the same evaluation snapshot. See the paper’s scoring method.

A direct head-to-head win rate compares just two models: 0.5 indicates parity, above 0.5 favors the named model, and below 0.5 favors its opponent. When interpreting a gap between headline scores, consider its size alongside the models’ task-level results. A difference of 0.005 equals 0.5 percentage points.

As a rule of thumb, headline scores less than 0.5 percentage points apart are close enough that we would not put much weight on their ordering without uncertainty estimates. A gap of 0.5 to less than 1 point is a small lead: potentially useful, but not a large separation. A lead of 1 to less than 3 points is notable, particularly among already strong models. Gaps of 3 points or more are substantial relative to the high-performing models in the published v1 results. For example, 0.640 versus 0.643 is a close result, while 0.640 versus 0.660 represents a notable two-point lead.

This mini update adds the following models to the original set:

Medmarks-Verified
Size:
Showing 9 models across 18 configurations
4Gemma 4 31B Thinking NewGoogle31.27BMedium60.7%
8Qwen3.5 27B Thinking NewAlibaba27.78BMedium59.3%
9Qwen3.6 35B-A3B Thinking NewAlibaba35.95BMedium59.3%
10Gemma 4 31B Instruct NewGoogle31.27BMedium59.1%
11Qwen3.5 35B-A3B Thinking NewAlibaba35.95BMedium58.6%
12Qwen3.6 27B Thinking NewAlibaba27.78BMedium58.5%
13Qwen3.8 27B Thinking (xhigh) NewAlibaba27.78BMedium58.5%
15Qwen3.8 27B Thinking (medium) NewAlibaba27.78BMedium58.1%
16Gemma 4 26B-A4B Thinking NewGoogle26.54BMedium57.9%
18Qwen3.8 27B Thinking (low) NewAlibaba27.78BMedium57.8%
19Muse Glimmer 30B (xhigh) NewMeta29.78BMedium57.6%
20Muse Glimmer 30B (high) NewMeta29.78BMedium57.4%
24Qwen3.6 35B-A3B Instruct NewAlibaba35.95BMedium56.6%
25Muse Glimmer 30B (medium) NewMeta29.78BMedium56.6%
26Gemma 4 26B-A4B Instruct NewGoogle26.54BMedium55.8%
27Muse Glimmer 30B (low) NewMeta29.78BMedium55.8%
29Qwen3.8 27B Instruct NewAlibaba27.78BMedium55.7%
41Nemotron 3.5 Lightning 30B-A3B NewNVIDIA31.58BMedium52.5%

Ranks and weighted mean win rates use the full 89-configuration snapshot. Each row is one evaluated configuration. Bars show relative scores across that snapshot; detailed scores show per-dataset win rates.

Several models have multiple reasoning modes, so there’s 18 new configurations to increase the total to 89.

Gemma 4 31B: the New 20B-40B Pareto Frontier

Unlike the Gemma 3 family, the Gemma 4 models perform extremely well on Medmarks-Verified. Gemma 4 31B Thinking slots into 4th place, in between Grok 4 and Sonnet 4.5 on the updated leaderboard.

In head-to-head pairings, Gemma 4 31B Thinking's win rate is 0.502 against Sonnet 4.5, 0.506 against GPT-5.2, and 0.510 against GLM-4.7 FP8. These are essentially ties for practical purposes, which is an impressive result given Gemma 4's size. Furthermore, Gemma 4 achieves this using a respectable 1.7K tokens per task. Significantly more than GPT-5.2's 285, but far less than some of the local thinking models in this update.

Model Mode Win rate Rank Mean tokens
Gemma 3 27B Instruct 0.4418 69 266
MedGemma 27B Instruct/medical 0.4829 54 1,037
Gemma 4 26B-A4B Instruct 0.5579 26 411
Gemma 4 31B Instruct 0.5905 10 317
Gemma 4 26B-A4B Thinking 0.5786 16 3,258
Gemma 4 31B Thinking 0.6075 4 1,727

The instruct model is also interesting. Gemma 4 31B Instruct ranks 10th while averaging only 317 output tokens. It has a higher average score than both Gemma 3 27B and MedGemma 27B on all datasets, with a direct winrate of 0.648 and 0.608, respectively.

Gemma 4 31B Thinking is second only to Gemini 3 Pro on both MedXpert Reasoning and MedXpert Understanding. Its rank isn't being carried by easy medical multiple choice.

Selected models in the 20B–40B update: overall win rate versus mean output tokens. Hollow markers indicate instruct configurations; filled markers indicate thinking or reasoning configurations.
Figure 1. Selected models in the 20B–40B update: overall win rate versus mean output tokens. Hollow markers indicate instruct configurations; filled markers indicate thinking or reasoning configurations.

The Curious Case of Qwen 3.8 27B

If you were aware of the hype surrounding Qwen 3.8 27B, you might come away with the impression that it was an Opus class model you could run on your local machine. We won't take any positions on Qwen 3.8 27B's coding or agentic abilities, but on health, the story is much less exciting.

Let's start with the positives. Like Gemma 4, every new Qwen Thinking model outperforms the prior generation’s Qwen3 30B-A3B Thinking. If Gemma 4 didn’t exist, Qwen 3.5 27B would be the best performing model in this weight class. Qwen3.6 also substantially reduced the average number of tokens per generation.

Model Architecture Winrate Rank Mean tokens Δ vs Qwen3
Qwen3 30B-A3B MoE 0.5392 35 1,849
Qwen3.5 27B Dense 0.5933 8 3,797 +5.41 pp
Qwen3.5 35B-A3B MoE 0.5856 11 4,263 +4.64 pp
Qwen3.6 27B Dense 0.5850 12 1,994 +4.58 pp
Qwen3.6 35B-A3B MoE 0.5932 9 2,399 +5.40 pp
Qwen3.8 27B low Dense 0.5781 18 809 +3.89 pp
Qwen3.8 27B medium Dense 0.5811 15 1,034 +4.19 pp
Qwen3.8 27B xhigh Dense 0.5847 13 3,288 +4.55 pp

Now for the curious part. Agentic training appears to have taken its toll on the dense 27B series with medical performance peaking with Qwen3.5. The 27B winrate drops from 0.5933 for Qwen3.5 to 0.5850 for Qwen3.6, then sits at 0.5847 for Qwen3.8 xhigh. Head-to-head, Qwen3.6 loses to Qwen3.5 on 16 of 27 datasets. And while Qwen3.8 vs 3.6 is an overall tie, 3.8 loses more individual datasets than it wins.

The Qwen MoE models don’t follow this pattern. Qwen3.6 35B-A3B’s win rate increases from 0.5856 to 0.5932 while its mean output length falls from 4,263 to 2,399 tokens, a 44% reduction. Head-to-head, it records a narrow 0.507 win rate against Qwen3.5 35B-A3B, improving on 11 of the 27 Medmarks-Verified datasets. Overall, Qwen3.6 offers a better quality/efficiency trade-off, even if its improvement is not consistent across tasks.

Qwen generations compared on win rate and output length. Arrows connect successive releases within each architecture.
Figure 2. Qwen generations compared on win rate and output length. Arrows connect successive releases within each architecture.

This result might show a benefit of mixture of expert models over dense models during agentic post-training: perhaps the agentic coding updates ignored more of the medical experts and improved overall reasoning, while the same post-training on the Qwen 27B series caused a mild case of catastrophic forgetting on medical tasks.

Last, Qwen3.8 27B still can't catch Sonnet. At xhigh reasoning level, its direct win rate against Sonnet 4.5 is 0.474. Far short of an Opus class model.

More Reasoning Still Mostly Buys More Tokens

Qwen3.8 introduces multiple reasoning levels to the open-weight Qwen models in this weight class, matching behavior previously limited to gpt-oss 20b.

Medium to xhigh raises Qwen3.8's score from 0.5811 to 0.5847. That's 0.36 pp. Mean output length goes from 1,034 to 3,288 tokens, or 3.18x.

Moving from medium to xhigh gains 0.36 percentage points for 3.18 times the mean output tokens.
Figure 3. Moving from medium to xhigh gains 0.36 percentage points for 3.18 times the mean output tokens.

While the median token usage barely moves, from 786 to 801, the mean explodes because xhigh will wander off on long, unsuccessful reasoning detours. Incorrect Qwen3.8 xhigh answers average 7,254 output tokens while correct answers average 2,215, replicating one of our original Medmarks findings.

Muse Glimmer: Efficient But Not Quite a Qwen Killer

Muse Glimmer is Meta’s first open-source model in this weight class since Llama 3, and for a first release in over two years it’s impressive. Especially since Glimmer offers four efficient reasoning levels out of the box with linear performance scaling before flattening at xhigh.

Effort Winrate Rank Mean tokens
Low 0.5575 27 309
Medium 0.5659 25 625
High 0.5743 20 1,249
Xhigh 0.5759 19 1,501

Muse Glimmer’s edge is efficiency, not raw score. While Qwen3.8 wins at all matched effort levels, Muse uses 62% fewer tokens at low reasoning, 40% fewer at medium, and 54% fewer at xhigh. At xhigh, Muse trails Qwen by only 0.88 percentage points while using less than half the tokens.

Muse Glimmer and Qwen 3.8 across reasoning levels. At xhigh, Muse uses 54% fewer output tokens and trails Qwen by 0.88 percentage points.
Figure 4. Muse Glimmer and Qwen 3.8 across reasoning levels. At xhigh, Muse uses 54% fewer output tokens and trails Qwen by 0.88 percentage points.

One interesting Muse Glimmer result from MedBullets: while adding the irrelevant fifth answer option drops Glimmer’s score from 0.785 to 0.737 at low reasoning, 0.817 to 0.768 at medium, and from 0.860 to 0.810 at high, at xhigh, the score falls only from 0.844 to 0.825.

Nemotron 3.5 Lightning: Slow but Steady Progress

Nvidia has been making steady progress improving Nemotron across generations, with Nemotron 3.5 Lightning narrowly beating gpt-oss 20b high using a similar number of tokens (head-to-head winrate of 0.511; higher on 18/27 datasets), but not matching Qwen3 30B-A3B Thinking while using double the tokens (head-to-head winrate 0.487; higher on 11/27 datasets).

Nemotron generations compared with gpt-oss 20b at low, medium, and high reasoning effort, and Qwen3 30B-A3B Thinking. Arrows connect successive Nemotron releases.
Figure 5. Nemotron generations compared with gpt-oss 20b at low, medium, and high reasoning effort, and Qwen3 30B-A3B Thinking. Arrows connect successive Nemotron releases.

At the dataset level, Lightning is one of the stranger models we’ve tested. It ranks 12th on Med-HALT NOTA and 13th on SCTPublic, both of which test uncertainty and resistance to bad answers. It also scores 0.865 on Med-HALT’s false-confidence test, up from just 0.229 for Nano V3. But it scores terribly on MedHallu, ranking 81st, 83rd, and 84th across the easy, medium, and hard splits. Whatever improved between Nemotron Nano V3 and 3.5 Lightning, it did not produce good medical hallucination-detection capabilities.

Medmarks Fall 2026 Update

We're finishing up evals on the recent frontier and near-frontier models, including planned evaluations of Claude Fable 5.1 and GPT 6 Astra. Look for the full update in the coming weeks.