Updating Medmarks with recent 20B–40B parameter models
We're finishing our fall 2026 update to Medmarks-Verified, but before the full release we wanted to highlight some consumer/prosumer models in the 20-40 billion parameter range. The models you can run on a single 5090 or a well spec'd Mac.
With 70 models evaluated across 89 configurations, Medmarks is the largest open-source benchmark suite for testing how well language models perform on a variety of medical tasks. For this update, we’re focusing on Medmarks-Verified which spans medical knowledge, expert-level clinical reasoning, medical calculations, and reasoning under uncertainty. Most tasks are multiple choice or similarly objective, and we combine the per-benchmark scores into a single win rate. The datasets, prompts, and grading code are all public, so the numbers can be reproduced on local open-weight models and proprietary APIs alike. Read the launch blog post or paper for more details.
A win rate summarizes how a model performs relative to other models in the evaluation. Comparisons award a per question win 1 point, a tie half a point, and a loss 0 points. The headline score combines comparisons against the other models into a weighted mean across datasets. A score of 0.640 (64.0%) is a relative comparison score, not 64% accuracy on medical questions. It also depends on which models are included, so compare headline scores within the same evaluation snapshot. See the paper’s scoring method.
A direct head-to-head win rate compares just two models: 0.5 indicates parity, above 0.5 favors the named model, and below 0.5 favors its opponent. When interpreting a gap between headline scores, consider its size alongside the models’ task-level results. A difference of 0.005 equals 0.5 percentage points.
As a rule of thumb, headline scores less than 0.5 percentage points apart are close enough that we would not put much weight on their ordering without uncertainty estimates. A gap of 0.5 to less than 1 point is a small lead: potentially useful, but not a large separation. A lead of 1 to less than 3 points is notable, particularly among already strong models. Gaps of 3 points or more are substantial relative to the high-performing models in the published v1 results. For example, 0.640 versus 0.643 is a close result, while 0.640 versus 0.660 represents a notable two-point lead.
This mini update adds the following models to the original set:
| CareQA (EN) | HeadQA v2 | LongHealth T1 | LongHealth T2 | M-ARC | MedHALT R-FCT | MedHALT R-NOTA | MedMCQA | MedBullets O4 | MedBullets O5 | MedCalc-Bench | MedConcepts E | MedConcepts H | MedConcepts M | MedHallu E | MedHallu H | MedHallu M | MedQA | MedXpert R | MedXpert U | MetaMedQA | MMLU-Pro Hlth | PubHealth Rev | PubMedQA | SCT Public | SuperGPQA E | SuperGPQA H | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3 Pro Preview | Undisclosed | API | 64.5% | 96.8% | 94.0% | 91.3% | 91.0% | 71.3% | 94.6% | 74.4% | 86.9% | 90.9% | 87.7% | 84.1% | 100.0% | 99.2% | 99.0% | 86.5% | 74.5% | 85.5% | 95.4% | 70.8% | 68.7% | 81.7% | 83.7% | 94.8% | 81.3% | 68.8% | 76.9% | 60.7% | |
| 2 | GPT-5.1 (medium) | OpenAI | Undisclosed | API | 62.1% | 94.9% | 92.3% | 90.8% | 90.3% | 65.0% | 93.4% | 69.9% | 84.7% | 89.5% | 87.2% | 79.8% | 100.0% | 94.4% | 97.1% | 80.1% | 60.8% | 74.1% | 89.7% | 45.5% | 43.6% | 82.7% | 82.2% | 93.6% | 77.7% | 78.0% | 68.9% | 52.4% |
| 3 | Grok 4 | xAI | Undisclosed | API | 61.6% | 95.1% | 91.7% | 90.4% | 86.5% | 80.0% | 93.5% | 78.5% | 83.8% | 87.8% | 84.0% | 78.0% | 100.0% | 99.5% | 99.8% | 70.4% | 46.5% | 62.6% | 93.5% | 45.6% | 47.1% | 81.7% | 82.4% | 91.7% | 78.2% | 73.6% | 66.5% | 46.9% |
| 4 | Gemma 4 31B Thinking New | 31.27B | Medium | 60.7% | 93.6% | 91.1% | 91.2% | 89.7% | 76.3% | 90.6% | 66.5% | 80.1% | 87.7% | 85.0% | 67.3% | 98.9% | 69.9% | 74.7% | 83.8% | 63.9% | 79.2% | 94.2% | 59.8% | 53.5% | 79.4% | 80.8% | 93.0% | 80.5% | 73.6% | 66.8% | 62.5% | |
| 5 | Claude Sonnet 4.5 | Anthropic | Undisclosed | API | 60.7% | 93.5% | 91.4% | 90.9% | 90.6% | 71.0% | 89.7% | 56.0% | 80.9% | 84.0% | 80.5% | 75.9% | 99.9% | 95.9% | 97.4% | 83.8% | 60.0% | 73.9% | 92.8% | 40.7% | 36.8% | 79.0% | 83.9% | 91.5% | 76.3% | 72.3% | 67.5% | 47.9% |
| 6 | GPT-5.2 (medium) | OpenAI | Undisclosed | API | 60.5% | 94.7% | 92.0% | 89.1% | 88.7% | 60.0% | 90.0% | 69.8% | 83.7% | 90.0% | 87.2% | 81.4% | 99.9% | 93.2% | 95.9% | 57.7% | 37.7% | 49.5% | 90.5% | 42.3% | 39.9% | 82.9% | 81.4% | 94.1% | 75.9% | 74.3% | 67.5% | 48.2% |
| 7 | GLM 4.7 FP8 | Z.ai | 355B | Large | 60.1% | 94.4% | 91.3% | 91.2% | 88.8% | 58.3% | 86.2% | 69.0% | 80.2% | 88.0% | 84.5% | 74.5% | 99.9% | 86.6% | 93.2% | 72.2% | 52.0% | 66.5% | 94.8% | 43.7% | 42.1% | 79.8% | 80.1% | 90.7% | 77.4% | 68.1% | 66.1% | 48.4% |
| 8 | Qwen3.5 27B Thinking New | Alibaba | 27.78B | Medium | 59.3% | 93.4% | 91.0% | 91.4% | 90.2% | 55.3% | 90.9% | 68.7% | 78.5% | 87.4% | 83.9% | 57.7% | 99.3% | 75.1% | 82.3% | 76.5% | 53.1% | 73.4% | 94.6% | 44.4% | 39.9% | 79.7% | 79.3% | 91.5% | 80.9% | 71.6% | 67.1% | 54.7% |
| 9 | Qwen3.6 35B-A3B Thinking New | Alibaba | 35.95B | Medium | 59.3% | 92.4% | 90.2% | 90.6% | 90.8% | 52.0% | 82.8% | 63.7% | 77.5% | 86.9% | 82.9% | 78.7% | 99.1% | 75.7% | 83.1% | 82.7% | 62.3% | 75.3% | 91.9% | 41.8% | 37.8% | 78.7% | 79.4% | 91.0% | 80.5% | 73.4% | 67.3% | 55.5% |
| 10 | Gemma 4 31B Instruct New | 31.27B | Medium | 59.1% | 92.3% | 89.7% | 90.5% | 89.5% | 69.3% | 78.2% | 53.3% | 78.7% | 84.7% | 81.4% | 64.5% | 98.8% | 68.6% | 72.9% | 84.1% | 64.2% | 80.3% | 92.9% | 56.6% | 50.5% | 71.1% | 79.1% | 93.1% | 79.7% | 74.1% | 62.0% | 60.2% | |
| 11 | Qwen3.5 35B-A3B Thinking New | Alibaba | 35.95B | Medium | 58.6% | 92.8% | 90.5% | 91.8% | 90.8% | 45.3% | 85.5% | 65.5% | 78.9% | 87.8% | 83.7% | 50.7% | 99.3% | 77.5% | 84.8% | 77.4% | 53.0% | 72.0% | 93.7% | 43.7% | 39.0% | 78.6% | 79.1% | 89.9% | 80.6% | 69.4% | 65.3% | 49.8% |
| 12 | Qwen3.6 27B Thinking New | Alibaba | 27.78B | Medium | 58.5% | 92.8% | 90.6% | 91.1% | 91.3% | 59.7% | 87.6% | 65.5% | 61.0% | 88.0% | 82.6% | 81.6% | 98.0% | 71.1% | 76.3% | 82.7% | 63.0% | 78.6% | 86.5% | 43.0% | 37.1% | 79.2% | 78.3% | 91.9% | 81.1% | 67.5% | 67.8% | 57.5% |
| 13 | Qwen3.8 27B Thinking (xhigh) New | Alibaba | 27.78B | Medium | 58.5% | 91.9% | 90.2% | 92.1% | 90.6% | 56.7% | 88.2% | 59.4% | 76.5% | 88.6% | 85.9% | 71.4% | 97.9% | 64.7% | 68.5% | 78.5% | 60.8% | 72.1% | 90.0% | 44.7% | 36.7% | 81.1% | 79.8% | 92.9% | 75.3% | 75.4% | 65.7% | 50.1% |
| 14 | Qwen3 235B-A22B Thinking | Alibaba | 235B | Large | 58.4% | 93.1% | 90.0% | 90.7% | 90.2% | 66.0% | 90.1% | 68.3% | 78.2% | 85.7% | 81.8% | 69.8% | 99.5% | 79.3% | 86.0% | 69.9% | 51.7% | 66.7% | 93.0% | 36.9% | 36.4% | 78.0% | 78.7% | 87.4% | 77.1% | 64.4% | 66.7% | 45.9% |
| 15 | Qwen3.8 27B Thinking (medium) New | Alibaba | 27.78B | Medium | 58.1% | 91.7% | 89.5% | 91.8% | 90.2% | 68.7% | 82.5% | 59.7% | 75.7% | 85.4% | 83.3% | 76.0% | 97.7% | 63.5% | 66.8% | 71.7% | 48.5% | 62.7% | 93.1% | 40.6% | 35.9% | 78.5% | 77.3% | 92.9% | 79.5% | 71.0% | 65.2% | 51.6% |
| 16 | Gemma 4 26B-A4B Thinking New | 26.54B | Medium | 57.9% | 92.1% | 89.8% | 90.7% | 89.3% | 53.7% | 84.8% | 64.6% | 77.5% | 86.6% | 83.7% | 65.5% | 98.2% | 67.8% | 70.5% | 78.3% | 61.8% | 73.0% | 93.5% | 53.1% | 44.2% | 78.7% | 79.0% | 69.9% | 79.0% | 75.3% | 60.3% | 54.8% | |
| 17 | Baichuan M3 235B | Baichuan | 235B | Large | 57.9% | 91.6% | 88.9% | 88.6% | 89.5% | 68.3% | 86.5% | 66.6% | 77.9% | 85.5% | 82.9% | 67.1% | 98.9% | 73.9% | 80.9% | 71.2% | 53.4% | 68.7% | 92.7% | 40.9% | 39.0% | 74.2% | 77.2% | 85.6% | 78.7% | 63.9% | 61.9% | 47.8% |
| 18 | Qwen3.8 27B Thinking (low) New | Alibaba | 27.78B | Medium | 57.8% | 91.3% | 89.2% | 91.7% | 90.0% | 63.0% | 85.4% | 59.4% | 74.8% | 87.3% | 83.7% | 80.5% | 97.6% | 62.9% | 67.5% | 77.7% | 52.9% | 67.2% | 83.4% | 40.8% | 34.0% | 76.8% | 77.3% | 92.2% | 76.9% | 74.4% | 65.3% | 52.1% |
| 19 | Muse Glimmer 30B (xhigh) New | Meta | 29.78B | Medium | 57.6% | 91.5% | 89.4% | 90.4% | 86.5% | 86.7% | 91.6% | 71.6% | 76.1% | 84.4% | 82.5% | 72.0% | 99.2% | 83.7% | 88.1% | 56.8% | 40.5% | 52.0% | 74.9% | 35.5% | 31.9% | 79.9% | 78.0% | 88.7% | 69.1% | 70.9% | 60.1% | 39.8% |
| 20 | Muse Glimmer 30B (high) New | Meta | 29.78B | Medium | 57.4% | 91.4% | 89.2% | 90.2% | 86.4% | 86.7% | 91.5% | 70.8% | 76.0% | 86.0% | 81.0% | 67.7% | 99.3% | 82.2% | 88.4% | 57.4% | 40.4% | 52.5% | 77.0% | 36.4% | 32.4% | 79.1% | 77.2% | 89.0% | 68.7% | 68.9% | 59.3% | 38.9% |
| 21 | Qwen3 Next 80B-A3B Thinking | Alibaba | 80B | Large | 57.0% | 92.1% | 89.1% | 89.2% | 87.1% | 75.0% | 86.9% | 78.9% | 76.3% | 82.4% | 78.8% | 68.3% | 98.9% | 71.2% | 75.6% | 62.5% | 45.4% | 55.9% | 91.1% | 32.8% | 30.6% | 75.4% | 76.9% | 84.6% | 75.9% | 68.2% | 62.6% | 49.5% |
| 22 | MiniMax M2.1 | MiniMax | 230B | Large | 56.8% | 91.8% | 88.0% | 88.8% | 88.2% | 41.7% | 73.3% | 61.1% | 76.8% | 83.9% | 77.3% | 66.5% | 99.0% | 76.7% | 83.8% | 57.6% | 38.3% | 49.9% | 91.1% | 34.4% | 34.2% | 77.3% | 83.0% | 90.6% | 75.9% | 70.0% | 59.2% | 47.8% |
| 23 | gpt-oss 120b (high) | OpenAI | 117B | Medium | 56.7% | 90.8% | 88.2% | 89.3% | 88.4% | 37.3% | 87.7% | 59.0% | 75.8% | 85.4% | 82.6% | 71.7% | 97.6% | 69.1% | 73.8% | 73.5% | 48.9% | 66.5% | 93.3% | 39.2% | 34.7% | 79.4% | 76.1% | 88.9% | 76.7% | 67.6% | 54.5% | 46.7% |
| 24 | Qwen3.6 35B-A3B Instruct New | Alibaba | 35.95B | Medium | 56.6% | 92.2% | 88.8% | 88.1% | 89.5% | 45.0% | 69.2% | 49.2% | 76.9% | 85.3% | 82.4% | 82.0% | 99.4% | 76.7% | 83.3% | 81.2% | 55.4% | 75.0% | 92.6% | 39.0% | 36.0% | 72.1% | 78.6% | 58.2% | 79.1% | 63.0% | 65.2% | 53.8% |
| 25 | Muse Glimmer 30B (medium) New | Meta | 29.78B | Medium | 56.6% | 91.3% | 88.8% | 89.8% | 86.6% | 86.0% | 90.7% | 69.3% | 74.8% | 81.7% | 76.8% | 54.5% | 99.1% | 81.6% | 86.7% | 57.9% | 39.8% | 53.7% | 67.3% | 35.1% | 34.5% | 77.9% | 77.8% | 88.6% | 68.4% | 71.5% | 58.1% | 38.2% |
| 26 | Gemma 4 26B-A4B Instruct New | 26.54B | Medium | 55.8% | 91.1% | 88.6% | 86.8% | 85.9% | 42.3% | 59.1% | 53.0% | 75.2% | 84.3% | 78.2% | 63.5% | 97.8% | 67.6% | 69.4% | 78.6% | 59.4% | 70.4% | 91.7% | 46.6% | 42.0% | 66.9% | 76.7% | 88.9% | 78.1% | 66.9% | 57.8% | 51.5% | |
| 27 | Muse Glimmer 30B (low) New | Meta | 29.78B | Medium | 55.8% | 89.6% | 88.3% | 89.3% | 87.8% | 80.0% | 88.3% | 66.5% | 74.1% | 78.5% | 73.7% | 45.9% | 98.6% | 78.5% | 85.2% | 54.6% | 36.7% | 51.9% | 76.3% | 33.5% | 32.2% | 75.2% | 77.8% | 89.5% | 72.3% | 71.9% | 57.4% | 37.9% |
| 28 | gpt-oss 120b (medium) | OpenAI | 117B | Medium | 55.7% | 90.2% | 87.4% | 88.3% | 87.6% | 34.7% | 88.0% | 60.6% | 74.3% | 84.3% | 81.0% | 68.9% | 96.1% | 65.3% | 67.9% | 67.9% | 44.5% | 59.6% | 91.5% | 35.6% | 31.9% | 77.3% | 76.0% | 88.2% | 78.1% | 71.4% | 52.7% | 45.0% |
| 29 | Qwen3.8 27B Instruct New | Alibaba | 27.78B | Medium | 55.7% | 91.4% | 87.9% | 89.3% | 89.7% | 49.3% | 71.5% | 49.5% | 69.4% | 83.3% | 81.2% | 80.7% | 97.0% | 63.5% | 66.9% | 82.5% | 59.5% | 73.5% | 89.2% | 39.2% | 32.8% | 65.8% | 76.2% | 91.7% | 68.5% | 67.1% | 61.3% | 51.9% |
| 30 | MiniMax M2 | MiniMax | 230B | Large | 55.1% | 91.3% | 87.8% | 88.9% | 87.1% | 46.3% | 75.6% | 59.9% | 76.3% | 81.4% | 76.4% | 25.2% | 98.2% | 76.7% | 84.0% | 61.5% | 40.9% | 54.1% | 90.7% | 33.8% | 31.4% | 76.5% | 78.4% | 85.1% | 76.6% | 69.5% | 55.4% | 40.6% |
| 31 | Qwen3 Next 80B-A3B Instruct | Alibaba | 80B | Large | 54.9% | 91.6% | 86.4% | 87.4% | 90.3% | 53.0% | 69.9% | 55.2% | 76.1% | 80.2% | 76.9% | 64.0% | 97.8% | 71.1% | 72.6% | 50.2% | 26.5% | 41.5% | 90.1% | 29.1% | 28.7% | 66.4% | 77.8% | 87.7% | 76.9% | 65.9% | 62.7% | 46.2% |
| 32 | Intellect 3 | Prime Intellect | 106B | Large | 54.6% | 90.6% | 88.3% | 88.7% | 88.5% | 52.3% | 55.9% | 59.0% | 74.0% | 80.6% | 76.8% | 63.6% | 98.0% | 67.2% | 74.1% | 70.7% | 42.7% | 62.5% | 90.5% | 31.4% | 30.5% | 75.3% | 76.7% | 86.1% | 75.5% | 64.2% | 56.7% | 43.3% |
| 33 | Qwen3 30B-A3B Thinking FP8 | Alibaba | 30.5B | Medium | 54.0% | 89.2% | 86.6% | 88.2% | 88.3% | 63.3% | 78.9% | 66.7% | 72.7% | 78.2% | 73.9% | 62.1% | 94.8% | 59.4% | 57.4% | 70.1% | 47.5% | 63.4% | 88.0% | 26.4% | 27.0% | 73.4% | 75.9% | 83.4% | 76.6% | 64.1% | 58.4% | 43.6% |
| 34 | Qwen3 30B-A3B Thinking 8-bit | Alibaba | 30.5B | Medium | 54.0% | 89.5% | 86.6% | 88.6% | 87.5% | 60.3% | 78.6% | 67.2% | 72.4% | 78.6% | 73.9% | 62.2% | 94.7% | 59.4% | 57.7% | 67.2% | 46.2% | 63.2% | 88.6% | 26.4% | 26.1% | 72.8% | 76.4% | 83.7% | 76.8% | 66.9% | 58.0% | 43.2% |
| 35 | Qwen3 30B-A3B Thinking | Alibaba | 30.5B | Medium | 53.9% | 89.2% | 86.6% | 87.6% | 87.7% | 61.3% | 78.2% | 67.2% | 72.8% | 77.1% | 73.8% | 61.8% | 95.0% | 59.6% | 57.6% | 67.5% | 47.1% | 62.6% | 88.8% | 26.3% | 26.0% | 73.5% | 75.1% | 83.4% | 77.5% | 62.6% | 58.2% | 46.4% |
| 36 | AntAngelMed 100B | Ant Group | 100B | Large | 53.3% | 88.6% | 87.0% | 88.7% | 86.7% | 60.3% | 78.2% | 70.4% | 73.9% | 79.3% | 73.1% | 50.1% | 82.1% | 48.0% | 60.4% | 67.6% | 48.4% | 63.6% | 90.0% | 28.0% | 26.1% | 76.4% | 77.0% | 86.2% | 67.8% | 74.1% | 57.4% | 45.5% |
| 37 | gpt-oss 120b (low) | OpenAI | 117B | Medium | 53.2% | 88.4% | 85.8% | 87.4% | 85.9% | 29.3% | 86.7% | 58.9% | 71.9% | 81.6% | 76.1% | 66.1% | 89.9% | 53.0% | 60.8% | 61.5% | 40.4% | 59.0% | 88.6% | 28.9% | 28.0% | 73.5% | 73.7% | 87.9% | 77.4% | 70.4% | 49.6% | 41.9% |
| 38 | Baichuan M2 32B | Baichuan | 32B | Medium | 53.2% | 87.3% | 84.8% | 88.8% | 88.0% | 42.7% | 80.4% | 66.1% | 72.1% | 76.5% | 72.7% | 67.5% | 93.4% | 55.3% | 52.7% | 60.6% | 36.8% | 52.6% | 89.4% | 27.5% | 25.1% | 74.8% | 75.9% | 85.7% | 69.9% | 71.2% | 57.2% | 38.7% |
| 39 | Qwen3 VL 30B-A3B Thinking | Alibaba | 30.5B | Medium | 53.1% | 90.1% | 87.5% | 87.8% | 87.8% | 44.3% | 72.4% | 67.0% | 73.0% | 78.6% | 73.4% | 24.9% | 94.9% | 59.5% | 56.9% | 66.5% | 46.9% | 61.9% | 89.1% | 27.8% | 24.8% | 73.5% | 76.0% | 85.4% | 78.5% | 65.2% | 58.8% | 43.6% |
| 40 | Qwen3 30B-A3B Thinking 4-bit | Alibaba | 30.5B | Small | 52.9% | 88.7% | 86.4% | 87.2% | 87.0% | 49.0% | 76.4% | 68.5% | 71.5% | 76.1% | 71.6% | 59.7% | 94.1% | 55.4% | 54.7% | 66.8% | 45.6% | 60.8% | 87.7% | 25.1% | 24.7% | 71.3% | 74.7% | 82.5% | 77.0% | 65.6% | 55.9% | 44.5% |
| 41 | Nemotron 3.5 Lightning 30B-A3B New | NVIDIA | 31.58B | Medium | 52.5% | 90.0% | 87.5% | 87.4% | 88.9% | 38.0% | 86.5% | 69.2% | 67.8% | 79.7% | 73.5% | 58.7% | 86.0% | 59.3% | 54.5% | 51.6% | 30.4% | 42.8% | 85.7% | 29.1% | 27.4% | 74.9% | 74.3% | 86.5% | 70.7% | 71.9% | 53.2% | 42.5% |
| 42 | GLM 4.5 Air | Z.ai | 106B | Large | 52.1% | 88.3% | 88.5% | 84.5% | 87.0% | 33.7% | 65.2% | 58.3% | 66.1% | 78.6% | 74.6% | 59.1% | 94.9% | 67.6% | 75.4% | 65.3% | 39.2% | 55.1% | 85.7% | 16.0% | 16.8% | 74.6% | 74.7% | 80.4% | 65.5% | 64.9% | 58.2% | 40.4% |
| 43 | Llama 3.3 70B Instruct | Meta | 70B | Large | 51.7% | 86.8% | 82.4% | 85.8% | 88.6% | 36.0% | 58.2% | 35.5% | 74.1% | 71.6% | 68.9% | 49.3% | 98.5% | 68.0% | 76.2% | 74.3% | 47.1% | 59.7% | 85.0% | 22.4% | 24.6% | 63.3% | 72.5% | 88.2% | 79.3% | 60.4% | 49.2% | 33.5% |
| 44 | Qwen3 14B (Thinking) | Alibaba | 14B | Small | 51.7% | 87.2% | 85.2% | 88.2% | 87.8% | 45.0% | 71.3% | 73.8% | 69.7% | 73.7% | 68.0% | 47.5% | 93.2% | 55.2% | 53.0% | 53.7% | 35.9% | 48.0% | 83.8% | 22.0% | 22.2% | 64.8% | 72.6% | 83.7% | 76.0% | 66.4% | 53.6% | 36.7% |
| 45 | gpt-oss 20b (high) | OpenAI | 21B | Small | 51.6% | 86.3% | 84.9% | 87.8% | 88.0% | 30.3% | 80.5% | 49.9% | 69.9% | 77.1% | 74.2% | 60.6% | 84.6% | 50.5% | 52.3% | 68.0% | 46.3% | 65.8% | 88.6% | 28.5% | 25.7% | 74.3% | 72.7% | 81.9% | 76.3% | 68.3% | 46.0% | 40.7% |
| 46 | Qwen3 30B-A3B Instruct 8-bit | Alibaba | 30.5B | Medium | 51.2% | 88.7% | 82.6% | 84.5% | 88.6% | 42.7% | 57.7% | 51.4% | 71.5% | 75.8% | 70.7% | 56.8% | 92.1% | 52.7% | 52.5% | 59.2% | 31.3% | 47.7% | 86.8% | 25.0% | 25.2% | 61.9% | 75.0% | 83.8% | 76.4% | 64.2% | 55.7% | 41.6% |
| 47 | Qwen3 30B-A3B Instruct | Alibaba | 30.5B | Medium | 51.2% | 88.8% | 82.6% | 84.1% | 88.6% | 46.3% | 58.4% | 51.6% | 71.6% | 77.2% | 71.0% | 57.5% | 92.3% | 52.0% | 52.9% | 58.8% | 29.0% | 43.6% | 86.2% | 25.1% | 24.6% | 61.8% | 74.4% | 83.9% | 76.6% | 65.7% | 55.7% | 40.1% |
| 48 | Qwen3 30B-A3B Instruct FP8 | Alibaba | 30.5B | Medium | 51.1% | 88.7% | 82.4% | 84.3% | 88.6% | 43.0% | 58.8% | 51.1% | 71.5% | 74.1% | 71.4% | 55.7% | 91.2% | 52.2% | 53.1% | 58.1% | 30.9% | 46.8% | 86.6% | 24.5% | 24.7% | 61.9% | 74.7% | 83.8% | 76.3% | 66.9% | 56.1% | 39.8% |
| 49 | Qwen3 30B-A3B Instruct 4-bit | Alibaba | 30.5B | Small | 50.5% | 88.2% | 82.4% | 83.1% | 87.8% | 39.7% | 54.0% | 54.3% | 70.5% | 74.9% | 70.2% | 53.3% | 91.4% | 46.4% | 51.3% | 63.5% | 35.8% | 51.4% | 85.8% | 24.0% | 24.1% | 60.1% | 74.2% | 84.2% | 76.7% | 64.2% | 53.6% | 40.7% |
| 50 | gpt-oss 20b (medium) | OpenAI | 21B | Small | 50.0% | 85.3% | 82.4% | 88.3% | 87.4% | 26.0% | 75.7% | 49.8% | 68.6% | 72.3% | 70.2% | 55.9% | 81.0% | 46.0% | 47.1% | 69.3% | 46.8% | 63.2% | 85.1% | 24.2% | 23.1% | 71.6% | 71.4% | 81.7% | 75.6% | 65.6% | 44.3% | 38.7% |
| 51 | Ling Flash 2.0 | Ant Group | 100B | Large | 49.7% | 86.9% | 83.7% | 84.4% | 85.5% | 26.7% | 65.9% | 61.5% | 68.0% | 73.2% | 67.3% | 46.2% | 86.2% | 47.9% | 48.7% | 76.5% | 54.4% | 72.2% | 83.7% | 20.2% | 21.6% | 53.1% | 71.7% | 86.5% | 76.5% | 52.9% | 54.7% | 34.9% |
| 52 | Nemotron Nano V3 30B-A3B | NVIDIA | 30.5B | Medium | 49.2% | 87.2% | 84.9% | 88.0% | 87.2% | 29.3% | 22.9% | 51.5% | 69.6% | 77.3% | 73.7% | 35.5% | 86.3% | 52.6% | 47.6% | 61.6% | 42.4% | 58.0% | 86.9% | 27.9% | 25.2% | 73.6% | 72.9% | 87.3% | 72.2% | 63.7% | 48.6% | 32.9% |
| 53 | Olmo 3.1 32B Think | Ai2 | 32B | Medium | 48.7% | 84.2% | 82.4% | 86.1% | 81.6% | 40.3% | 81.9% | 72.1% | 63.7% | 67.9% | 62.9% | 16.4% | 83.5% | 48.7% | 47.7% | 50.4% | 34.7% | 42.5% | 80.6% | 21.2% | 22.6% | 60.6% | 70.9% | 82.9% | 70.9% | 65.7% | 45.5% | 33.3% |
| 54 | MedGemma 27B | 27B | Medium | 48.3% | 87.0% | 83.9% | 82.8% | 80.2% | 39.3% | 48.2% | 35.6% | 65.0% | 79.4% | 75.4% | 48.6% | 77.5% | 49.2% | 43.6% | 74.7% | 50.7% | 68.8% | 87.4% | 26.4% | 24.2% | 36.2% | 72.3% | 84.7% | 71.5% | 59.2% | 50.1% | 31.6% | |
| 55 | Qwen3 8B (Thinking) | Alibaba | 8B | Small | 48.3% | 84.2% | 82.4% | 87.2% | 87.9% | 26.7% | 49.4% | 67.8% | 65.8% | 66.6% | 62.2% | 45.2% | 86.8% | 48.0% | 38.8% | 55.9% | 37.6% | 51.5% | 79.9% | 17.8% | 21.1% | 62.1% | 69.6% | 82.4% | 74.8% | 67.8% | 48.1% | 32.9% |
| 56 | Olmo 3 32B Think | Ai2 | 32B | Medium | 48.2% | 83.8% | 81.7% | 86.5% | 84.3% | 29.7% | 71.7% | 63.5% | 63.4% | 64.9% | 62.3% | 7.8% | 82.5% | 50.1% | 47.2% | 59.5% | 40.3% | 53.1% | 80.5% | 20.6% | 22.2% | 65.4% | 69.0% | 84.0% | 73.0% | 67.6% | 45.5% | 34.3% |
| 57 | Qwen2.5 32B Instruct | Alibaba | 32.5B | Medium | 48.1% | 82.0% | 80.0% | 86.1% | 88.5% | 20.7% | 70.6% | 68.8% | 64.4% | 61.9% | 54.1% | 45.4% | 93.0% | 48.6% | 52.1% | 46.3% | 22.2% | 35.0% | 75.4% | 15.8% | 16.7% | 57.6% | 71.2% | 85.6% | 73.8% | 50.2% | 48.5% | 25.3% |
| 58 | Nemotron Nano 12B V2 | NVIDIA | 12.31B | Small | 48.1% | 85.0% | 82.3% | 85.7% | 86.7% | 23.3% | 47.1% | 60.2% | 64.1% | 70.6% | 62.7% | 47.2% | 88.9% | 51.8% | 40.6% | 60.0% | 42.9% | 50.9% | 81.5% | 20.6% | 20.4% | 66.2% | 71.1% | 85.0% | 66.7% | 63.7% | 44.8% | 31.2% |
| 59 | Qwen3 4B Thinking | Alibaba | 4B | Tiny | 47.0% | 83.1% | 81.7% | 86.8% | 86.7% | 57.7% | 33.6% | 59.4% | 63.0% | 63.6% | 57.4% | 36.9% | 75.6% | 50.8% | 38.4% | 67.9% | 47.9% | 62.1% | 76.5% | 17.3% | 17.3% | 63.4% | 67.8% | 78.9% | 74.6% | 61.5% | 48.0% | 28.0% |
| 60 | Trinity Mini | Arcee AI | 26B | Medium | 46.6% | 83.6% | 81.4% | 79.0% | 80.5% | 21.0% | 30.0% | 49.1% | 65.0% | 66.1% | 60.0% | 0.0% | 95.4% | 54.5% | 65.9% | 64.5% | 42.0% | 56.9% | 78.3% | 18.5% | 19.9% | 64.9% | 69.8% | 82.8% | 74.6% | 58.6% | 45.8% | 29.0% |
| 61 | Hermes 4 70B | Nous Research | 70B | Large | 46.5% | 72.4% | 79.9% | 82.5% | 86.9% | 23.0% | 44.8% | 38.2% | 65.0% | 58.2% | 54.8% | 35.4% | 95.6% | 50.0% | 66.6% | 69.2% | 52.0% | 62.6% | 75.1% | 15.6% | 19.0% | 56.6% | 68.5% | 86.3% | 75.6% | 53.5% | 43.4% | 27.8% |
| 62 | gpt-oss 20b (low) | OpenAI | 21B | Small | 46.2% | 81.1% | 78.6% | 85.5% | 85.4% | 19.0% | 57.9% | 49.4% | 64.4% | 66.3% | 58.7% | 47.9% | 69.6% | 41.3% | 40.5% | 65.4% | 47.3% | 61.1% | 78.7% | 19.3% | 18.7% | 65.5% | 66.4% | 80.0% | 74.8% | 64.0% | 40.4% | 28.0% |
| 63 | Phi 4 Reasoning | Microsoft | 14B | Small | 46.1% | 69.7% | 82.4% | 87.8% | 88.0% | 39.0% | 60.8% | 58.3% | 63.4% | 65.2% | 58.9% | 53.0% | 61.6% | 43.8% | 43.4% | 64.8% | 48.4% | 59.8% | 73.0% | 20.2% | 21.7% | 59.5% | 73.8% | 85.7% | 55.6% | 64.1% | 48.1% | 36.7% |
| 64 | Ministral 3 14B Reasoning | Mistral AI | 14B | Small | 46.1% | 84.0% | 76.0% | 83.8% | 87.9% | 23.3% | 34.6% | 43.4% | 64.8% | 65.2% | 60.5% | 47.0% | 89.3% | 47.1% | 50.2% | 59.9% | 39.8% | 52.9% | 77.2% | 19.5% | 20.1% | 51.7% | 69.0% | 81.8% | 69.8% | 49.0% | 42.9% | 24.9% |
| 65 | Hermes 4 14B | Nous Research | 14B | Small | 45.9% | 78.6% | 78.2% | 80.8% | 83.0% | 33.7% | 33.0% | 54.2% | 63.7% | 62.4% | 56.7% | 42.0% | 91.5% | 48.4% | 52.2% | 67.7% | 50.8% | 59.3% | 73.4% | 17.2% | 17.3% | 46.5% | 64.3% | 82.6% | 76.1% | 53.1% | 46.6% | 29.0% |
| 66 | Mirothinker 1.5 30B | MiroMind | 30.5B | Medium | 45.7% | 83.4% | 81.6% | 82.8% | 84.4% | 26.0% | 51.8% | 60.0% | 67.6% | 66.1% | 62.7% | 15.8% | 72.3% | 37.8% | 36.4% | 62.3% | 43.1% | 55.0% | 78.8% | 19.5% | 20.7% | 65.9% | 66.1% | 79.1% | 70.9% | 63.7% | 47.1% | 36.1% |
| 67 | Ministral 3 14B Instruct | Mistral AI | 14B | Small | 45.6% | 85.6% | 76.8% | 83.2% | 88.9% | 30.7% | 8.8% | 9.3% | 67.6% | 68.4% | 62.7% | 49.5% | 84.6% | 53.6% | 50.4% | 60.8% | 37.3% | 53.6% | 80.6% | 21.5% | 21.4% | 53.8% | 70.9% | 83.4% | 73.2% | 28.8% | 46.5% | 27.6% |
| 68 | Magistral Small | Mistral AI | 24B | Medium | 45.5% | 85.3% | 78.4% | 85.1% | 89.0% | 20.7% | 35.9% | 32.2% | 66.8% | 69.6% | 65.2% | 48.0% | 77.2% | 34.0% | 43.9% | 51.5% | 35.7% | 47.9% | 79.5% | 18.3% | 21.8% | 54.8% | 68.5% | 85.0% | 75.5% | 52.8% | 44.6% | 27.6% |
| 69 | Gemma 3 27B | 27B | Medium | 44.2% | 81.5% | 77.1% | 80.0% | 76.1% | 24.3% | 17.8% | 18.9% | 63.5% | 62.4% | 56.0% | 43.5% | 85.4% | 50.1% | 45.8% | 53.2% | 37.3% | 51.2% | 77.1% | 16.4% | 19.1% | 54.8% | 66.6% | 82.2% | 68.0% | 61.1% | 41.7% | 21.4% | |
| 70 | Olmo 3.1 32B Instruct | Ai2 | 32B | Medium | 43.7% | 80.3% | 69.1% | 78.8% | 77.7% | 29.3% | 46.1% | 41.2% | 53.6% | 59.7% | 56.4% | 38.0% | 77.5% | 40.9% | 42.8% | 42.6% | 32.9% | 38.2% | 74.4% | 17.5% | 19.7% | 43.6% | 65.8% | 83.9% | 65.9% | 51.6% | 41.5% | 23.8% |
| 71 | DASD 4B Thinking | Alibaba-Apsara | 4B | Tiny | 43.3% | 78.5% | 77.1% | 84.7% | 85.5% | 30.0% | 27.4% | 44.6% | 58.1% | 59.7% | 55.8% | 29.4% | 57.0% | 40.5% | 32.2% | 62.9% | 47.8% | 60.0% | 73.2% | 19.5% | 17.0% | 60.1% | 61.2% | 80.0% | 73.3% | 59.2% | 42.0% | 31.8% |
| 72 | Jamba2 Mini 52B | AI21 Labs | 52B | Large | 42.7% | 78.8% | 69.3% | 82.1% | 78.5% | 18.7% | 17.9% | 23.0% | 59.8% | 61.5% | 52.9% | 32.5% | 76.3% | 48.3% | 49.1% | 69.0% | 53.4% | 61.2% | 71.7% | 15.4% | 16.2% | 44.5% | 60.7% | 82.1% | 73.9% | 48.5% | 39.8% | 22.6% |
| 73 | Granite 4.0H Small | IBM | 32B | Medium | 42.7% | 74.6% | 75.7% | 77.0% | 52.6% | 21.3% | 9.9% | 38.0% | 59.8% | 57.5% | 52.7% | 26.8% | 90.2% | 53.0% | 62.1% | 66.3% | 50.9% | 59.0% | 69.6% | 14.7% | 16.1% | 51.6% | 63.7% | 83.6% | 61.5% | 46.2% | 39.4% | 23.0% |
| 74 | Ministral 3 8B Instruct | Mistral AI | 8B | Small | 42.5% | 83.2% | 73.3% | 83.6% | 89.3% | 20.7% | 3.8% | 3.0% | 65.9% | 63.1% | 59.1% | 47.3% | 73.6% | 43.9% | 38.5% | 50.0% | 31.5% | 43.0% | 77.6% | 17.4% | 20.3% | 52.3% | 68.3% | 83.6% | 73.7% | 4.0% | 43.5% | 26.4% |
| 75 | Gemma 3 12B | 12B | Small | 41.9% | 76.9% | 73.0% | 76.4% | 74.8% | 20.0% | 20.5% | 20.2% | 58.6% | 57.1% | 50.5% | 26.9% | 84.9% | 47.4% | 42.0% | 52.1% | 37.5% | 45.5% | 68.7% | 15.4% | 17.0% | 50.8% | 59.3% | 80.9% | 67.8% | 51.7% | 36.5% | 20.6% | |
| 76 | Ministral 3 8B Reasoning | Mistral AI | 8B | Small | 41.2% | 80.0% | 62.6% | 82.8% | 85.4% | 25.7% | 23.3% | 45.3% | 59.4% | 59.2% | 52.9% | 36.2% | 47.0% | 33.6% | 31.6% | 57.5% | 34.6% | 46.7% | 71.8% | 15.7% | 18.2% | 46.8% | 61.6% | 80.1% | 74.4% | 48.4% | 38.2% | 24.3% |
| 77 | Olmo 3 7B Think | Ai2 | 7B | Small | 39.5% | 71.5% | 71.0% | 79.8% | 81.6% | 16.3% | 51.8% | 56.2% | 51.6% | 50.0% | 41.7% | 0.7% | 32.3% | 36.1% | 31.1% | 59.4% | 44.8% | 51.4% | 61.6% | 12.9% | 14.7% | 47.1% | 53.9% | 76.9% | 67.9% | 51.5% | 34.4% | 25.8% |
| 78 | DASD 30B-A3B | Alibaba-Apsara | 30.5B | Medium | 39.0% | 82.8% | 79.3% | 87.8% | 86.9% | 34.3% | 57.7% | 41.0% | 14.1% | 75.4% | 68.0% | 47.0% | 19.0% | 10.6% | 12.2% | 67.4% | 46.1% | 63.3% | 11.9% | 2.5% | 1.8% | 72.2% | 71.2% | 65.8% | 16.6% | 58.4% | 53.9% | 37.6% |
| 79 | Llama 3.1 8B Instruct | Meta | 8B | Small | 37.9% | 72.1% | 68.1% | 76.4% | 70.1% | 30.0% | 28.1% | 25.3% | 49.0% | 57.6% | 52.3% | 21.0% | 30.6% | 21.9% | 25.5% | 67.2% | 49.4% | 60.4% | 57.0% | 12.1% | 14.8% | 46.6% | 55.7% | 81.8% | 74.2% | 43.5% | 36.3% | 22.6% |
| 80 | Ministral 3 3B Instruct | Mistral AI | 3B | Tiny | 37.5% | 76.8% | 65.4% | 79.8% | 84.5% | 15.3% | 0.2% | 0.2% | 53.4% | 52.3% | 46.0% | 29.6% | 52.8% | 42.9% | 32.1% | 64.4% | 43.2% | 56.0% | 64.8% | 12.9% | 16.1% | 39.1% | 58.6% | 80.0% | 70.6% | 0.0% | 35.5% | 20.6% |
| 81 | MedGemma 4B | 4B | Tiny | 36.1% | 64.7% | 60.6% | 67.3% | 33.0% | 23.7% | 17.3% | 16.9% | 55.4% | 52.9% | 46.5% | 6.0% | 32.6% | 31.7% | 30.2% | 56.0% | 50.4% | 52.7% | 64.9% | 14.8% | 13.1% | 43.1% | 43.5% | 78.2% | 72.7% | 45.5% | 31.6% | 19.0% | |
| 82 | MedGemma 4B 1.5 | 4B | Tiny | 35.8% | 66.5% | 66.2% | 56.5% | 36.4% | 22.0% | 13.4% | 37.5% | 51.1% | 60.4% | 53.2% | 24.7% | 34.7% | 30.7% | 28.9% | 56.8% | 36.2% | 48.3% | 67.9% | 17.0% | 14.4% | 44.6% | 51.8% | 52.1% | 61.1% | 57.6% | 33.3% | 22.9% | |
| 83 | Ministral 3 3B Reasoning | Mistral AI | 3B | Tiny | 35.6% | 67.0% | 60.5% | 77.7% | 77.8% | 14.3% | 9.4% | 26.7% | 46.0% | 43.2% | 37.4% | 17.5% | 45.9% | 31.9% | 32.7% | 52.7% | 40.6% | 48.6% | 53.2% | 12.1% | 11.7% | 37.0% | 45.2% | 73.9% | 68.9% | 43.7% | 31.7% | 18.3% |
| 84 | Trinity Nano Preview | Arcee AI | 6B | Tiny | 34.2% | 62.7% | 60.8% | 63.2% | 40.8% | 12.7% | 21.5% | 18.5% | 50.2% | 40.9% | 38.2% | 10.9% | 32.6% | 25.6% | 25.4% | 57.7% | 44.4% | 54.1% | 54.0% | 12.1% | 13.5% | 42.6% | 41.6% | 80.8% | 71.3% | 43.4% | 30.9% | 16.1% |
| 85 | SmolLM3 3B | Hugging Face | 3B | Tiny | 33.7% | 52.9% | 63.7% | 57.2% | 40.0% | 22.0% | 12.9% | 48.1% | 47.0% | 44.4% | 39.6% | 1.6% | 35.1% | 36.1% | 29.1% | 51.7% | 44.0% | 51.3% | 50.0% | 13.9% | 14.7% | 40.6% | 43.1% | 73.9% | 46.1% | 49.6% | 29.4% | 14.1% |
| 86 | Granite 4.0H Tiny | IBM | 7B | Small | 33.5% | 56.0% | 63.8% | 58.8% | 37.6% | 15.7% | 13.0% | 18.5% | 48.5% | 41.5% | 38.1% | 0.9% | 43.4% | 29.0% | 30.1% | 52.2% | 44.9% | 51.7% | 52.0% | 11.9% | 10.3% | 39.6% | 37.5% | 75.9% | 69.5% | 39.1% | 27.3% | 15.4% |
| 87 | Olmo 3 7B Instruct | Ai2 | 7B | Small | 33.4% | 64.7% | 55.6% | 65.2% | 37.7% | 25.7% | 24.1% | 44.8% | 42.9% | 43.2% | 37.6% | 0.1% | 8.0% | 18.3% | 11.4% | 57.9% | 45.9% | 54.3% | 53.7% | 12.6% | 13.6% | 32.5% | 42.2% | 75.1% | 65.3% | 38.0% | 29.0% | 20.9% |
| 88 | Gemma 3 4B | 4B | Tiny | 30.5% | 59.7% | 57.2% | 62.0% | 31.9% | 11.0% | 8.8% | 12.8% | 46.4% | 43.3% | 37.3% | 12.8% | 31.5% | 34.0% | 30.3% | 13.7% | 9.4% | 12.7% | 51.6% | 11.1% | 11.8% | 36.7% | 39.4% | 73.3% | 63.6% | 42.2% | 26.2% | 18.6% | |
| 89 | AFM 4.5B | Arcee AI | 4.5B | Tiny | 30.3% | 48.5% | 58.1% | 63.2% | 48.3% | 16.3% | 12.3% | 22.3% | 44.8% | 37.6% | 33.5% | 7.2% | 16.2% | 12.9% | 20.3% | 33.9% | 27.6% | 29.1% | 49.8% | 16.5% | 16.4% | 34.4% | 28.3% | 78.2% | 64.5% | 28.3% | 29.2% | 18.4% |
No models match these filters.
Ranks and weighted mean win rates use the full 89-configuration snapshot. Each row is one evaluated configuration. Bars show relative scores across that snapshot; detailed scores show per-dataset win rates.
Several models have multiple reasoning modes, so there’s 18 new configurations to increase the total to 89.
Unlike the Gemma 3 family, the Gemma 4 models perform extremely well on Medmarks-Verified. Gemma 4 31B Thinking slots into 4th place, in between Grok 4 and Sonnet 4.5 on the updated leaderboard.
In head-to-head pairings, Gemma 4 31B Thinking's win rate is 0.502 against Sonnet 4.5, 0.506 against GPT-5.2, and 0.510 against GLM-4.7 FP8. These are essentially ties for practical purposes, which is an impressive result given Gemma 4's size. Furthermore, Gemma 4 achieves this using a respectable 1.7K tokens per task. Significantly more than GPT-5.2's 285, but far less than some of the local thinking models in this update.
| Model | Mode | Win rate | Rank | Mean tokens |
|---|---|---|---|---|
| Gemma 3 27B | Instruct | 0.4418 | 69 | 266 |
| MedGemma 27B | Instruct/medical | 0.4829 | 54 | 1,037 |
| Gemma 4 26B-A4B | Instruct | 0.5579 | 26 | 411 |
| Gemma 4 31B | Instruct | 0.5905 | 10 | 317 |
| Gemma 4 26B-A4B | Thinking | 0.5786 | 16 | 3,258 |
| Gemma 4 31B | Thinking | 0.6075 | 4 | 1,727 |
The instruct model is also interesting. Gemma 4 31B Instruct ranks 10th while averaging only 317 output tokens. It has a higher average score than both Gemma 3 27B and MedGemma 27B on all datasets, with a direct winrate of 0.648 and 0.608, respectively.
Gemma 4 31B Thinking is second only to Gemini 3 Pro on both MedXpert Reasoning and MedXpert Understanding. Its rank isn't being carried by easy medical multiple choice.
If you were aware of the hype surrounding Qwen 3.8 27B, you might come away with the impression that it was an Opus class model you could run on your local machine. We won't take any positions on Qwen 3.8 27B's coding or agentic abilities, but on health, the story is much less exciting.
Let's start with the positives. Like Gemma 4, every new Qwen Thinking model outperforms the prior generation’s Qwen3 30B-A3B Thinking. If Gemma 4 didn’t exist, Qwen 3.5 27B would be the best performing model in this weight class. Qwen3.6 also substantially reduced the average number of tokens per generation.
| Model | Architecture | Winrate | Rank | Mean tokens | Δ vs Qwen3 |
|---|---|---|---|---|---|
| Qwen3 30B-A3B | MoE | 0.5392 | 35 | 1,849 | — |
| Qwen3.5 27B | Dense | 0.5933 | 8 | 3,797 | +5.41 pp |
| Qwen3.5 35B-A3B | MoE | 0.5856 | 11 | 4,263 | +4.64 pp |
| Qwen3.6 27B | Dense | 0.5850 | 12 | 1,994 | +4.58 pp |
| Qwen3.6 35B-A3B | MoE | 0.5932 | 9 | 2,399 | +5.40 pp |
| Qwen3.8 27B low | Dense | 0.5781 | 18 | 809 | +3.89 pp |
| Qwen3.8 27B medium | Dense | 0.5811 | 15 | 1,034 | +4.19 pp |
| Qwen3.8 27B xhigh | Dense | 0.5847 | 13 | 3,288 | +4.55 pp |
Now for the curious part. Agentic training appears to have taken its toll on the dense 27B series with medical performance peaking with Qwen3.5. The 27B winrate drops from 0.5933 for Qwen3.5 to 0.5850 for Qwen3.6, then sits at 0.5847 for Qwen3.8 xhigh. Head-to-head, Qwen3.6 loses to Qwen3.5 on 16 of 27 datasets. And while Qwen3.8 vs 3.6 is an overall tie, 3.8 loses more individual datasets than it wins.
The Qwen MoE models don’t follow this pattern. Qwen3.6 35B-A3B’s win rate increases from 0.5856 to 0.5932 while its mean output length falls from 4,263 to 2,399 tokens, a 44% reduction. Head-to-head, it records a narrow 0.507 win rate against Qwen3.5 35B-A3B, improving on 11 of the 27 Medmarks-Verified datasets. Overall, Qwen3.6 offers a better quality/efficiency trade-off, even if its improvement is not consistent across tasks.
This result might show a benefit of mixture of expert models over dense models during agentic post-training: perhaps the agentic coding updates ignored more of the medical experts and improved overall reasoning, while the same post-training on the Qwen 27B series caused a mild case of catastrophic forgetting on medical tasks.
Last, Qwen3.8 27B still can't catch Sonnet. At xhigh reasoning level, its direct win rate against Sonnet 4.5 is 0.474. Far short of an Opus class model.
Qwen3.8 introduces multiple reasoning levels to the open-weight Qwen models in this weight class, matching behavior previously limited to gpt-oss 20b.
Medium to xhigh raises Qwen3.8's score from 0.5811 to 0.5847. That's 0.36 pp. Mean output length goes from 1,034 to 3,288 tokens, or 3.18x.
While the median token usage barely moves, from 786 to 801, the mean explodes because xhigh will wander off on long, unsuccessful reasoning detours. Incorrect Qwen3.8 xhigh answers average 7,254 output tokens while correct answers average 2,215, replicating one of our original Medmarks findings.
Muse Glimmer is Meta’s first open-source model in this weight class since Llama 3, and for a first release in over two years it’s impressive. Especially since Glimmer offers four efficient reasoning levels out of the box with linear performance scaling before flattening at xhigh.
| Effort | Winrate | Rank | Mean tokens |
|---|---|---|---|
| Low | 0.5575 | 27 | 309 |
| Medium | 0.5659 | 25 | 625 |
| High | 0.5743 | 20 | 1,249 |
| Xhigh | 0.5759 | 19 | 1,501 |
Muse Glimmer’s edge is efficiency, not raw score. While Qwen3.8 wins at all matched effort levels, Muse uses 62% fewer tokens at low reasoning, 40% fewer at medium, and 54% fewer at xhigh. At xhigh, Muse trails Qwen by only 0.88 percentage points while using less than half the tokens.
One interesting Muse Glimmer result from MedBullets: while adding the irrelevant fifth answer option drops Glimmer’s score from 0.785 to 0.737 at low reasoning, 0.817 to 0.768 at medium, and from 0.860 to 0.810 at high, at xhigh, the score falls only from 0.844 to 0.825.
Nvidia has been making steady progress improving Nemotron across generations, with Nemotron 3.5 Lightning narrowly beating gpt-oss 20b high using a similar number of tokens (head-to-head winrate of 0.511; higher on 18/27 datasets), but not matching Qwen3 30B-A3B Thinking while using double the tokens (head-to-head winrate 0.487; higher on 11/27 datasets).
At the dataset level, Lightning is one of the stranger models we’ve tested. It ranks 12th on Med-HALT NOTA and 13th on SCTPublic, both of which test uncertainty and resistance to bad answers. It also scores 0.865 on Med-HALT’s false-confidence test, up from just 0.229 for Nano V3. But it scores terribly on MedHallu, ranking 81st, 83rd, and 84th across the easy, medium, and hard splits. Whatever improved between Nemotron Nano V3 and 3.5 Lightning, it did not produce good medical hallucination-detection capabilities.
We're finishing up evals on the recent frontier and near-frontier models, including planned evaluations of Claude Fable 5.1 and GPT 6 Astra. Look for the full update in the coming weeks.