Real numbers from actually running each model on one NVIDIA DGX Spark (128GB unified memory) — not vendor benchmarks. Every row below sent real requests to a real server and measured what happened.
Models benchmarked
15
on one GB10 box
🏆 No.1 overall quality (LocalScore)
nvidia/Qwen3.6-35B-A3B-NVFP4 (reasoning: OFF)
🤖 No.1 for a personal-agent harness (Hermes Score)
nvidia/Qwen3.6-27B-NVFP4 (reasoning: OFF)
Which one should you use? For the most broadly accurate model on general tasks (coding, instructions, safety, long documents), go with nvidia/Qwen3.6-35B-A3B-NVFP4 (reasoning: OFF). For a personal-agent harness specifically — tool calls, memory across turns, serving many concurrent sessions — nvidia/Qwen3.6-27B-NVFP4 (reasoning: OFF) is the stronger pick; see the Hermes Benchmark section below for why.
⚠️ Quality and Responsiveness are reused column names with genuinely different definitions in the Overall table vs. the Hermes table — a model can score very differently on each. Hover the ⓘ on any column header below for a quick reminder.
Different from the No.1 picks above on purpose — each card here picks a winner for one specific thing you might care about, straight from the same raw measurements, instead of one composite number trying to represent everything at once. A model can win a category here and rank modestly elsewhere, or vice versa.
Fastest (with real answers, not just fast garbage)
RedHatAI/Muse-Glimmer-30B-NVFP4
210 tok/s · 30B
Highest peak decode speed among models that still cleared a basic quality bar (Quality ≥ 50/100) -- the pick when speed is what you care about most, as long as it's not getting things wrong to get there.
Most accurate (that isn't painfully slow)
unsloth/Qwen3.6-35B-A3B-NVFP4 (reasoning: ON)
87/100 quality · 35B
Highest Quality score among models still doing at least 20 tokens/sec -- the pick when correctness matters most and you just need it to not crawl.
Best at coding
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 (reasoning: ON)
100% coding accuracy · 30B
Highest code-generation accuracy (exec-verified against test cases), speed not considered at all -- the pick for a coding assistant specifically.
Best at tool calling
unsloth/Qwen3.6-35B-A3B-NVFP4 (reasoning: ON)
100/100 tool-use quality · 35B
Highest tool-use accuracy from the graded eval suite -- the pick for anything agentic that lives or dies on correctly calling functions.
Best for a Hermes-style personal agent
nvidia/Qwen3.6-27B-NVFP4 (reasoning: OFF)
82/100 Hermes Score · 27B
Highest Hermes Score -- see the dedicated section below for the full breakdown.
Best at agentic/autonomous planning
nvidia/Gemma-4-31B-IT-NVFP4
55/100 planning quality · 31B
Highest score on the eval suite's autonomous_planning domain -- not told the exact steps, just a goal and a toolbox; the pick for a harness that needs the model to figure out its own plan, not just follow one.
Most concurrent agent sessions
unsloth/Qwen3.6-35B-A3B-NVFP4 (reasoning: ON)
32 concurrent · 35B
Highest orchestrator capacity ceiling -- the pick when you need to share one box across the most simultaneous tool-chain agent sessions, not just serve one well.
Best for long documents
unsloth/Qwen3.6-35B-A3B-NVFP4 (reasoning: ON)
100/100 recall · 35B
Highest score on the eval suite's long_context (needle-in-haystack) domain -- the pick for summarizing or answering questions about long documents, not just short chats.
Best for structured/JSON output
unsloth/Qwen3.6-35B-A3B-NVFP4 (reasoning: ON)
100/100 structured quality · 35B
Highest score on the eval suite's structured domain -- the pick for a harness that parses the model's output programmatically and can't tolerate malformed JSON.
A faster-to-scan companion to the detail tables below, not a replacement — each panel ranks every model with data for that one metric and shows the top 5, bar length and the number at the tip both carrying the same value. Hover or tab through a bar for the full model name if it's cut off.
Overall quality
LocalScore, 0-100
Agent-harness fitness
Hermes Score, 0-100
Peak decode speed
best-case tok/s, single stream
Average decode speed
blended across mixed workloads
Prompt-processing speed
peak prefill tok/s
Coding accuracy
exec-verified, lowest concurrency tested
Tool-use quality
graded eval domain, 0-100
Concurrent tool-chain sessions
orchestrator-shaped traffic
General-purpose accuracy across 22 graded tasks (tool use, coding, safety, instruction-following, and more). LocalScore is one 0-100 number combining how often the model got the task right (Quality), how consistent it was across repeats (Reliability), how good its accuracy-per-token was (Efficiency), and how fast it started responding (Responsiveness). A different 0-100 number from Hermes Score below — see that section for what that one means. Avg tok/s here is a blended average across every decode-speed reading recorded for this model (short quick tasks and long-context runs together), so it can land well below the model's best case — that peak number is what Best model for → Fastest and the Decode peak column further down report instead.
| model | label | status | LocalScore | QualityMean graded score across the 22-task eval suite. A different task set from the Hermes table's "Quality" below. | ReliabilityHow consistent the score was across 3 repeats — 100 = zero variance. | EfficiencyAverage decode speed during the eval run, scored against an 80 tok/s target. | ResponsivenessMedian TOTAL reply time (first token through the full answer) vs. a 10s cap. Not the same measurement as the Hermes table's "Responsiveness." | avg tok/s | last run |
|---|---|---|---|---|---|---|---|---|---|
| nvidia/Qwen3.6-35B-A3B-NVFP4 | nvidia-qwen36-35b-a3b-nvfp4-boosted-nothink | complete | 92.0 | 86.7 | 100.0 | 97.0 | 93.1 | 316.9 | 2026-08-14 |
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | complete | 83.7 | 87.1 | 97.9 | 100.0 | 48.9 | 747.4 | 2026-08-15 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | complete | 80.7 | 63.6 | 99.7 | 100.0 | 85.6 | 62.0 | 2026-08-16 |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | complete | 79.9 | 87.0 | 100.0 | 39.6 | 88.9 | 31.7 | 2026-08-14 |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | complete | 78.1 | 81.9 | 99.6 | 100.0 | 31.8 | 79.1 | 2026-08-16 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | complete | 77.1 | 63.4 | 99.7 | 100.0 | 67.9 | 190.1 | 2026-08-16 |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | complete | 76.8 | 87.7 | 97.0 | 100.0 | 14.0 | 528.4 | 2026-08-15 |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | complete | 71.1 | 77.5 | 100.0 | 24.8 | 81.6 | 19.9 | 2026-08-14 |
| google/gemma-4-12B-it | google-gemma-4-12b-it | complete | 69.8 | 86.7 | 99.8 | 11.3 | 67.5 | 8.0 | 2026-08-14 |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | complete | 69.1 | 86.7 | 99.8 | 10.2 | 65.1 | 7.9 | 2026-08-15 |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | complete | 68.7 | 87.7 | 99.6 | 14.5 | 56.9 | 13.9 | 2026-08-16 |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | complete | 66.9 | 78.7 | 97.3 | 15.9 | 68.3 | 21.7 | 2026-08-15 |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | complete | 63.0 | 57.9 | 96.8 | 21.3 | 90.9 | 11.6 | 2026-08-15 |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | complete | 55.5 | 42.9 | 98.5 | 42.1 | 64.8 | 22.8 | 2026-08-14 |
| RedHatAI/Muse-Glimmer-30B-NVFP4*No vLLM-compatible tool-call/reasoning parser exists for this model's custom "Onyx ATEM" chat template -- confirmed 2026-08-16, no matching vLLM plugin available. A harness-side workaround (spark_bench_plus.py's chat_stream) strips its raw reasoning text before grading, which should help plain-text domains, but its tool calls are also raw text that never reaches vLLM's tool_calls field, so tool-use-dependent numbers specifically may still understate this model's real capability. See METHODOLOGY.md. | muse-glimmer-30b-nvfp4 | complete | 54.2 | 63.9 | 100.0 | 15.3 | 37.2 | 22.2 | 2026-08-16 |
The peak number of simultaneous requests of that traffic type this box was
actually tested against — orchestrator is multi-step tool-chain traffic (the shape an
autonomous agent sends), coding_agent is code-generation requests, chat_agent
is casual back-and-forth conversation. All three are swept over the same concurrency levels (1, 2,
4, 8, 16, 32), so the numbers are directly comparable to each other.
| model | label | tool-chain agents | coding agents | chat sessions | tool-calling works? |
|---|---|---|---|---|---|
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | 32 | 32 | 32 | ✅ |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | 32 | 2 | 32 | ✅ |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | 32 | 32 | 32 | ✅ |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | 32 | 32 | 2 | ✅ |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | 32 | 32 | 16 | ✅ |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | 32 | 32 | 32 | ✅ |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | 16 | 32 | 32 | ✅ |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | 16 | 32 | 32 | ✅ |
| google/gemma-4-12B-it | google-gemma-4-12b-it | 16 | 32 | 32 | ✅ |
| nvidia/Qwen3.6-35B-A3B-NVFP4 | nvidia-qwen36-35b-a3b-nvfp4-boosted-nothink | 8 | 8 | — | ✅ |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | 1 | 32 | 16 | ❌ |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | 1 | 32 | 16 | ❌ |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | 1 | 32 | 16 | ❌ |
| RedHatAI/Muse-Glimmer-30B-NVFP4 | muse-glimmer-30b-nvfp4 | 1 | 32 | 16 | ❌ |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | 1 | 32 | 16 | ❌ |
One 0-100 number for "how good would this model be behind a Hermes-style personal agent": 50% how well it actually completes real agent tasks (tool use, web search, remembering things across a conversation), 30% how many concurrent sessions it sustains, 20% how quickly it starts responding.
| model | label | Hermes Score | QualityMean score on a separate 7-task hermes-shaped task set (tool chains, web search, memory). Not the 22-task eval suite in the Overall table. | CapacityConcurrent hermes-shaped sessions handled vs. a 16-session target. Can mean "tested this far without breaking," not always a proven ceiling. | ResponsivenessBaseline time-to-first-token only, single user, vs. a 1.5s target. Not the same measurement as the Overall table's "Responsiveness." | approx size |
|---|---|---|---|---|---|---|
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | 82.1 | 64.1 | 100.0 | 100.0 | 27B |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | 80.6 | 61.1 | 100.0 | 100.0 | 31B |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | 79.8 | 59.6 | 100.0 | 100.0 | 27B |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | 79.0 | 58.1 | 100.0 | 100.0 | 26B |
| google/gemma-4-12B-it | google-gemma-4-12b-it | 79.0 | 58.1 | 100.0 | 100.0 | 12B |
| nvidia/Qwen3.6-35B-A3B-NVFP4 | nvidia-qwen36-35b-a3b-nvfp4-boosted-nothink | 69.3 | 68.7 | 50.0 | 100.0 | 35B |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | 62.9 | 59.1 | 100.0 | 16.7 | 27B |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | 54.3 | 53.5 | 25.0 | 100.0 | 4B |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | 50.5 | 46.1 | 25.0 | 100.0 | — |
| RedHatAI/Muse-Glimmer-30B-NVFP4 | muse-glimmer-30b-nvfp4 | 50.2 | 45.5 | 25.0 | 100.0 | 30B |
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | 49.1 | 58.1 | 50.0 | 25.5 | 35B |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | 46.6 | 38.2 | 25.0 | 100.0 | 3B |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | 46.5 | 45.5 | 12.5 | 100.0 | 30B |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | 37.2 | 45.5 | 25.0 | 35.0 | 30B |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | 37.0 | 64.1 | 12.5 | 5.8 | 27B |
Smallest model that's actually good enough (70+/100): google/gemma-4-12B-it (~12B, score 79.0) — the one to reach for if you're tight on memory. Largest model tested: nvidia/Gemma-4-31B-IT-NVFP4 (~31B, score 80.6) — the highest-quality option this box can run.
Same-model decode speed at its smallest vs. largest tested context, both
from the same speed sweep run. Stability is
100 × (tok/s at largest context) / (tok/s at smallest context) — 100% means no slowdown at all as
the prompt grows; a low number means the model falls off hard on long documents even though it
might look fine on a short prompt. Peak is the single fastest decode_tps ever recorded for
this model at any context size, no quality filter. reasoning? is derived from whether any
measured reasoning time was ever recorded for this label, not a name guess — note that a
reasoning-capable model's thinking-ON and thinking-OFF entries can have very different stability
curves, so check the label, not just the model name.
| model | label | reasoning? | @ smallest ctx | @ largest ctx | stability | peak (any context) |
|---|---|---|---|---|---|---|
| RedHatAI/Muse-Glimmer-30B-NVFP4 | muse-glimmer-30b-nvfp4 | no | 12 tok/s @3K tokens | 12 tok/s @60K tokens | 96% | 210 |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | no | 17 tok/s @3K tokens | 16 tok/s @15K tokens | 96% | 17 |
| google/gemma-4-12B-it | google-gemma-4-12b-it | no | 7 tok/s @4K tokens | 7 tok/s @129K tokens | 91% | 8 |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | no | 7 tok/s @4K tokens | 6 tok/s @64K tokens | 91% | 7 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | 🧠 yes | 78 tok/s @4K tokens | 70 tok/s @133K tokens | 90% | 87 |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | no | 33 tok/s @4K tokens | 29 tok/s @66K tokens | 87% | 33 |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | no | 19 tok/s @4K tokens | 16 tok/s @64K tokens | 86% | 19 |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | no | 29 tok/s @4K tokens | 25 tok/s @129K tokens | 84% | 30 |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | 🧠 yes | 11 tok/s @3K tokens | 9 tok/s @124K tokens | 81% | 11 |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | 🧠 yes | 12 tok/s @3K tokens | 9 tok/s @124K tokens | 80% | 12 |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | 🧠 yes | 10 tok/s @3K tokens | 2 tok/s @124K tokens | 22% | 10 |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | 🧠 yes | 10 tok/s @3K tokens | 2 tok/s @124K tokens | 20% | 10 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | 🧠 yes | 64 tok/s @4K tokens | 12 tok/s @133K tokens | 19% | 85 |
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | 🧠 yes | 58 tok/s @3K tokens | 8 tok/s @124K tokens | 13% | 58 |
Reasoning models average stability: 46% (n=7) vs. non-reasoning: 90% (n=7). Most stable at long context: RedHatAI/Muse-Glimmer-30B-NVFP4 (muse-glimmer-30b-nvfp4) (96%). Fastest single-context peak (no quality filter): RedHatAI/Muse-Glimmer-30B-NVFP4 (muse-glimmer-30b-nvfp4) (210 tok/s).
Best-case single-stream timing at the smallest context tested, the
largest context that actually returned a real answer, and the GPU memory budget each model was
served at. TTFT = time to first token, TPOT = time per output token (decode latency), prefill =
how fast it reads the prompt before answering. Max document size is the actual measured
prompt token count of the largest successful run, not the round number passed to
--contexts — different tokenizers turn the same target size into different real
token counts, so this can land a bit above or below what was requested.
| model | label | GPU util used | Load time | Max document size | TTFT (ms) | Prefill (tok/s) | Decode peak (tok/s) | TPOT (ms) |
|---|---|---|---|---|---|---|---|---|
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | 0.90 | — | 124K tokens | 21692 | 80109 | 12 | 85 |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | 0.90 | 735s | 64K tokens | 2601 | 1593 | 7 | 148 |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | 0.90 | — | 124K tokens | 23415 | 95030 | 11 | 92 |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | 0.90 | 292s | 129K tokens | 736 | 5876 | 30 | 34 |
| google/gemma-4-12B-it | google-gemma-4-12b-it | 0.90 | 292s | 129K tokens | 1704 | 2864 | 8 | 134 |
| nvidia/Qwen3.6-35B-A3B-NVFP4 | nvidia-qwen36-35b-a3b-nvfp4-boosted-nothink | 0.90 | 121s | — | — | — | — | — |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | 0.90 | 364s | 124K tokens | 24700 | 2381 | 10 | 97 |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | 0.90 | 251s | 64K tokens | 633 | 6954 | 19 | 54 |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | 0.90 | 925s | 30K tokens | 1288 | 4529 | 17 | 59 |
| RedHatAI/Muse-Glimmer-30B-NVFP4 | muse-glimmer-30b-nvfp4 | 0.90 | 282s | 120K tokens | 20536 | 3954 | 210 | 5 |
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | 0.90 | 221s | 124K tokens | 4436 | 8193 | 58 | 17 |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | 0.90 | 181s | 66K tokens | 403 | 15618 | 33 | 30 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | 0.90 | 231s | 133K tokens | 622 | 12755 | 87 | 17 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.90 | 201s | 133K tokens | 2986 | 12066 | 85 | 12 |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | 0.90 | 131s | 124K tokens | 24645 | 1935 | 10 | 97 |
Real GPU behavior sampled every 5s for the duration of each model's benchmark run (not the box's own hermes-vllm production traffic) — how hot, how loaded, and how much memory it actually used, versus the --gpu-memory-utilization flag we asked vLLM to target. Only present for runs after telemetry capture was added; older runs show —.
| model | label | Peak GPU util | Avg GPU util | Peak memory used | Peak temp | Peak power |
|---|---|---|---|---|---|---|
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4-nothink | 96% | 96% | — | 71°C | 69W |
| nvidia/Gemma-4-31B-IT-NVFP4 | nvidia-gemma-4-31b-it-nvfp4 | 96% | 96% | — | 72°C | 76W |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4-nothink | 96% | 96% | — | 74°C | 69W |
| nvidia/Gemma-4-26B-A4B-NVFP4 | nvidia-gemma-4-26b-a4b-nvfp4 | 96% | 96% | — | 77°C | 73W |
| google/gemma-4-12B-it | google-gemma-4-12b-it | 96% | 96% | — | 78°C | 87W |
| nvidia/Qwen3.6-35B-A3B-NVFP4 | nvidia-qwen36-35b-a3b-nvfp4-boosted-nothink | 96% | 91% | — | 61°C | 37W |
| unsloth/Qwen3.8-27B-NVFP4 | unsloth-qwen38-27b-nvfp4 | 96% | 96% | — | 81°C | 88W |
| google/gemma-4-E4B-it | google-gemma-4-e4b-it | 96% | 96% | — | 72°C | 74W |
| poolside/Laguna-S-2.1-NVFP4 | poolside-laguna-s-21-nvfp4 | 96% | 96% | — | 76°C | 78W |
| RedHatAI/Muse-Glimmer-30B-NVFP4 | muse-glimmer-30b-nvfp4 | 96% | 96% | — | 77°C | 80W |
| unsloth/Qwen3.6-35B-A3B-NVFP4 | unsloth-qwen36-35b-a3b-nvfp4 | 96% | 96% | — | 76°C | 80W |
| LiquidAI/LFM2.5-2.6B | liquidai-lfm25-26b | 96% | 96% | — | 79°C | 92W |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 | nvidia-nvidia-nemotron-3-nano-30b-a3b-nvfp4 | 96% | 96% | — | 82°C | 83W |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nemotron-3.5-lightning-30b-a3b-nvfp4 | 96% | 96% | — | 80°C | 90W |
| nvidia/Qwen3.6-27B-NVFP4 | nvidia-qwen36-27b-nvfp4 | 96% | 96% | — | 82°C | 89W |