vllm-bench results explorer

Our September 2026 benchmark of open-weight models served by vLLM, one protocol for every model.

Two modded RTX 4090D 48GB (Ada, 96GB total), 330 W per card, vLLM 0.26.0, 26 Sep 2026, three runs per figure (medians). Full write-up ยท The repo

Read in your browser with FileReader. Nothing is uploaded. Any results.csv written by bench.py works; extraction, coding and tool_eval columns are optional.

Decode speed

Single-stream decode, tok/s. Whiskers show the lowest and highest of the runs where we kept them; otherwise the run-to-run spread is printed.

Prompt

Speed against tool-calling quality

Short-prompt decode tok/s against the tool-eval pass rate at temperature 0.2. Models without a score are left out. Hover or tab to a point for its name.

One request against four at once

Output tok/s for one stream, and the total across four simultaneous short requests.

Tokens per watt

Short-prompt decode tok/s divided by median board power during decode. One-card rows count the working card; two-card rows add both.

Qwen3.6-27B quantisation ladder

Same model, same card, five configurations. Single-stream decode, short prompt, tok/s.

One card or two: Qwen3.8-27B AWQ

The same model on one card, split across two with tensor parallel, and pipelined across two. Output tok/s, with cold 16k time to first token.

All figures as a table