Skip to content
Blog

Qwen3.8-27B is now live, also the winning recipes for 5090s are now public

· Mert Göksel

Qwen3.8-27B is now live on Astorias. The production deployment serves the FP8 weights, not the NVFP4 build we benchmarked here. Every number in this post comes from the NVFP4 runs on RTX 5090s.

Pricing for this model is now $0.10 Input / $1.60 Output / $0.05 Cache (per million for each).

Qwen3.8-27B is a strong open model, and it is small enough to run on a gaming card. That makes it tempting to self-host. Before we committed to it, we wanted to know how much serving one RTX 5090 gives you, and what changes when you add a second.

We spent a week benchmarking it, and this post covers what we found. We published the best configurations in our open repository at github.com/Astorias-AI/qwen-3.8-27B-public.

The setup

Everything in these benchmarks is open source under Apache-2.0.

The model is Qwen3.8-27B in NVFP4. NVFP4 is a 4-bit number format that Blackwell GPUs run natively. It shrinks the weights enough to fit on a single 32 GB card.

The draft model is DFlash2, which we use for speculative decoding. A small model guesses the next several tokens, and the large model checks all the guesses in one pass. When the guesses are right, you get several tokens for about the cost of one.

The server is SGLang with HiCache turned on. HiCache keeps the attention cache for prompts the server has already seen, and moves it to system RAM when GPU memory fills up. Requests that share a long prefix, like an agent's system prompt, skip most of their prefill work.

We ran every benchmark on the GPU host itself, so network latency does not show up in the numbers.

One RTX 5090

A 5090 has 32 GB of memory. After the model and the draft model load, only a few gigabytes remain. That space has to hold the attention cache for every active conversation, the per-request state for the model's linear-attention layers, and the precompiled GPU graphs that SGLang uses for decoding.

We hit that limit right away. We asked for a 65,536-token context pool, and the server allocated 28,192 tokens. Asking for more does not help, because the memory is not there.

The biggest gain came from one setting. Qwen3.8 is a hybrid model, and each active request needs a few slots of saved linear-attention state. With 8 slots, the server ran only two requests at a time. With 12, it runs four. On a benchmark with eight clients at once, output rose from about 256 to 422 tokens per second on the same card and software.

Workload Output tokens/s Time to first token
One request, 8k tokens in / 1k out 122 775 ms
8 concurrent, 2k in / 256 out 422 2.7 s
8 concurrent, shared prompt prefix 356 479 ms
Multi-turn chat, 4 concurrent 402 212 ms

Speculative decoding averaged about three accepted tokens per step. On the shared-prefix test, 83% of each prompt came from the cache, and the first token arrived in 479 ms.

One card is enough for a small team, a single agent pipeline, or a development endpoint, as long as prompts fit in about 28k tokens of context.

Two RTX 5090s

With two cards, SGLang uses tensor parallelism to split the model across both GPUs, and each card holds half the weights. With that memory freed, the server allocated the model's full 262,144-token context window.

The extra memory also let us raise almost every capacity setting. The server now runs eight concurrent requests instead of four, keeps more saved-state slots, drafts six tokens per step instead of four, and precompiles larger GPU graphs. We tested each change on its own, then checked that the best values worked together before we adopted them.

Workload Output tokens/s Time to first token
One request, 8k tokens in / 1k out 126 2.4 s
8 concurrent, 2k in / 1k out 537 2.8 s
8 concurrent, shared prompt prefix 747 1.2 s
Multi-turn chat, 4 concurrent 425 340 ms

The best result is 747 tokens per second on shared-prefix traffic, which is the pattern agents and chat products produce. Speculative decoding also accepted more tokens here, 3.7 to 3.9 per step, because the draft model proposes six tokens instead of four.

The two tables use different output lengths. The two-card load tests generate 1,024 output tokens per request, and the one-card load tests generate 256. Read each table as its own measurement, and don't divide one by the other.

Two cards make sense for long documents, coding agents with large contexts, or more concurrent users on one endpoint.

What is still slow

We tuned the serving settings. The GPU kernels for this card still need work.

The 5090 uses the SM120 variant of Blackwell, and the data-center parts use SM100. The fastest prefill kernels for Qwen3.8's linear-attention layers only target SM100, so the 5090 runs a slower general-purpose path. Long prefill chunks also delay token generation for other active requests. We cap prefill chunks at 2,048 tokens to keep that delay short.

A prefill kernel written for the 5090 is where we plan to look next.

The configurations

The best one-card and two-card configurations are open source at github.com/Astorias-AI/qwen-3.8-27B-public. If you run them on your own hardware, send us your numbers.