Skip to main content
Qwen3.8 27B (qwen-3.8-27b) is a 27-billion-parameter text LLM served on PolarGrid edge nodes via Triton’s vllm_backend. Weights ship pre-quantized to FP8 (~28 GB VRAM) and load directly on PolarGrid’s Blackwell edge GPUs without runtime requantization.

Headline benchmark

We publish two numbers side by side. End-to-end is what your application actually experiences (request → response, network included). Server-only is what the GPU spends on inference (apples-to-apples vs centralized providers’ published “inference-only” figures). The gap is the latency PolarGrid’s yvr-02 PoP eliminates by being at the edge. Bench: 60 streaming chat-completion runs (5 warmup, concurrency 1, max_tokens 48) against https://api.yvr-02.edge.polargrid.ai, captured 2026-08-19 from a macOS host over the public internet, paced at 1.2 req/s to stay under the per-key rate limit. Server-only is read from the gateway’s pg_metadata SSE event (inference_ttft_ms); end-to-end is client wall-clock. Reasoning mode off (default). CUDA graphs on (enforce_eager=false). Raw runs: benchmarks/fleet-2026-08-18-qwen-3.8/yvr-02-c1.json. Read the server-only row, not the end-to-end row, when comparing against the Qwen3.5 27B card. The two captures ran from different client locations, so their network legs differ (210 ms here vs 106 ms there) even though server-side latency is effectively identical. The end-to-end number tells you what that benchmark host saw, not what the model got slower at. Fleet-wide, not one node. All 13 production nodes were measured (780 runs, zero errors). The eleven 2-GPU nodes land at 87–91 ms server TTFT p50 with 30.8–31.4 tok/s decode — a 4 ms spread across eleven geographically separate machines. The two 4-GPU nodes (yto-01, yvr-02, which also carry telephony/livekit workloads) sit at 72 ms with 29.4–29.5 tok/s: lower TTFT, slightly lower decode. That is a machine-class difference, not a per-node regression.
Apples-to-apples disclaimer. Other providers usually publish only their server-side number; comparing it to our server-only row is the fair baseline. Our end-to-end row is what you’ll see from a customer-side request because PolarGrid runs at the edge. The network row above shows exactly how much that’s worth in milliseconds.

How this compares

PolarGrid wins on TTFT against frontier-reasoning providers because of edge proximity — the server does the work in tens of milliseconds and the client is close to it. PolarGrid remains 4 to 70 times behind specialty silicon on raw throughput; RTX 6000 Pro Blackwell workstation FLOPS are below H100 and H200 datacenter FLOPS. CUDA graphs are enabled (enforce_eager=false, the fleet default). Against the Qwen3.5 27B it replaced, the swap is latency-neutral at concurrency 1 and better under load. Measured on the staging canary 2026-08-17: 88 / 121 / 118 / 143 ms at c=1 / 8 / 16 / 32 for 3.8, versus 88 / 129 / 144 / 173 ms for 3.5. Same first-token latency for a single caller; the gap opens as concurrency rises.

Quickstart

Capabilities

Function calling

Pass OpenAI-shape tools and the model returns a tool_calls array on the assistant message (or as a delta.tool_calls chunk when streaming). The gateway speaks Qwen’s Hermes tool-call template under the hood.
tool_choice accepts "auto" (model decides), "none" (force plain text), "required" (force a tool call), or { "type": "function", "function": { "name": "<tool>" } } to force a specific tool. After invoking the tool yourself, append a role: "tool" message containing the result and re-call the model:

Structured output (JSON mode)

Use response_format to force the model to emit valid JSON. Backed server-side by vLLM’s structured_outputs constrained decoding, so the output is guaranteed to parse.
{"type": "json_object"} accepts any valid JSON; json_schema constrains it to your schema.

Reasoning mode

Qwen3.8 ships with a “thinking” mode that emits a <think>...</think> reasoning trace before the user-visible answer. PolarGrid’s /v1/chat/completions endpoint disables this by default to keep first-token latency low. The 27B variant runs the same toggle, with deeper reasoning quality at the cost of longer generation time. To enable thinking on a per-request basis:

Model identifier

Call this model with the canonical id qwen-3.8-27b at all inference endpoints (/v1/chat/completions, /v1/completions). The HuggingFace repo id Qwen/Qwen3.8-27B-FP8 is accepted at /v1/models/load for hot-loading purposes but does not resolve at inference time — use the canonical id for chat and completions calls.

Notes

  • License: Apache 2.0 (no auth required to pull weights).
  • Native FP8 — no runtime quantization step at load.
  • VRAM is tight: a single 46 GB L40S can host this model or the voice stack, not both. Multi-GPU edges pin 27B to its own GPU; see backend/edge-production-setup/CLAUDE.md for the layout matrix.

See also