modular.

Serve a model (it’s harder than you think)

01

Model

Selected
Qwen3 8B
Hugging FaceQwen/Qwen3-8B
Kind of model
Hugging Face profile

Public models only.

Model reference
Source
02

Engines

Weight format / quantization
03

Configuration

Context window
GPU
Memory per GPU
GPUs

Scheduling

Tokens per batch
Sequences per step
Chunked prefillSplit long prompts across scheduler steps
Eager executionUseful when debugging; may reduce throughput
Memory and precision Cache, offload, dtype
CPU weight offload per GPU

Moves some weights through system RAM each step; speed depends on the CPU–GPU link.

KV cache allocation

A fixed cache allocation replaces the GPU memory target below.

Compute dtype
KV cache dtype

FP8 cache needs hardware and checkpoint support; verify quality for your workload.

Request behavior Priority and logs
Priority schedulingServe requests by priority instead of arrival order
Request loggingLog request information at the server's log level
Advanced settings Network, memory, and flags
Host
90%
50%99%
Optional
Trust model codeAllow model repository code where the server supports it
Prefix cacheReuse computation for shared prompt prefixes

Deploy a full server (one command at a time)

01

Services

    Nothing stacked yet.

    Add the current configurationChange it
    02

    Startup

    Check the machineConfirm the GPU driver, the GPU count, and that containers can see the GPUs
    Wait until readyHold until every service answers, and stop with its log if one fails to start
    03

    Gateway

    HTTPS gatewayOne address for every service, behind an API key

    Themes

    Tracker

    Road to a finished release.