Serve a model (it’s harder than you think)
01
Model
Selected
Qwen3 8BHugging Face
Qwen/Qwen3-8BKind of model
Design
Checkpoint
Hugging Face profile
Public models only.
Model reference
Source
Ornith 1.5 serving profileUse its reasoning and tool parsers; check your variant’s chat template
02
Engines
Weight format / quantization
vLLM recipe
Variables: {{model}}, {{port}}, {{host}}, {{context}}, {{quant}}. Quote values in your template as needed.
03
Configuration
Context window
512262K
GPU
Memory per GPU
GPUs
Tensor parallel
Model behavior
Ornith 1.5Reasoning parserSeparate reasoning from the final answer
Tool-call parserExpose model tool calls in the API
Scheduling
Tokens per batch
Sequences per step
Chunked prefillSplit long prompts across scheduler steps
Eager executionUseful when debugging; may reduce throughput
Memory and precision Cache, offload, dtype
CPU weight offload per GPU
Moves some weights through system RAM each step; speed depends on the CPU–GPU link.
KV cache allocation
A fixed cache allocation replaces the GPU memory target below.
Compute dtype
KV cache dtype
FP8 cache needs hardware and checkpoint support; verify quality for your workload.
Request behavior Priority and logs
Priority schedulingServe requests by priority instead of arrival order
Request loggingLog request information at the server's log level
Advanced settings Network, memory, and flags
Pipeline stages
Host
50%99%
Optional
Trust model codeAllow model repository code where the server supports it
Prefix cacheReuse computation for shared prompt prefixes