the model
Qwen/Qwen3-8B
Served by Native Inference · vLLM. Point a task at it and the audit prices the ceiling before a token is spent.
Vendored from the released binary · supernovae-st/nika@5d61b7e13f1f (v0.107.0) · digest-verified, re-derived at build, gated in CI
Point a task at it
tasks:
ask:
infer:
model: "native/qwen"
prompt: "…"The seats · 2
- native/qwen128k context · 8k out
Native Inference
- vllm/qwen128k context · 8k out
vLLM · the provider room
The price
No exact-match row at this pin · the engine resolves pattern rules at run time and nika audit prints the real ceiling. A local seat is unpriced, never free.