OpenAI-compatible /v1
Drop-in for any client that already talks to OpenAI.
vLLM Warden is the operator console for serving large language models from a single GPU host. It manages model lifecycle, exposes an OpenAI-compatible API at /v1, schedules across multiple GPUs, and gives you per-consumer API tokens with rate limits — all from one self-hosted appliance, no external control plane.
Drop-in for any client that already talks to OpenAI.
Let vLLM Warden pick the right GPU for each model.
Per-consumer keys, TPS caps, usage telemetry.
One Docker compose stack, no external control plane.