Models and pricing
One model today, more as the fleet grows. Prices in USD per million tokens, read live from the gateway document.
One model today, more as the fleet grows. Prices in USD per million tokens, read live from the gateway document.
| Model id | qwen/qwen3.8-27b |
|---|---|
| Endpoint | https://api.auxilio.ai/v1, OpenAI-compatible |
| Context | 128k tokens, prompt up to 124k |
| Max output | 32k tokens |
| Weights | FP8 on vLLM, own hardware |
| Input | $0.30 per million tokens |
| Cached input | $0.08 per million tokens |
| Output | $1.80 per million tokens, reasoning tokens billed as output |
| Capacity | 19.2M input and 800k output tokens per minute, 512 concurrent streams |
| Region | EU, datacenters in RO, own hardware |
| Data retention | None for prompts and completions. Policy |
| Status | Serving |
Over the per-model capacity the gateway answers 429 at once; back off and retry. One request up to 32 MB.
Per token used. Terms are agreed when your key is issued.
Every number on this page comes from api.auxilio.ai/openrouter/models, the same document the gateway publishes to marketplaces.
Need the whole card? The same fleet is available as virtual machines with one to eight GPUs. GPU Cloud