Models and pricing

One model today, more as the fleet grows. Prices in USD per million tokens, read live from the gateway document.

Qwen3.8 27B

FP8 weights on vLLM, on our own GPU hardware in Romania.

Model id qwen/qwen3.8-27b
Endpoint https://api.auxilio.ai/v1, OpenAI-compatible
Context 128k tokens, prompt up to 124k
Max output 32k tokens
Weights FP8 on vLLM, own hardware
Input $0.30 per million tokens
Cached input $0.08 per million tokens
Output $1.80 per million tokens, reasoning tokens billed as output
Capacity 19.2M input and 800k output tokens per minute, 512 concurrent streams
Region EU, datacenters in RO, own hardware
Data retention None for prompts and completions. Policy
Status Serving

Features

  • Streaming, with usage in every chunk
  • Tool calls with automatic tool choice, JSON schema and structured outputs
  • Thinking on by default, switched off per request (see the docs)
  • Sampling controls: temperature, top_p, top_k, penalties, seed, stop

Limits

Over the per-model capacity the gateway answers 429 at once; back off and retry. One request up to 32 MB.

Billing

Per token used. Terms are agreed when your key is issued.

Request an API key

Live document

Every number on this page comes from api.auxilio.ai/openrouter/models, the same document the gateway publishes to marketplaces.

GPU servers

Need the whole card? The same fleet is available as virtual machines with one to eight GPUs. GPU Cloud