API reference
Model ids, limits, headers and usage
What every endpoint shares: which model ids are accepted, what happens to large output limits, the rate-limit headers on each answer, the usage fields, per-key spend limits and the errors to expect.
- Base URL
- https://api.vani.ai/v1
- Output limit
- Per model: top_provider.max_completion_tokens (higher values are lowered)
- Rate-limit headers
- x-ratelimit-{limit,remaining,reset}-{requests,tokens}
- Model list
- GET https://api.vani.ai/v1/models
On this pageModel ids
Model ids
Vani model ids include the vendor, for example anthropic/claude-haiku-4.5, openai/gpt-6-sol or zhipu/glm-5.3. GET /models lists every id your key can use, with context_length, top_provider.max_completion_tokens, created and, for prepaid keys, pricing in USD per token.
- A vendor's own id is accepted when it names exactly one model your key can use:
claude-haiku-4-5,claude-opus-4-8-20260115,gpt-5.5-2026-02-01,models/gemini-3.8-flash, Bedrock'santhropic.claude-haiku-4-5-20251001-v1:0and Vertex'sclaude-haiku-4-5@20251001all resolve. Case, dots and underscores, a date or-latestsuffix and a-previewsuffix do not matter. - The request is served, billed, logged and answered as the Vani model; the response carries
x-vani-model-routed-from: <the id you sent>and itsmodelfield is the Vani id. - An id is never resolved to a model your key cannot use (a plan that does not include it, or a model without a published API price). That request gets the usual 404
model_not_availablefor the id you sent. - Retired ids keep working the same way (for example
zhipu/glm-5-turbois served byzhipu/glm-5.3-flash).
Output limits
Each model has its own output limit (reasoning included): top_provider.max_completion_tokens in GET /models is the largest value a request for that model can get.
- A larger
max_tokens/max_completion_tokens(Chat Completions),max_output_tokens(Responses) ormax_tokens(Messages) is lowered to the limit instead of refused. The answer is otherwise unchanged and carriesx-vani-max-output-clamped: <the limit used>. - 0, a negative number or a non-integer is still refused with 400.
- Messages: when the lowered
max_tokensleaves no room for yourthinking.budget_tokens(at least 1,024 below it), the request is refused with 400; lower the budget or disable thinking. - With no limit in the request, Chat Completions and Responses use 4,096 output tokens. Messages requires
max_tokens.
Rate limits and headers
Prepaid accounts get requests-per-second, requests-per-minute, tokens-per-minute and concurrent-request limits from their tier (raised automatically by lifetime top-ups; see Limits in the console). Every answer to an API key carries OpenAI's rate-limit headers:
| Header | Meaning |
|---|---|
| x-ratelimit-limit-requests | Requests per minute for this key's account. |
| x-ratelimit-remaining-requests | Requests left in the current minute window. |
| x-ratelimit-reset-requests | Time until the window is full again, e.g. 850ms, 1.5s, 6m0s. |
| x-ratelimit-limit-tokens | Tokens per minute (prepaid keys). |
| x-ratelimit-remaining-tokens | Tokens left: the limit minus the last 60 seconds of usage and reservations in flight. |
| x-ratelimit-reset-tokens | Time until the oldest of those leaves the window. |
Note
Retry-After in seconds.Usage fields
- Cached input:
usage.prompt_tokens_details.cached_tokens(Chat Completions),usage.input_tokens_details.cached_tokens(Responses),usage.cache_read_input_tokens(Messages): the prompt tokens billed at the cached-input price. - Reasoning:
usage.completion_tokens_details.reasoning_tokens(Chat Completions),usage.output_tokens_details.reasoning_tokens(Responses) andusage.output_tokens_details.thinking_tokens(Messages) say how much of the output was reasoning. When the provider reports no count but billed output the answer does not show (Claude 5.x models think adaptively and the provider returns neither the thinking nor its size), the figure is Vani's estimate: the billed output minus the visible answer. Billing always uses the provider's total. - The console's usage page and CSV exports show cached input per day, model and key; the request log shows cached and reasoning tokens per request.
Reasoning effort and thinking
- Chat Completions and Responses:
reasoning_effort/reasoning.effort(noneandminimalmeanlow;xhighmeansextra_high). Models without an effort setting ignore it. - Messages:
output_config.effortsets the effort. Without it,thinking: {"type": "disabled"}asks for the lowest effort,{"type": "enabled", "budget_tokens": N}for the effort of that budget, andadaptivekeeps the model's default. Thinking blocks are not returned. - Claude Sonnet 5.5, Opus 5.5 and Fable 5.1 cannot turn thinking off. Their current route also does not apply effort or thinking settings yet, so they decide how much to think; the reasoning fields above show what it cost.
- A small output limit can be spent on reasoning alone: the answer then ends with
finish_reason: "length"/stop_reason: "max_tokens"and no text, and the reasoning count says why. - A model's safety refusal is an answer, not an error:
finish_reason: "content_filter"(Chat Completions) orstop_reason: "refusal"(Messages), with the provider's short message as the text.
Keys and spend limits
- Create, rename and revoke keys under API keys in the console (
POST /v1/keys,PATCH /v1/keys/{id},DELETE /v1/keys/{id}with your signed-in session; API keys cannot manage keys). - A prepaid key can have a monthly spend limit ($1 to $100,000 per UTC calendar month). Settled charges and requests in flight count. A request the rest of the limit cannot cover is refused with 402
api_key_spend_limit_reached(details:limit_micro,spent_micro,resets_at) until the next month or until you raise the limit. - Every rename and limit change, and the first refusal each month, is recorded in the key's history (
GET /v1/keys/{id}/events).
Set a $25 monthly limit (signed-in session)
curl -X PATCH "https://api.vani.ai/v1/keys/$KEY_ID" \
-H "Authorization: Bearer $SESSION_TOKEN" \
-H "Content-Type: application/json" \
-d '{"monthly_spend_limit_micro":"25000000"}'
# Remove it again with {"monthly_spend_limit_micro":null}.Errors to expect
| Status · code | What to do |
|---|---|
| 400 bad_request | A field is unsupported or out of range; the message names it. Nothing was billed. |
| 401 invalid_token | The key is wrong, revoked or expired. |
| 402 api_insufficient_balance | Add funds or turn on automatic top-up. |
| 402 api_key_spend_limit_reached | This key reached its monthly limit; raise it or use another key. |
| 404 model_not_available | Use an id from GET /models (vendor ids are accepted when they name one model). |
| 429 rate_limited / api_concurrency_limit | Wait for Retry-After, or for an active request to finish. |
| 502 / stream error event | The provider failed. A request that delivered nothing is not charged. |
More reference: MCP tools on the API · Media API: images and videos · Artifacts API · Responses API: scope and limits · API overview