API reference

API reference

Model ids, limits, headers and usage

What every endpoint shares: which model ids are accepted, what happens to large output limits, the rate-limit headers on each answer, the usage fields, per-key spend limits and the errors to expect.

Base URL
https://api.vani.ai/v1
Output limit
Per model: top_provider.max_completion_tokens (higher values are lowered)
Rate-limit headers
x-ratelimit-{limit,remaining,reset}-{requests,tokens}
Model list
GET https://api.vani.ai/v1/models
On this pageModel ids

Model ids

Vani model ids include the vendor, for example anthropic/claude-haiku-4.5, openai/gpt-6-sol or zhipu/glm-5.3. GET /models lists every id your key can use, with context_length, top_provider.max_completion_tokens, created and, for prepaid keys, pricing in USD per token.

  • A vendor's own id is accepted when it names exactly one model your key can use: claude-haiku-4-5, claude-opus-4-8-20260115, gpt-5.5-2026-02-01, models/gemini-3.8-flash, Bedrock's anthropic.claude-haiku-4-5-20251001-v1:0 and Vertex's claude-haiku-4-5@20251001 all resolve. Case, dots and underscores, a date or -latest suffix and a -preview suffix do not matter.
  • The request is served, billed, logged and answered as the Vani model; the response carries x-vani-model-routed-from: <the id you sent> and its model field is the Vani id.
  • An id is never resolved to a model your key cannot use (a plan that does not include it, or a model without a published API price). That request gets the usual 404 model_not_available for the id you sent.
  • Retired ids keep working the same way (for example zhipu/glm-5-turbo is served by zhipu/glm-5.3-flash).

Output limits

Each model has its own output limit (reasoning included): top_provider.max_completion_tokens in GET /models is the largest value a request for that model can get.

  • A larger max_tokens / max_completion_tokens (Chat Completions), max_output_tokens (Responses) or max_tokens (Messages) is lowered to the limit instead of refused. The answer is otherwise unchanged and carries x-vani-max-output-clamped: <the limit used>.
  • 0, a negative number or a non-integer is still refused with 400.
  • Messages: when the lowered max_tokens leaves no room for your thinking.budget_tokens (at least 1,024 below it), the request is refused with 400; lower the budget or disable thinking.
  • With no limit in the request, Chat Completions and Responses use 4,096 output tokens. Messages requires max_tokens.

Rate limits and headers

Prepaid accounts get requests-per-second, requests-per-minute, tokens-per-minute and concurrent-request limits from their tier (raised automatically by lifetime top-ups; see Limits in the console). Every answer to an API key carries OpenAI's rate-limit headers:

HeaderMeaning
x-ratelimit-limit-requestsRequests per minute for this key's account.
x-ratelimit-remaining-requestsRequests left in the current minute window.
x-ratelimit-reset-requestsTime until the window is full again, e.g. 850ms, 1.5s, 6m0s.
x-ratelimit-limit-tokensTokens per minute (prepaid keys).
x-ratelimit-remaining-tokensTokens left: the limit minus the last 60 seconds of usage and reservations in flight.
x-ratelimit-reset-tokensTime until the oldest of those leaves the window.

Note

Plan-billed keys (created in Vani Chat) report the per-user request window; their output is bounded by the plan's allowance, so no token headers are sent. A 429 also carries Retry-After in seconds.

Usage fields

  • Cached input: usage.prompt_tokens_details.cached_tokens (Chat Completions), usage.input_tokens_details.cached_tokens (Responses), usage.cache_read_input_tokens (Messages): the prompt tokens billed at the cached-input price.
  • Reasoning: usage.completion_tokens_details.reasoning_tokens (Chat Completions), usage.output_tokens_details.reasoning_tokens (Responses) and usage.output_tokens_details.thinking_tokens (Messages) say how much of the output was reasoning. When the provider reports no count but billed output the answer does not show (Claude 5.x models think adaptively and the provider returns neither the thinking nor its size), the figure is Vani's estimate: the billed output minus the visible answer. Billing always uses the provider's total.
  • The console's usage page and CSV exports show cached input per day, model and key; the request log shows cached and reasoning tokens per request.

Reasoning effort and thinking

  • Chat Completions and Responses: reasoning_effort / reasoning.effort (none and minimal mean low; xhigh means extra_high). Models without an effort setting ignore it.
  • Messages: output_config.effort sets the effort. Without it, thinking: {"type": "disabled"} asks for the lowest effort, {"type": "enabled", "budget_tokens": N} for the effort of that budget, and adaptive keeps the model's default. Thinking blocks are not returned.
  • Claude Sonnet 5.5, Opus 5.5 and Fable 5.1 cannot turn thinking off. Their current route also does not apply effort or thinking settings yet, so they decide how much to think; the reasoning fields above show what it cost.
  • A small output limit can be spent on reasoning alone: the answer then ends with finish_reason: "length" / stop_reason: "max_tokens" and no text, and the reasoning count says why.
  • A model's safety refusal is an answer, not an error: finish_reason: "content_filter" (Chat Completions) or stop_reason: "refusal" (Messages), with the provider's short message as the text.

Keys and spend limits

  • Create, rename and revoke keys under API keys in the console (POST /v1/keys, PATCH /v1/keys/{id}, DELETE /v1/keys/{id} with your signed-in session; API keys cannot manage keys).
  • A prepaid key can have a monthly spend limit ($1 to $100,000 per UTC calendar month). Settled charges and requests in flight count. A request the rest of the limit cannot cover is refused with 402 api_key_spend_limit_reached (details: limit_micro, spent_micro, resets_at) until the next month or until you raise the limit.
  • Every rename and limit change, and the first refusal each month, is recorded in the key's history (GET /v1/keys/{id}/events).
Shell

Set a $25 monthly limit (signed-in session)

curl -X PATCH "https://api.vani.ai/v1/keys/$KEY_ID" \
  -H "Authorization: Bearer $SESSION_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"monthly_spend_limit_micro":"25000000"}'

# Remove it again with {"monthly_spend_limit_micro":null}.

Errors to expect

Status · codeWhat to do
400 bad_requestA field is unsupported or out of range; the message names it. Nothing was billed.
401 invalid_tokenThe key is wrong, revoked or expired.
402 api_insufficient_balanceAdd funds or turn on automatic top-up.
402 api_key_spend_limit_reachedThis key reached its monthly limit; raise it or use another key.
404 model_not_availableUse an id from GET /models (vendor ids are accepted when they name one model).
429 rate_limited / api_concurrency_limitWait for Retry-After, or for an active request to finish.
502 / stream error eventThe provider failed. A request that delivered nothing is not charged.

More reference: MCP tools on the API · Media API: images and videos · Artifacts API · Responses API: scope and limits · API overview

We use cookies for analytics and to measure how well our campaigns are working. None of it is needed to run the site, and your conversations are never included. See our Privacy Policy.