Base URL
https://api.vani.ai/v1Include /v1 once. Select the API format documented in your client's guide.
Vani documentation
Connect your preferred client, build with the API, or get more from chat. Start with configuration examples and a clear view of what each integration supports.
Connect with three values: a base URL, your API key, and an enabled model identifier. Vani exposes bounded Chat Completions, stateless Responses and Anthropic Messages adapters at the endpoint below.
https://api.vani.ai/v1Include /v1 once. Select the API format documented in your client's guide.
Create a key in Developer settings after verifying your account. Keep it in your client's secret field or environment.
openai/gpt-6-solA starting example. GET /models returns the models available to your account.
Important
For pay-as-you-go access, open the API platformto manage prepaid keys, check your balance, and see API usage by model and key for today, 7, 30 or 90 days. Eligible API models cost 30% below their verified official rates; private models priced directly by Vani are labelled “Vani price”. Rates include source references and context tiers; unavailable rates are never treated as free.
export VANI_BASE_URL="https://api.vani.ai/v1"
export VANI_MODEL="openai/gpt-6-sol"
# Set VANI_API_KEY in your shell secret store or environment.
# Use the API origin shown here, ending in /v1 once.
# The website /chat and /api/chat routes are not this API base.
# Start by fetching /models; that check does not invoke inference.curl --fail-with-body "$VANI_BASE_URL/models" \
-H "Authorization: Bearer $VANI_API_KEY"
# Use an exact data[].id from this account-scoped response.Use an exact data[].id from the authenticated model list as VANI_MODEL. The examples start with openai/gpt-6-sol; change it when needed. Model access is account-specific; disabled inference can still prevent a request.
curl --fail-with-body "$VANI_BASE_URL/chat/completions" \
-H "Authorization: Bearer $VANI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-6-sol","messages":[{"role":"user","content":"Hello"}],"max_completion_tokens":256,"stream":false}'Note
A curated set of 15 popular clients and coding tools. Each guide separates supported API settings from features that need an adapter or extra validation.
15 of 15 guides · Popular coding and chat clients.
Configure a named custom provider using Hermes' documented Chat Completions transport.
Use the Codex CLI with Vani’s stateless Responses adapter, including local function and namespaced custom tools.
Add Vani as a custom provider with the OpenAI-compatible Chat Completions package.
Connect Claude Code to Vani’s Anthropic Messages adapter for text, streaming and client-executed tools.
Cursor's current official BYOK guide documents provider keys, but does not verify this custom Vani endpoint or GPT-6 Sol.
Use Cline's OpenAI Compatible provider and the exact Vani model identifier.
Configure a Chat-only OpenAI-compatible model and explicitly disable automatic Responses routing.
Point Aider's OpenAI-compatible transport at Vani and preserve the exact Vani model ID.
A fast code editor with its own agent and configurable OpenAI-compatible model providers.
Connect the OpenClaw agent runtime to Vani through a custom Chat Completions provider.
Use an admin-managed OpenAI-compatible connection with the Vani model allowlist.
Add Vani as a custom OpenAI-compatible endpoint using a minimal request baseline.
Configure the OpenAI-compatible provider and an exact custom Vani model id.
Use a custom OpenAI endpoint for ordinary chat; preserve the Vani model prefix.
Choose the Generic OpenAI LLM provider for text chat while keeping document indexing separate.
Note
Media library is an archive of generated results, while creation stays in chat.
Choose Search in the composer when you need current information. The submitted message becomes the search query; typing alone does not start a search.
Start a saved evidence job from Research history or confirm a research action proposed in chat. Follow its progress, cancel it if needed, and return to saved sources. The composer's Research mode can also add evidence to a message; it does not by itself promise a separate saved job.
Source links and citations help you check the evidence. Search results and web pages are untrusted reference material, and generated summaries can be incomplete or wrong. Open the cited source before relying on a claim.
Note
Create a Project for related conversations and reusable instructions. Changes apply to future sends; they do not rewrite saved messages or a request already being processed.
Resolve active work before branching. A missing or expired original attachment can block a draft. After an uncertain send, refresh saved results instead of repeatedly generating the same draft.
Your subscription allowance is shared across available models and conversations. The app applies both billing-period and weekly limits, along with request-rate and concurrency limits. A higher plan changes configured capacity; it does not guarantee a fixed number of messages, tokens or images.
Your usage dashboard offers Today, 7 days and 30 days as UTC calendar windows including today. Request counts use their start date; token and allowance records use settlement time, so activity spanning midnight can appear on different days.
| Feature | Supported behavior |
|---|---|
| Public QA endpoints | API-key access at https://api.vani.ai/v1: GET /models; POST /chat/completions, /responses, /messages and /messages/count_tokens. These fixed gateway routes require a Vani API key; browser session cookies do not authenticate API calls. |
| Account and keys | Create a key in Developer settings after verifying your local account. Keep keys in client secret storage or environment variables. QA verification prepares a private local message; no email is sent. |
| Chat Completions | Text strings or text-only content parts; user, assistant, system and tool messages. developer is normalized to system. Use max_completion_tokens or max_tokens, never both. Up to 64 function schemas and matching tool results are accepted. |
| Common client controls | stop accepts one string or up to four nonempty sequences totalling at most 1,024 UTF-8 bytes. parallel_tool_calls is forwarded as a boolean. frequency_penalty and presence_penalty accept finite values from -2 to 2. Model support still varies. n may be omitted or 1; store must be false or null. user is bounded metadata, ignored for authentication and not forwarded. |
| Stateless Responses | Text input and instructions, full explicit message history, function/custom tool calls and results, namespaced tools, streaming and non-streaming output. Use max_output_tokens and store:false. Resend complete input for each call; no previous_response_id, stored conversation, background mode or response retrieval. Custom grammar hints are advisory, not a guarantee of grammar-constrained generation. |
| Anthropic Messages | Text/system blocks, tool_use/tool_result, tool definitions and tool_choice, max_tokens, stop_sequences, streaming and non-streaming output. Use thinking.type=disabled. For Claude models, ephemeral cache_control hints on system and user text blocks are passed to the model provider where it takes them (response header x-vani-cache-hints: forwarded); hints on assistant blocks, tools, tool results and the request itself are accepted and ignored. A hint never guarantees a cache hit. top_k and extended thinking are unsupported. |
| Token count estimate | POST /messages/count_tokens returns a local token estimate and x-vani-token-count: local-estimate. It does not invoke a model or replace provider-reported usage used for billing. |
| Streaming and errors | Read complete SSE frames and terminal events. HTTP 200 can be followed by an error. Reported usage depends on the provider; local estimates are labelled. Disconnecting requests cancellation, but completed output can still be settled. |
| Fast models | Where a model offers priority processing, GET /models lists a separate fast model, for example openai/gpt-6-astra-fast, with its own published rate. While one is listed, sending service_tier "priority" (or "fast") with the standard model is the same request at the same rate. auto and default are always accepted; other service tiers are refused. |
| Learning mode | Add "vani_mode": "learning" to a Chat Completions, Responses or Messages request and the model tutors instead of answering outright: it asks what the learner already knows, gives one hint at a time and checks understanding, and still gives a worked solution when asked or answers directly when the request is not a learning task. Vani adds the instruction as a system message after yours; it is billed as input tokens. "standard", null or no field is a normal request; any other value is refused with 400. |
| Billing and persistence | Each completion is billed by its key: keys created in Vani Chat use the subscription allowance and limits; API platform keys use the prepaid balance at the published API rates. API calls are independent runs; retries do not promise Idempotency-Key replay protection. store:false does not erase operational, usage or security records. |
| Prompt caching | Prompt caching is automatic where the model's provider supports it. Cache hits are not guaranteed: the provider may serve a request on shared or busy hardware. usage.prompt_tokens_details.cached_tokens (Responses: usage.input_tokens_details.cached_tokens; Messages: usage.cache_read_input_tokens) shows the prompt tokens billed at the cached-input price. Models marked “No cache discount” in the API price list bill cached input at the input price. |
| Browser media and research | Chat's own media and research workflows are separate from client function tools. Tool settings can let Vani create requested media right away; web tool proposals still ask for confirmation. These adapters do not add raw image/video-generation endpoints or execute arbitrary client functions. |
| Unsupported capabilities | No raw image, audio or file input; embeddings; hosted web/computer tools; stateful Responses; encrypted reasoning continuation; background jobs; or extended Anthropic thinking. An optional encrypted-content projection can be omitted by the Responses adapter; no encrypted state is fabricated. JSON-schema output modes and unpriced extra completions remain unsupported. |
Important
Chat Completions accepts function definitions, but your API client executes them. Vani does not automatically run arbitrary client functions. Validate the model’s proposed arguments, apply your own permissions, execute the function, then send the assistant tool call and a matching tool result in the next request.
{
"model": "openai/gpt-6-sol",
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"max_completion_tokens": 256,
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
],
"additionalProperties": false
}
}
}
],
"tool_choice": "auto"
}Every assistant tool call must have a matching result before another message. Built-in web research and media actions in the Vani chat UI are a separate workflow with saved action state. In Tool settings, Vani can create requested images and videos right away or ask first. Web tool proposals always require confirmation.
Vani Chat can show charts, diagrams, formulas and clickable choices in a reply. Each is a plain-text block inside the assistant message, so the API returns it exactly as written, and nothing is added to API requests. To get these blocks from the API, put the instruction below in your system message, then render the blocks in your client, or show them as code.
Rich replies: this chat draws the fenced blocks below. Use one only when it clearly helps
(comparing numbers, a trend, a process, a formula, a choice the user must make); most replies
need none. Use the exact language tag, never nest a block inside another fence, and explain in
prose around it.
Math: LaTeX inline as $…$ and display equations as $$…$$ on their own lines; in a reply with
math, write a literal dollar sign as \$.
Charts: a vani-chart block holding one JSON object, e.g. {"type":"bar","title":"Revenue by
quarter","x":{"label":"Quarter","values":["Q1","Q2","Q3"]},"y":{"label":"Revenue","unit":"$"},"series":[{"name":"2025","values":[120,135,160]}]}.
type is line, bar, area, pie or scatter; each series has one number or null per x value;
optional "description", "source", "stacked":true (bar, area), "horizontal":true (bar). pie:
x.values are the slice labels and there is exactly one series. scatter: no x.values; each series
has "points":[[x,y],…]. Up to 8 series and 500 values each. Chart only real numbers from the
conversation or the user's files, never invented ones; aggregate large data and say so in
"description".
Diagrams: a mermaid block using only flowchart (TD or LR), sequenceDiagram, stateDiagram-v2,
mindmap or pie, with up to 40 nodes and short labels; no styling, classDef, click or HTML.
Choices: when you need the user to pick between concrete options before you can continue, end
the reply with one vani-input block holding one JSON object, e.g.
{"type":"single","question":"Which plan should I
detail?","options":["Basic","Pro","Enterprise"]}. type is single (one click), multiple
(checkboxes; optional "min" and "max"), rank (the user orders the options) or form with 2 to 5
"fields":[{"name":"budget","label":"Budget","type":"number","required":true}] (field types text,
textarea, number, select with "options", date, email). Give 2 to 8 options with short labels; an
option may be {"label":"…","description":"…"}. The answer arrives as the user's next message
(e.g. "Chose: Pro") and the user may type something else instead. At most one vani-input per
reply, never to confirm a tool action.
Charts and diagrams show data and structure only: they never replace a requested picture or
video, which only the image and video tools may make, and neither do SVG or drawing code.| Block | Format | Render with |
|---|---|---|
| Math | LaTeX: $…$ inline, $$…$$ on their own lines for display (\(…\) and \[…\] are accepted too). A literal dollar sign in a reply with math is written \$. | KaTeX or MathJax. Vani Chat uses KaTeX with trust off, so \href, \url, \includegraphics and the \html… commands never become links, images or markup. |
| Charts | A ```vani-chart fence holding one JSON object: type (line, bar, area, pie, scatter), title, optional description and source, x { label, unit, values }, y { label, unit, min, max }, series [{ name, values }] with one number or null per x value, optional stacked (bar, area) and horizontal (bar). Pie: one series, x.values are slice labels. Scatter: no x.values; each series has points [[x, y], …]. Up to 8 series and 500 values per series; unknown keys are ignored. | Any chart library: the JSON maps directly onto Recharts, Chart.js or Vega-Lite. Show the block as code when it does not validate. |
| Diagrams | A ```mermaid fence using the flowchart / graph (TD, LR), sequenceDiagram, stateDiagram-v2, mindmap and pie syntax, without styling, classDef, click or HTML. | Mermaid.js renders these blocks unchanged. Vani Chat draws this subset with its own renderer and shows any other diagram type as code. |
| Choices | A ```vani-input fence holding one JSON object: type single (one click), multiple (checkboxes; min, max), rank (order the options) or form (2 to 5 fields: name, label, type text | textarea | number | select | date | email, required, options for select); question; 2 to 8 options as text or { label, description }. | Buttons or form controls. Send the answer as the next user message, in words: "Chose: Pro", "Chose: EU, Asia", "Ranked:\n1. Speed\n2. Cost" or "Answered:\n- Budget: 5000". Users can always type something else instead. |
```vani-chart
{
"type": "bar",
"title": "Revenue by quarter",
"description": "Quarterly revenue, 2025 vs 2026.",
"x": {
"label": "Quarter",
"values": ["Q1", "Q2", "Q3", "Q4"]
},
"y": {
"label": "Revenue",
"unit": "$"
},
"series": [
{
"name": "2025",
"values": [120, 135, 160, 172]
},
{
"name": "2026",
"values": [140, 151, 189, null]
}
],
"source": "Finance export, October 2026"
}
```Treat blocks as untrusted model output
Use https://api.vani.ai/v1 as the base URL and exactly one /v1 prefix. Choose the endpoint matching the client protocol: chat/completions, responses or messages. Other paths are not exposed by the fixed gateway.
Verify the local account before creating API keys or signing uploads. Native signup remains unverified; the private QA spool does not send email.
Choose the documented protocol for your client and use text/tool messages. Disable raw image/audio/file input, provider-only fields, extended thinking and stateful features. A custom base URL alone does not establish full compatibility.
Check key expiry or revocation. Create or manage keys in Developer settings.
Check text-only content, output limits, stop-sequence bounds, complete tool-call/result pairing and the adapter-specific field list. Validation failures are rejected before a billable run.
Check the account model list and plan eligibility. A catalogue entry does not prove that inference is configured.
Wait for Retry-After when supplied. Avoid bursts; read and write requests have separate budgets and account limits still apply across clients.
Check both weekly and billing-period progress and reset times. 402 usage_limit_exceeded means the billing period; 429 weekly_limit_reached means the weekly window. Active work may reserve allowance before settlement.
Check the status page and save the request reference. An HTTP 200 stream can still contain an error; inspect the stream and usage before retrying.
Before you retry
Idempotency-Key. A timeout can occur after work starts. Inspect saved usage and the request reference before issuing a new completion.Closing the compatibility stream requests cancellation; already produced output can still be settled. Read error frames even after HTTP 200. The status pageseparates recent model outcomes from low-sample or unknown observations. It shows Vani's recorded attempts, not official global provider uptime, and does not invent recovery estimates.
Saved chats, projects, uploaded files, research evidence and in-app updates use this deployment’s database and storage. Selected model providers and web sources may still receive the request data needed for inference or search. Self-hosted application storage does not mean inference runs on this server.
Temporary chat is not added to saved conversation history, but operational run, usage and security records can still exist. Uploaded files have their own retention lifecycle. This is not a zero-retention promise.
Use Report a bug in the sidebar, or Report this error beside a failed message. The report contains the text you enter, an optional page path and an optional request reference. Conversations and files are not attached automatically.
A request reference helps staff correlate operational metadata; it is not proof of request ownership. Keep secrets and sensitive content out of bug reports. Manage keys in Developer settings and review the privacy policy.
We use cookies for analytics and to measure how well our campaigns are working. None of it is needed to run the site, and your conversations are never included. See our Privacy Policy.