Vani documentation

Your models. Your tools.
One place to get started.

Connect your preferred client, build with the API, or get more from chat. Start with configuration examples and a clear view of what each integration supports.

On this pageAPI quickstart

API quickstart

Connect with three values: a base URL, your API key, and an enabled model identifier. Vani exposes bounded Chat Completions, stateless Responses and Anthropic Messages adapters at the endpoint below.

Base URL

https://api.vani.ai/v1

Include /v1 once. Select the API format documented in your client's guide.

Your API key

Create a key in Developer settings after verifying your account. Keep it in your client's secret field or environment.

Model identifier

openai/gpt-6-sol

A starting example. GET /models returns the models available to your account.

Important

A key is billed where it was created. Keys from Vani Chat's Developer settings use your plan's allowance, the same allowance your chats use. Keys from the API platform draw down its separate prepaid balance at the published API rates. The API key is your account credential; the base URL chooses Vani, and the model identifier chooses the model. These API routes accept Vani API keys, not browser session cookies. The browser's /chat and /api/chat routes are not API base URLs.

For pay-as-you-go access, open the API platformto manage prepaid keys, check your balance, and see API usage by model and key for today, 7, 30 or 90 days. Eligible API models cost 30% below their verified official rates; private models priced directly by Vani are labelled “Vani price”. Rates include source references and context tiers; unavailable rates are never treated as free.

Set your environment

Shell

Environment variables

export VANI_BASE_URL="https://api.vani.ai/v1"
export VANI_MODEL="openai/gpt-6-sol"
# Set VANI_API_KEY in your shell secret store or environment.
# Use the API origin shown here, ending in /v1 once.
# The website /chat and /api/chat routes are not this API base.
# Start by fetching /models; that check does not invoke inference.

Find a model identifier

Shell

Discover account model identifiers

curl --fail-with-body "$VANI_BASE_URL/models" \
  -H "Authorization: Bearer $VANI_API_KEY"

# Use an exact data[].id from this account-scoped response.

Use an exact data[].id from the authenticated model list as VANI_MODEL. The examples start with openai/gpt-6-sol; change it when needed. Model access is account-specific; disabled inference can still prevent a request.

Send a request

Shell

cURL: text completion

curl --fail-with-body "$VANI_BASE_URL/chat/completions" \
  -H "Authorization: Bearer $VANI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-6-sol","messages":[{"role":"user","content":"Hello"}],"max_completion_tokens":256,"stream":false}'

Note

These are generic SDK configuration examples for this endpoint contract. The wire adapter has local fixture tests; every SDK version and third-party application has not been certified. Set automatic retries to zero when a repeated request could consume balance twice.

Find your client

A curated set of 15 popular clients and coding tools. Each guide separates supported API settings from features that need an adapter or extra validation.

15 of 15 guides · Popular coding and chat clients.

Coding agents

Hermes Agent

Configure a named custom provider using Hermes' documented Chat Completions transport.

Chat Completions setup
Coding agents

Codex

Use the Codex CLI with Vani’s stateless Responses adapter, including local function and namespaced custom tools.

Responses adapter
Coding agents

OpenCode

Add Vani as a custom provider with the OpenAI-compatible Chat Completions package.

Chat Completions setup
Coding agents

Claude Code

Connect Claude Code to Vani’s Anthropic Messages adapter for text, streaming and client-executed tools.

Messages adapter
Editors

Cursor

Cursor's current official BYOK guide documents provider keys, but does not verify this custom Vani endpoint or GPT-6 Sol.

Check client version
Editors

Cline

Use Cline's OpenAI Compatible provider and the exact Vani model identifier.

Chat Completions setup
Continue logoEditors

Continue

Configure a Chat-only OpenAI-compatible model and explicitly disable automatic Responses routing.

Chat Completions setup
Aider logoCoding agents

Aider

Point Aider's OpenAI-compatible transport at Vani and preserve the exact Vani model ID.

Chat Completions setup
Zed logoEditors

Zed

A fast code editor with its own agent and configurable OpenAI-compatible model providers.

Custom provider · validate tools
Coding agents

OpenClaw

Connect the OpenClaw agent runtime to Vani through a custom Chat Completions provider.

Text chat setup · configuration checked
Chat workspaces

Open WebUI

Use an admin-managed OpenAI-compatible connection with the Vani model allowlist.

Text chat setup · configuration checked
LibreChat logoChat workspaces

LibreChat

Add Vani as a custom OpenAI-compatible endpoint using a minimal request baseline.

Text chat setup · configuration checked
Chat workspaces

LobeHub / LobeChat

Configure the OpenAI-compatible provider and an exact custom Vani model id.

Text chat setup · configuration checked
Chat workspaces

Cherry Studio

Use a custom OpenAI endpoint for ordinary chat; preserve the Vani model prefix.

Text chat setup · configuration checked
Chat workspaces

AnythingLLM

Choose the Generic OpenAI LLM provider for text chat while keeping document indexing separate.

Text chat setup · configuration checked

One conversation for text and media

  1. Open New chat, choose an available model, and send your message. You can change models as you continue.
  2. Use the paperclip to attach a file. Turn on Image or Videounder the message box to create media from your words; choose the model (and a video's length) in the bar that appears. The line under the message box always says what Send will do and what it uses from your allowance.
  3. Open Tool settings to choose whether Vani creates requested images and videos right away or asks first, how prompts are handled (Enhance prompt lets your chat model add visual details; Use my words sends your message unchanged), and your preferred models. Creating right away uses your allowance, up to one media generation per reply. Web tool proposals always ask first. Completed images and videos are ready to use without a continuation step.
  4. Use Stop to request cancellation. Work already completed can still count toward usage.

Note

Media eligibility and availability are separate. Image/video tools remain unavailable until their provider connection, capabilities and prices are configured. A model listed in the picker does not guarantee a successful generation.

Media library is an archive of generated results, while creation stays in chat.

Search, research and sources

Web search

Choose Search in the composer when you need current information. The submitted message becomes the search query; typing alone does not start a search.

Research

Start a saved evidence job from Research history or confirm a research action proposed in chat. Follow its progress, cancel it if needed, and return to saved sources. The composer's Research mode can also add evidence to a message; it does not by itself promise a separate saved job.

Source links and citations help you check the evidence. Search results and web pages are untrusted reference material, and generated summaries can be incomplete or wrong. Open the cited source before relying on a claim.

Note

Native web tools are browser-session features. Supplying an API key to Chat Completions does not grant access to Vani’s built-in search/research workflow.

Projects, instructions and branches

Create a Project for related conversations and reusable instructions. Changes apply to future sends; they do not rewrite saved messages or a request already being processed.

  • Archive a project to stop new assignments and new project chats. Existing conversations keep the project instructions.
  • Delete a project to detach its conversations. The saved chats remain available.
  • If another tab updates the project first, reload its latest version while preserving your unsaved edits, review them, then save again.
  • Use Edit, Regenerate or Branch on a saved message to create a separate conversation. The original stays intact; edited or regenerated responses preserve a saved branch draft. If a draft is waiting, use Generate response to start it.

Resolve active work before branching. A missing or expired original attachment can block a draft. After an uncertain send, refresh saved results instead of repeatedly generating the same draft.

Plans, pacing and usage

Your subscription allowance is shared across available models and conversations. The app applies both billing-period and weekly limits, along with request-rate and concurrency limits. A higher plan changes configured capacity; it does not guarantee a fixed number of messages, tokens or images.

Your usage dashboard offers Today, 7 days and 30 days as UTC calendar windows including today. Request counts use their start date; token and allowance records use settlement time, so activity spanning midnight can appear on different days.

  • Provider-reported tokens exclude local estimates. Records without provider reports are labelled separately.
  • Active work can reserve allowance. Reset times are scheduled dates, not a prediction of when capacity will run out.
  • The model-rate calculator explains subscription allowance debits per million input/output tokens. On Claude Sonnet 5.5, Claude Fable 5.1, MiMo V2.6 Flash, MiMo V2.6 Pro, Kimi K2.8 Preview and Claude Opus 5.5, the part of a prompt the provider serves from its cache uses a fraction of the usual allowance; on other models cached input counts at the full input rate.
  • API keys created in Vani Chat use this same allowance; API platform keys use a separate prepaid balance. Chat allowance values are not cash balances or provider invoices. Automatic API top-ups require configured payment processing; check the API platform for funding availability.

Supported API and client settings

FeatureSupported behavior
Public QA endpointsAPI-key access at https://api.vani.ai/v1: GET /models; POST /chat/completions, /responses, /messages and /messages/count_tokens. These fixed gateway routes require a Vani API key; browser session cookies do not authenticate API calls.
Account and keysCreate a key in Developer settings after verifying your local account. Keep keys in client secret storage or environment variables. QA verification prepares a private local message; no email is sent.
Chat CompletionsText strings or text-only content parts; user, assistant, system and tool messages. developer is normalized to system. Use max_completion_tokens or max_tokens, never both. Up to 64 function schemas and matching tool results are accepted.
Common client controlsstop accepts one string or up to four nonempty sequences totalling at most 1,024 UTF-8 bytes. parallel_tool_calls is forwarded as a boolean. frequency_penalty and presence_penalty accept finite values from -2 to 2. Model support still varies. n may be omitted or 1; store must be false or null. user is bounded metadata, ignored for authentication and not forwarded.
Stateless ResponsesText input and instructions, full explicit message history, function/custom tool calls and results, namespaced tools, streaming and non-streaming output. Use max_output_tokens and store:false. Resend complete input for each call; no previous_response_id, stored conversation, background mode or response retrieval. Custom grammar hints are advisory, not a guarantee of grammar-constrained generation.
Anthropic MessagesText/system blocks, tool_use/tool_result, tool definitions and tool_choice, max_tokens, stop_sequences, streaming and non-streaming output. Use thinking.type=disabled. For Claude models, ephemeral cache_control hints on system and user text blocks are passed to the model provider where it takes them (response header x-vani-cache-hints: forwarded); hints on assistant blocks, tools, tool results and the request itself are accepted and ignored. A hint never guarantees a cache hit. top_k and extended thinking are unsupported.
Token count estimatePOST /messages/count_tokens returns a local token estimate and x-vani-token-count: local-estimate. It does not invoke a model or replace provider-reported usage used for billing.
Streaming and errorsRead complete SSE frames and terminal events. HTTP 200 can be followed by an error. Reported usage depends on the provider; local estimates are labelled. Disconnecting requests cancellation, but completed output can still be settled.
Fast modelsWhere a model offers priority processing, GET /models lists a separate fast model, for example openai/gpt-6-astra-fast, with its own published rate. While one is listed, sending service_tier "priority" (or "fast") with the standard model is the same request at the same rate. auto and default are always accepted; other service tiers are refused.
Learning modeAdd "vani_mode": "learning" to a Chat Completions, Responses or Messages request and the model tutors instead of answering outright: it asks what the learner already knows, gives one hint at a time and checks understanding, and still gives a worked solution when asked or answers directly when the request is not a learning task. Vani adds the instruction as a system message after yours; it is billed as input tokens. "standard", null or no field is a normal request; any other value is refused with 400.
Billing and persistenceEach completion is billed by its key: keys created in Vani Chat use the subscription allowance and limits; API platform keys use the prepaid balance at the published API rates. API calls are independent runs; retries do not promise Idempotency-Key replay protection. store:false does not erase operational, usage or security records.
Prompt cachingPrompt caching is automatic where the model's provider supports it. Cache hits are not guaranteed: the provider may serve a request on shared or busy hardware. usage.prompt_tokens_details.cached_tokens (Responses: usage.input_tokens_details.cached_tokens; Messages: usage.cache_read_input_tokens) shows the prompt tokens billed at the cached-input price. Models marked “No cache discount” in the API price list bill cached input at the input price.
Browser media and researchChat's own media and research workflows are separate from client function tools. Tool settings can let Vani create requested media right away; web tool proposals still ask for confirmation. These adapters do not add raw image/video-generation endpoints or execute arbitrary client functions.
Unsupported capabilitiesNo raw image, audio or file input; embeddings; hosted web/computer tools; stateful Responses; encrypted reasoning continuation; background jobs; or extended Anthropic thinking. An optional encrypted-content projection can be omitted by the Responses adapter; no encrypted state is fabricated. JSON-schema output modes and unpriced extra completions remain unsupported.

Configure a generic OpenAI-compatible client

  1. Select the protocol supported by your client guide: Chat Completions, Responses, or Anthropic Messages.
  2. For OpenAI-compatible clients, use the base URL ending in /v1. Anthropic SDKs that append /v1/messages use the origin without /v1 instead. Follow your client guide, then set your API key and an enabled model identifier.
  3. Disable unsupported extras such as raw media input, response_format, extended thinking, stored Responses state and automatic provider fallbacks.
  4. Start with a short text request, then enable streaming, then test the client’s own function tools.

Important

Responses clients must send full stateless input with store:false; Anthropic clients must disable extended thinking. Raw uploads, audio and unsupported schema fields still need additional support. A configurable base URL alone is not proof of end-to-end compatibility.

Client function tools versus Vani actions

Chat Completions accepts function definitions, but your API client executes them. Vani does not automatically run arbitrary client functions. Validate the model’s proposed arguments, apply your own permissions, execute the function, then send the assistant tool call and a matching tool result in the next request.

JSON

Function definition request

{
  "model": "openai/gpt-6-sol",
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    }
  ],
  "max_completion_tokens": 256,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get weather for a city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string"
            }
          },
          "required": [
            "city"
          ],
          "additionalProperties": false
        }
      }
    }
  ],
  "tool_choice": "auto"
}

Every assistant tool call must have a matching result before another message. Built-in web research and media actions in the Vani chat UI are a separate workflow with saved action state. In Tool settings, Vani can create requested images and videos right away or ask first. Web tool proposals always require confirmation.

Charts, diagrams, math and choices

Vani Chat can show charts, diagrams, formulas and clickable choices in a reply. Each is a plain-text block inside the assistant message, so the API returns it exactly as written, and nothing is added to API requests. To get these blocks from the API, put the instruction below in your system message, then render the blocks in your client, or show them as code.

Text

System message for rich blocks

Rich replies: this chat draws the fenced blocks below. Use one only when it clearly helps
(comparing numbers, a trend, a process, a formula, a choice the user must make); most replies
need none. Use the exact language tag, never nest a block inside another fence, and explain in
prose around it.

Math: LaTeX inline as $…$ and display equations as $$…$$ on their own lines; in a reply with
math, write a literal dollar sign as \$.

Charts: a vani-chart block holding one JSON object, e.g. {"type":"bar","title":"Revenue by
quarter","x":{"label":"Quarter","values":["Q1","Q2","Q3"]},"y":{"label":"Revenue","unit":"$"},"series":[{"name":"2025","values":[120,135,160]}]}.
type is line, bar, area, pie or scatter; each series has one number or null per x value;
optional "description", "source", "stacked":true (bar, area), "horizontal":true (bar). pie:
x.values are the slice labels and there is exactly one series. scatter: no x.values; each series
has "points":[[x,y],…]. Up to 8 series and 500 values each. Chart only real numbers from the
conversation or the user's files, never invented ones; aggregate large data and say so in
"description".

Diagrams: a mermaid block using only flowchart (TD or LR), sequenceDiagram, stateDiagram-v2,
mindmap or pie, with up to 40 nodes and short labels; no styling, classDef, click or HTML.

Choices: when you need the user to pick between concrete options before you can continue, end
the reply with one vani-input block holding one JSON object, e.g.
{"type":"single","question":"Which plan should I
detail?","options":["Basic","Pro","Enterprise"]}. type is single (one click), multiple
(checkboxes; optional "min" and "max"), rank (the user orders the options) or form with 2 to 5
"fields":[{"name":"budget","label":"Budget","type":"number","required":true}] (field types text,
textarea, number, select with "options", date, email). Give 2 to 8 options with short labels; an
option may be {"label":"…","description":"…"}. The answer arrives as the user's next message
(e.g. "Chose: Pro") and the user may type something else instead. At most one vani-input per
reply, never to confirm a tool action.

Charts and diagrams show data and structure only: they never replace a requested picture or
video, which only the image and video tools may make, and neither do SVG or drawing code.
BlockFormatRender with
MathLaTeX: $…$ inline, $$…$$ on their own lines for display (\(…\) and \[…\] are accepted too). A literal dollar sign in a reply with math is written \$.KaTeX or MathJax. Vani Chat uses KaTeX with trust off, so \href, \url, \includegraphics and the \html… commands never become links, images or markup.
ChartsA ```vani-chart fence holding one JSON object: type (line, bar, area, pie, scatter), title, optional description and source, x { label, unit, values }, y { label, unit, min, max }, series [{ name, values }] with one number or null per x value, optional stacked (bar, area) and horizontal (bar). Pie: one series, x.values are slice labels. Scatter: no x.values; each series has points [[x, y], …]. Up to 8 series and 500 values per series; unknown keys are ignored.Any chart library: the JSON maps directly onto Recharts, Chart.js or Vega-Lite. Show the block as code when it does not validate.
DiagramsA ```mermaid fence using the flowchart / graph (TD, LR), sequenceDiagram, stateDiagram-v2, mindmap and pie syntax, without styling, classDef, click or HTML.Mermaid.js renders these blocks unchanged. Vani Chat draws this subset with its own renderer and shows any other diagram type as code.
ChoicesA ```vani-input fence holding one JSON object: type single (one click), multiple (checkboxes; min, max), rank (order the options) or form (2 to 5 fields: name, label, type text | textarea | number | select | date | email, required, options for select); question; 2 to 8 options as text or { label, description }.Buttons or form controls. Send the answer as the next user message, in words: "Chose: Pro", "Chose: EU, Asia", "Ranked:\n1. Speed\n2. Cost" or "Answered:\n- Budget: 5000". Users can always type something else instead.
Text

vani-chart block

```vani-chart
{
  "type": "bar",
  "title": "Revenue by quarter",
  "description": "Quarterly revenue, 2025 vs 2026.",
  "x": {
    "label": "Quarter",
    "values": ["Q1", "Q2", "Q3", "Q4"]
  },
  "y": {
    "label": "Revenue",
    "unit": "$"
  },
  "series": [
    {
      "name": "2025",
      "values": [120, 135, 160, 172]
    },
    {
      "name": "2026",
      "values": [140, 151, 189, null]
    }
  ],
  "source": "Finance export, October 2026"
}
```

Treat blocks as untrusted model output

Validate the JSON before drawing, render text as text (never as HTML), keep KaTeX trust off, and send a choice only when the user clicks it. A block that does not validate can always be shown as code.

Errors, retries and cancellation

404 / wrong API path

Use https://api.vani.ai/v1 as the base URL and exactly one /v1 prefix. Choose the endpoint matching the client protocol: chat/completions, responses or messages. Other paths are not exposed by the fixed gateway.

403 / email_unverified

Verify the local account before creating API keys or signing uploads. Native signup remains unverified; the private QA spool does not send email.

Client adds unsupported fields

Choose the documented protocol for your client and use text/tool messages. Disable raw image/audio/file input, provider-only fields, extended thinking and stateful features. A custom base URL alone does not establish full compatibility.

401 / invalid_token

Check key expiry or revocation. Create or manage keys in Developer settings.

400 / bad_request

Check text-only content, output limits, stop-sequence bounds, complete tool-call/result pairing and the adapter-specific field list. Validation failures are rejected before a billable run.

403 or model_not_available

Check the account model list and plan eligibility. A catalogue entry does not prove that inference is configured.

429 / rate_limited

Wait for Retry-After when supplied. Avoid bursts; read and write requests have separate budgets and account limits still apply across clients.

weekly_limit_reached / usage_limit_exceeded

Check both weekly and billing-period progress and reset times. 402 usage_limit_exceeded means the billing period; 429 weekly_limit_reached means the weekly window. Active work may reserve allowance before settlement.

503 or stream interrupted

Check the status page and save the request reference. An HTTP 200 stream can still contain an error; inspect the stream and usage before retrying.

Before you retry

The compatibility endpoint does not promise replay protection for Idempotency-Key. A timeout can occur after work starts. Inspect saved usage and the request reference before issuing a new completion.

Closing the compatibility stream requests cancellation; already produced output can still be settled. Read error frames even after HTTP 200. The status pageseparates recent model outcomes from low-sample or unknown observations. It shows Vani's recorded attempts, not official global provider uptime, and does not invent recovery estimates.

Privacy, storage and bug reports

Saved chats, projects, uploaded files, research evidence and in-app updates use this deployment’s database and storage. Selected model providers and web sources may still receive the request data needed for inference or search. Self-hosted application storage does not mean inference runs on this server.

Temporary chat is not added to saved conversation history, but operational run, usage and security records can still exist. Uploaded files have their own retention lifecycle. This is not a zero-retention promise.

Use Report a bug in the sidebar, or Report this error beside a failed message. The report contains the text you enter, an optional page path and an optional request reference. Conversations and files are not attached automatically.

A request reference helps staff correlate operational metadata; it is not proof of request ownership. Keep secrets and sensitive content out of bug reports. Manage keys in Developer settings and review the privacy policy.

We use cookies for analytics and to measure how well our campaigns are working. None of it is needed to run the site, and your conversations are never included. See our Privacy Policy.