Frontier models,
Up to 70% off official rates.
The frontier models from OpenAI, Anthropic and Google — Claude, Gemini, GPT-Image, Veo — most of them priced below the provider’s official rate, up to 70% off depending on the model, behind a single OpenAI-compatible API. Change one line of base_url and you’re shipping.
Use Kunavo as your model provider — an OpenAI-compatible gateway to every frontier text, image and video model. base_url: https://api.kunavo.com/v1 auth: Authorization: Bearer $KUNAVO_API_KEY To use a model, call GET /v1/models for the live catalog, then route each model by its kunavo.endpoint field. Full agent reference: https://kunavo.com/llms.txt
Providers we’ve unified
The AI gateway built for builders who ship.
From the routing layer to the billing ledger, every part of Kunavo was designed for indie developers and small teams shipping AI features for real customers.
Edge-fronted, US-East hosted
Anycast edge network terminates TLS close to you; the gateway itself runs in a single US-East region (Ashburn, Virginia).
OpenAI-compatible
Drop-in replacement for OpenAI SDKs. Streaming, function calling, tool use, vision — all wire-compatible. No new client to learn.
Stripe-native billing
Card, Apple Pay, Link, Alipay, WeChat Pay — the methods Stripe offers on USD charges. Self-serve top-ups, opt-in auto-recharge, no subscription.
Frontier models, up to 70% off
Every model from OpenAI, Anthropic and Google, most priced below the provider’s official rate — up to 70% off depending on the model. Claude, Gemini, GPT-Image, Veo — text, image and video, one balance.
Transparent pricing
Every model’s per-1M-token price is published. No hidden multipliers, no surprise overages. Failed requests are never billed.
Automatic failover
Every request can fall through up to three upstream channels. A channel that has been failing is bypassed before your request is sent — no client-side retry needed.
First-class streaming
Native SSE pass-through. Time-to-first-token matches the upstream provider — no buffering, no batching, no delay.
Granular usage data
Per-call analytics by model, key and IP. Webhook deliveries when an image, video or music task finishes. Export everything as CSV when you need it.
Prompt caching, up to 90% off
Cache reads bill at 10% of input on Claude, GPT and Gemini — pass cache_control on your system prompt and long context becomes a near-free re-read. Hit rate and savings are shown live in your dashboard.
What to build with Kunavo.
- Customer Support
AI customer support
The fastest-ROI AI deployment in any B2C SaaS — automate ticket triage, draft 80% of responses, and escalate the rest cleanly. Production code, real cost numbers, and the compliance pitfalls that catch teams off-guard.
Explore - Knowledge Base
RAG chatbot API
Most internal knowledge bases are dead documentation — nobody finds anything. A Claude-backed RAG chatbot turns them into a real assistant that cites sources and refuses when it doesn't know. Here's the production pattern.
Explore - Trust & Safety
AI content moderation
Modern moderation isn't just regex — it's nuance: sarcasm, dog whistles, brand-context misuse, image+text combinations. LLMs do this far better than rule-based systems, at a price that scales.
Explore - Developer Tools
AI code assistant
Cursor, Aider, Cline, Continue.dev — they're all powered by the same handful of frontier LLMs. If you're building a coding tool (or a co-pilot inside your own dev product), here's the architecture and the cost reality.
Explore - Data Processing
AI data extraction
The boring, valuable use case. Invoices, receipts, contracts, leads, resumes — anywhere you'd previously have built a parser, an LLM with JSON-mode does it in 30 lines, more accurately, and you can ship in a day instead of a quarter.
Explore
Frontier models, up to 70% off official.
Claude Fable 5
Frontier reasoning and long-horizon agents — the model Fable 5.1 succeeds.
Claude Opus 5
Near-flagship Opus reasoning at half the price of Fable 5 — vision and agentic coding.
Claude Opus 4.8
Anthropic Opus 4.8 — stronger agentic coding and honesty.
Claude Opus 4.7
Anthropic Opus 4.7 — previous-generation Opus; reasoning and vision.
Claude Opus 4.6
Anthropic Opus 4.6 — deep reasoning, exceptional agentic ability.
Claude Sonnet 5
Near-Opus coding and agentic quality at Sonnet cost.
Claude Sonnet 4.6
Balanced speed/quality — the everyday production workhorse, elite coding.
Claude Haiku 4.5
Anthropic Haiku 4.5 — fast and cost-efficient.
Gemini 3.8 Flash
Google's newest Flash — long-horizon software engineering and agentic execution.
Gemini 3.7 Flash
Google's newest Flash — stronger coding and agentic execution at half the 3.6 official rate.
Gemini 3.6 Flash
Gemini 3.6 Flash — thinking-by-default at Flash latency, with native audio input.
Gemini 3.1 Pro
Gemini 3.1 Pro — Google's flagship for coding, agents, and cross-modal analysis.
Gemini 2.5 Flash
Previous-gen Gemini Flash — extreme value.
GPT-6 Astra
OpenAI's newest flagship — frontier reasoning and long-horizon agentic coding.
GPT-5.6 Sol
OpenAI's newest flagship — top-tier reasoning and agentic coding.
GPT-5.6 Terra
GPT-5.6 mid tier — the everyday workhorse of the 5.6 family.
GPT-5.6 Luna
GPT-5.6 fast tier — high-volume, low-latency execution at mini-class cost.
GPT-5.5
OpenAI GPT-5.5 — flagship reasoning and agentic tool use.
Point your agent at llms.txt
It uses every model itself.
Hand one instruction to Claude Code, Cursor, Cline — or any OpenAI-compatible agent. It reads the live model catalog from Kunavo and drives text, image and video models on its own. No SDK, no glue code.
- OpenAI-wire compatible — agents need no custom integration
- GET /v1/models is the live catalog — never hardcode model names
- One key for every modality: text, image, video, audio
Use Kunavo as your model provider — an OpenAI-compatible gateway to every frontier text, image and video model. base_url: https://api.kunavo.com/v1 auth: Authorization: Bearer $KUNAVO_API_KEY To use a model, call GET /v1/models for the live catalog, then route each model by its kunavo.endpoint field. Full agent reference: https://kunavo.com/llms.txt
The more you pre-pay, the more you save.
Pre-paid wallet. $10 starts you up. No subscription, no minimum, balance never expires.
Starter
Just exploring
- Access to every model and endpoint
- Per-call usage analytics
- Email support
- No subscription, no expiry — top up from $10
Builder
Limited · +$10Shipping a product
- $100 deposit = $110 credit
- 10% bonus, limited time
- Priority email support
- Everything every account gets
Scale
Limited · +$250Running production traffic
- $1000 deposit = $1250 credit
- 25% bonus, limited time
- Priority email support
- Everything every account gets
Enterprise
Limited · +$2000High-volume scale
- $5000 deposit = $7000 credit
- 40% bonus, limited time
- Priority email support
- Everything every account gets
Included with every account
- Unlimited API keys
- Per-key monthly spend limits
- Per-key IP allowlists
- Auto-recharge from a saved card — any whole-dollar amount from $10, opt-in
- Webhooks for image, video and music tasks
- Usage API and CSV export
- OpenAI-compatible and Anthropic Messages endpoints
Start with the popular guides.
- Setup
Gemini API key
Get a Gemini key and call it with the OpenAI SDK — or use one Kunavo key for everything.
Read - Setup
Claude API key
Get a Claude key and call Claude via the OpenAI SDK or the native Messages API.
Read - Pricing
Claude API pricing 2026
Per-model rates ~60% under Anthropic, worked cost examples, and prompt-caching savings.
Read - Pricing
Gemini API pricing 2026
Gemini 3.8 / 3.7 Flash and 2.5 rates per 1M tokens, set against Google's intro pricing, with worked examples.
Read - Pricing
GPT API pricing 2026
GPT-5.5, 5.6 Sol, Terra and Luna rates ~60% under OpenAI's list, with worked examples.
Read - Concept
LLM gateway
One API for every model — routing, fallback, billing, observability and security, explained.
Read - API
OpenAI-compatible API
Point any OpenAI SDK at hosted Claude, Gemini and GPT by changing only base_url.
Read - Tools
LLM cost calculator
Estimate per-request and monthly API cost for Claude, Gemini and GPT — with savings vs list.
Read - Video
Text-to-video API
Generate video from a prompt on an OpenAI-style endpoint — live today on Google Veo 3.1.
Read - Video
Veo 3.1 API
Google Veo 3.1 text-to-video with native audio, ~50–70% under Google's list.
Read - Audio
Suno API
Generate full songs from a prompt via /v1/audio/music — metered per request, no subscription.
Read - Image
Nano Banana API
Google's image models through an OpenAI-compatible endpoint, ~25–50% under Google's list.
Read - Image
GPT-Image-2 API
OpenAI's image model on the OpenAI-compatible images endpoint at about half of list.
Read - Compare
OpenRouter alternative
Kunavo vs OpenRouter — text breadth vs multimodal, pricing, API keys and payments.
Read - Compare
OpenRouter alternatives 2026
Seven honest alternatives by motivation — cheaper, multimodal, BYO-key or open source.
Read - Models
Gemini 3 API
Gemini 3.8 and 3.7 Flash are live — what they cost against Google's introductory rate, and how to call them.
Read - Setup
CC Switch setup
Add Kunavo to CC Switch for Claude Code and Codex — which API format to pick, and why.
Read
Recent deep dives.
- 实战·5 min
在中国稳定调用 Claude / GPT / Gemini — Kunavo 中国友好路由实测
从北京/上海/深圳调用 Kunavo:无需代理即可直连。支付宝/微信支付/双币卡都可用。健壮重试代码 + 何时真的需要代理。
- 教學·6 min
用 Veo 3 為台灣品牌做短影片廣告(Sora 同端點待上線)— 5 分鐘完整教學
從文字 prompt 到 9:16 直式 Reels / TikTok / IG 短影片,全流程教學。圖生影片用既有產品照片動起來。5 種台灣品牌實際應用、繁中文字渲染注意事項、商業可用品質的進階技巧。
- 実装ガイド·8 min
日本語 RAG チャットボットを Claude で構築 — 5,000 文書のナレッジベースを 30 行で
社内ドキュメント 5,000 件を Claude Sonnet 4.6 で検索可能にする RAG 完全実装。1 クエリ約 0.9 円(prompt caching 適用後)。埋め込みは Kunavo では提供しておらず、その工程のみ OpenAI に直接課金されます。日本語特有のトークン消費・ハルシネーション対策・本番投入チェックリスト含む。
Everything you’re
wondering about.
Didn’t answer your question? Email us at contact@kunavo.com — we reply within 24 hours.
Kunavo is purpose-built for indie developers and small teams shipping production AI features. Three real differences: (1) we cover text, image, video and music on one balance; (2) Stripe-native checkout, Alipay, Apple Pay, WeChat Pay all included — no off-platform invoices; (3) full transparency on routing — we never silently swap your model to a cheaper one.
Most models are priced below the provider’s official list price — up to 70% off depending on the model; a few sit at list, and each model page says so. Bigger top-ups add a bonus on top. You also save operationally: one account, one balance, one SDK, no commitment minimums. The per-1M-token price for every model is published on /pricing — easy to compare against the upstream listing anytime.
Yes. We implement the full set of OpenAI endpoints: /v1/chat/completions, /v1/embeddings, /v1/images/generations, /v1/models and /v1/video/generations. Streaming, function calling, vision and tool use all behave identically. Projects using the OpenAI SDK migrate by changing base_url — that’s it.
No. Kunavo is a pre-paid wallet. Top-ups stay in your account forever — no subscriptions, no monthly minimums, no expiration. Account closure refunds remaining balance to your original payment method.
Never. 4xx and 5xx responses are not billed. Streaming responses that disconnect mid-flight are billed only for the tokens actually delivered. Every charge is visible per-call in the usage dashboard, exportable as CSV for accounting.
Cards (Visa, Mastercard, Amex, JCB, UnionPay), Apple Pay, Link, Cash App Pay, Klarna, Amazon Pay, Alipay and WeChat Pay. Local rails follow the buyer's country — Pix in Brazil, Bancontact in Belgium, BLIK in Poland, EPS in Austria, MB WAY in Portugal — with Stripe converting the USD price to local currency at checkout. SEPA Direct Debit, BACS and BECS are not among the methods we accept; a card issued in those countries works normally. Auto-recharge is opt-in and card-only: save a card once and the wallet refills itself by an amount you choose, from $10.
Kunavo runs in one US-East region (Ashburn, Virginia) behind an anycast edge that terminates TLS near you. Accounts, billing and usage data live in a single primary database with daily encrypted volume snapshots.
Three minutes to your first call.
One OpenAI-compatible API for Claude, Gemini, GPT-Image, Veo and Suno — $10 minimum top-up, pay only for what you call.