# Cheaper Inference > A wallet-backed, OpenAI- and Anthropic-compatible inference API. Point an existing SDK at it, keep your code, and pay less per request through routing, caching and prompt optimization. The API is served from https://api.cheaperinference.com/v1. Authenticate with `Authorization: Bearer ci_live_...` or `x-api-key: ci_live_...`; both are accepted on every endpoint. Requests are billed against a prepaid wallet, and every response carries a `cheaper_inference` object saying what it cost and what it saved. This file is the index. The same documentation is available three other ways: - Full text, every section inline: https://cheaperinference.com/llms-full.txt - Searchable JSON: `GET https://api.cheaperinference.com/v1/docs?q=...`; any valid API key, no scope required. Omit `q` for the table of contents. - Human documentation: https://cheaperinference.com/docs, and the API reference at https://cheaperinference.com/api-reference Entries below read `- [Title](link) `id`: summary; keywords`. The `id` is the same handle `/v1/docs` returns and llms-full.txt marks each section with. ## Getting started - [Quickstart](https://cheaperinference.com/docs#quick-start) `quickstart`: CheaperInference provides OpenAI-compatible API endpoints. Point your SDK at https://api.cheaperinference.com/v1 and use a CheaperInference key created in the dashboard under API keys.; start, setup, base url, first request, curl, swap - [API keys](https://cheaperinference.com/docs#quick-start) `api-keys`: Keys are created in the dashboard under API keys. New keys begin ci_live_; existing customer keys beginning ir_live_ remain valid after the platform cutover.; key, token, auth, bearer, revoke, scopes, ci_live, ir_live ## API - [Chat completions](https://cheaperinference.com/docs#api-reference) `chat-completions`: POST https://api.cheaperinference.com/v1/chat/completions accepts OpenAI-style model, messages, max_tokens and supported optional controls. Check each model's endpoint, supported_endpoints and capabilities in GET /v1/models.; chat, completions, messages, openai, stream, sse - [Claude reasoning and thinking](https://cheaperinference.com/docs#claude-reasoning) `claude-reasoning`: To request Claude thinking through Chat Completions, send reasoning_effort: "high" or reasoning: { "effort": "high" }, with a sufficient output limit such as max_tokens: 8192. These spellings are aliases; conflicting values return 400.; claude, reasoning_effort, reasoning.effort, thinking, budget_tokens, reasoning_content, include_reasoning, output_config, haiku, sonnet, opus - [Claude prompt caching](https://cheaperinference.com/docs#prompt-caching) `prompt-caching`: Claude requests can carry cache_control on the request, tools, messages and supported content blocks. The Chat-to-Messages bridge preserves supported cache controls during translation. x-ci-prompt-cache: on adds breakpoints to an unmarked Claude request at the system/tools prefix and the last eligible conversation…; claude, cache_control, prompt cache, cache hit, cache miss, x-ci-prompt-cache, cache read, cache write, affinity, session - [Text completions](https://cheaperinference.com/docs#api-reference) `text-completions`: POST https://api.cheaperinference.com/v1/completions serves the legacy prompt-style interface for older clients. New integrations should use chat completions instead; this endpoint exists for compatibility with code that has not migrated.; completions, prompt, legacy - [Anthropic Messages API](https://cheaperinference.com/docs#messages-api) `messages`: POST https://api.cheaperinference.com/v1/messages implements the Anthropic Messages surface, and POST https://api.cheaperinference.com/v1/messages/count_tokens returns token counts without running inference. This is what makes Claude Code and other Anthropic-native clients work against CheaperInference with no code…; anthropic, claude, messages, count_tokens, x-api-key - [Responses API](https://cheaperinference.com/docs#api-reference) `responses`: POST https://api.cheaperinference.com/v1/responses implements the OpenAI Responses API, which is what Codex speaks. Configure Codex with wire_api = "responses" and point base_url at https://api.cheaperinference.com/v1.; responses, codex, openai responses, wire_api - [Pro reasoning](https://cheaperinference.com/docs#pro-reasoning) `pro-reasoning`: To use Pro, send a POST request to /v1/responses with store: false and reasoning.mode: "pro". Set reasoning.effort separately to control how much reasoning the model does.; pro, reasoning.mode, luna, terra, sol, astra, gpt-5.5-pro, aggregate tokens - [Listing models](https://cheaperinference.com/docs#models-and-pricing) `models`: GET https://api.cheaperinference.com/v1/models returns every model currently available with its id, capabilities and current pricing. Use it to discover model ids rather than hardcoding them, since the catalog moves.; models, catalog, list, available, ids - [Pricing change feed](https://cheaperinference.com/docs#models-and-pricing) `pricing-changes`: GET https://api.cheaperinference.com/v1/pricing/changes?since=... returns the latest customer-facing model object for each catalog row changed at or after the required RFC 3339 since timestamp; a removal has change_type "removed" and model null. previous_pricing and current_pricing each carry input_per_million,…; pricing, pricing/changes, price feed, poll, pricing_version, since, cache pricing - [Price ceiling](https://cheaperinference.com/docs#price-ceiling) `price-ceiling`: Add min_discount_percent (0 to 99.99) to a chat completions, completions, messages or responses request to set the minimum discount off list price the request will accept. The gateway checks the value against our published list price, not a price a source claims, and strips the field before forwarding so no upstream…; price ceiling, min_discount_percent, discount, max price, cap, models/supply, min_discount_unavailable - [Zero data retention](https://cheaperinference.com/docs#zero-data-retention) `zdr`: Add zdr: true to a chat completions, completions, messages or responses request to route only through providers that support a zero-data-retention policy. They do not keep your prompts or responses and do not train on them.; zdr, zero data retention, privacy, no retention, data protection, zdr_capacity_unavailable, PII, masking, anonymization, contract coverage, provider agreement - [Ranking: speed, discount, balance](https://cheaperinference.com/docs#ranking) `ranking`: Add ranking to a chat completions, completions, messages or responses request to choose the order the gateway tries the discounted sources: "discount" (cheapest first, the historic default), "speed" (highest recent speed score first, cost breaking ties), or "balance" (the default: speed score times discount, highest…; ranking, speed, discount, balance, fastest, cheapest, route order, invalid_ranking - [Image generation and editing](https://cheaperinference.com/docs#images-and-vision) `image-generation`: POST https://api.cheaperinference.com/v1/images/generations creates images from a prompt (not Chat Completions); each data entry carries a URL or base64 image data. POST https://api.cheaperinference.com/v1/images/edits edits images, and only works with a model whose capabilities.image_edit is true.; image, images, generate image, edit image, images/generations, images/edits, mask, nano-banana - [Video generation](https://cheaperinference.com/docs#images-and-vision) `video-generation`: POST https://api.cheaperinference.com/v1/videos/generations creates video from a prompt. Generation is buffered while the provider completes the job, so use a client timeout of at least five minutes; the response carries a videos array with URLs or base64 MP4 data.; video, videos, generate video, videos/generations, seedance, audio - [Response feedback](https://cheaperinference.com/docs#api-reference) `feedback`: POST https://api.cheaperinference.com/v1/feedback attaches a 1-to-5 score and an optional comment to a previous request, identified by its request id. Feedback is what turns an A/B experiment from a cost comparison into a quality comparison, so send it if you are running experiments.; feedback, score, rating, quality, experiment - [Request and response headers](https://cheaperinference.com/docs#api-reference) `headers`: Every response carries attribution headers: x-ci-request-id identifies the request for feedback and support, x-ci-tokens-saved is the number of tokens removed before billing, x-ci-saved-usd is the dollar amount saved against list price, and x-ci-techniques lists which optimizations fired. On the request side,…; headers, x-ci, savings, attribution, request id - [Usage reporting](https://cheaperinference.com/docs#usage-and-billing) `usage-api`: GET https://api.cheaperinference.com/v1/usage/requests returns per-request usage with cursor pagination, and GET https://api.cheaperinference.com/v1/usage/daily returns daily aggregates. Both accept a date range and api_key_id to restrict usage to one key.; usage, spend, reporting, daily, requests, export, cache, cache hit, cache percentage, cached tokens, prompt cache - [Checking your balance](https://cheaperinference.com/docs#account) `balance-api`: GET https://api.cheaperinference.com/v1/account/balance returns the workspace wallet: balance_usd, available_usd, reserved_usd, plus auto_recharge_enabled, threshold_usd and recharge_amount_usd. Spend against available_usd, not balance_usd; the difference is money already held for requests that have not settled, so…; balance, credit, wallet, funds, top up, auto-recharge, available, reserved, how much left - [Error reference](https://cheaperinference.com/docs#errors) `errors`: OpenAI-compatible endpoints (chat/completions, completions, responses, images, videos) return errors as {"error": {"message", "type", "param", "code"}}, so existing OpenAI error handling works unchanged. /v1/messages and /v1/messages/count_tokens return Anthropic's {"type": "error", "error": {"type", "message"}}…; error, 400, 402, 401, 429, 500, 502, 503, insufficient, failed, error shape, envelope, unsupported_parameter, unexpected_parameter, invalid_parameter, request_rejected, context_length_exceeded, output_limit_exceeded, gateway_error, param, end_user_blocked, safety_identifier - [Which endpoint a model uses](https://cheaperinference.com/docs#api-reference) `endpoint-formats`: Text models are served in four request formats: OpenAI Chat Completions (POST https://api.cheaperinference.com/v1/chat/completions), OpenAI legacy Completions (POST https://api.cheaperinference.com/v1/completions), OpenAI Responses (POST https://api.cheaperinference.com/v1/responses) and Anthropic Messages (POST…; endpoint, endpoints, request format, wrong endpoint, model_endpoint_mismatch, supported_endpoints, image model, video model, responses model, chat completions, messages, four formats - [Cache token fields in usage](https://cheaperinference.com/docs#token-usage) `cache-token-fields`: Usage field names follow the endpoint you called, not the provider that served the request, so a field name says nothing about the serving provider or architecture. On OpenAI-compatible endpoints (chat/completions, completions, responses), prompt-cache reads are reported as cached_tokens: read…; cached_tokens, cache_read_input_tokens, cache_creation_input_tokens, prompt_tokens_details, input_tokens_details, cache tokens, usage fields, cache read, cache write, field name - [Reasoning tokens and completion_tokens](https://cheaperinference.com/docs#token-usage) `reasoning-tokens`: CheaperInference does not compute, pad or estimate output tokens: usage reports the counts the serving model returns. For most reasoning models, completion_tokens (output_tokens on Responses and Messages) includes the tokens the model spent reasoning as well as the visible answer, so a short reply can carry a large…; reasoning tokens, reasoning_tokens, completion_tokens, completion_tokens_details, output_tokens_details, thinking tokens, short answer, high token count, too many tokens, token count, billed tokens - [Timeouts](https://cheaperinference.com/docs#timeouts) `timeouts`: Streaming (SSE) and non-streaming requests have the same request limit. They run through the same endpoints and are not capped differently.; timeout, time limit, how long, max duration, deadline, sse timeout, streaming timeout, long request, 504, request took too long, slow model, per request limit, maximum request time - [Rate limits and retries](https://cheaperinference.com/docs#errors) `rate-limits`: Eligible upstream failures are retried before the error reaches your code. Retry 429 and 5xx responses with exponential backoff; do not retry 4xx other than 429, since the request will fail the same way.; rate limit, 429, retry, backoff, timeout, concurrency ## Billing - [How billing works](https://cheaperinference.com/dashboard/billing) `wallet`: Billing is a prepaid wallet, not a subscription and not an invoice. Each request reserves its worst-case cost up front, then refunds the unused part once the actual token counts are known, so a request can never overdraw the wallet.; wallet, balance, prepaid, credit, charge, cost - [Refunds and promotional credits in practice](https://cheaperinference.com/legal/terms#fees) `refunds-and-credits`: How support applies the refund clause of the Terms of Service (section 7, Fees, wallet credits and taxes, which is indexed separately under Legal). This is operational guidance, not part of the Terms.; refund, money back, return payment, cancel, final, non-refundable, chargeback, dispute, bonus, promotional credit, welcome credit, crypto transfer, wrong network, wrong token, terms - [Topping up](https://cheaperinference.com/dashboard/billing) `topping-up`: Top up on the Billing page. The card form is collected in-app by Stripe; there is no redirect to a hosted checkout page.; top up, topup, add funds, card, payment, stripe - [Auto-recharge](https://cheaperinference.com/dashboard/billing) `auto-recharge`: Auto-recharge charges your saved card once when the settled wallet balance reaches a threshold you set. It is the way to stop production traffic failing with 402 when the balance runs out.; auto recharge, automatic, threshold, top up, 402 - [Pricing](https://cheaperinference.com/#pricing) `pricing`: Curated model pricing and discounts vary by model and can change. Check the Models page or GET /v1/models for current prices before sending requests.; price, cost, discount, markup, list price ## Account - [Ask us anything: Docs and My account](https://cheaperinference.com/docs#ask-us-anything) `ask-us-anything`: The Ask us anything button in the dashboard header opens a panel with two tabs, split by what they can see. The Docs tab searches only this documentation: ask about setup, API behavior, billing rules, models, error codes or troubleshooting; it never reads your account, keys or balance, and if the docs do not cover a…; ask, assistant, help, question, docs tab, my account tab, panel, ask us anything, screenshot, attach - [Connecting Slack, Discord and Telegram](https://cheaperinference.com/docs#chat-platforms) `chat-platforms`: Connect Slack, Discord or Telegram from the Reports page in the dashboard. Slack and Discord connect through the platform's own authorization flow; Telegram opens a direct chat with the bot through your personal link, and a separate link adds the bot to a group.; slack, discord, telegram, chat bot, connect, integration, group, channel - [Chat notifications and scheduled reports](https://cheaperinference.com/docs#chat-notifications) `chat-notifications`: A connected Slack, Discord or Telegram channel can receive the savings report on a schedule (daily, weekly or monthly, at the hour and day you choose) and event notifications, on by default and adjustable per channel: low wallet balance while auto-recharge is off, spend anomalies, an organization spend limit being…; notifications, alerts, low balance, balance alert, spend anomaly, report schedule, digest, send report, subscribe, unsubscribe, daily, weekly, monthly - [Asking the bot about your account](https://cheaperinference.com/docs#chat-assistant) `chat-assistant`: The bot on every connected platform is the account assistant. Type balance, usage or info for an instant answer, or ask anything in plain language: it reads your account (balance, usage, savings, experiments, account details, report settings) and searches this same documentation, so it can answer billing, profile,…; ask, assistant, bot, balance, usage, info, question, slack, discord, telegram - [Closing your account](https://cheaperinference.com/docs#account-closure) `account-closure`: Close your account yourself from the dashboard: Settings, then Profile, under "Permanently close account". Closure requires a sign-in within the last 10 minutes (if the request is refused, sign out, sign back in and return), typing your account email exactly, and confirming that any positive prepaid balance is…; delete account, delete my account, close account, cancel account, remove account, deactivate, erase my data, gdpr, account deletion ## Optimization - [Choose your starting point](https://cheaperinference.com/docs/optimization#optimization) `optimization`: Optimization brings Prompt Studio, Experiments, and Techniques together. They are connected starting points, not mandatory onboarding steps.; optimization, overview, start, workflow, workspace, scope - [Prompt Studio: prepare and run](https://cheaperinference.com/docs/optimization#prompt-studio) `prompt-studio`: Use Prompt Studio when you have a prompt to improve for a specific task and model. Start with a representative task and define what a useful answer must include.; prompt studio, rewrite, system prompt, user prompt, target model, candidate evaluations, billing - [Review, reuse, and test a Studio result](https://cheaperinference.com/docs/optimization#prompt-studio-results) `prompt-studio-results`: A completed run supplies a candidate prompt to review. Its optimizer scores describe the optimization run; they are not measured production savings or a guarantee of quality on your traffic.; prompt studio, scores, history, copy, handoff, optimized prompt, variables - [Experiments: create control and treatments](https://cheaperinference.com/docs/optimization#experiments) `experiments`: Compare variants using test requests or application traffic sent to an experiment. The workflow is create variants, send requests, review cost, tokens, and feedback, then choose a winner.; experiments, ab test, variants, control, treatment, weights, sticky, random - [Send test requests and application traffic](https://cheaperinference.com/docs/optimization#experiment-traffic) `experiment-traffic`: An experiment needs requests before it can provide evidence. Start with the detail page's Test a variant panel, then use its integration examples to connect the application traffic you want to compare.; experiments, traffic, headers, integration, test request, X-CI-Experiment, X-CI-End-User - [Interpret results and choose a winner](https://cheaperinference.com/docs/optimization#experiment-results) `experiment-results`: Compare cost and token usage alongside task quality. Zero requests means no evidence yet.; experiments, feedback, quality, cost, tokens, promote, winner, pause, resume - [Techniques: review and configure defaults](https://cheaperinference.com/docs/optimization#techniques) `techniques`: Techniques contains optional request optimizations. Available means an implementation can run on eligible requests.; techniques, available, enabled, default, workspace, permissions, owner, admin, safe compression - [Understand technique tradeoffs](https://cheaperinference.com/docs/optimization#technique-tradeoffs) `technique-tradeoffs`: Savings depend on the workload, request eligibility, and technique configuration. Use the live catalog for current availability.; techniques, json, whitespace, context, output budget, concise, routing, cache, risk - [Exact-match response caching](https://cheaperinference.com/docs/optimization#caching) `caching`: Exact-match response caching can reuse an eligible identical deterministic request from the account's private cache instead of making another provider call. It is an optional technique, separate from a provider's prompt-cache usage.; cache, cached, repeat, identical, deterministic, techniques - [Research radar and troubleshooting](https://cheaperinference.com/docs/optimization#technique-research-radar) `technique-research-radar`: Research radar contains source-linked ideas that are not available to run. Discovery or review does not enable a technique.; research, radar, evidence, sources, implementation, unavailable, troubleshooting ## Integrations - [Pi: Claude reasoning and prompt caching](https://cheaperinference.com/docs#pi-claude) `pi`: Pi does not automatically add cache markers for a custom cheaper-inference provider. In Pi builds supporting compat.cacheControlFormat, set it to "anthropic" in the model entry of ~/.pi/agent/models.json.; pi, pi-mono, models.json, cacheControlFormat, cacheRetention, supportsCacheControl, thinkingFormat, custom provider - [Cursor](https://cheaperinference.com/docs/integrations#cursor) `cursor`: In Cursor, open Settings → Models → OpenAI API Key → Override OpenAI Base URL. Set the base URL to https://api.cheaperinference.com/v1 and the key to your CheaperInference API key.; cursor, editor, ide, override base url - [Claude Code](https://cheaperinference.com/docs/integrations#claude-code) `claude-code`: Set ANTHROPIC_BASE_URL to exactly https://api.cheaperinference.com, with no /v1 or /v1/messages suffix. Set ANTHROPIC_AUTH_TOKEN to your CheaperInference API key, then run claude from that terminal.; claude code, anthropic, cli, ANTHROPIC_BASE_URL, setup, base url, 404, not found, extra v1, double v1, /v1/v1/messages - [Codex](https://cheaperinference.com/docs/integrations#codex) `codex`: Add CheaperInference as a model provider in ~/.codex/config.toml with wire_api = "responses", base_url = "https://api.cheaperinference.com/v1", and env_key = "CHEAPER_INFERENCE_API_KEY". Then export that variable with your CheaperInference API key. env_key holds the NAME of the environment variable, not the key…; codex, config.toml, model_provider, responses, env_key, missing environment variable - [Obsidian (Claudian)](https://cheaperinference.com/docs/integrations#obsidian) `obsidian`: Use YishenTu's Claudian plugin in desktop Obsidian with Claude Code installed. In Claudian's Claude provider Environment settings, set ANTHROPIC_BASE_URL=https://api.cheaperinference.com (no /v1 suffix), ANTHROPIC_AUTH_TOKEN to your Cheaper Inference key, and ANTHROPIC_MODEL to an exact catalog ID such as…; obsidian, claudian, vault, notes, fable, MCP, local REST API - [OpenClaw](https://cheaperinference.com/dashboard/integrations#openclaw) `openclaw`: OpenClaw talks to CheaperInference as a custom Chat Completions provider. In your OpenClaw configuration, add a "cheaper-inference" provider with baseUrl set to https://api.cheaperinference.com/v1, api set to "openai-completions", and apiKey read from CHEAPER_INFERENCE_API_KEY.; openclaw, agent platform, openai-completions - [Hermes Agent](https://cheaperinference.com/dashboard/integrations#hermes) `hermes`: Run "hermes model", choose "Custom endpoint (self-hosted / VLLM / etc.)", enter https://api.cheaperinference.com/v1 as the base URL, paste your CheaperInference API key, and pick "Chat Completions" as the API mode. For a manual setup, put CHEAPER_INFERENCE_API_KEY in ~/.hermes/.env and add a custom_providers block in…; hermes, hermes agent, custom endpoint, vllm - [OpenCode](https://cheaperinference.com/dashboard/integrations#opencode) `opencode`: Configure OpenCode with an OpenAI-compatible custom provider. Open /connect, choose Other, enter cheaper-inference, and add your key.; opencode, openai-compatible, @ai-sdk/openai-compatible - [OpenWork](https://cheaperinference.com/dashboard/integrations#openwork) `openwork`: OpenWork is an OpenCode-powered desktop agent that reads OpenCode provider config from each workspace's opencode.json. Add a "cheaper-inference" provider there with baseURL pointing at https://api.cheaperinference.com/v1 and env: ["CHEAPER_INFERENCE_API_KEY"] (the env array declares the key; OpenWork stores the…; openwork, desktop agent, opencode schema - [Open WebUI](https://cheaperinference.com/dashboard/integrations#open-webui) `open-webui`: In Open WebUI, open Admin Settings → Connections → OpenAI → Add Connection. Enter https://api.cheaperinference.com/v1 as the base URL and your CheaperInference API key, then save.; open webui, openwebui, chat interface - [Cline](https://cheaperinference.com/dashboard/integrations#cline) `cline`: In Cline (VS Code), choose OpenAI Compatible as the API provider, set the Base URL to https://api.cheaperinference.com/v1, paste your CheaperInference API key, and enter an exact model id such as gpt-5.4. Advanced model settings (image support, tool use, context/output limits) are model-specific; enable them only…; cline, vs code, openai compatible - [Aider](https://cheaperinference.com/dashboard/integrations#aider) `aider`: Aider is a terminal pair programmer that reads OpenAI-compatible envs. Export OPENAI_API_BASE=https://api.cheaperinference.com/v1 and OPENAI_API_KEY to your CheaperInference API key, then run "aider --model openai/gpt-5.4".; aider, pair programmer, openai/, openai_api_base - [LibreChat](https://cheaperinference.com/dashboard/integrations#librechat) `librechat`: Add CHEAPER_INFERENCE_API_KEY to LibreChat's .env, then add a custom endpoint in librechat.yaml with name "Cheaper Inference", baseURL https://api.cheaperinference.com/v1, apiKey read from that env var, and models.fetch true so it discovers ids automatically. Restart LibreChat.; librechat, team chat, custom endpoint - [Continue](https://cheaperinference.com/dashboard/integrations#continue) `continue`: Add a model to Continue's config.yaml with provider: openai, apiBase pointing at https://api.cheaperinference.com/v1, and apiKey set to your CheaperInference API key. Keep useResponsesApi: false so Continue uses the broadly compatible Chat Completions path; the Responses endpoint is targeted at Codex workloads, not…; continue, continue.dev, ide assistant - [OpenAI Node SDK](https://cheaperinference.com/dashboard/integrations#framework-node) `framework-node`: Use the official OpenAI Node SDK with new OpenAI({ apiKey: process.env.CHEAPER_INFERENCE_API_KEY, baseURL: "https://api.cheaperinference.com/v1" }). Check GET /v1/models for the selected model's endpoint, supported_endpoints and capabilities.; openai, node, javascript, typescript, sdk - [OpenAI Python SDK](https://cheaperinference.com/dashboard/integrations#framework-python) `framework-python`: Use the official OpenAI Python SDK with OpenAI(api_key=os.environ["CHEAPER_INFERENCE_API_KEY"], base_url="https://api.cheaperinference.com/v1"). Check GET /v1/models for the selected model's endpoint, supported_endpoints and capabilities.; openai, python, sdk - [Vercel AI SDK](https://cheaperinference.com/dashboard/integrations#framework-vercel) `framework-vercel`: Use @ai-sdk/openai-compatible. Call createOpenAICompatible({ name: "cheaper-inference", apiKey: process.env.CHEAPER_INFERENCE_API_KEY, baseURL: "https://api.cheaperinference.com/v1" }), then pass cheaperInference("gpt-5.4") to generateText or streamText.; vercel, ai sdk, @ai-sdk/openai-compatible - [LangChain](https://cheaperinference.com/dashboard/integrations#framework-langchain) `framework-langchain`: Use langchain_openai's ChatOpenAI. Construct ChatOpenAI(model="gpt-5.4", api_key=os.environ["CHEAPER_INFERENCE_API_KEY"], base_url="https://api.cheaperinference.com/v1") and invoke it as normal.; langchain, langchain_openai, chatopenai - [OpenAPI spec](https://cheaperinference.com/api-reference) `openapi`: A machine-readable OpenAPI 3.1 document is served at https://api.cheaperinference.com/openapi.json and checked in at docs/api/openapi.json. Use it to generate a typed client rather than hand-writing request shapes.; openapi, spec, schema, swagger, generate client ## Troubleshooting - [Requests failing with 402](https://cheaperinference.com/dashboard/billing) `troubleshoot-402`: A 402 means the wallet could not cover the request's worst-case reservation. Top up on the Billing page, and turn on auto-recharge so it does not recur.; 402, insufficient, balance, declined, stopped working - [Requests failing with 401](https://cheaperinference.com/dashboard/keys) `troubleshoot-401`: A 401 means the key is missing, malformed or revoked. Check the key is sent as "Authorization: Bearer ..." and has not been revoked on the API keys page.; 401, unauthorized, invalid key, revoked, auth - [Getting 404 from an agent or editor](https://cheaperinference.com/docs/integrations#verify) `troubleshoot-404`: For a 404 Not Found during agent or editor setup, check the base URL first. Clients differ: Cursor and Codex need https://api.cheaperinference.com/v1, while Claude Code needs ANTHROPIC_BASE_URL=https://api.cheaperinference.com with no /v1 or /v1/messages suffix.; 404, not found, base url, v1, double v1, extra v1, /v1/v1/messages, ANTHROPIC_BASE_URL - [Savings look lower than expected](https://cheaperinference.com/dashboard/requests) `troubleshoot-no-savings`: Open the request on the Requests page and read its savings receipt, which itemises the provider list price, the marketplace discount and every optimization that fired. Two common explanations: the model is a marketplace-only model, which carries no list-price discount because the official list price is unknown; or…; savings, no savings, attribution, expected, lower ## Legal - [Terms of Service, section 1: Agreement, authority, and eligibility](https://cheaperinference.com/legal/terms#agreement) `terms-agreement`: You must be at least 18 years old and legally capable of entering into these Terms. The Services are intended for business and professional use and are not directed to children.; agreement, authority, eligibility, legal, terms, terms of service - [Terms of Service, section 2: The Services and inference routing](https://cheaperinference.com/legal/terms#services) `terms-services`: Cheaper Inference provides an OpenAI-compatible interface and related tools that route eligible requests to third-party model providers. Available models, capabilities, routes, and prices may change.; inference, legal, routing, services, terms, terms of service - [Terms of Service, section 3: Accounts, workspaces, and API keys](https://cheaperinference.com/legal/terms#accounts) `terms-accounts`: You must provide accurate account information and keep it current. You are responsible for activity under your account, workspace, and API keys, including activity by users you invite or authorize.; accounts, keys, legal, terms, terms of service, workspaces - [Terms of Service, section 4: Customer content and instructions](https://cheaperinference.com/legal/terms#content) `terms-content`: Customer Content means prompts, instructions, images, files, data, and other material submitted through the Services, and outputs generated for Customer. As between the parties and to the extent permitted by law, Customer retains its rights in Customer Content.; content, customer, instructions, legal, terms, terms of service - [Terms of Service, section 5: AI outputs and human review](https://cheaperinference.com/legal/terms#outputs) `terms-outputs`: Models can produce inaccurate, incomplete, offensive, or non-unique output. Similar or identical output may be generated for other users.; accuracy, hallucination, human, human review, legal, outputs, review, terms, terms of service - [Terms of Service, section 6: Acceptable use](https://cheaperinference.com/legal/terms#acceptable-use) `terms-acceptable-use`: You must not use or help others use the Services to: Violate law, sanctions, export controls, or third-party rights. Generate, distribute, or facilitate malware, credential theft, unauthorized surveillance, exploitation, fraud, or unlawful harm.; abuse, acceptable, gateway, jailbreak, legal, prohibited, resale, resell, reseller, shared access, terms, terms of service - [Terms of Service, section 7: Fees, wallet credits, and taxes](https://cheaperinference.com/legal/terms#fees) `terms-fees`: Usage charges. The Services use wallet-based, usage-priced billing.; auto-recharge, bonus, credits, dispute, fees, final, legal, money back, non-refundable, promotional, refund, tax, taxes, terms, terms of service, top-up, wallet - [Terms of Service, section 8: Availability, support, and changes](https://cheaperinference.com/legal/terms#availability) `terms-availability`: We may update, limit, suspend, or discontinue models, routes, features, or the Services to address provider changes, security, law, maintenance, or business needs. We do not guarantee that any model, provider, feature, response time, capacity, price, or region will remain available.; availability, changes, deprecate, legal, outage, sla, support, support hours, terms, terms of service, uptime - [Terms of Service, section 9: Third-party models and services](https://cheaperinference.com/legal/terms#third-parties) `terms-third-parties`: The Services depend on third-party model, cloud, payment, communications, analytics, and infrastructure providers. Their services may be unavailable, changed, or discontinued, and their lawful terms or use policies may apply to Customer's use.; legal, models, services, terms, terms of service, third-party - [Terms of Service, section 10: Privacy, security, and data processing](https://cheaperinference.com/legal/terms#privacy) `terms-privacy`: Our Privacy Policy explains how Keak processes personal data as a controller or business. When Keak processes Customer Personal Data on Customer's behalf, the Data Processing Addendum applies if incorporated into the parties' agreement or executed by both parties.; data, gdpr, legal, personal data, privacy, processing, processor, retention, security, terms, terms of service, zdr - [Terms of Service, section 11: Ownership, documentation, and feedback](https://cheaperinference.com/legal/terms#ownership) `terms-ownership`: Keak and its licensors own the Services, software, websites, documentation, designs, trademarks, and related technology, excluding Customer Content. No rights are granted except the limited right to use the Services under these Terms.; documentation, feedback, legal, ownership, terms, terms of service - [Terms of Service, section 12: Suspension and termination](https://cheaperinference.com/legal/terms#suspension) `terms-suspension`: You may stop using the Services at any time. Contact support to request account closure, subject to outstanding obligations and required record retention.; ban, banned, close account, closure, legal, suspended, suspension, terminate, termination, terms, terms of service - [Terms of Service, section 13: Disclaimers](https://cheaperinference.com/legal/terms#disclaimers) `terms-disclaimers`: TO THE MAXIMUM EXTENT PERMITTED BY LAW, THE SERVICES, MODELS, PROVIDER ROUTES, DOCUMENTATION, AND OUTPUTS ARE PROVIDED "AS IS" AND "AS AVAILABLE." KEAK AND ITS AFFILIATES AND LICENSORS DISCLAIM ALL EXPRESS, IMPLIED, AND STATUTORY WARRANTIES, INCLUDING MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE,…; disclaimers, legal, terms, terms of service - [Terms of Service, section 14: Limitation of liability](https://cheaperinference.com/legal/terms#liability) `terms-liability`: TO THE MAXIMUM EXTENT PERMITTED BY LAW, NEITHER PARTY WILL BE LIABLE FOR INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, CONSEQUENTIAL, OR PUNITIVE DAMAGES, OR FOR LOST PROFITS, REVENUE, GOODWILL, BUSINESS OPPORTUNITY, OR DATA, EVEN IF ADVISED THAT THE DAMAGES WERE POSSIBLE. TO THE MAXIMUM EXTENT PERMITTED BY LAW, KEAK'S…; legal, liability, limitation, terms, terms of service - [Terms of Service, section 15: Indemnity](https://cheaperinference.com/legal/terms#indemnity) `terms-indemnity`: Customer will defend, indemnify, and hold harmless Keak, its affiliates, and their officers, directors, employees, and agents from third-party claims, damages, losses, liabilities, and reasonable legal fees arising from Customer Content, Customer's products or services, Customer's breach of these Terms, or Customer's…; indemnity, legal, terms, terms of service - [Terms of Service, section 16: Governing law and courts](https://cheaperinference.com/legal/terms#law) `terms-law`: Delaware law governs these Terms, without regard to conflict-of-law rules. The state and federal courts located in Delaware have exclusive jurisdiction over disputes arising out of or relating to these Terms or the Services, and each party consents to their jurisdiction and venue.; courts, governing, legal, terms, terms of service - [Terms of Service, section 17: Changes, notices, and general terms](https://cheaperinference.com/legal/terms#general) `terms-general`: Changes. We may update these Terms.; assignment, changes, changes to terms, general, legal, notice, notices, terms, terms of service - [Terms of Service, section 18: Contact us](https://cheaperinference.com/legal/terms#contact) `terms-contact`: Keak AI, Inc. 651 North Broad Street Middletown, Delaware 19709 United States Legal and support contact form; address, company, contact, keak ai, legal, terms, terms of service