# compareai.today — Full AI Model Pricing Data > Compare API pricing across 9 leading AI providers. Updated August 12, 2026. > Built by [Digital Five](https://www.digitalfive.com). ## Providers ### Anthropic (Claude) - Website: https://anthropic.com - Docs: https://platform.claude.com/docs/en/about-claude/pricing - Strengths: Deep reasoning, Long context (1M tokens), Safety & alignment, Code generation, Vision capabilities - Best for: Enterprise research, complex analysis, and tasks requiring deep contextual understanding ### OpenAI (GPT) - Website: https://openai.com - Docs: https://platform.openai.com/docs/pricing - Strengths: Versatile NLP, Content generation, Code assistance, Data analysis, Largest ecosystem - Best for: General-purpose AI tasks, content creation, and building AI-powered applications ### Google (Gemini) - Website: https://ai.google.dev - Docs: https://ai.google.dev/gemini-api/docs/pricing - Strengths: Native multimodality, Google Search grounding, STEM reasoning, Huge context (1M tokens), Free tier - Best for: Multimodal applications, enterprise research, and STEM problem-solving ### xAI (Grok) - Website: https://x.ai - Docs: https://docs.x.ai/developers/models - Strengths: Real-time data from X, Advanced reasoning, Low hallucination, 1M context window, Cost-effective - Best for: Real-time analysis, trend monitoring, and applications needing current information ### Perplexity AI (Sonar) - Website: https://perplexity.ai - Docs: https://docs.perplexity.ai/guides/pricing - Strengths: Cited answers, Real-time web search, Research synthesis, Source verification, Deep research mode - Best for: Research, fact-checking, and tasks requiring cited, up-to-date information ### Microsoft Azure (Azure OpenAI) - Website: https://azure.microsoft.com/en-us/products/ai-services/openai-service - Docs: https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/ - Strengths: Microsoft 365 integration, Enterprise security, Azure scalability, Compliance, Custom deployments - Best for: Enterprise productivity, Microsoft ecosystem integration, and custom AI solutions ### Mistral AI - Website: https://mistral.ai - Docs: https://mistral.ai/pricing - Strengths: Open-source models, Cost-efficient, European data sovereignty, Code generation, Edge AI - Best for: Cost-conscious enterprises, code generation, and European data compliance ### Meta (Llama) - Website: https://llama.meta.com - Docs: https://www.together.ai/pricing - Strengths: Fully open-source, Strong reasoning, Customizable, Community-driven, Multiple hosting options - Best for: Custom AI applications, self-hosted deployments, and open-source innovation ### DeepSeek - Website: https://deepseek.com - Docs: https://platform.deepseek.com/api-docs/pricing - Strengths: Ultra-low pricing, Strong coding ability, Mathematical reasoning, Open-source, Efficient architecture - Best for: Budget-conscious developers, coding tasks, and mathematical reasoning ## Complete Model Pricing All prices in USD per million tokens (MTok). Sorted by provider. ### Anthropic | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Claude Fable 5 | Premium | $10.00 | $50.00 | 1M tokens | Yes | Top tier above Opus (June 9, 2026). Public Mythos-class model at double Opus pricing. | | Claude Opus 5 | Premium | $5.00 | $25.00 | 1M tokens | Yes | New Opus flagship (July 24, 2026). Near-Fable-5 capability at half the price, same $5/$25 as Opus 4.8. Cached read $0.50. Fast mode $10/$50. Prompt-cache minimum halves to 512 tokens. | | Claude Opus 4.8 | Premium | $5.00 | $25.00 | 1M tokens | Yes | Previous Opus flagship (May 2026), superseded by Opus 5 at the same $5/$25 price. | | Claude Opus 4.7 | Premium | $5.00 | $25.00 | 1M tokens | Yes | Older Opus release. The 4.7-generation tokenizer bills roughly 30% more tokens than 4.6 for the same text. | | Claude Sonnet 5 | Standard | $2.00 | $10.00 | 1M tokens | Yes | New Standard flagship (GA). The $2/$10 launch rate is now the standard price — Anthropic confirmed Aug 11, 2026 that the scheduled Sep 1 rise to $3/$15 will not occur. Cached read $0.20, 5m cache write $2.50, 1h cache write $4.00. New tokenizer bills ~30% more tokens than Sonnet 4.6 (1.0-1.35x). Full 1M context at standard rates. Adaptive thinking on by default. | | Claude Sonnet 4.6 | Standard | $3.00 | $15.00 | 1M tokens | Yes | Previous-generation Sonnet; superseded by Sonnet 5. Still strong for production workloads. | | Claude Haiku 4.5 | Budget | $1.00 | $5.00 | 200k tokens | Yes | Fast and affordable for high-volume tasks. | ### OpenAI | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | GPT-5.6 Sol | Premium | $5.00 | $30.00 | 1.05M | Yes | Latest flagship (GA Jul 9, 2026). Cached read $0.50, cache write $6.25 (1.25x input). 2x in/1.5x out above 272K. | | GPT-5.6 Terra | Standard | $2.00 | $12.00 | 1.05M | Yes | Mid tier (GA Jul 9, 2026). Cut 20% from $2.50/$15.00 on Jul 30, 2026. Cached read $0.20, cache write $2.50 (1.25x input). 2x in/1.5x out above 272K. | | GPT-5.6 Luna | Budget | $0.20 | $1.20 | 1.05M | Yes | Budget tier (GA Jul 9, 2026). Cut 80% from $1.00/$6.00 on Jul 30, 2026 — OpenAI's cheapest model. Cached read $0.02, cache write $0.25 (1.25x input). 2x in/1.5x out above 272K. | | GPT-5.5 | Premium | $5.00 | $30.00 | 1M | Yes | Previous flagship (Apr 2026), superseded by GPT-5.6 Sol. 1M context. Cached input $0.50. 2x in/1.5x out above 272K. | | GPT-5.5 Pro | Premium | $30.00 | $180.00 | 1M | Yes | Highest-precision flagship variant (GA). Most expensive model tracked. | | GPT-5.4 | Premium | $2.50 | $15.00 | 1M | Yes | Previous flagship. Cached input $0.25. 2x in/1.5x out above 272K. | | GPT-5.4 Mini | Standard | $0.75 | $4.50 | 400K | Yes | Balanced performance and cost. | | GPT-5.4 Nano | Budget | $0.20 | $1.25 | 400K | Yes | Most affordable OpenAI model. | | o4-mini | Standard | $1.10 | $4.40 | 200K | Yes | Reasoning model. Cached input $0.275. Deprecated — API shutdown Oct 23, 2026; OpenAI points migrations at GPT-5.6 Terra. | ### Google (Gemini) | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Gemini 3.1 Pro Preview | Premium | $2.00 | $12.00 | 1M | Yes | Preview release. 1M context. $4/$18 above 200K tokens. Cached input $0.20. | | Gemini 3.6 Flash | Standard | $1.50 | $7.50 | 1M | Yes | Newest GA Flash (July 21, 2026). Same $1.50 input as 3.5 Flash with output cut from $9 to $7.50. Cached input $0.15. | | Gemini 3 Flash Preview | Standard | $0.50 | $3.00 | 1M | Yes | Preview release. Faster successor to 2.5 Flash. | | Gemini 3.5 Flash | Standard | $1.50 | $9.00 | 1M | Yes | Previous GA workhorse (Google I/O, May 2026), superseded by 3.6 Flash. Cached input $0.15. | | Gemini 3.5 Flash-Lite | Budget | $0.30 | $2.50 | 1M | Yes | Newest Flash-Lite (GA, July 21, 2026). Higher-capability tier above 3.1 Flash-Lite. Cached input $0.03. | | Gemini 3.1 Flash-Lite | Budget | $0.25 | $1.50 | 1M | Yes | Newest Flash-Lite (GA, May 2026). Native multimodal input. | | Gemini 2.5 Pro | Premium | $1.25 | $10.00 | 1M | Yes | GA flagship. 1M context. $2.50/$15 above 200K tokens. | | Gemini 2.5 Flash | Standard | $0.30 | $2.50 | 1M | Yes | Output price includes thinking tokens. | | Gemini 2.5 Flash-Lite | Budget | $0.10 | $0.40 | 1M | Yes | Lightweight 2.5 variant. Native multimodal input. | ### xAI (Grok) | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Grok 4.5 | Premium | $2.00 | $6.00 | 500K tokens | No | Flagship (Jul 8, 2026). $4/$12 above 200K prompt tokens. Cached input $0.30. No batch discount. Not yet available in the EU. | | Grok 4.3 | Premium | $1.25 | $2.50 | 1M tokens | Yes | Previous flagship (May 2026). Cheaper than Grok 4.5 with a larger 1M context. Cached input $0.20. Batch 20% off. | | Grok 4.20 | Standard | $1.25 | $2.50 | 1M tokens | Yes | Repriced to match 4.3. Lowest hallucination rate; reasoning and multi-agent variants. | | Grok Build 0.1 | Budget | $1.00 | $2.00 | 256K tokens | No | No batch discount. Coding specialist. grok-code-fast-1 is now an alias for this model. Cached input $0.20. | ### Perplexity AI | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Sonar Deep Research | Premium | $2.00 | $8.00 | N/A | No | Exhaustive multi-search research mode. Token rates above are only part of the bill: citation tokens $2/MTok, reasoning tokens $3/MTok, plus $5 per 1,000 searches. | | Sonar Pro | Premium | $3.00 | $15.00 | N/A | No | Enhanced search with multi-step reasoning. Per-request search fees apply on top: $6-$14 per 1,000 requests depending on search context size. | | Sonar Reasoning Pro | Standard | $2.00 | $8.00 | N/A | No | Reasoning with search capabilities. Per-request search fees apply on top: $6-$14 per 1,000 requests depending on search context size. | | Sonar | Budget | $1.00 | $1.00 | N/A | No | Basic search-augmented generation. Per-request search fees apply on top: $5-$12 per 1,000 requests depending on search context size. | ### Microsoft Azure | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Azure GPT-5.5 | Premium | $5.00 | $30.00 | 1M | Yes | Azure-hosted GPT-5.5 (Foundry Global Standard). Cached input $0.50. 2x in/1.5x out above 272K. | | Azure GPT-5.4 | Premium | $2.50 | $15.00 | 1M | Yes | Foundry Global Standard. $5/$22.50 above 272K. Cached input $0.25. | | Azure GPT-5.2 | Standard | $1.75 | $14.00 | 400K | Yes | Previous Azure flagship. | | Azure GPT-5 Mini | Budget | $0.25 | $2.00 | 400K | Yes | Budget-friendly Azure option. | | Azure o4-mini | Standard | $1.10 | $4.40 | 200K | Yes | Reasoning model on Azure. Deprecated — OpenAI retires the API on Oct 23, 2026. | ### Mistral AI | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Mistral Large 3 | Premium | $0.50 | $1.50 | 256K | Yes | Flagship 675B open-weight MoE. Context expanded from 32K to 256K. | | Mistral Medium 3.5 | Premium | $1.50 | $7.50 | 256K | Yes | Frontier multimodal/agentic model (commercial tier). | | Magistral Small | Standard | $0.50 | $1.50 | 128K | Yes | Open-weight reasoning model. Replaces Magistral Medium 1.2, retired Jul 31, 2026, at a quarter of its $2.00/$5.00 rate. | | Mistral Small 4 | Standard | $0.15 | $0.60 | 256K | Yes | Latest small generation (v26.03). 256K context. | | Codestral | Standard | $0.30 | $0.90 | 256K | Yes | Specialized for code generation. Price cut; context expanded to 256K. | | Devstral 2 | Standard | $0.40 | $2.00 | 256K | Yes | Open-weight agentic coding model (123B dense). 72.2% on SWE-bench Verified. Ships with the Mistral Vibe CLI. | ### Meta (Llama) | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | Muse Spark 1.1 | Standard | $1.25 | $4.25 | 1M tokens | No | Meta's first paid first-party API (Meta Model API, public preview Jul 9, 2026). Multimodal agentic model with computer use. US developers only; $20 free credits on signup. | | Llama 4 Scout | Budget | $0.10 | $0.30 | 10M tokens | No | 10M context window. Hosted on DeepInfra/Together. Pricing varies by host. | | Llama 4 Maverick | Standard | $0.20 | $0.80 | 1M tokens | No | Hosted rates rose to $0.20/$0.80 on DeepInfra. Varies by host. | | Llama 3.3 70B | Standard | $0.10 | $0.32 | 128K | No | Commodity host rates $0.10/$0.32 (DeepInfra); Together is $1.04. Varies widely by host. | | Llama 3.1 8B | Budget | $0.02 | $0.04 | 128K | No | Llama 3.1 8B on hosted providers ($0.02/$0.04 on DeepInfra) — 128K context, the cheapest rate tracked. Supersedes the original Llama 3 8B (8K). | ### DeepSeek | Model | Tier | Input/MTok | Output/MTok | Context | Batch Discount | Notes | |-------|------|-----------|-------------|---------|---------------|-------| | DeepSeek V4 Flash | Budget | $0.14 | $0.28 | 1M tokens | No | Replaces V3/R1. Cache hits drop to $0.0028/MTok input. DeepSeek has announced 2x peak-hour pricing (09:00-12:00 and 14:00-18:00 Beijing time); start date not yet confirmed. | | DeepSeek V4 Pro | Standard | $0.435 | $0.87 | 1M tokens | No | 75% price cut made permanent (May 2026). Cache hits $0.003625/MTok input. | ## Blog Articles | Article | URL | Description | |---------|-----|-------------| | AI API Pricing August 2026: What Changed Across 10 Providers | https://compareai.today/blog/ai-api-pricing-august-2026 | OpenAI cut GPT-5.6 Luna by 80%, Anthropic cancelled the Claude Sonnet 5 price rise, and DeepSeek warned of an increase. Every AI API price change in August 2026, verified against vendor pricing pages. | | Claude Opus 5: What's New, Pricing and Who It's For | https://compareai.today/blog/claude-opus-5-whats-new-pricing-who-its-for | Claude Opus 5 launched on 24 July 2026 at $5/$25 per million tokens - the same price as Opus 4.8. Here's what actually changed, how it compares to Fable 5 and GPT-5.6 Sol, and which teams should switch. | | Cheapest AI API 2026: Cost Per Million Tokens Ranked | https://compareai.today/blog/cheapest-ai-api-2026 | The cheapest AI APIs in 2026, ranked by real blended cost per million tokens. Amazon Nova, GPT-5 nano, Ministral 3, DeepSeek and Gemini Flash-Lite compared — then price your own workload free. | | AI News July 2026: The Biggest Stories So Far | https://compareai.today/blog/ai-news-july-2026 | Claude Sonnet 5, GPT-5.6 and Grok 4.5 all landed in a fortnight, Gemini 3.5 Pro slipped, and the EU AI Act gets its teeth in August. Here's July 2026's AI news so far - and what each story does to your cost per million tokens. | | Fable 5 Returns: Inside Anthropic's Export-Control Standoff | https://compareai.today/blog/fable-5-returns-export-control-standoff | Claude Fable 5 is back after a 19-day suspension. Here's what triggered the export controls, what changed, and how it now compares on compareai.today. | | AI News Today: Sonnet 5, Export Rules, Black Market | https://compareai.today/blog/ai-news-today-sonnet-5-export-rules-black-market | Claude Sonnet 5 launched at roughly half the price of Opus 4.8, the US tightened its AI chip export rules to reach Chinese firms' overseas subsidiaries, and reporting exposed a black market reselling API tokens at up to 93 percent off. Here is today's AI news roundup, and what each story means if you build with AI. | | The Future of Operating Systems: How AI Rewrites the OS | https://compareai.today/blog/future-of-operating-systems-ai | AI is turning the operating system from passive plumbing into an agentic layer that reads your intent and acts for you. See what an AI-native OS means, who's building it, and the benefits, risks and UK implications. | | What Is a Token in AI? How AI Pricing Really Works | https://compareai.today/blog/what-is-a-token-in-ai | A token is the unit AI APIs bill you for — roughly four characters of text. See how tokens are counted, why output costs more than input, and how to estimate and cut your AI bill. | | How Reliable Is Claude? Uptime and Integration Risk | https://compareai.today/blog/claude-uptime-integration-risk | Claude is dependable for everyday chat, but its API has no default uptime guarantee, and 2026 brought a run of outages. See why wiring Claude into tooling carries extra risk, and how to build resilient AI integrations. | | Claude Fable 5 and Mythos 5 Suspended: Alternatives | https://compareai.today/blog/claude-fable-5-mythos-5-suspended-alternatives | The US government has suspended Claude Fable 5 and Mythos 5 for all users. Here's what happened, what it means, and the best alternatives to switch to now. | | SpaceX IPO 2026: What It Means for AI, Grok & Starlink | https://compareai.today/blog/spacex-ipo-2026-ai-grok-starlink | SpaceX's $1.75T IPO is also an AI story: it now owns xAI and Grok and is building orbital data centers. See what it means for AI buyers and Grok pricing. | | Claude Fable 5 Explained: Pricing and Capabilities | https://compareai.today/blog/claude-fable-5-explained-pricing-capabilities | Claude Fable 5 is Anthropic's most powerful public model yet. See its benchmarks, $10/$50 pricing and how its API cost compares to GPT, Gemini and Grok. | | AI Image Generator Pricing Compared (2026): Cost Per Image | https://compareai.today/blog/ai-image-generator-pricing-compared | Midjourney, DALL-E/GPT Image, Firefly, Stable Diffusion, Imagen and Flux — converted into real cost per image, from about $0.005 to $0.08. See which is cheapest for your workload. | | Claude Opus 4.8: New Features, Benchmarks, and Pricing | https://compareai.today/blog/claude-opus-4-8-whats-new-benchmarks-pricing | Anthropic released Claude Opus 4.8 on May 28. See verified benchmarks vs GPT-5.5 and Gemini 3.1 Pro, new pricing, and what to pin if you build on the API. | | Claude Opus 4.7 Review: Same Price, 35% Bigger Bills (Here's Why) | https://compareai.today/blog/claude-opus-4-7-review-same-price-bigger-bills | Claude Opus 4.7 keeps $5/$25 per MTok, but a new tokenizer counts up to 35% more tokens. See workload cost impact, mitigations, and a clean upgrade decision. | | DeepSeek V4 Just Halved AI Inference Costs Again — What It Means For Your Stack | https://compareai.today/blog/deepseek-v4-halved-ai-inference-costs | DeepSeek V4 launched in April 2026 at roughly half V3's per-token rate, undercutting GPT-5 and Claude on cost. Here's what to do in the next 90 days. | | ChatGPT vs Claude vs Gemini (2026): Real Output Comparison | https://compareai.today/blog/chatgpt-vs-claude-vs-gemini-2026 | Practical comparison of ChatGPT, Claude, and Gemini for writing, coding, and research with scoring rubric. | | The Best AI Image Generator For Realistic Photos | https://compareai.today/blog/best-ai-image-generator-realistic-photos | Side-by-side comparison of Midjourney, DALL-E, Adobe Firefly, Stable Diffusion, and Google Imagen. | | I Ran the Same Prompt on 7 AI Models — The Results Will Surprise You | https://compareai.today/blog/same-prompt-7-ai-models | Running the same prompt across seven AI models reveals surprising differences in reasoning and safety. | | Anthropic Holds Claude Opus 4.6 Pricing Steady at $5/$25 per MTok | https://compareai.today/blog/claude-opus-4-6-pricing | Anthropic's latest flagship model maintains competitive pricing while expanding the context window to 1M tokens at standard rates. | | OpenAI GPT-5.4 Series: New Nano Model at $0.20/MTok Input | https://compareai.today/blog/openai-gpt-5-4-launch | OpenAI expands its model lineup with GPT-5.4 Nano, the most affordable model yet at $0.20 input / $1.25 output per million tokens. | | Google Drops Gemini 2.0 Flash to $0.075/MTok — Cheapest Major Model | https://compareai.today/blog/gemini-2-0-flash-price-drop | Google aggressively cuts pricing on Gemini 2.0 Flash, now the cheapest model from a major provider. | ## FAQ Browse all questions at https://compareai.today/faq. ### About compareai.today (https://compareai.today/faq/about) **Q: Is compareai.today free to use?** A: Yes, completely free. No sign-up, no email capture, no credit card. All tools — the cost calculator, comparison table, charts, and pricing history — are available immediately without an account. **Q: Which AI providers does compareai.today cover?** A: We track 10 providers: Anthropic (Claude), OpenAI (GPT), Google (Gemini), xAI (Grok), Perplexity, Microsoft Azure, Mistral, Meta (Llama), DeepSeek, and Manus — 51 token-priced models plus Manus credit plans. **Q: How often are prices updated?** A: We update pricing data monthly and whenever a provider announces a price change. Exchange rates for currency conversion are updated daily. Every provider page links to the official pricing docs so you can verify the source. **Q: How accurate is the pricing data?** A: Prices are taken from official provider pricing pages and re-verified in a monthly sweep against provider documentation. Prices can vary by region, usage tier, and contract, so always confirm with the provider before committing to volume. **Q: Where does the pricing data come from?** A: Directly from each provider's official pricing documentation — the same pages linked in our footer. For open-weight models like Meta Llama that are served by third-party hosts, we show representative rates from major hosts such as DeepInfra and Together, and note that pricing varies by host. **Q: Which currencies can I see prices in?** A: Prices can be displayed in 18 currencies, including USD, GBP, EUR, JPY, INR, and AUD. Exchange rates refresh daily; USD is always the canonical source rate. **Q: Do you have an API or embed widget?** A: There's no embed widget, but the full pricing dataset is published in machine-readable form: JSON at compareai.today/data/pricing.json, CSV at /data/pricing.csv, and markdown at /llms-full.txt — all free to use with attribution (CC BY 4.0) and refreshed with every pricing update. If you need something custom, contact Digital Five via digitalfive.com. **Q: Who built compareai.today?** A: compareai.today is built and maintained by Digitalfive.com — a UK digital agency specialising in Website Design, SEO, PPC and Agents. ### AI API Pricing Basics (https://compareai.today/faq/pricing-basics) **Q: What is a token in AI pricing?** A: A token is the unit AI models use to process text — roughly 4 characters, or about three-quarters of an English word. APIs bill per token, so a 1,000-word prompt is around 1,300 input tokens. Every price on this site is quoted per million tokens (MTok). **Q: What does "per MTok" mean?** A: MTok is one million tokens, the standard unit for AI API pricing. If a model costs $2.50/MTok input, then one million tokens of input — roughly 750,000 words — costs $2.50. Input (your prompt) and output (the model's response) are billed at different rates. **Q: Why do output tokens cost more than input tokens?** A: Generating text is far more compute-intensive than reading it, so most providers price output 3–6x above input — Claude Opus 5, for example, charges $5.00/MTok in but $25.00/MTok out. If your workload produces long responses, output pricing will dominate your bill. **Q: What is a context window?** A: The context window is the maximum amount of text (in tokens) a model can consider at once — your prompt, any documents, and the conversation so far. Current flagships like Claude Fable 5 and Gemini 2.5 Pro support 1M tokens, and OpenAI's GPT-5.6 family reaches 1.05M; Meta's Llama 4 Scout reaches 10M. Bigger windows let you process whole codebases or document sets in one call. **Q: What is batch processing and why is it cheaper?** A: Batch processing queues API requests for asynchronous completion (typically within 24 hours) instead of answering in real time. Because providers can schedule the work on spare capacity, Anthropic, OpenAI, Google, Mistral, and Azure discount batch jobs by 50% on both input and output tokens. The rate is not universal: xAI discounts by only 20%, and then only on Grok 4.3 and the Grok 4.20 variants, while DeepSeek has no batch discount at all. Anything that doesn't need an instant response — ETL, evals, backfills — belongs in a batch. **Q: What is prompt caching?** A: Prompt caching lets you reuse a repeated prompt prefix (system prompts, shared documents) at a fraction of the normal input rate. OpenAI bills cached GPT-5.6 Sol input at $0.50/MTok versus $5.00 list, and DeepSeek V4 Flash cache hits drop to under a cent per MTok. Check whether your model charges to write the cache as well as read it: the GPT-5.6 family bills cache writes at 1.25x the input rate, so prompts that are refreshed constantly rather than reused can cost more, not less. If many requests share the same preamble, caching is often the single biggest saving available. **Q: How many tokens are in typical text?** A: Rules of thumb for English: 1 token ≈ 4 characters, 1,000 tokens ≈ 750 words, and a typical A4 page is 500–800 tokens. Code and non-English text usually tokenize less efficiently. Tokenizers also differ per model family — Anthropic's current tokenizer, for instance, can produce noticeably more tokens for the same text than older Claude versions. **Q: How is API pricing different from ChatGPT Plus or Claude Pro?** A: Consumer subscriptions (ChatGPT Plus, Claude Pro) are flat monthly fees for using a chat app, with usage limits. API pricing is pay-per-token for building the models into your own products and workflows — there's no monthly fee, you pay exactly for what you process. Manus is a third pattern: credit-based plans where tasks consume credits. ### Claude API Pricing FAQs (https://compareai.today/faq/claude-pricing) **Q: How much does Claude Fable 5 cost?** A: Claude Fable 5 is priced at $10.00 per million input tokens and $50.00 per million output tokens — double Opus 5. With batch processing, this drops to $5.00/$25.00. It launched June 9, 2026 as the top tier above Opus. **Q: How much does Claude Opus 5 cost?** A: Claude Opus 5 is priced at $5.00 per million input tokens and $25.00 per million output tokens. With batch processing, this drops to $2.50/$12.50. It launched July 24, 2026 at exactly the same rate as the Opus 4.8 it replaces, so the upgrade costs nothing extra per token. Fast mode, if you enable it, is billed separately at $10.00/$50.00 per MTok. **Q: How much does Claude Opus 4.8 cost?** A: Claude Opus 4.8 is priced at $5.00 per million input tokens and $25.00 per million output tokens. With batch processing, this drops to $2.50/$12.50. It was superseded by Opus 5 in July 2026 but remains available at the same rate, as does Opus 4.7. **Q: How much does Claude Sonnet 5 cost?** A: Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens. The $2.00/$10.00 launch rate, originally introductory through August 31, 2026, is now the standard price: Anthropic confirmed on August 11, 2026 that the scheduled increase to $3/$15 will not take place. With batch processing, the rate drops to $1.00/$5.00. It's the current mid-range Claude model, succeeding Sonnet 4.6. **Q: What is the cheapest Claude model?** A: Claude Haiku 4.5 is the most affordable at $1.00/MTok input and $5.00/MTok output. With the batch discount, effective rates are $0.50/$2.50. **Q: Does Claude offer batch processing discounts?** A: Yes, all Claude models support batch processing with a 50% discount on both input and output tokens. **Q: What context window does Claude support?** A: Claude Fable 5, Opus 5, and Sonnet 5 support a 1M-token context window at standard per-token pricing, while Haiku 4.5 supports 200K — well beyond most legacy models. **Q: How do I estimate my Claude API costs?** A: Use our free calculator at compareai.today. Select a Claude model, enter your monthly token volume, and get an instant cost estimate in your preferred currency — all 18 currencies are supported. ### ChatGPT & OpenAI API Pricing FAQs (https://compareai.today/faq/chatgpt-pricing) **Q: How much does GPT-5.6 cost per token?** A: GPT-5.6 went GA on July 9, 2026 in three tiers: Sol at $5.00 per million input tokens and $30.00 per million output tokens, Terra at $2.00/$12.00, and Luna at $0.20/$1.20. On July 30, 2026 — three weeks after launch — OpenAI cut Luna by 80% (from $1.00/$6.00) and Terra by 20% (from $2.50/$15.00), leaving Sol unchanged. That makes Luna OpenAI's cheapest model and pushes Terra below the older GPT-5.4 on both input and output. All three carry a 1.05M-token context window. Above 272K tokens, input doubles and output rises 1.5x. GPT-5.6 also introduces a cache-write charge at 1.25x the input rate, where cached reads stay at 10% of input. **Q: How much does GPT-5.5 cost per token?** A: GPT-5.5 is priced at $5.00 per million input tokens and $30.00 per million output tokens. It was OpenAI's flagship until GPT-5.6 Sol superseded it in July 2026 at the same $5.00/$30.00 rate. The older GPT-5.4 is $2.50/$15.00. Batch processing offers a 50% discount on both. **Q: How much does GPT-5.5 Pro cost?** A: GPT-5.5 Pro, the highest-precision flagship variant, costs $30.00 per million input tokens and $180.00 per million output tokens — 6x standard GPT-5.5 and the most expensive model we track. Batch processing halves it to $15.00/$90.00. **Q: What is the cheapest ChatGPT API model in 2026?** A: GPT-5.4 Nano is the most affordable at $0.20 per million input tokens and $1.25 per million output tokens. With batch processing, this drops to $0.10/$0.625. **Q: Does OpenAI offer batch processing discounts?** A: Yes — the Batch API covers the GPT family, including the new GPT-5.6 tiers and GPT-5.5 Pro, at a 50% discount on both input and output tokens. Batch jobs are processed within 24 hours. **Q: How do I calculate my ChatGPT API costs?** A: Use our free calculator at compareai.today. Enter your estimated monthly token volume, select your model, and get an instant cost estimate in 18 currencies. **Q: What context window sizes do GPT models support?** A: The GPT-5.6 family (Sol, Terra, and Luna) supports a 1.05M-token context window — the largest in the GPT family — with 128K max output. GPT-5.5, GPT-5.5 Pro, and GPT-5.4 support 1M tokens. GPT-5.4 Mini and GPT-5.4 Nano support 400K tokens, and the reasoning-focused o4-mini supports 200K, though o4-mini is deprecated and OpenAI shuts down its API on October 23, 2026. ### Gemini, Grok & Other Provider FAQs (https://compareai.today/faq/other-providers) **Q: How much does Google Gemini cost?** A: Gemini 2.5 Pro, the GA flagship, is $1.25/$10.00 per MTok, and the newer Gemini 3.1 Pro Preview is $2.00/$12.00. At the budget end, Gemini 2.5 Flash-Lite starts at $0.10/$0.40. Gemini models offer 1M-token context windows, 50% batch discounts, and a generous free tier. **Q: How much does xAI Grok cost?** A: Grok 4.5, the current flagship, costs $2.00/$6.00 per MTok with a 500K context, rising to $4.00/$12.00 above 200K prompt tokens. It is xAI's priciest model, not its cheapest: the previous flagship Grok 4.3 is $1.25/$2.50 with a full 1M context and an unusually small output premium (2x input, where most providers charge 5x or more). Grok Build 0.1, the coding specialist, is $1.00/$2.00. xAI's batch discount is 20%, not the 50% common elsewhere, and it only covers Grok 4.3 and the Grok 4.20 variants. **Q: Why is DeepSeek so cheap?** A: DeepSeek V4 Flash costs $0.14/$0.28 per MTok and V4 Pro $0.435/$0.87 — after a 75% price cut was made permanent in May 2026. Aggressive inference-efficiency work lets DeepSeek undercut Western flagships by an order of magnitude, and cache hits drop input costs to fractions of a cent. There's no batch discount because the list price effectively is the discount. **Q: How is Meta Llama priced?** A: Llama models are open-weight, so you pay a hosting provider (DeepInfra, Together, and others) rather than Meta — and rates vary by host. Representative pricing runs from Llama 3.1 8B at $0.02/$0.04 per MTok to Llama 4 Maverick at $0.20/$0.80. Llama 4 Scout offers a 10M-token context window, the largest we track, at $0.10/$0.30. **Q: How does Perplexity Sonar pricing work?** A: Perplexity's Sonar models bundle live web search into the per-token price: Sonar is $1.00/$1.00 per MTok, Sonar Reasoning Pro $2.00/$8.00, and Sonar Pro $3.00/$15.00. There are no batch discounts, but you're paying for retrieval plus generation in one call rather than running your own search stack. **Q: What does Mistral cost?** A: Mistral Large 3, the open-weight 675B flagship, is $0.50/$1.50 per MTok — premium capability at budget-tier pricing. Mistral Small 4 is $0.15/$0.60 and the code-focused Codestral is $0.30/$0.90. Mistral doesn't offer batch discounts. **Q: Is Azure OpenAI cheaper than OpenAI direct?** A: No — list prices match. Azure GPT-5.5 costs $5.00/$30.00 per MTok, the same as OpenAI direct. Teams choose Azure for enterprise agreements, regional data residency, and compliance rather than price. Azure also carries some models at different lifecycle stages, like GPT-5.2 at $1.75/$14.00. **Q: How does Manus credit pricing work?** A: Manus is credit-based rather than per-token: the free plan includes a 1,000-credit starter pack plus 300 daily credits, and paid plans run from $20/month for 4,000 credits to $200/month for 40,000. A typical data-analysis task consumes about 200 credits; a full app build can use 900+. ### Comparing AI Models & Providers (https://compareai.today/faq/comparisons) **Q: Is GPT cheaper than Claude?** A: On input tokens, at list price, yes at most tiers: GPT-5.4 ($2.50) is 2x cheaper than Opus 5 ($5.00) and GPT-5.4 Nano ($0.20) is 5x cheaper than Haiku 4.5 ($1.00). Sonnet 5 ($2.00) now undercuts GPT-5.4 on input permanently, after Anthropic made its launch rate standard. Opus 5 ($25.00) is also cheaper on output than the flagship GPT-5.6 Sol ($30.00). **Q: Why would anyone choose Claude if GPT is cheaper?** A: Claude flagships now match GPT on context (1M tokens for Opus 5 and Sonnet 5), Opus 5 undercuts GPT-5.6 Sol on output ($25.00 vs $30.00), and Anthropic's Constitutional AI safety approach plus output quality on tasks like creative writing and nuanced analysis keep it competitive. **Q: Can I switch from Claude to GPT easily?** A: Yes, both use similar API formats. Check our migration guides for step-by-step instructions. Most switches require minimal code changes. **Q: Which has better context windows — GPT or Claude?** A: The flagships are now tied: GPT-5.5, GPT-5.4, Claude Opus 5, and Claude Sonnet 5 all offer 1M tokens, and the GPT-5.6 family edges ahead at 1.05M. Claude Haiku 4.5 has 200K, while GPT-5.4 Mini and Nano have 400K. **Q: How do I decide between GPT and Claude for my project?** A: Use our comparison table and calculator to model your specific workload. Consider: monthly token volume, required context window, quality requirements, and budget constraints. **Q: Which AI model is best for coding on a budget?** A: Strong budget coding options include Grok Build 0.1 ($1.00/$2.00 per MTok, a dedicated coding specialist), Mistral's Codestral ($0.30/$0.90), and Claude Haiku 4.5 ($1.00/$5.00). For the hardest problems, teams still reach for premium models like Claude Opus 5 or GPT-5.6 Sol — a common pattern is routing routine edits to a budget model and escalating complex work. ### Cutting AI Costs (https://compareai.today/faq/cost-optimization) **Q: What is the absolute cheapest AI API right now?** A: By input cost, Meta Llama 3.1 8B at ~$0.02/MTok input and ~$0.04/MTok output is the cheapest, with Llama 4 Scout next at $0.10 — though both are offered through third-party hosts, so pricing varies by provider. The cheapest from a major first-party API is Google Gemini 2.5 Flash-Lite at $0.10/MTok input and $0.40/MTok output, dropping to $0.05/$0.20 with batch processing. **Q: Is the cheapest AI API also the worst quality?** A: Not necessarily. Gemini 2.5 Flash-Lite and GPT-5.4 Nano perform well on most standard tasks. Quality differences mainly show up in complex reasoning, creative writing, and nuanced analysis. **Q: Can I mix budget and premium models?** A: Absolutely. Many teams use a routing strategy — budget models for simple queries and premium models for complex ones. This can cut costs by 60-80% while maintaining quality where it matters. **Q: Do budget models support batch processing?** A: Many do. Gemini 2.5 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5 all offer 50% batch discounts. DeepSeek, Meta Llama, and Mistral budget models don't — but their list prices are already among the lowest we track. **Q: How do I switch from a premium to a budget model?** A: Check our migration guides for step-by-step instructions on switching between providers. Most API formats are similar and require minimal code changes. **Q: Are there hidden costs with budget AI APIs?** A: Watch for output token costs (often 3-4x input costs), rate limits on free tiers, and potential quality trade-offs on complex tasks. Our comparison table shows all costs transparently. **Q: What's the fastest way to cut AI API costs?** A: Three levers, in order of typical impact: move non-urgent workloads to batch processing (an instant 50% off with most major providers), cache repeated prompt prefixes, and route simple requests to budget models. Our Spend Analyzer shows what you'd save by switching providers, and the Strategy Builder models multi-model routing. ### Using Our Tools (https://compareai.today/faq/tools) **Q: How do I use the cost calculator?** A: Pick one or more models, enter your expected monthly input and output token volumes, and toggle batch processing if your workload allows it. The calculator shows a monthly cost estimate per model, side by side, in any of 18 currencies. **Q: Can I compare specific models side by side?** A: Yes, use our comparison table to filter by provider, sort by price, and compare any models side by side. You can also use the visual charts for a graphical view. **Q: Do you track batch processing discounts?** A: Yes, we indicate which models support batch processing and factor each provider's actual discount into our calculator when you select the batch option — 50% at most providers, but 20% at xAI, and nothing at all on models their batch APIs exclude. **Q: Can I use this for enterprise budgeting?** A: Absolutely. Our calculator supports custom token volumes, batch processing toggles, and 18 currencies — perfect for building business cases and budget proposals. **Q: What does the Cost Per Seat tool do?** A: It estimates what AI costs per team member per month: choose a usage profile (light to heavy), a model, and team size, and it converts per-token prices into a per-seat monthly figure — useful for comparing API costs against flat-rate subscription seats. **Q: What does the Spend Analyzer do?** A: Enter your current provider, model, and monthly spend, and it shows what the same workload would cost on comparable models from other providers — including how much you'd save (or lose) by switching. **Q: What does the Strategy Builder do?** A: It models a multi-model routing strategy: you describe your workload mix (simple vs complex requests), and it estimates costs for routing each slice to an appropriate model tier instead of sending everything to one flagship. **Q: How do price alerts work?** A: The Price Alerts page tracks recent pricing changes across all providers and shows which potential changes the community is watching. There's no sign-up — it's a live changelog of the AI pricing market you can check any time. **Q: Can I see historical AI pricing trends?** A: Yes — the Pricing History explorer charts how model prices have moved month by month since early 2026, including launches and price cuts. It's the easiest way to see the deflation trend in AI inference costs. ## Site Information - URL: https://compareai.today - Built by [Digital Five](https://www.digitalfive.com). - Data updated: August 12, 2026 - Deployment: Cloudflare Pages (global edge) - No tracking cookies, no user accounts required