Anthropic Pricing & API Models
Claude series developer, known for safety, long context, and coding excellence
35 paid models · 12 free · Price range: $0.25 - $30.00 /1M
Last updated: Aug 8, 2026
About Anthropic
Anthropic pricing helps developers compare model availability, token billing, context limits, and use-case fit. Anthropic was founded by former OpenAI researchers with a focus on AI safety. Their Claude models feature industry-leading 200K context windows, exceptional coding abilities, and the Constitutional AI safety framework. Claude 3.5 Sonnet is widely regarded as one of the best models for coding tasks.
Key Highlights
- 200K ultra-long context window
- Industry-leading code generation quality
- Constitutional AI safety framework
- Native PDF and image understanding
- Excellent instruction following
Pricing Features
- Pay-per-token billing
- Prompt Caching (discounted)
Supports Prompt Caching (up to 90% discount on cached tokens) and Batch API (50% off). Message Batches API for async processing.
API Features
Common Use Cases
- • Enterprise Chat Applications
- • Code Generation & Review
- • Long Document Analysis
- • Research Assistance
- • Content Writing
Anthropic pricing guide
Focused notes for developers comparing official pricing, API docs, token billing, and model fit.
Anthropic API pricing for Claude Sonnet, Opus, and Haiku
Use this page to compare Anthropic API pricing across Claude model families by input price, output price, context length, and supported capabilities. Sonnet-family models are commonly evaluated for coding, agent workflows, and long-document analysis, while Haiku-family models can be a better fit when latency and high-volume cost matter more than peak reasoning quality.
- Compare input and output token prices separately before choosing a Claude model.
- Use Sonnet models for code generation, software agents, and complex analysis.
- Use Haiku models for high-volume chat, classification, and lightweight extraction tasks.
- Check Opus-family models when maximum reasoning quality matters more than unit cost.
How to reduce Anthropic Claude API costs
Anthropic pricing can change materially when you use prompt caching or batch processing. Prompt caching is useful when the same system prompt, policy document, or knowledge-base context is reused across many requests. Batch processing is better for offline workloads that can wait for asynchronous completion. Verify current rates on the official Anthropic pricing documentation before launch.
- Use prompt caching for repeated instructions, long reference documents, and stable agent context.
- Use batch requests for offline analysis, enrichment, and non-real-time content workflows.
- Estimate real spend with your own input/output token ratio instead of comparing only headline prices.
How to read Anthropic models pricing
Anthropic models pricing should be compared across input rate, output rate, context length, and capability fit. A lower input price is not automatically cheaper when responses are long, prompts are repeated, or the workload needs a larger context window. Use the live model rows to compare the complete input/output mix.
- Compare Sonnet, Haiku, and Opus families by the workload they need to serve.
- Check output pricing carefully for code generation, long summaries, and agent traces.
- Use context length and caching eligibility as cost-planning variables, not just model features.
Anthropic API pricing: a quick budgeting example
For a simple forecast, multiply input tokens by the input rate and output tokens by the output rate, then add any applicable cached, batch, or enterprise adjustments. For example, compare a 500k-input/200k-output chatbot workload with the model rows above before choosing between Haiku, Sonnet, and Opus families.
This page covers API and token pricing. Claude consumer plans, Claude Code subscriptions, and regional contract terms are separate buying intents; verify those on Anthropic's official pages rather than mixing them into the API estimate.
Related search terms
📊 Anthropic Model Comparison
Compare all models side by side. Sorted by total price (input + output).
| Model | Tier | Input /1M | Output /1M | Total /1M | Context | Best For |
|---|---|---|---|---|---|---|
| Claude 3 Haiku | Budget | $0.25 | $1.25 | $1.50 | 200k | Image analysis |
| Claude Haiku 4.5 (batch) | Budget | $0.50 | $2.50 | $3.00 | 200k | Complex reasoning, math |
| Claude 3.5 Haiku | Budget | $0.80 | $4.00 | $4.80 | 200k | Image analysis |
| Claude Sonnet 5 (batch) | Budget | $1.00 | $5.00 | $6.00 | 1.0M | Complex reasoning, math |
| Anthropic Claude Haiku Latest | Budget | $1.00 | $5.00 | $6.00 | 200k | Complex reasoning, math |
| Claude Haiku 4.5 | Budget | $1.00 | $5.00 | $6.00 | 200k | Complex reasoning, math |
🎯 Which Anthropic Model Should You Choose?
Quick recommendations based on your use case.
💰 Anthropic Monthly Cost Examples
Estimated monthly costs for common use cases.
| Use Case | Monthly Usage | Claude 3 Haiku (Budget) |
Claude Opus 5 (batch) (Flagship) |
|---|---|---|---|
|
Customer Service Bot
1000 conversations/day
|
500k input 200k output |
$0.38/mo | $3.75/mo |
|
Code Assistant
200 requests/day
|
1.0M input 500k output |
$0.88/mo | $8.75/mo |
|
Data Analysis
500 analyses/day
|
2.0M input 300k output |
$0.88/mo | $8.75/mo |
⚔️ Anthropic vs Competitors
How does {brand} compare to other major AI providers?
| Brand | Model | Input /1M | Output /1M | Total /1M | Context | vs {brand} |
|---|---|---|---|---|---|---|
Anthropic
|
Claude Opus 5 (batch) Current | $2.50 | $12.50 | $15.00 | 1.0M | — |
OpenAI
|
Codex Mini | $1.50 | $6.00 | $7.50 | 200k | 50% cheaper |
OpenAI
|
GPT-5.2 Pro | $21.00 | $168.00 | $189.00 | 400k | 1160% more |
OpenAI
|
GPT-5.2 | $1.75 | $14.00 | $15.75 | 400k | 5% more |
OpenAI
|
GPT-5.1-Codex-Max | $1.25 | $10.00 | $11.25 | 400k | 25% cheaper |
OpenAI
|
GPT-5.1 | $1.25 | $10.00 | $11.25 | 400k | 25% cheaper |
OpenAI
|
GPT-5.1-Codex | $1.25 | $10.00 | $11.25 | 400k | 25% cheaper |
All Models
❓ Anthropic Pricing FAQ
What is the cheapest Anthropic model?
The cheapest Anthropic model is Claude 3 Haiku at $1.50 per 1M tokens (input + output combined).
What is the maximum context length for Anthropic models?
Anthropic models support up to 1.0M context length, allowing you to process large documents and maintain long conversations.
How do I choose between Anthropic models?
For budget projects, choose the cheapest model. For code generation, prioritize low output price. For complex reasoning, choose models with reasoning capability. Use our scenario guide above.
Does Anthropic support prompt caching?
Yes, Anthropic supports prompt caching which can significantly reduce costs for repeated prompts. Check individual model pages for caching prices.
How does Anthropic pricing work?
Anthropic API pricing is usually evaluated separately for input and output tokens, with model family, context usage, caching, batching, and deployment channel affecting the final estimate. Use the model table and calculator for a workload-specific comparison.
Does Anthropic API pricing include separate input and output token prices?
Yes. For budgeting, treat input and output tokens separately. Long prompts or documents increase input cost, while code generation, summaries, and agent traces usually increase output cost.
When should I use Anthropic prompt caching?
Use prompt caching when many requests share the same long prompt, policy text, examples, or retrieval context. It can reduce repeated-context cost, but it is less useful for one-off prompts.
Does Anthropic offer regional or enterprise pricing?
Regional availability, enterprise terms, and contract discounts can differ from public API rates. Treat the public model rows as a planning baseline, then verify the applicable region, account terms, and official pricing page before committing spend.
Does Anthropic pricing cover Claude Code or consumer plans?
This page focuses on public Anthropic API model pricing. Claude Code, consumer subscriptions, regional offers, and enterprise contracts can use separate plans or terms, so verify those products on the official Anthropic pages instead of mixing them into an API token estimate.