Haiku 5.5 Pricing: Token Rates & Cost Examples
Haiku 5.5 costs $0.10 per million input tokens and $0.50 for output up to 100K input tokens. Above 100K, rates rise to $0.50 input and $2.50 output.
Prices checked October 10, 2026. USD per million tokens unless stated otherwise. Rates below are for Anthropic's standard Claude API with global routing.
Haiku 5.5 pricing at a glance
| Token category | Prompt ≤100K | Prompt >100K |
|---|---|---|
| Standard input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache read | $0.01 | $0.05 |
| Cache write: 5 minutes | $0.125 | $0.625 |
| Cache write: 1 hour | $0.20 | $1.00 |
How the 100K threshold works
The threshold is evaluated separately for each request. Count all input tokens, including cache reads and cache writes. At exactly 100,000 input tokens, use the lower tier. Above 100,000, use the higher tier for that request's token categories, including output—not just the input tokens above the threshold. Earlier requests are not repriced.
For example, 90,000 cached input tokens plus 20,000 uncached input tokens gives a 110,000-token prompt. A cache hit does not move that request into the lower tier. Official long-context rules.
Calculate your API cost
For a request without caching:
Cost = input tokens / 1,000,000 × input rate + output tokens / 1,000,000 × output rate.
These are worked estimates, not measured API runs. They exclude tools, regional surcharges, taxes, retries, and negotiated discounts.
Short request: $0.0007
With 2,000 input tokens and 1,000 output tokens:
- Input: 2,000 / 1,000,000 × $0.10 = $0.0002.
- Output: 1,000 / 1,000,000 × $0.50 = $0.0005.
- Total: $0.0007. At identical usage, 10,000 requests cost $7.
Long request: $0.065
With 120,000 input tokens and 2,000 output tokens, use the higher tier:
- Input: 120,000 / 1,000,000 × $0.50 = $0.06.
- Output: 2,000 / 1,000,000 × $2.50 = $0.005.
- Total: $0.065. At identical usage, 1,000 requests cost $65.
Cached long request: $0.017
For the 110,000-token prompt described above, assume 90,000 cache-read tokens, 20,000 uncached input tokens, no new cache writes, and 1,000 output tokens:
- Cache read: 90,000 / 1,000,000 × $0.05 = $0.0045.
- Uncached input: 20,000 / 1,000,000 × $0.50 = $0.01.
- Output: 1,000 / 1,000,000 × $2.50 = $0.0025.
- Total: $0.017 for this request. The earlier request that created the cache has its own cost.
Keep cache reads, cache writes, and uncached input in separate buckets when estimating charges; do not charge the same token under multiple input categories.
Haiku 5.5 Batch API pricing
Batch processing is asynchronous and offers a 50% discount on input and output token rates:
| Batch token category | Prompt ≤100K | Prompt >100K |
|---|---|---|
| Input | $0.05 | $0.25 |
| Output | $0.25 | $1.25 |
For the short-request example above, batch token charges are $0.00035 per request, or $3.50 for 10,000 identical requests. Batch pricing reference.
Haiku 5.5 vs. Haiku 4.5
Haiku 4.5's standard input/output rates are $1/$5 per million tokens. Haiku 5.5's lower tier is 90% cheaper per token; its higher tier is 50% cheaper. Task-level savings can differ because tokenization and usage change. Anthropic's comparison.
What else should you budget for?
The examples assume a fixed request and output size. A production workflow may repeat a request, retain more conversation history, or produce longer answers than expected. For planning, track the full cost of a successfully completed task rather than the cost of a single attempt.
Test representative jobs and record token usage by category, accepted results, latency, retries, and manual corrections. For a script-writing workflow, divide the total expense by the number of drafts you can actually use. A cheaper request is not necessarily a cheaper finished script.
Image and video generation need their own budget. This article estimates the text-processing portion of a workflow; it does not claim that Kavello provides the Haiku API.
Common pricing questions
Does a cache hit avoid the higher tier?
No. Cached input still counts toward prompt length. Use the cache-read rate for the applicable tier.
Are these prices the same on every provider?
Check the provider that invoices you. Partner platforms and regional settings can have different rates; these examples use standard global Claude API pricing. Platform pricing details.
Is $7 a guaranteed monthly cost?
No. It is the arithmetic result for 10,000 requests containing exactly 2,000 uncached input tokens and 1,000 output tokens each, using the standard lower tier. Your workload and additional charges determine your bill.