Pricing
Pay for tokens. That's the whole bill.
Every model on this page costs 20–70% less through Inferray than it does at list price. No subscription, no seat count, no monthly minimum — you top up a balance and spend it down.
The math
Same tokens, smaller number at the bottom.
Nothing about your workload has to change to get this. Move the slider to see what your current spend looks like on Inferray prices.
Estimate only. Your actual saving is the per-model discount in the table below, and the higher end assumes a cache-heavy workload.
Every model, every price
The full list
Prices are per 1M tokens. The struck-through number is what the model costs direct; the bold one is what you pay.
Anthropic
4 models| Model | Context | Input | Cached input | Output | You save |
|---|---|---|---|---|---|
| Claude Fable 5claude-fable-5 | 1M | $10$7.50 | $1$0.75 | $50$37.50 | −25% |
| Claude Opus 4.8claude-opus-4-8 | 1M | $5$3.75 | $0.50$0.38 | $25$18.75 | −25% |
| Claude Sonnet 5claude-sonnet-5 | 1M | $2$1.50 | $0.20$0.15 | $10$7.50 | −25% |
| Claude Haiku 4.5claude-haiku-4-5 | 200K | $1$0.75 | $0.10$0.08 | $5$3.75 | −25% |
OpenAI
3 models| Model | Context | Input | Cached input | Output | You save |
|---|---|---|---|---|---|
| GPT-5.5gpt-5.5 | 1.1M | $5$4.00 | $0.50$0.40 | $30$24.00 | −20% |
| GPT-5.4gpt-5.4 | 1.1M | $2.50$2.00 | $0.25$0.20 | $15$12.00 | −20% |
| GPT-5.2gpt-5.2 | 400K | $1.75$1.40 | $0.17$0.14 | $14$11.20 | −20% |
| Model | Context | Input | Cached input | Output | You save |
|---|---|---|---|---|---|
| Gemini 2.5 Progemini-2.5-pro | 1.0M | $1.25$1.06 | $0.13$0.11 | $10$8.50 | −15% |
What's not on the invoice
Four things you never get charged for
No seats
Invite the whole team. You pay for tokens, not headcount.
No subscription
Nothing renews and nothing lapses. Top up when you want to.
No minimum
Spend $5 this month and $5,000 the next. Same price either way.
No support tier
The discount is the same whether you send one request or a million.
Before you ask
Questions worth answering
Is it the same model I'd get direct?
Yes. Same model, same version, same answers. The only thing that changes is what it costs you.
Where does the discount come from?
We buy inference in volume and hand the difference back. There's no trial window and no tier to reach — the discount is on your first request and every one after it.
Why is the range so wide?
The flat discount per model is the floor. Cached input is priced far below fresh input, so workloads that reuse a long system prompt or a big document land near the top of the range.
What happens when my balance runs out?
Requests stop rather than silently overspending. Turn on auto top-up in the console and your balance refills before that happens.
Do I get an invoice?
Every top-up produces a receipt, and the console breaks spend down by project, key, and model so you can see where it went.
Start on the lower price today
Create a key, top up, and your very next request is billed at the Inferray rate.