Skip to content

Pricing

Pay for tokens. That's the whole bill.

Every model on this page costs 20–70% less through Inferray than it does at list price. No subscription, no seat count, no monthly minimum — you top up a balance and spend it down.

The math

Same tokens, smaller number at the bottom.

Nothing about your workload has to change to get this. Move the slider to see what your current spend looks like on Inferray prices.

$2,000/mo
You keep / month$400–$1,400
You keep / year$4,800–$16,800

Estimate only. Your actual saving is the per-model discount in the table below, and the higher end assumes a cache-heavy workload.

Every model, every price

The full list

Prices are per 1M tokens. The struck-through number is what the model costs direct; the bold one is what you pay.

Anthropic

4 models
ModelContextInputCached inputOutputYou save
Claude Fable 5claude-fable-51M$10$7.50$1$0.75$50$37.5025%
Claude Opus 4.8claude-opus-4-81M$5$3.75$0.50$0.38$25$18.7525%
Claude Sonnet 5claude-sonnet-51M$2$1.50$0.20$0.15$10$7.5025%
Claude Haiku 4.5claude-haiku-4-5200K$1$0.75$0.10$0.08$5$3.7525%

OpenAI

3 models
ModelContextInputCached inputOutputYou save
GPT-5.5gpt-5.51.1M$5$4.00$0.50$0.40$30$24.0020%
GPT-5.4gpt-5.41.1M$2.50$2.00$0.25$0.20$15$12.0020%
GPT-5.2gpt-5.2400K$1.75$1.40$0.17$0.14$14$11.2020%

Google

1 model
ModelContextInputCached inputOutputYou save
Gemini 2.5 Progemini-2.5-pro1.0M$1.25$1.06$0.13$0.11$10$8.5015%

What's not on the invoice

Four things you never get charged for

No seats

Invite the whole team. You pay for tokens, not headcount.

No subscription

Nothing renews and nothing lapses. Top up when you want to.

No minimum

Spend $5 this month and $5,000 the next. Same price either way.

No support tier

The discount is the same whether you send one request or a million.

Before you ask

Questions worth answering

Is it the same model I'd get direct?

Yes. Same model, same version, same answers. The only thing that changes is what it costs you.

Where does the discount come from?

We buy inference in volume and hand the difference back. There's no trial window and no tier to reach — the discount is on your first request and every one after it.

Why is the range so wide?

The flat discount per model is the floor. Cached input is priced far below fresh input, so workloads that reuse a long system prompt or a big document land near the top of the range.

What happens when my balance runs out?

Requests stop rather than silently overspending. Turn on auto top-up in the console and your balance refills before that happens.

Do I get an invoice?

Every top-up produces a receipt, and the console breaks spend down by project, key, and model so you can see where it went.

Start on the lower price today

Create a key, top up, and your very next request is billed at the Inferray rate.