A team routing $1,000 a month of tokens through most AI gateways gives up $600 to $2,400 a year on a charge that never appears as a line item. The rate the provider publishes is one layer of what you pay; the markup layer between you and that rate is the other, and for most teams it is the larger and the less visible of the two. OrcaRouter quotes providers’ published rates directly and adds nothing on top of token costs, which is why any cheapest llm api pricing comparison only means something once both layers are on the table.
All prices and latency figures read 2026-09-15.
Two numbers decide every token bill
When you compare LLM providers, you are usually comparing a single number: the price per million tokens for the model you care about. That number is real, it is published, and it is how the market lets you shop. It is just not the whole price.
The whole price is the provider’s rate plus whatever layer sits between you and the provider. On most AI gateways that layer is a margin on every token that passes through. Our own pricing page characterises the typical range at 5% to 20% of token spend — wider, for most teams, than the entire gap between the cheapest and the mid-priced models in the catalog. A cheaper model saves you fractions of a cent per request. A 20% margin costs you a fifth of your bill, whatever the model.
The two numbers are also easy to confuse, which is how the layer hides. If a gateway quotes you $0.18 for a model a provider sells at $0.15, you see one number and reasonably assume it is the model price. It is the model price plus margin, but nothing on the page tells you the split. The quoted price is doing double duty: it is both the number you compare and the number that conceals the margin, and nobody has an incentive to separate them for you.
What 5–20% adds up to over a year
The markup reads as a small percentage and stops reading that way the moment you annualise it.
| Monthly token spend | At a 5% markup | At a 10% markup | At a 20% markup |
| $250 | $150 a year | $300 a year | $600 a year |
| $1,000 | $600 a year | $1,200 a year | $2,400 a year |
| $5,000 | $3,000 a year | $6,000 a year | $12,000 a year |
Those figures are the markup alone, on top of the token price. A team spending $1,000 a month is not paying $12,000 a year for tokens; it is paying that plus between $600 and $2,400 a year for the privilege of passing those tokens through a middle layer. Because the margin is quoted into the rate rather than listed separately, it is almost impossible to recover later by shopping — you cannot compare a number you cannot see.
The point is not that every gateway marks up 20%. The point is that you have no way to tell, and the range you are guessing inside is larger than the difference between most cheap models. That is what makes the markup layer the biggest invisible line in the bill: it is larger than the optimisations you can see, and it is the one you cannot price-check.
Why does the markup layer stay invisible?
The markup is invisible for a structural reason, not an accidental one. A gateway that adds a margin has no incentive to show the provider’s rate beside its own; the difference is its business model. So the published rate absorbs the margin, and every comparison table on the internet — including the ones that found you the cheap model — is comparing prices that each quietly include a different, unpublished add-on.
That is the specific failure this whole class of comparison questions runs into. Model shopping is a solved problem: model rates are published, comparable and indexed, so the market has compressed them. The gateway layer is not, because the margin is neither published nor comparable. The money sits in the layer nobody lists.

How to check what you are actually paying
You do not have to take anyone’s word for what their margin is, because the check takes about ninety seconds.
- Pick a model you already use and note its input and output rate on the provider’s own pricing page.
- Find the same model on the gateway you are evaluating.
- Compare the two numbers — not “is the gateway competitive”, but whether they are identical.
A gap of any size is the gateway’s margin, expressed as a rate. An identical pair means the gateway is genuinely pass-through, and the claim is worth trusting because you can verify it again whenever a provider reprices: a table generated live from the same catalog the model pages read will move with it.
Here is a concrete example from our own catalog, read on the same day. DeepSeek’s official pricing page publishes off-peak rates of $0.15 per million input and $0.60 per million output for the model it serves as `deepseek-flash`; our model page for `deepseek-v4.1-flash` lists exactly $0.15 / $0.60. The two match to the cent. That is what pass-through looks like — not a competitive price that happens to be close, but the provider’s own number reproduced.

Zero is a business model, not a discount
When a provider says it adds nothing on top of token costs, the useful question is not whether you believe it but how it can be true — because the answer decides whether the zero survives contact with reality.
On OrcaRouter, token usage is pass-through: the platform is funded by the optional Team and Enterprise subscriptions, which sell seats, compliance enforcement and reporting, rather than by a margin on your spend. The free Hacker tier is the full gateway — free forever, three API keys, no token markup. That structure matters more than the number. A discount is a promotion: it exists at the vendor’s discretion and gets withdrawn when the unit economics change. A structural zero is a business model: it lasts only as long as the thing being sold is genuinely something else, and it is checkable, which is the property you actually want in a billing partner.
What to take from this
The rate you compare is the visible half of your bill; the markup layer is the invisible half, and it is where a 5% to 20% swing lives. For most teams that swing is larger than anything achievable by switching from one cheap model to another — so the practical order is to check the layer first, then the rate, then the latency. The first check takes ninety seconds, and most teams never run it.
Sourcing note: The 5–20% markup range is how OrcaRouter’s own pricing page characterises common gateway practice; it is a statement about the market, not an independent survey, and the comparison check above is how to establish the actual figure for any specific gateway. Model rates and median time-to-first-token figures are OrcaRouter’s own catalog and production telemetry, read 2026-09-15; latency is a 7-day rolling window reflecting one platform’s traffic mix, regions and live provider load, not a controlled benchmark. DeepSeek’s off-peak rates are from DeepSeek’s official pricing documentation (api-docs.deepseek.com/quick_start/pricing), vendor-reported, read 2026-09-15; the exact match between that page and the rate shown for deepseek-v4.1-flash is our own observation on the same date.



