The short version. OpenRouter passes provider pricing through with no markup on inference and charges 5.5% when you buy credits, which is not the same thing. Vercel AI Gateway charges no markup and no platform fee on tokens at all, and puts its money in optional add-ons. Cloudflare AI Gateway gives away analytics, caching and rate limiting, and charges 5% only if you use its unified billing. Portkey is free as open source and $49 a month hosted for 100,000 logs. LiteLLM is $0 and open source, and you are the one running it.
Every price and fee here came off the vendor’s own pricing or documentation page on 16 August 2026.
Calling three model providers from one codebase? Our directory keeps shortlists for teams running AI in production and reviews them on a schedule.
Nobody charges for the routing, so find where the fee actually sits
The pitch for every gateway is the same and it is a good pitch: one API instead of four, failover when a provider has a bad afternoon, caching, a spend cap, and logs. What surprises people is that the routing itself is essentially free everywhere. The money is somewhere else in each product, and it is somewhere different each time.
Read these four sentences carefully, because they are the whole comparison:
- OpenRouter takes 5.5% with a $0.80 minimum when you buy credits by card, and explicitly no markup on inference.
- Cloudflare gives the core gateway away and takes 5% on credit purchases only if you opt into unified billing.
- Vercel takes nothing on tokens, including with your own keys, and bills separately for custom reporting, allowlists and zero data retention.
- Portkey takes nothing on tokens and charges $49 a month for logs.
A fee on credit purchases is not a fee on your bill, and the difference matters. Five and a half percent applies when you top up, so it is one charge on the amount you load, not a tax on every request. It is still real money and it is not the same shape as a per-request markup.
The other thing worth naming before the list: what you are mostly paying for here is logs and governance, not routing. Every product on this page will route requests for nothing. Not one of them will store a month of logs for nothing beyond a small allowance.
1. OpenRouter, for reaching every model without a procurement round
No markup on inference; provider pricing is passed through. Credit purchases carry 5.5% with a $0.80 minimum by card, or 5% by crypto. Bringing your own provider keys is free up to a monthly allowance and then carries 5% on usage above $25,000 on pay-as-you-go. Verified 16 August 2026.
The practical value is not routing, it is access. One account and one key reaches hundreds of models from dozens of providers, including the ones you would otherwise be creating accounts and adding cards for on a Tuesday afternoon to try something.
For evaluation work that is transformative. Comparing six models on your own prompts stops being a week of account setup and becomes an afternoon, and the pricing model means you pay list price for the tokens either way.
Do the arithmetic on the credit fee rather than reacting to the percentage. Loading $1,000 costs $55. That is meaningful at scale and trivial while you are experimenting, which is exactly the phase most people are in when they reach for this.
Who it is wrong for: a team settled on one provider at volume, where you are adding a hop and a fee for a flexibility you have decided not to use. Also anyone whose compliance position requires a direct contract with the model provider.
2. Vercel AI Gateway takes nothing on tokens, including your own keys
No markup and no platform fee on tokens, on either the free or the paid tier, and no fee when bringing your own key. The free tier covers a subset of models with lower rate limits. Add-ons are charged separately: custom reporting at $0.075 per 1,000 tag, user ID or quota entity writes and $5 per 1,000 reporting queries; a team-wide provider allowlist at $0.10 per 1,000 successful requests; team-wide zero data retention at $0.10 per 1,000 requests. Per-request allowlist filtering and per-request zero data retention cost nothing extra. Verified 16 August 2026.
Zero markup with your own keys is the cleanest position on this page. You keep your provider relationship and your negotiated rates, and the gateway takes nothing for standing in the middle. Compare that to a 5% BYOK fee elsewhere and it is a genuine difference rather than marketing.
The add-on pricing rewards reading. The same capability costs nothing per request and $0.10 per 1,000 requests team-wide, for both the allowlist and zero data retention. Setting it per request instead of team-wide is free, and that is the sort of thing you find in the docs rather than on a comparison table.
The obvious catch is gravity. This is most natural if you are already on Vercel, and adopting it otherwise means a new vendor in your critical path for a feature the others also provide.
Who it is wrong for: teams with no other Vercel footprint, and anyone whose model catalogue needs are broad, since the free tier covers a subset rather than everything.
3. Cloudflare gives away the parts people expected to buy
The core gateway features are free, including dashboard analytics, caching and rate limiting. Persistent logs are free with limits by plan: 100,000 logs on Free and 10 million on Paid. Data loss prevention scanning is included on all plans. Unified billing carries a 5% fee on credit purchases with provider rates passed through at no markup. Guardrails bill as Workers AI token-based inference. Logpush is on the Workers Paid plan at 10 million a month plus $0.05 per million. Verified 16 August 2026.
Caching, rate limiting and analytics for nothing is a strong opening, and 100,000 persistent logs on a free plan is more than most teams generate in a month of a new feature.
The structure is honest in a way worth pointing out. The 5% is attached to unified billing, which is an optional convenience, so if you keep your own provider accounts you are using a capable gateway at no cost at all. That is a real free tier rather than a trial.
It ranks third rather than first because it is infrastructure rather than a product for the AI team. The dashboard thinks in requests and cache hit rates, not in prompts and model versions, and if you want to know why an answer got worse this is not where you will find out.
Who it is wrong for: teams wanting prompt-level debugging and evaluation. That is an observability purchase and this is not one, though it pairs well with something that is.
4. Portkey sells the logs, and gives away the gateway
Open source and self-hosted is free with no request limit and no overage permitted. Developer is free forever with 10,000 recorded logs a month, and requests above the threshold simply are not recorded. Production is $49 a month with 100,000 recorded logs and $9 a month for every additional 100,000, up to 3 million requests. Enterprise is custom. Verified 16 August 2026.
The pricing tells you exactly what the product is. Requests are unlimited and free; recorded logs are the meter. Portkey is selling observability and governance with routing attached, which is the opposite orientation from the three above.
For a team that wants one purchase rather than two, that is the argument. Guardrails, prompt management, virtual keys and per-team budgets sit alongside the routing, and $49 for 100,000 logs is competitive against buying a gateway and an observability tool separately.
The Developer tier’s behaviour is worth understanding before you rely on it: past 10,000 logs, requests keep working and stop being recorded. That fails quietly, which is the failure mode you least want in the tool you bought to see what is happening.
Who it is wrong for: teams that already run a dedicated observability tool, where you are paying for the half you have. Route through something free instead.
5. LiteLLM is free, and you are the operator
Open source at $0 with no credit card, including 140-plus provider integrations, logging to Langfuse, Arize Phoenix, LangSmith and OTEL, virtual keys, budgets and teams, load balancing with request and token rate limits, and guardrails. Enterprise adds support with custom SLAs, JWT auth, SSO and audit logs at pricing that is not published. Verified 16 August 2026.
Read that feature list again and note that it is the free tier. Virtual keys, per-team budgets, load balancing and rate limits are the governance features other products put behind a plan, and here they are in the open source project.
What you are taking on is operating it. A proxy in front of every model call is now a service you deploy, monitor, patch and page someone about at three in the morning. That cost is real and it is not on any pricing page.
It also logs to the observability tools rather than competing with them, which is the right architecture and makes it the natural pairing with self-hosted Langfuse or Phoenix if your requirement is that nothing leaves your network.
Who it is wrong for: small teams without an engineer who wants to own it. Free software with an on-call rotation is not free, and one of the hosted options above will cost you less than the first outage.
Situations this list does not cover
One provider, one model, modest volume. Call the API directly. A gateway is a dependency you have not earned yet.
Enterprise API management. If the requirement includes an existing API estate, developer portals and formal governance, this is a Kong or Apigee conversation and these are the wrong shortlist.
Self-hosted models. Serving your own weights is an inference infrastructure problem, and routing is the small part of it.
Anyone with no spend cap. Every product here can stop runaway spend and most teams have not configured it. Do that before comparing anything else on this page.
Also weighed
Helicone. Observability first with gateway features attached, at $79 a month on Pro. Worth comparing against Portkey if you want both jobs from one vendor.
Kong AI Gateway. The right answer when the AI traffic has to sit inside an API gateway you already run, and overkill when it does not.
Writing your own router. Two hundred lines and a weekend, and then failover, retries, streaming edge cases, token counting and a year of maintenance. Teams underestimate the tail of this reliably.
The provider SDKs directly. Still the fastest path for a prototype, and there is no shame in staying there until switching providers is a thing you actually need to do.
Questions we get
Is a 5% credit fee the same as a 5% markup?
No, and the difference is worth understanding. OpenRouter’s 5.5% applies when you top up your balance, not to each inference call, and inference itself is passed through at provider list price. It is one charge on the amount you load rather than a tax on usage.
Which gateway is genuinely free?
Cloudflare’s core gateway, if you keep your own provider accounts and skip unified billing, with 100,000 persistent logs on the free plan. LiteLLM is free as software, and Portkey’s open source edition is free with unlimited requests. Vercel takes nothing on tokens on either tier.
Do I need a gateway and an observability tool?
Sometimes not. Portkey and Helicone each cover both to a useful standard. Buy two only when you have a specific requirement neither one meets, because the overlap in caching, key management and logging is where teams quietly pay twice.
What breaks when the gateway goes down?
Everything downstream of it, which is the honest cost of the convenience. You have put a dependency in front of every model call. Check the failover behaviour and whether you can bypass it in an incident before you route production traffic through anything here.
Does bringing my own key avoid the fees?
On Vercel, yes, with no markup or fee on BYOK. On OpenRouter, BYOK is free up to a monthly allowance and then carries 5% on usage above $25,000 on pay-as-you-go plans. Read your own volume against that number before assuming BYOK is free.
What is the first thing to configure?
A spend cap, then caching. Most teams reach for a gateway after a surprising invoice and then spend their first week on model routing, which is the interesting problem rather than the expensive one.
Our maintained directory picks for AI cost and reliability cover what happens after the routing works, including attributing model spend to the customer who caused it.
Last reviewed: 16 August 2026. Every price and fee was read off the vendor’s own pricing or documentation page on that date. LiteLLM does not publish enterprise pricing and we have said so rather than estimating. Fees on credit purchases are stated as such and are not markups on inference. We review this page quarterly.
Leave a Reply