The short version. Langfuse is $29 a month for the Core plan and free forever if you run it yourself, which makes it the sensible default. Arize Phoenix is open source and local-first, and its hosted sibling starts free at 25,000 spans with a $50 Pro tier. LangSmith is $39 a seat and is the obvious answer if you already build on LangChain. Helicone at $79 has the easiest integration on this page and doubles as a gateway. Braintrust at $249 is priced for teams whose problem is evaluation quality rather than cost visibility.
Every price here came off the vendor’s own pricing page on 16 August 2026.
Running a model call on every request? Our directory keeps shortlists for teams running AI in production, reviewed on a schedule.
Five products, five different billing units
The reason this category is hard to shop for is not that the products differ. It is that no two of them charge for the same thing, so the prices are not comparable until you decide what your own unit is.
- Langfuse bills in “units”: 50,000 free a month, then $8 per 100,000.
- LangSmith bills per seat at $39, plus traces, plus its own compute units at $1.50 and storage units at $1.00.
- Braintrust bills in processed gigabytes at $4, plus scores at $2.50 per 1,000, plus model credits.
- Helicone bills in requests and gigabytes of storage.
- Arize bills in spans and gigabytes.
The spread that produces is wider than any feature difference. Langfuse Core is $29 a month. Braintrust Pro is $249. That is not a premium, it is a different question being answered, and if you compare them on a feature grid you will conclude one of them is overpriced by eight times, which is not what is happening.
So we ranked on cost at the volume a small team actually generates, on whether there is a free self-hosted escape route, and on whether the tool helps you answer the only question that matters at two in the morning: this answer got worse, and what changed.
1. Langfuse is the default, and the escape hatch is real
Hobby is free with 50,000 units a month. Core is $29 a month with 100,000 units included. Pro is $199 a month. Enterprise is $2,499 a month. Additional units run $8 per 100,000, falling to $7 above a million, $6.50 above ten million and $6 above fifty million. Langfuse is open source and can be self-hosted for free. Verified 16 August 2026.
Two things put this first. The hosted price is the lowest serious option here at $29, and the open source version is not a crippled demo. You can run the whole thing on your own infrastructure at no licence cost, which changes the negotiation permanently.
That matters more than it sounds in a category this young. Observability tools accumulate your entire history of prompts, outputs and traces, and a vendor holding that has leverage over you at renewal. The self-host path means the worst case is an infrastructure project rather than a migration under pressure.
The graduated overage rates are also published in full, which is rarer here than it should be. You can model a year of growth on the pricing page without contacting anyone.
Who it is wrong for: teams whose problem is evaluation rigour rather than visibility. Langfuse traces and scores well, and if your day is spent building test suites and comparing model versions, Braintrust is built around that and this is not.
2. Arize Phoenix, if you would rather it never left your machine
Phoenix is the open-source, local-first platform for tracing, evaluation, experimentation and prompt iteration, described by Arize as somewhere to “stay local, stay open, and move to Arize AX only when you need it”. AX Free is free with 25,000 spans and 1 GB a month. AX Pro is $50 a month with 50,000 spans and 10 GB. AX Enterprise is custom. Verified 16 August 2026.
Phoenix is the answer to a question the rest of this page mostly ignores, which is whether your prompts and outputs should leave your network at all. For anyone handling regulated data, that is not a preference, it is the deciding constraint, and a local-first tool removes an entire vendor review.
The hosted ladder is priced sensibly too. AX Free at 25,000 spans is enough to instrument a real feature and see whether the data changes any decisions, and AX Pro at $50 is half of Helicone’s Pro tier.
What you give up relative to Langfuse is a smaller hosted free tier, and relative to LangSmith a weaker story if your application is built on LangChain primitives.
Who it is wrong for: a small team with no appetite for running anything. Local-first is a benefit only if somebody is willing to own the local part, and if nobody is, you are buying the hosted tier anyway and the ranking above applies.
3. LangSmith charges per seat, which is the thing to check
Developer is $0 per seat with up to 5,000 base traces a month, then pay-as-you-go. Plus is $39 a seat a month with up to 10,000 base traces, then pay-as-you-go. Enterprise is custom. LangChain Compute Units are $1.50 each and LangChain Storage Units are $1.00. Per-trace overage rates are not stated on the pricing page. Verified 16 August 2026.
If your application is already built on LangChain, this is the least friction available. The tracing is native, the abstractions line up, and nobody has to instrument anything by hand.
The per-seat model is the part to think through, because it is the only one here that scales with your team rather than your traffic. Six engineers on Plus is $234 a month before a single trace of overage, against $29 for Langfuse Core at any headcount. If everybody needs access, this becomes the most expensive option on the page other than Braintrust.
The Developer tier at $0 with 5,000 traces is a genuinely useful free plan for one person prototyping, and it is worth knowing the compute and storage units exist before you model costs, because they are separate line items from the seat.
Who it is wrong for: larger teams, and anyone not using LangChain. Outside that ecosystem the native advantage disappears and you are paying per seat for parity.
4. Helicone is one line of code away
Hobby is free with 10,000 requests and 1 GB of storage. Pro is $79 a month with the same 10,000 requests and 1 GB free, then usage-based. Team is $799 a month. Enterprise is contact us. The pricing page references a calculator but does not publish per-request or per-token overage rates. Verified 16 August 2026.
Helicone works as a proxy, so integration is changing a base URL rather than instrumenting your code. When the goal is finding out today what your model spend looks like broken down by user, that difference is worth real money.
It is also a gateway, with caching, rate limiting and key management, which means for some teams it replaces two purchases. That is the strongest argument for it and it is why it is on this list despite the price.
It ranks fourth because $79 is steep next to Langfuse at $29 and Arize AX Pro at $50, and because the overage rates are not published. For a usage-based product, a pricing page with a calculator instead of a rate is a page that has decided you should talk to somebody, and we do not estimate numbers a vendor has chosen not to print.
Who it is wrong for: teams that need to model costs precisely before committing, and teams already running a separate gateway, where you are paying for an overlap.
5. Braintrust is an evaluation product with observability attached
Starter is $0 with $10 of credits, 1 GB of processed data and $4 per GB after, 10,000 scores and $2.50 per 1,000 after, and 14-day retention. Pro is $249 a month with $249 of credits, 5 GB of processed data and $3 per GB after, 50,000 scores and $1.50 per 1,000 after, and 30-day retention plus $0.50 per GB per month. Enterprise is custom. Verified 16 August 2026.
Note what it meters. Scores, not traces. That tells you what the product is for: running evaluations systematically, comparing model and prompt versions against a test set, and knowing whether a change made things better before it ships.
That is a real discipline and most teams are not doing it. If yours is, and if a regression in answer quality costs you customers rather than embarrassment, $249 is defensible and the cheaper tools on this page will not replace it.
If yours is not, this is eight times the price of Langfuse for a capability you will admire and not use. The honest test is whether anyone on the team currently maintains an evaluation set. If nobody does, buying this will not create the habit.
Who it is wrong for: teams whose actual question is what the model is costing and where the latency went. That is cost and trace visibility, and everything above does it for less.
Teams this ranking does not serve
Anyone with one prompt and no traffic. Log the request and response to your existing logging. You do not have an observability problem yet.
Regulated environments with a hard data boundary. The decision starts and ends with what can be self-hosted, which narrows this to Langfuse, Phoenix or LiteLLM’s logging and skips the comparison entirely.
Large platform teams. Above a certain scale this becomes an OpenTelemetry conversation inside the observability stack you already run, not a new vendor.
Anyone who will not look at the traces. Instrumenting everything and reading none of it is the most common outcome in this category. Pick one question you cannot currently answer, and buy the cheapest thing that answers it.
Considered, not ranked
Weights and Biases Weave. Strong on experiment tracking and a natural fit for teams already using W&B for model training. A different centre of gravity from the five above.
OpenTelemetry with your existing tools. Most of these emit OTEL. If you already run Grafana or Datadog, cheap instrumentation into what you have may beat a new subscription, and it is the option nobody selling you a dashboard will mention.
Your provider’s own dashboard. Free, already there, and shows spend by API key. If your only question is what you are spending in total, stop here.
Portkey. A gateway with observability attached rather than the reverse, and worth comparing if you want both jobs in one purchase.
Questions engineers ask us
What is the cheapest real option?
Self-hosted Langfuse or Phoenix, both free and both open source, if you are willing to run them. Hosted, Langfuse Core at $29 is the lowest serious price, with Arize AX Pro at $50 next.
Why is Braintrust eight times Langfuse?
Because it is not selling the same thing. It meters scores rather than traces, which is a tell: it is an evaluation platform, and evaluation is a discipline you either practise or do not. If nobody on your team maintains a test set, the cheaper tools cover what you actually need.
How do I compare these on price?
Decide your own unit first, then convert. Pick a representative month of traffic, count requests, traces or spans depending on the product, and price each one against that. Comparing plan tiers directly will mislead you, because no two of these meter the same event.
Does per-seat pricing matter?
Only LangSmith charges that way, and it matters a lot above three or four people. Six engineers is $234 a month before overage, where every other tool here is priced on traffic and indifferent to headcount.
Should I self-host?
If you have a data residency requirement, yes, and the decision is already made. Otherwise treat it as leverage rather than a plan. Knowing you could move is worth something at renewal even if you never do.
Do I need this and a gateway?
Often not. Helicone and Portkey each do both jobs to a useful standard, and buying one of each is a common way to pay twice for caching and key management. Work out which half is your real problem before buying either.
See our maintained directory picks for AI cost and reliability, including tools that attribute model spend to a specific customer and plan.
Last reviewed: 16 August 2026. Every price was read off the vendor’s own pricing page on that date. Helicone does not publish per-request overage rates and LangSmith does not publish per-trace overage rates; we have said so rather than estimating. We review this page quarterly.
Leave a Reply