Comparison · LLM Observability

LLM Observability & Evaluation Platforms Compared

An open, side-by-side comparison of LLM observability and evaluation platforms, the tracing-and-scoring layer that tells you whether your agents actually work in production. Capability facts are public and drawn from each provider; the rank is decided by the engineers who live in the dashboards, never by payment.

The comparison facts on this page are open and drawn from each provider. The rank is decided by community votes, never by payment.
12 of 12 providers
How to read thisEvery value comes straight from the provider's own site.Contact sales means they do not publish a price.Not disclosed or means the spec is not stated on their site.We never fill blanks with guesses.
#ProviderFocusOpen sourceSelf-hostableFree tierRatingVote
1BR
Braintrust
US
Evals
No reviews yet
2GA
Galileo
US
Evals
No reviews yet
3CA
Confident AI
US
Evals
No reviews yet
4LA
LangSmith
US
Observability
No reviews yet
5LA
Langfuse
DE
Observability
No reviews yet
6AP
Arize Phoenix
US
Observability
No reviews yet
7HE
Helicone
US
Observability
No reviews yet
8W&
Weights & Biases Weave
US
Observability
No reviews yet
9CO
Comet Opik
US
Observability
No reviews yet
10TR
Traceloop
US
Observability
No reviews yet
11HO
HoneyHive
US
Observability
No reviews yet
12PR
PromptLayer
US
Prompt management
No reviews yet
How ranking works: facts are open; the order settles on community votes. See the full standings → Browse the directory →