LLM Observability
10 providers in this category.
LLM observability platforms trace, evaluate, and monitor applications built on language models: logging each request with cost and latency, running quality evals, versioning prompts, and sometimes routing traffic through a model gateway. Compare tracing depth, eval workflows, prompt management, gateway support, open-source and self-hosting options, and free tier. Pricing is often seat-plus-usage, so confirm rates.
Companies
Neutral ordering. Not a recommendation.
Arize
arize.com
Arize is an AI observability company whose AX platform monitors, traces, and evaluates both traditional ML models and LLM applications.
Braintrust
www.braintrust.dev
Braintrust is an evaluation-first platform for teams shipping AI products, centered on systematic evals with datasets and scoring functions.
Confident AI
www.confident-ai.com
Confident AI is the commercial evaluation and observability platform from the creators of DeepEval, the Apache 2.0 open-source LLM evaluation framework.
Galileo
galileo.ai
Galileo is an evaluation and observability platform for generative AI applications and agents, oriented toward testing, monitoring, and guardrailing model behavior in production.
Helicone
www.helicone.ai
Helicone is an open-source LLM observability platform built around a logging proxy and AI gateway, Apache 2.0 licensed and free to self-host.
HoneyHive
www.honeyhive.ai
HoneyHive is an observability and evaluation platform for LLM applications and agents, unifying tracing, offline evaluation against test datasets, human review queues, and experimentation.
Langfuse
langfuse.com
Langfuse is an open-source, MIT-licensed LLM engineering platform covering tracing, evaluation, prompt management, and usage analytics.
LangSmith
www.langchain.com
LangSmith is the observability and evaluation platform from LangChain for teams building LLM applications, with or without the LangChain framework.
Portkey
portkey.ai
Portkey is an AI gateway with observability built in: applications route model calls through it to gain retries, fallbacks, load balancing, caching, and guardrails across 250+ providers.
