Score quality with repeatable tests
Systematic evaluation of model and agent outputs against datasets and judges.
Who it's for
AI teams shipping changes who need signal on quality beyond vibes.
Keep exploring
10 in catalog
Evals is featured by
-
Opik
Open-source LLM evaluation and observability
LLM Ops -
Weights & Biases
The AI developer platform
LLM Ops -
Langfuse
The open-source LLM engineering platform
LLM Ops -
LangSmith
Build trustworthy AI agents
LLM Ops -
Braintrust
Enterprise AI evaluation
LLM Ops -
Arize Phoenix
Open-source AI observability
LLM Ops -
DeepEval
The open-source LLM evaluation framework
LLM Ops -
promptfoo
Test your LLM app before you ship
LLM Ops -
W&B Weave
Observability and continuous improvement for production agents
LLM Ops -
Agenta
The open-source workspace for your agents
LLM Ops