AI Observability
Run AI quality workflows
Use evaluations, scores, sessions, and users to improve AI behavior.
Quality work starts from production traces, then moves through repeatable checks and controlled model changes.
Quality surfaces
- Scores: review score trends, failed checks, evaluator output, and quality drift across models, prompts, or applications.
- Sessions and Users: investigate AI behavior by session or end user when debugging journeys rather than isolated requests.
- Evaluations: run repeatable checks and review pass/fail patterns before changing production flows.
Review loop
- Open Trace Explorer when a request is slow, failed, expensive, or user-visible.
- Capture feedback and comments on the request instead of keeping notes outside the product.
- Turn repeated failures into evaluation cases or score checks so future prompt or model changes can be tested.
- Use Scores and Evaluations to confirm the fix improved quality without creating a new regression.
When to use each tool
| Need | Use |
|---|---|
| Debug one bad answer | Trace detail, spans, comments, and scores |
| Compare model behavior | Playground and Evaluations |
| Create repeatable checks | Scores and Evaluations |
| Tune prompt behavior | Prompt Insights and Playground |
| Investigate by customer journey | Sessions and Users |
| Prepare review evidence | Trace metadata, comments, scores, and action history |
Need more help?
Keep moving with the right next guide
If your team still runs into setup issues, empty dashboards, billing delays, callback failures, or alerting problems, continue with troubleshooting before re-running the entire onboarding flow.