AI Observability

Run AI quality workflows

Use evaluations, scores, sessions, and users to improve AI behavior.

Quality work starts from production traces, then moves through repeatable checks and controlled model changes.

Quality surfaces

  • Scores: review score trends, failed checks, evaluator output, and quality drift across models, prompts, or applications.
  • Sessions and Users: investigate AI behavior by session or end user when debugging journeys rather than isolated requests.
  • Evaluations: run repeatable checks and review pass/fail patterns before changing production flows.

Review loop

  • Open Trace Explorer when a request is slow, failed, expensive, or user-visible.
  • Capture feedback and comments on the request instead of keeping notes outside the product.
  • Turn repeated failures into evaluation cases or score checks so future prompt or model changes can be tested.
  • Use Scores and Evaluations to confirm the fix improved quality without creating a new regression.

When to use each tool

NeedUse
Debug one bad answerTrace detail, spans, comments, and scores
Compare model behaviorPlayground and Evaluations
Create repeatable checksScores and Evaluations
Tune prompt behaviorPrompt Insights and Playground
Investigate by customer journeySessions and Users
Prepare review evidenceTrace metadata, comments, scores, and action history

Need more help?

Keep moving with the right next guide

If your team still runs into setup issues, empty dashboards, billing delays, callback failures, or alerting problems, continue with troubleshooting before re-running the entire onboarding flow.