Total latency
Shows whether the request is slow.
Stage profiler
Shows which stage is slow and how its share of the total changed.
Why one end-to-end number cannot tell an AI operations team where the bottleneck moved.
Shows whether the request is slow.
Shows which stage is slow and how its share of the total changed.
Often compared only to a global threshold.
Compares stage and end-to-end behavior against prior profiles.
Teams still need to investigate the source.
Bottleneck movement and SLO budget guide the next diagnostic action.
Suitable for dashboards.
Suitable for GO/WARN/BLOCK deployment gates when thresholds are explicit.
Track total latency for user experience. Use stage profiling when the objective is diagnosis and controlled deployment.
Profile RAG traces with p50/p90/p95/p99, error rates, SLO breaches, bottleneck ranking and instrumentation gaps across arbitrary retrieval, reranking and generation stages.
See the Actor