Find the bottleneck: the critical path of a request, from your traces.
A slow request is rarely slow everywhere. Some calls run in parallel and some wait on each other; only the chain of waits sets the end-to-end time. That chain is the critical path, and making anything off it faster changes nothing the user sees.
Updated
Total time misleads; self time does not
A span's duration includes the time it spent waiting on its children. Rank spans by duration and the root always wins, which tells you nothing. Self time — the part of a span's duration not covered by the children it waited on — is the time that hop itself costs, and it is what a fix can remove.
Parallel calls complicate it further: of two calls made at once, only the one that finishes last holds the request up. The critical path follows that one.
Entry points first
A critical path belongs to an entry point — where traffic enters your services, such as a server route or a queue a consumer reads. Ritele names entry points by service and http.route, RPC method or topic, so a path is for "POST /transfers", not for one of a thousand raw URLs.
Which hop is worth improving
For each entry point, Ritele's Paths view shows the hop-by-hop path requests take, the self time at every hop, and its share of the whole. It can also rank the hops by how much time each would save if its p99 self time were 20% shorter, so the hop most worth working on comes first.
The paths are built hourly from kept traces. Because tail sampling keeps every error, those paths include the failing requests, not only a random sample of the healthy ones.
What the traces need to carry
http.route on server spans rather than the raw URL; the messaging conventions on consumers; and a span for every outbound call, so each hop appears on the path and none of its time is charged to its caller. A database call with no client span looks like slow application code.