A billing service that usually answers in 180 milliseconds starts taking about 4 seconds. The API gateway in front of checkout gives up after 1 second and tries twice more. Within a few minutes the gateway's connection pool to the billing service is saturated, the database behind billing is busy with invoice queries whose callers have already disconnected, and checkout is listed as down. The slow query started it, and every timed-out checkout after that became extra queries against the same invoices table.
The retry helper is a few lines in the HTTP client, added after checkout surfaced one error and someone wanted the next error swallowed. It treats a timeout, a connection reset, and a 503 as the same event, and it has no count of how many other callers are doing the same thing in the same second.
The work that outlives the caller
When the gateway stops waiting at 1 second, the billing process often keeps going. Closing the client side of the HTTP connection cancels server-side work only if the server is watching that connection and the handler can abandon what it is doing. At the moment of the timeout, the handler is blocked in the database driver. The worker holds its connection until the query returns. The plan that needed 4 seconds still needs 4 seconds. The bytes come back to a handler that can no longer write them anywhere the user will see, and the database connection returns to its pool only then.
By that point the gateway has opened the next attempt. With a 1 second timeout and a handler that runs for 4 seconds, the first attempt is still executing when the second arrives, and the second is still executing when the third arrives. The client sends the retry because the attempt timed out, which happens well before the handler completes. One checkout has become three invoice queries, each holding a database connection for the full runtime.
Those checkouts share one database. At 50 checkouts a second with three overlapping attempts each, the database sees about 150 new queries a second. Held for 4 seconds apiece, that is about 600 queries in flight, against a database that was already missing its latency target when it was serving the original 50. As the runtime grows, the number in flight grows with it, and the extra concurrency pushes the runtime up again. If max_connections is 200, those 600 sessions never all run. The surplus waits in the application's database pool or is rejected at the database. A rejected connection is a failure the HTTP client has been told to retry, so the gateway schedules more attempts into a database that is already refusing them. The database spends its remaining capacity on responses for a gateway that has already returned an error to the user.
An idempotency key on the checkout POST stops a retry from inserting a second invoice. The lookup query still runs, and the connection stays checked out until the statement finishes. Adding the key removes duplicate ledger lines and leaves the three-attempt policy in place. The number of in-flight queries stays the same.
Before the status page changes
In-flight requests on the billing service climb while that service's own success rate is still high. The handlers that reach completion do so after the gateway has stopped waiting, and those successes never appear in the gateway's error budget. Gateway errors rise next, mostly timeouts and 502s, while the billing service's error rate stays low. Billing's logs show completed responses. The gateway's logs show timeouts, so the gateway is the process that gets restarted. Restarting it drains the gateway pool. The database queries already in progress keep running. The next wave of retries fills the gateway pool again.
Server-side p95 and client-side p95 are percentiles over different requests. Server-side p95 might sit around 6 seconds. Client-side p95 sits on the timeout, at 1 second, because the client never records the slow successes. A trace sampled from a call that completed shows one slow success and omits the attempts the gateway abandoned. Graph the number of handlers that ran past the caller's deadline, next to a retry counter.
That counter lives on the client, so a dashboard of the billing service's request rate, errors, and duration can stay calm through the climb. A gateway metric such as upstream_retries, or a client metric labeled with the attempt number, shows how many attempts each checkout became. When the series sits near 1.0 through a normal week and then holds around 2.4, the extra attempts are adding load on top of the slow query. Pool wait belongs on the same chart. Once wait time is a large fraction of the 1 second timeout, new attempts fail in the gateway pool before they reach the invoice query. A policy that retries on timeout or on connection failure then schedules another attempt for a failure the pool produced.
Timeouts, cancellation, and a shared retry budget
Set the client timeout from the latency checkout can tolerate, and keep it above the healthy runtime of the call. If invoice lookup p99 is 400 milliseconds when the database is inside normal load, a 1 second client timeout absorbs a cold cache page without scheduling a second attempt. A 200 millisecond timeout on that same call will open retries on an ordinary afternoon. A report endpoint whose healthy runtime is 8 seconds should stay off a path whose timeout is 2 seconds and whose client retries when that timeout fires. The shorter deadline starts a second copy of the query while the first copy is still running.
The billing service's call into the database often has no deadline of its own. PostgreSQL ships with statement_timeout at 0, which disables the timeout. The gateway will give up at 1 second. The statement runs until the plan finishes, or until something else ends the session. On client disconnect, the billing handler should cancel the statement. Drivers that accept a cancellation token will do that when the handler passes the token through to the driver. If that path is incomplete, set statement_timeout a little under the client timeout so the database ends the query before the gateway gives up. The connection returns to the database pool, and the gateway receives an error response before its own timer fires. If retries are reserved for timeouts and connection failures, and that error is returned as a normal response, the second attempt is never sent. If the policy retries the error as well, a later attempt still goes out, but the previous statement has ended, so the attempts no longer overlap.
A per-request cap of three attempts is how 50 checkouts a second become 150 query starts a second. A budget shared across callers caps that expansion. A token bucket used only for retries, or a circuit breaker that opens when the timeout rate on the billing service crosses a threshold chosen ahead of time, refuses further attempts while first attempts are still allowed to finish or fail. The user sees an error on checkout. The database pool holds the concurrency of those first attempts, which is the concurrency it was sized for.
Backoff and jitter change when attempts arrive. The shared budget is what limits how many are admitted. Exponential backoff spreads attempts after a short blip, and those attempts still run if the billing service is saturated when the delay ends. The same delay on every caller delivers them together. Jitter spreads that arrival across a window. A hedge, a second attempt sent to another replica before the first attempt has reached its timeout, draws on the same budget with a shorter delay. The extra query is a reasonable cost when it is uncommon and the attempt that loses is cancelled on the database. Once the path is already past its latency target, the hedge is one more query holding a connection.