A Retry Is a Second Request
A client sends a request. The server accepts it, updates a record, and starts preparing a response. The connection drops before that response reaches the client.
From the client's perspective, the operation failed. From the server's perspective, it may already be complete. Retrying without accounting for that difference can turn a network problem into duplicate work.
Give the intended action an identity
Consider creating a service request. Clicking submit once expresses one intention, even if the client needs several attempts to communicate it.
An idempotency key can represent that intention across attempts. The client reuses the key when retrying the same operation; a genuinely new submission receives a new one. The server needs an explicit contract for how it recognises and responds to repeats.
The Amazon Builders' Library discusses this approach in Making retries safe with idempotent APIs, including caller-provided identifiers and the problem of reusing an identifier for different intent.
Define the uncomfortable cases
For an API I were designing, I would write down these cases before implementing the happy path:
- The same caller repeats the same key and payload.
- The same caller repeats the key with a different payload.
- Two requests with the same key arrive concurrently.
- The first operation is still running when the retry arrives.
- A request arrives after the retention window for its key has ended.
Those are different situations. A key should be scoped to the caller or another appropriate boundary. A changed payload should not silently inherit the outcome of an unrelated request.
The retention window matters too. An idempotency contract that remembers keys for a limited period cannot promise duplicate detection forever. The client and the documentation need to reflect that boundary.
The storage boundary is part of the design
A lookup followed by an insert is not automatically safe under concurrency. Two workers can both observe that no record exists and then both proceed.
I would want uniqueness enforced at the storage layer, with the claim on the operation and its database changes coordinated transactionally where possible. External actions need their own treatment. Sending a message through another service cannot simply be rolled back with a local database transaction.
An outbox or durable job can help coordinate the handoff, but it does not remove the need to think about redelivery and the downstream provider's contract. The important question is what prevents a second attempt from creating a second effect.
Retry with a limit and a reason
I would distinguish a temporary transport problem from a validation failure. Repeating invalid input does not make it valid. Where retries are appropriate, they should have a bounded budget and avoid immediate synchronised bursts.
The interface should also preserve uncertainty honestly. A lost response may require checking the operation's status rather than telling the person to create another request.
Retries are a useful recovery tool. They become trustworthy when the system can distinguish another attempt at the same action from an entirely new action.