# Webhook Retry Best Practices: Signature Verification, Idempotency and Dead-Letter Queues
TL;DR: Solid webhook retry best practices cover both directions of the connection: as a receiver, verify the HMAC signature against the raw request body before parsing it, return a 2xx response within a few seconds, and dedupe on the event ID before you do anything with side effects. As a sender, retry failed deliveries with exponential backoff plus jitter, cap the number of attempts, and move exhausted events to a dead-letter queue instead of dropping them.
Webhooks turn a request-response API into a lossy, asynchronous message bus that runs over plain HTTP. The provider fires an HTTP POST, your endpoint might be slow, down, or mid-deploy, and the provider has to decide whether and how to retry. Get either side wrong and you end up with duplicate charges, missed order updates, or a security hole where anyone can POST a fake "payment succeeded" event to your endpoint.
This guide walks through both sides: what to check when you receive a webhook, and what to build when you send one, with the actual retry schedules that Stripe, Shopify, and GitHub document today.
How to verify a webhook signature
To verify a webhook signature, recompute an HMAC over the raw, unparsed request body using a shared secret, then compare it to the signature the provider sent in a header — using a constant-time comparison, never === or ==. If the computed and received signatures don't match, or the payload is stale, reject the request with a 4xx status before any business logic runs.
Three things trip people up in practice:
1. The raw body, not the parsed body. Frameworks like Express re-serialize JSON when you call JSON.stringify on a parsed object, and the re-serialized bytes (key order, whitespace, number formatting) almost never match what the provider originally sent and signed. You must capture the exact bytes off the wire before any body-parsing middleware touches them.
2. Constant-time comparison. A naive string === short-circuits on the first mismatched byte, which leaks timing information an attacker can use to guess the signature byte-by-byte. Use crypto.timingSafeEqual in Node.js (or the language's equivalent, e.g. Ruby's Rack::Utils.secure_compare).
3. Timestamp tolerance to stop replay attacks. Stripe and similar providers include a timestamp in the signed payload. Reject anything outside a tolerance window (Stripe's own libraries default to 300 seconds / 5 minutes) even if the signature itself is technically valid — otherwise a captured request can be replayed indefinitely.
How the major providers sign their payloads
| Provider | Header | Scheme | Algorithm |
|---|---|---|---|
| Stripe | Stripe-Signature | t=<timestamp>,v1=<signature> — signs "{timestamp}.{raw_body}" | HMAC-SHA256 |
| Shopify | X-Shopify-Hmac-SHA256 | Base64-encoded digest of the raw body | HMAC-SHA256 |
| GitHub | X-Hub-Signature-256 | sha256=<hex digest> of the raw body | HMAC-SHA256 |
All three land on the same underlying primitive — HMAC-SHA256 over the raw body with a shared secret — they just differ in how the signature is packaged into the header.
HMAC webhook verification in Node.js
Here's a complete, correct Express example that verifies a Stripe-style signature (the same shape applies to Shopify and GitHub with the header name and comparison string swapped). The critical detail is capturing the raw body with express.raw() before any JSON-parsing middleware runs on that route.
const express = require('express');
const crypto = require('crypto');
const app = express();
const WEBHOOK_SECRET = process.env.WEBHOOK_SECRET;
const TOLERANCE_SECONDS = 300; // reject events older than 5 minutes
function verifySignature(rawBody, signatureHeader, secret) {
// Stripe-style header: "t=1730210000,v1=8f1a3c...".
const parts = Object.fromEntries(
signatureHeader.split(',').map((pair) => pair.split('='))
);
const timestamp = parts.t;
const receivedSignature = parts.v1;
if (!timestamp || !receivedSignature) {
return { valid: false, reason: 'malformed signature header' };
}
const age = Math.abs(Date.now() / 1000 - Number(timestamp));
if (age > TOLERANCE_SECONDS) {
return { valid: false, reason: 'timestamp outside tolerance' };
}
const signedPayload = `${timestamp}.${rawBody}`;
const expectedSignature = crypto
.createHmac('sha256', secret)
.update(signedPayload, 'utf8')
.digest('hex');
const expected = Buffer.from(expectedSignature, 'utf8');
const received = Buffer.from(receivedSignature, 'utf8');
const valid =
expected.length === received.length &&
crypto.timingSafeEqual(expected, received);
return { valid, reason: valid ? null : 'signature mismatch' };
}
// express.raw() keeps req.body as a Buffer for this route only,
// so the bytes we verify match exactly what was signed.
app.post(
'/webhooks/provider',
express.raw({ type: 'application/json' }),
async (req, res) => {
const signatureHeader = req.headers['x-provider-signature'];
const rawBody = req.body.toString('utf8');
const { valid, reason } = verifySignature(
rawBody,
signatureHeader,
WEBHOOK_SECRET
);
if (!valid) {
return res.status(400).json({ error: reason });
}
const event = JSON.parse(rawBody);
const alreadyProcessed = await eventStore.exists(event.id);
if (alreadyProcessed) {
// Same event ID, already handled — acknowledge and stop.
return res.status(200).json({ received: true, duplicate: true });
}
await queue.enqueue('process-webhook-event', event);
return res.status(200).json({ received: true });
}
);
module.exports = app;Note what this handler does not do: it does not update a database record, charge a card, or send an email inline. It verifies, dedupes, enqueues, and returns. The actual business logic runs in a worker pulling off that queue, decoupled from the provider's timeout budget.
The receiver path: verify first, dedupe second, acknowledge fast, process later.
Return 2xx fast, process async
Providers time out a webhook request if it takes too long — Shopify, for example, treats anything other than a fast 2xx as a failure and queues a retry. If your handler holds the connection open while it emails a customer, calls three downstream APIs, and writes to two tables, you're gambling that all of it finishes inside the provider's timeout window every single time, under load, during a deploy, during a GC pause.
The fix is the classic decouple-with-a-queue pattern:
1. Verify the signature.
2. Check whether the event ID has already been processed; if so, return 200 immediately without re-running side effects.
3. Push the event onto a queue (SQS, a Postgres-backed job table, Redis, whatever you already run).
4. Return 200 as soon as the enqueue call succeeds.
5. A separate worker process dequeues and does the actual work, with its own retry/backoff logic that's independent of the provider's retry schedule.
This also means a slow downstream dependency (a flaky third-party API, a locked database row) degrades your worker queue's latency, not your webhook endpoint's success rate.
Idempotency key in payments: why event IDs aren't optional
An idempotency key in payments is a unique identifier — the provider's event ID, or a key you generate yourself on outbound requests — that lets you safely process the same operation twice without double-charging or double-shipping. Providers explicitly retry on any failure signal, including a timeout on your side after you actually succeeded internally, so "delivered exactly once" is not a guarantee any webhook provider makes. Duplicate delivery is expected, normal behavior, not an edge case.
Practically, this means:
- Store processed event IDs (Stripe's
evt_..., GitHub's delivery GUID, Shopify'sX-Shopify-Webhook-Id) in a table with a unique constraint, and check it before running side effects, not after. - Make the side effect itself idempotent where you can — e.g.,
INSERT ... ON CONFLICT DO NOTHINGon the order-fulfillment record — as defense in depth, since checking the event ID and enqueuing the job aren't atomic. - If you're the one issuing requests that must not double-fire (charging a card via a payment API, for instance), generate your own client-side idempotency key and pass it on the outbound request too — this is the same concept applied to the sending side, not just the receiving side.
Webhook retry mechanism: what senders should actually build
A webhook retry mechanism is what runs on the sending side after a delivery attempt fails: it schedules another attempt with an increasing delay, caps the total number of attempts, and eventually gives up and records the event somewhere durable instead of losing it. If you're building a product that sends webhooks to customers, you need the sender-side equivalent of everything above.
Exponential backoff with jitter
Retrying at a fixed interval synchronizes all your failed deliveries into a thundering herd against a recovering endpoint. Exponential backoff spaces retries out (1s, 2s, 4s, 8s, …), and adding random jitter to each delay prevents many simultaneous failures from retrying in lockstep and re-overwhelming the same endpoint the moment it comes back up.
function nextDelayMs(attempt, baseMs = 1000, maxMs = 15 * 60 * 1000) {
const exponential = Math.min(maxMs, baseMs * 2 ** attempt);
const jitter = Math.random() * exponential * 0.5; // full jitter range
return exponential / 2 + jitter;
}Cap attempts, then dead-letter
After a bounded number of attempts or a bounded time window, stop retrying and write the event to a dead-letter queue (DLQ) — a durable store of "deliveries we gave up on" that a human or a backfill job can inspect and manually replay later. Silently dropping the event after the last retry is the single most common webhook reliability bug; a DLQ turns permanent failures into a queryable, replayable backlog instead of a support ticket with no data behind it.
Each failed attempt backs off further and adds jitter; once attempts are exhausted, the event lands in the dead-letter queue instead of disappearing.
Disable chronically failing endpoints
If an endpoint has been failing for hours or days, continuing to retry it wastes resources on both sides. Shopify's own policy reflects this directly.
Stripe webhook retry policy vs. Shopify vs. GitHub
Here's what each provider's current documentation actually says, so you can stop guessing:
| Provider | Retry window | Backoff | On exhaustion |
|---|---|---|---|
| Stripe | Up to 3 days in live mode; 3 attempts over a few hours in test mode | Exponential | Event delivery stops; visible in the Event destinations dashboard |
| Shopify | Up to 8 attempts over a 4-hour period | Exponential, interval increases each failed attempt | Delivery stops; if failures persist across events, the webhook subscription itself is removed and must be re-created |
| GitHub | No automatic retry at all — GitHub does not automatically redeliver failed deliveries | N/A | You must poll the Deliveries API and manually (or via a scheduled script) redeliver failures within a 3-day lookback |
The GitHub row is the one people get wrong most often: unlike Stripe and Shopify, GitHub's default behavior is to attempt delivery once and then leave the failure sitting in the delivery log. If you need retry coverage for a GitHub webhook integration, you're expected to build a small polling script against the deliveries API yourself — GitHub's docs ship a reference implementation for exactly this.
If you're integrating several providers like this into one system, that's the kind of plumbing that's easy to get subtly wrong per-provider; see Wise Hustlers' API integration services if you'd rather have someone build and test it once.
Comparison: signature verification quick reference
| Provider | Header to read | What you HMAC | Comparison |
|---|---|---|---|
| Stripe | Stripe-Signature | "{timestamp}.{raw_body}" | crypto.timingSafeEqual on hex digests |
| Shopify | X-Shopify-Hmac-SHA256 | raw body only | crypto.timingSafeEqual on base64-decoded digests |
| GitHub | X-Hub-Signature-256 | raw body only | constant-time compare on sha256=-prefixed hex digest |
FAQ
How do you verify a webhook signature without a raw body?
You can't reliably — signature verification requires the exact bytes the provider signed, so any middleware that parses, re-serializes, or reformats the body before your verification code runs will break it. Configure your raw-body-capturing middleware (express.raw() in Express, disabling Next.js's default body parser, etc.) specifically on the webhook route, before any JSON parser touches that request.
What's the difference between a webhook retry mechanism and a message queue?
A webhook retry mechanism is the sender's logic for re-attempting HTTP delivery to your endpoint on failure (backoff, attempt caps, DLQ); a message queue is usually what you, the receiver, put the event onto internally after accepting it, so your own processing has independent retry semantics from the provider's delivery retries. They solve related but distinct problems on opposite ends of the same webhook.
Why do I need an idempotency key if webhook signatures already guarantee authenticity?
A signature proves the event is genuinely from the provider; it says nothing about whether you've already processed this exact event before. Providers explicitly resend events after network timeouts and ambiguous responses, so an idempotency key (the event ID) is what protects you from processing an authentic, duplicate delivery twice.
What should a dead-letter queue actually store?
At minimum: the full raw payload, the headers (including the signature, for later re-verification), the endpoint or event type, the number of attempts made, and the final failure reason — enough for someone to inspect why it failed and manually replay it without going back to the provider's own (often time-limited) delivery logs.
Sources
- Stripe — Receive Stripe events in your webhook endpoint
- Stripe — Webhook signatures
- Shopify — Verify webhook deliveries
- Shopify — Troubleshoot webhooks
- GitHub — Validating webhook deliveries
- GitHub — Handling failed webhook deliveries
- GitHub — Automatically redelivering failed deliveries for a repository webhook