Wise Hustlers — Digital Product & App Development Studio Logo
Get Consultation
By Wise Hustler Admin•9/29/2026•12 min read

Webhook Retry Best Practices: Signature Verification, Idempotency and Dead-Letter Queues

Webhook Retry Best Practices: Signature Verification, Idempotency and Dead-Letter Queues

# Webhook Retry Best Practices: Signature Verification, Idempotency and Dead-Letter Queues

TL;DR: Solid webhook retry best practices cover both directions of the connection: as a receiver, verify the HMAC signature against the raw request body before parsing it, return a 2xx response within a few seconds, and dedupe on the event ID before you do anything with side effects. As a sender, retry failed deliveries with exponential backoff plus jitter, cap the number of attempts, and move exhausted events to a dead-letter queue instead of dropping them.

Webhooks turn a request-response API into a lossy, asynchronous message bus that runs over plain HTTP. The provider fires an HTTP POST, your endpoint might be slow, down, or mid-deploy, and the provider has to decide whether and how to retry. Get either side wrong and you end up with duplicate charges, missed order updates, or a security hole where anyone can POST a fake "payment succeeded" event to your endpoint.

This guide walks through both sides: what to check when you receive a webhook, and what to build when you send one, with the actual retry schedules that Stripe, Shopify, and GitHub document today.

How to verify a webhook signature

To verify a webhook signature, recompute an HMAC over the raw, unparsed request body using a shared secret, then compare it to the signature the provider sent in a header — using a constant-time comparison, never === or ==. If the computed and received signatures don't match, or the payload is stale, reject the request with a 4xx status before any business logic runs.

Three things trip people up in practice:

1. The raw body, not the parsed body. Frameworks like Express re-serialize JSON when you call JSON.stringify on a parsed object, and the re-serialized bytes (key order, whitespace, number formatting) almost never match what the provider originally sent and signed. You must capture the exact bytes off the wire before any body-parsing middleware touches them.

2. Constant-time comparison. A naive string === short-circuits on the first mismatched byte, which leaks timing information an attacker can use to guess the signature byte-by-byte. Use crypto.timingSafeEqual in Node.js (or the language's equivalent, e.g. Ruby's Rack::Utils.secure_compare).

3. Timestamp tolerance to stop replay attacks. Stripe and similar providers include a timestamp in the signed payload. Reject anything outside a tolerance window (Stripe's own libraries default to 300 seconds / 5 minutes) even if the signature itself is technically valid — otherwise a captured request can be replayed indefinitely.

How the major providers sign their payloads

ProviderHeaderSchemeAlgorithm
StripeStripe-Signaturet=<timestamp>,v1=<signature> — signs "{timestamp}.{raw_body}"HMAC-SHA256
ShopifyX-Shopify-Hmac-SHA256Base64-encoded digest of the raw bodyHMAC-SHA256
GitHubX-Hub-Signature-256sha256=<hex digest> of the raw bodyHMAC-SHA256

All three land on the same underlying primitive — HMAC-SHA256 over the raw body with a shared secret — they just differ in how the signature is packaged into the header.

HMAC webhook verification in Node.js

Here's a complete, correct Express example that verifies a Stripe-style signature (the same shape applies to Shopify and GitHub with the header name and comparison string swapped). The critical detail is capturing the raw body with express.raw() before any JSON-parsing middleware runs on that route.

const express = require('express');
const crypto = require('crypto');

const app = express();

const WEBHOOK_SECRET = process.env.WEBHOOK_SECRET;
const TOLERANCE_SECONDS = 300; // reject events older than 5 minutes

function verifySignature(rawBody, signatureHeader, secret) {
  // Stripe-style header: "t=1730210000,v1=8f1a3c...".
  const parts = Object.fromEntries(
    signatureHeader.split(',').map((pair) => pair.split('='))
  );
  const timestamp = parts.t;
  const receivedSignature = parts.v1;

  if (!timestamp || !receivedSignature) {
    return { valid: false, reason: 'malformed signature header' };
  }

  const age = Math.abs(Date.now() / 1000 - Number(timestamp));
  if (age > TOLERANCE_SECONDS) {
    return { valid: false, reason: 'timestamp outside tolerance' };
  }

  const signedPayload = `${timestamp}.${rawBody}`;
  const expectedSignature = crypto
    .createHmac('sha256', secret)
    .update(signedPayload, 'utf8')
    .digest('hex');

  const expected = Buffer.from(expectedSignature, 'utf8');
  const received = Buffer.from(receivedSignature, 'utf8');

  const valid =
    expected.length === received.length &&
    crypto.timingSafeEqual(expected, received);

  return { valid, reason: valid ? null : 'signature mismatch' };
}

// express.raw() keeps req.body as a Buffer for this route only,
// so the bytes we verify match exactly what was signed.
app.post(
  '/webhooks/provider',
  express.raw({ type: 'application/json' }),
  async (req, res) => {
    const signatureHeader = req.headers['x-provider-signature'];
    const rawBody = req.body.toString('utf8');

    const { valid, reason } = verifySignature(
      rawBody,
      signatureHeader,
      WEBHOOK_SECRET
    );

    if (!valid) {
      return res.status(400).json({ error: reason });
    }

    const event = JSON.parse(rawBody);

    const alreadyProcessed = await eventStore.exists(event.id);
    if (alreadyProcessed) {
      // Same event ID, already handled — acknowledge and stop.
      return res.status(200).json({ received: true, duplicate: true });
    }

    await queue.enqueue('process-webhook-event', event);

    return res.status(200).json({ received: true });
  }
);

module.exports = app;

Note what this handler does not do: it does not update a database record, charge a card, or send an email inline. It verifies, dedupes, enqueues, and returns. The actual business logic runs in a worker pulling off that queue, decoupled from the provider's timeout budget.

Diagram showing the webhook receiver flow from signature verification to async processing

The receiver path: verify first, dedupe second, acknowledge fast, process later.

Return 2xx fast, process async

Providers time out a webhook request if it takes too long — Shopify, for example, treats anything other than a fast 2xx as a failure and queues a retry. If your handler holds the connection open while it emails a customer, calls three downstream APIs, and writes to two tables, you're gambling that all of it finishes inside the provider's timeout window every single time, under load, during a deploy, during a GC pause.

The fix is the classic decouple-with-a-queue pattern:

1. Verify the signature.

2. Check whether the event ID has already been processed; if so, return 200 immediately without re-running side effects.

3. Push the event onto a queue (SQS, a Postgres-backed job table, Redis, whatever you already run).

4. Return 200 as soon as the enqueue call succeeds.

5. A separate worker process dequeues and does the actual work, with its own retry/backoff logic that's independent of the provider's retry schedule.

This also means a slow downstream dependency (a flaky third-party API, a locked database row) degrades your worker queue's latency, not your webhook endpoint's success rate.

Idempotency key in payments: why event IDs aren't optional

An idempotency key in payments is a unique identifier — the provider's event ID, or a key you generate yourself on outbound requests — that lets you safely process the same operation twice without double-charging or double-shipping. Providers explicitly retry on any failure signal, including a timeout on your side after you actually succeeded internally, so "delivered exactly once" is not a guarantee any webhook provider makes. Duplicate delivery is expected, normal behavior, not an edge case.

Practically, this means:

  • Store processed event IDs (Stripe's evt_..., GitHub's delivery GUID, Shopify's X-Shopify-Webhook-Id) in a table with a unique constraint, and check it before running side effects, not after.
  • Make the side effect itself idempotent where you can — e.g., INSERT ... ON CONFLICT DO NOTHING on the order-fulfillment record — as defense in depth, since checking the event ID and enqueuing the job aren't atomic.
  • If you're the one issuing requests that must not double-fire (charging a card via a payment API, for instance), generate your own client-side idempotency key and pass it on the outbound request too — this is the same concept applied to the sending side, not just the receiving side.

Webhook retry mechanism: what senders should actually build

A webhook retry mechanism is what runs on the sending side after a delivery attempt fails: it schedules another attempt with an increasing delay, caps the total number of attempts, and eventually gives up and records the event somewhere durable instead of losing it. If you're building a product that sends webhooks to customers, you need the sender-side equivalent of everything above.

Exponential backoff with jitter

Retrying at a fixed interval synchronizes all your failed deliveries into a thundering herd against a recovering endpoint. Exponential backoff spaces retries out (1s, 2s, 4s, 8s, …), and adding random jitter to each delay prevents many simultaneous failures from retrying in lockstep and re-overwhelming the same endpoint the moment it comes back up.

function nextDelayMs(attempt, baseMs = 1000, maxMs = 15 * 60 * 1000) {
  const exponential = Math.min(maxMs, baseMs * 2 ** attempt);
  const jitter = Math.random() * exponential * 0.5; // full jitter range
  return exponential / 2 + jitter;
}

Cap attempts, then dead-letter

After a bounded number of attempts or a bounded time window, stop retrying and write the event to a dead-letter queue (DLQ) — a durable store of "deliveries we gave up on" that a human or a backfill job can inspect and manually replay later. Silently dropping the event after the last retry is the single most common webhook reliability bug; a DLQ turns permanent failures into a queryable, replayable backlog instead of a support ticket with no data behind it.

Diagram of a sender retry timeline with exponential backoff and jitter ending in a dead-letter queue

Each failed attempt backs off further and adds jitter; once attempts are exhausted, the event lands in the dead-letter queue instead of disappearing.

Disable chronically failing endpoints

If an endpoint has been failing for hours or days, continuing to retry it wastes resources on both sides. Shopify's own policy reflects this directly.

Stripe webhook retry policy vs. Shopify vs. GitHub

Here's what each provider's current documentation actually says, so you can stop guessing:

ProviderRetry windowBackoffOn exhaustion
StripeUp to 3 days in live mode; 3 attempts over a few hours in test modeExponentialEvent delivery stops; visible in the Event destinations dashboard
ShopifyUp to 8 attempts over a 4-hour periodExponential, interval increases each failed attemptDelivery stops; if failures persist across events, the webhook subscription itself is removed and must be re-created
GitHubNo automatic retry at all — GitHub does not automatically redeliver failed deliveriesN/AYou must poll the Deliveries API and manually (or via a scheduled script) redeliver failures within a 3-day lookback

The GitHub row is the one people get wrong most often: unlike Stripe and Shopify, GitHub's default behavior is to attempt delivery once and then leave the failure sitting in the delivery log. If you need retry coverage for a GitHub webhook integration, you're expected to build a small polling script against the deliveries API yourself — GitHub's docs ship a reference implementation for exactly this.

If you're integrating several providers like this into one system, that's the kind of plumbing that's easy to get subtly wrong per-provider; see Wise Hustlers' API integration services if you'd rather have someone build and test it once.

Comparison: signature verification quick reference

ProviderHeader to readWhat you HMACComparison
StripeStripe-Signature"{timestamp}.{raw_body}"crypto.timingSafeEqual on hex digests
ShopifyX-Shopify-Hmac-SHA256raw body onlycrypto.timingSafeEqual on base64-decoded digests
GitHubX-Hub-Signature-256raw body onlyconstant-time compare on sha256=-prefixed hex digest

FAQ

How do you verify a webhook signature without a raw body?

You can't reliably — signature verification requires the exact bytes the provider signed, so any middleware that parses, re-serializes, or reformats the body before your verification code runs will break it. Configure your raw-body-capturing middleware (express.raw() in Express, disabling Next.js's default body parser, etc.) specifically on the webhook route, before any JSON parser touches that request.

What's the difference between a webhook retry mechanism and a message queue?

A webhook retry mechanism is the sender's logic for re-attempting HTTP delivery to your endpoint on failure (backoff, attempt caps, DLQ); a message queue is usually what you, the receiver, put the event onto internally after accepting it, so your own processing has independent retry semantics from the provider's delivery retries. They solve related but distinct problems on opposite ends of the same webhook.

Why do I need an idempotency key if webhook signatures already guarantee authenticity?

A signature proves the event is genuinely from the provider; it says nothing about whether you've already processed this exact event before. Providers explicitly resend events after network timeouts and ambiguous responses, so an idempotency key (the event ID) is what protects you from processing an authentic, duplicate delivery twice.

What should a dead-letter queue actually store?

At minimum: the full raw payload, the headers (including the signature, for later re-verification), the endpoint or event type, the number of attempts made, and the final failure reason — enough for someone to inspect why it failed and manually replay it without going back to the provider's own (often time-limited) delivery logs.

Sources

Related articles