All posts

What is a webhook? How webhooks work, and the six ways they break in production

A webhook is an HTTP POST one system sends another when something happens. Receiving one is easy. Receiving each one exactly once, verified, in an order you can use, is the actual job — here is how it works, and the six ways it breaks.

Vadym Kykalo12 min read
  • webhooks
  • guide
  • security
  • reliability
  • http

Every figure quoted from a provider was read from that provider's own documentation on September 19, 2026, and each one is linked. Providers change these without notice — check the link before you rely on a number.

A customer pays for order ord_8412. The shop sends your backend a webhook, and your handler awards loyalty points, emails a receipt and calls the warehouse — which takes eleven seconds. The sender's timeout is ten, so it marks the delivery failed and sends it again. The customer gets two receipts and double points, and your logs show two perfectly healthy 200s.

Nothing in that story is a bug in the usual sense. Every line of the handler did what it said. It is what happens when code written for a function call meets a protocol built on retries — and it is the first of six failures this guide walks through, after the basics every one of them depends on.

What is a webhook?

A webhook is an HTTP POST that one system sends to a URL another system registered, at the moment an event happens. The receiver registers once — "when an order is created, tell me at this URL" — and from then on the sender calls it every time. This is what one looks like, signed the way the open Standard Webhooks specification describes:

POST /webhooks/orders HTTP/1.1
Host: api.example-shop.com
Content-Type: application/json
Content-Length: 186
webhook-id: msg_2mQ8hZk3Xv9aR1cT
webhook-timestamp: 1789983612
webhook-signature: v1,pX/4CnjtyRIdKhtP4FEBLl0a8pgk0IErOOHRkgsduog=

{
  "type": "order.created",
  "timestamp": "2026-09-21T09:40:12Z",
  "data": {
    "order_id": "ord_8412",
    "customer_id": "cus_2291",
    "total": 42.50,
    "currency": "EUR"
  }
}
POST /webhooks/orders HTTP/1.1content-type: application/jsonwebhook-id: msg_2mQ8hZk3Xv9aR1cTwebhook-timestamp: 1789983612webhook-signature: v1,pX/4CnjtyRIdKhtP…{ "type": "order.created", "data": { "order_id": "ord_8412", … } }Wherea URL the receiver registered forthis eventIdentitythe same on every retry: yourdeduplication keyFreshnessolder than ~5 minutes? reject: thereplay windowProofHMAC-SHA256 of id.timestamp.body, inbase64What happenedthe event type, and its data
Three headers carry everything a receiver needs to trust the request: which event it is, when it was sent, and proof that nobody changed it on the way.

The signature is real. It was made with the secret whsec_DioEXswemwdu+ScHzpEImNzhtrASzXe0vxCXPBmwuTM=: paste the body, that secret and the three headers into the signature verifier and it will say the signature is valid — and that the timestamp is far too old to accept. Both answers are correct, and the second one is a security feature.

Webhook vs API polling

The alternative to being told is asking: calling the other system's API on a timer to see whether anything changed.

order createdPollingyou ask on a timer6 requests, 1 usefulWebhookthey tell you, once1 request, 1 usefulone intervalGET /ordersnothing newGET /ordersnothing newGET /ordersnothing newGET /ordersnothing newGET /ordersfound itGET /ordersnothing newup to one polling interval latethe moment it happensPOSTyour endpointtime →wasted requestuseful request
Polling pays for every empty answer and still hears about the order up to one interval late. A webhook costs one request and arrives when the event happens.
Polling an APIReceiving a webhook
Who starts the requestYou, on a timerThe sender, when something happens
DelayUp to one polling intervalUsually seconds
Requests when nothing happensOne per interval, foreverNone
What you must runA scheduled jobA public HTTPS endpoint
If you are downYou catch up on the next pollYou depend on the sender's retries

Shorten a polling interval and you pay in requests and rate limits; lengthen it and you pay in lag. A webhook costs one request per event. But read the last row twice: polling fails safe, because a missed poll is repeated by the next one, while a webhook fails however the sender decides. That is why polling keeps a job in a good webhook integration — as the slow reconciliation that catches whatever the webhooks missed.

Sending webhooks: what the sender decides

Event types. Conventionally resource.verb in the past tense: order.created, invoice.paid. The Standard Webhooks spec asks for a full-stop delimited type, with a timestamp and a data object next to it. Treat types as a public API: renaming one breaks every receiver that routes on it.

Payloads. A full snapshot of the resource saves the receiver a round trip; a thin notification carrying only an id makes it fetch the current state, which is never stale. Stripe offers both and calls them exactly that.

Subscriptions. Which endpoint wants which event types. Sending everything to everyone just makes every receiver filter traffic it never asked for.

Signing. Anyone can POST to a public URL, so the receiver needs proof the request came from you, unaltered. That proof is an HMAC — a hash of the message keyed with a secret only the two of you hold. Standard Webhooks signs webhook-id, webhook-timestamp and the raw body joined by full stops, with HMAC-SHA256, and sends the base64 result after a v1, prefix. The secret is 24 to 64 random bytes, base64-encoded, with a whsec_ prefix.

A replay window. The timestamp is signed so that a captured request cannot be sent again tomorrow: the receiver rejects anything older than a few minutes. The spec leaves the tolerance to you; Stripe's libraries default to five minutes, a sensible place to start.

Timeouts and status codes. The spec recommends a request timeout "somewhere between 15 and 30s". 2xx is success and anything else is retried — including redirects, which Stripe also counts as failures. Two codes carry extra meaning: 410 Gone asks the sender to disable the endpoint, 429 Too Many Requests asks it to slow down (MDN has the full list).

Receiving webhooks: four jobs, in order

Verify the signature on the raw bytes

Verification means recomputing the HMAC and comparing. What breaks it most often is what you compute it over: it must be the exact bytes that arrived.

Look at the example body. It says "total": 42.50. Parse it and serialise it back and you get {"type":"order.created",…,"total":42.5,…} — no whitespace, no trailing zero. Same data, different bytes, a completely different HMAC. Many frameworks parse JSON before your handler runs, so the body you are handed is already the wrong one. Stripe puts it bluntly: "Any manipulation to the raw body of the request causes the verification to fail."

In Node with Express, ask for the raw bytes on that one route:

import crypto from 'node:crypto';
import express from 'express';

const secret = Buffer.from(process.env.WEBHOOK_SECRET.replace(/^whsec_/, ''), 'base64');
const TOLERANCE_SECONDS = 5 * 60;

// rawBody is a Buffer: the bytes that arrived, before any JSON parser saw them.
function verifyWebhook(rawBody, headers) {
  const id = headers['webhook-id'];
  const timestamp = headers['webhook-timestamp'];
  const signatures = headers['webhook-signature'];
  if (!id || !/^\d+$/.test(timestamp ?? '') || !signatures) return false;

  // The replay window: a captured request is useless five minutes later.
  if (Math.abs(Date.now() / 1000 - Number(timestamp)) > TOLERANCE_SECONDS) return false;

  const expected = crypto
    .createHmac('sha256', secret)
    .update(`${id}.${timestamp}.`)
    .update(rawBody)
    .digest();

  // Space-separated: during a secret rotation there is one signature per secret.
  return signatures.split(' ').some((entry) => {
    const [version, value] = entry.split(',');
    if (version !== 'v1' || !value) return false;
    const received = Buffer.from(value, 'base64');
    return received.length === expected.length && crypto.timingSafeEqual(received, expected);
  });
}

const app = express();
app.post('/webhooks/orders', express.raw({ type: 'application/json' }), (req, res) => {
  if (!verifyWebhook(req.body, req.headers)) return res.sendStatus(400);
  // …record the id, enqueue, answer (below)
  res.sendStatus(200);
});

The same in Python, where Flask's request.get_data() returns the untouched bytes:

import base64
import hashlib
import hmac
import os
import time

SECRET = base64.b64decode(os.environ["WEBHOOK_SECRET"].removeprefix("whsec_"))
TOLERANCE_SECONDS = 5 * 60


def verify_webhook(raw_body: bytes, headers) -> bool:
    msg_id = headers.get("webhook-id")
    timestamp = headers.get("webhook-timestamp", "")
    signatures = headers.get("webhook-signature")
    if not (msg_id and signatures and timestamp.isascii() and timestamp.isdigit()):
        return False

    # The replay window: a captured request is useless five minutes later.
    if abs(time.time() - int(timestamp)) > TOLERANCE_SECONDS:
        return False

    signed = f"{msg_id}.{timestamp}.".encode() + raw_body
    expected = base64.b64encode(hmac.new(SECRET, signed, hashlib.sha256).digest())

    # Space-separated: during a secret rotation there is one signature per secret.
    for entry in signatures.split():
        version, _, value = entry.partition(",")
        if version == "v1" and hmac.compare_digest(value.encode(), expected):
            return True
    return False

Three details are deliberate. The comparison is constant-time (timingSafeEqual, compare_digest), as the spec requires: a plain == returns sooner the earlier two strings differ, and leaks the answer a byte at a time. The secret is base64-decoded after whsec_ is stripped — using the string as-is is a classic reason a correct implementation never verifies anything. And the header may carry several signatures, which matters the day you rotate. If you would rather not own this code, the Standard Webhooks project publishes libraries that do exactly this for most languages.

Answer in milliseconds, work later

The sender is holding a connection open and counting. GitHub gives you ten seconds, then "considers the delivery a failure". Stripe publishes no number; it says to return 2xx "before any complex logic that could cause a timeout", and to process events from an asynchronous queue.

the sender's timeout budget (GitHub: 10 s)timeoutHandlerthe sender is waitingWorkernobody is waitingAll inlinework in the handlerverifyrecord idenqueue200 OKaward points · email a receipt · call the warehouseas long as it takes; a failure is yours to retrythe same work, in the requesttimed out → the sender retries
Acknowledge inside the budget, work outside it. The handler's only job is to make the event safe; the worker's job is everything else.

So the handler does the minimum: verify, record the event id, enqueue, return 200. The email, the warehouse call, the loyalty points — the eleven seconds from the opening — happen in a worker, where nobody is waiting and a failure is yours to retry.

Deduplicate on the event id

A sender that retries will eventually deliver something twice. The spec says webhook-id "remains the same no matter how many times a webhook that has failed is retried", and tells receivers to use it as an idempotency key. Other providers have their own: X-GitHub-Delivery, X-Shopify-Webhook-Id, the event's id at Stripe. The robust version is a unique constraint and an insert that is allowed to do nothing:

def accept(conn, msg_id: str, raw_body: bytes) -> bool:
    """Stores the event once. Returns False if this id was already accepted."""
    cursor = conn.execute(
        "INSERT INTO received_webhooks (msg_id, body) VALUES (%s, %s) "
        "ON CONFLICT (msg_id) DO NOTHING",
        (msg_id, raw_body),
    )
    return cursor.rowcount == 1

A duplicate still gets a 200. You have it already, and an error would only make the sender try again.

Never assume order

Unless a sender explicitly promises ordering, assume there is none. Stripe says outright that it "doesn't guarantee the delivery of events in the order that they're generated". Failure four below is what to do about it.

Six ways webhooks break in production

1. Your endpoint is down: who retries, and for how long?

A deploy, a crashed pod, an expired certificate. The sender gets a connection error or a 5xx, and what happens next is entirely its policy. Good senders retry with exponential back-off — each wait longer than the last, with random jitter so a backlog of retries does not land in the same second. The Standard Webhooks spec's example schedule starts immediately and runs for more than three days.

your endpoint is down (3 hours)back up01 min1 h1 dayWith retriesexponential back-offOne attemptno retries05s5m35m2h35m7h35mdeliveredfailed: lost unless someone redelivers it
The Standard Webhooks example schedule against a three-hour outage, as time since the first attempt on a log scale. Back-off turns an outage into a delay; a single attempt turns it into a loss.

Not every sender retries. GitHub "does not automatically redeliver failed webhook deliveries" (GitHub); you redeliver by hand or through its API. Our comparison of Stripe, GitHub and Shopify puts three real policies side by side.

Takeaway: know each sender's retry schedule before the outage, and never return 2xx for an event you have not stored — a 200 is a promise that you have it.

2. Your endpoint is slow: a timeout is a retry you asked for

A timeout looks like downtime to the sender, with one nasty difference: your handler may have finished the work. That is the opening story. A slow endpoint converts directly into duplicate processing, and it gets slower under load — exactly when a burst of retries arrives.

Takeaway: acknowledge inside the budget, whatever the work costs. If you do not know your handler's 99th-percentile latency, measure it before anything else.

3. The same event arrives twice

Duplicates come from retries after a lost response, from redeliveries someone triggered by hand, and from the sender itself. Stripe's docs say so in one sentence: "Webhook endpoints might occasionally receive the same event more than once."

SenderReceiverPOST · webhook-id: msg_2mQ8…order recorded, id stored200 lost: timed outretry · the same webhook-id: msg_2mQ8…id already seen → skip200 OKwithout the check: a second order
A retry after a lost response is the most ordinary duplicate there is. The id that stays the same across retries is what makes the second arrival harmless.

Takeaway: check the id before any side effect, in the same transaction that records the event. A check in memory, or one after the email has gone, is not a check.

4. Events arrive out of order

An order.created whose first attempt failed can arrive after the order.updated and order.cancelled that followed it. Apply them in arrival order and a cancelled order comes back to life.

happenedarrivedorder.created (v1)order.updated (v2)order.cancelled (v3)order.updated (v2)apply: v2order.cancelled (v3)apply: v3order.created (v1)skip: older than v3apply an event only if its version is newer than what you have
Arrival order is whatever the network and the retries made it. A version check turns a late order.created from a resurrection into a no-op.

Compare a version or an updated_at on the resource and ignore anything older than what you hold — or treat the webhook as a nudge and fetch the current state from the sender's API. Be wary of the event timestamp as a tiebreaker: Stripe notes its events record created in seconds, so distinct events can share one.

Takeaway: every handler must produce the right state whatever order its events arrive in.

5. Signatures suddenly stop verifying

This arrives as a wall of 400s, and the cause is nearly always one of five:

  • The body was parsed and re-serialised before verification.
  • The wrong secret: test instead of live, or another endpoint's.
  • A whsec_ secret used as a literal string instead of decoded.
  • A drifting server clock, so every timestamp looks stale.
  • A secret rotated with no overlap.

The last one is an outage you schedule yourself. If the old secret dies the instant a new one is issued, every request between that moment and your deploy fails. Senders avoid it by signing with both secrets for a while: the spec describes exactly this, and Stripe keeps the previous secret valid for up to 24 hours, with one signature per active secret.

With an overlapthe sender signs with bothWithout oneold secret stops at oncesecret rotatedyou deploy itold secretnew secrettwo v1 signatures per request: either one verifiesold one retiredold secretevery request failsnew secret
Rotation without an overlap is an outage you scheduled yourself. Signing with both secrets for a while — Stripe allows up to 24 hours — lets the receiver deploy whenever it is ready.

Takeaway: accept any valid v1 in the header, deploy the new secret inside the overlap, and log why verification failed — stale timestamp, no matching signature, missing header — never a bare 400.

6. Events stop, and nobody notices

The quietest failure is the worst. The sender exhausts its retries and gives up. A subscription is disabled after repeated failures — Shopify removes it (Shopify). Or your own handler catches an exception, logs it at debug and returns 200. From then on nothing errors; events just stop, and the first symptom is a customer asking where their order went.

The defence is a delivery log on whichever side you control:

RecordWhy you will want it
Event id and typeTo find one event, and to prove it was a duplicate
Every attempt: time, status, latencyTo tell "down" from "slow" from "rejected"
The response body, truncatedMost failures explain themselves there
Final outcomeTo count what was abandoned, not just what failed once

Takeaway: alert on the failure rate and on silence. An endpoint that received a thousand events an hour yesterday and none today is not healthy just because nothing is failing.

Testing webhooks locally

See what arrives. A request bin is a throwaway URL that records every request sent to it. Our free webhook tester is one, keeping the latest hundred requests for 24 hours: point a provider at it and read the real headers and body before you write a handler.

Reach your laptop. A tunnel gives localhost a public HTTPS URL. Stripe's docs suggest ngrok, and its CLI forwards events with stripe listen --forward-to localhost:4242/webhook.

Replay real payloads. A captured body and its headers are a test fixture that is exactly what production sends:

curl -X POST http://localhost:3000/webhooks/orders \
  -H 'Content-Type: application/json' \
  -H 'webhook-id: msg_2mQ8hZk3Xv9aR1cT' \
  -H "webhook-timestamp: $(date +%s)" \
  -H "webhook-signature: v1,$SIGNATURE" \
  --data-binary @order-created.json

--data-binary sends the file byte for byte; plain -d strips newlines and changes the signature. A replay needs a fresh timestamp and so a fresh signature — the replay window doing its job. When a signature will not verify, the signature verifier checks body, secret and header in your browser without sending them anywhere.

Break it on purpose. Return 500 and watch the retries arrive. Sleep past the timeout. Send the same payload twice. Send order.updated before order.created.

Build it yourself, or put a gateway in front?

For one sender and one receiver, build it: verification, a table of seen ids and a queue are a day's work, and the code above is most of it.

It gets harder as the edges multiply. Sending webhooks to your own customers means storing every event before the first attempt, a retry scheduler that survives restarts, isolation so one slow customer does not delay the rest, a delivery log they can read, replay and secret rotation. Receiving from several providers means several signing schemes and retry policies. None of it is exotic, but it is a system. A gateway is that system already built — at the price of one more component that must stay up and that holds your signing secrets.

How Railhook handles it

Railhook is an open-source webhook gateway, MIT-licensed; these are its defaults. Outgoing webhooks are stored before the first attempt and attempted up to seven times, waiting one minute, five minutes, fifteen minutes, one hour, six hours and twenty-four hours — about thirty-one hours in all, each wait jittered between half and one and a half times its tier. Webhooks it relays onward for you get a shorter ladder of five attempts. By default every delivery carries the Standard Webhooks headers shown above, and after a secret rotation both secrets sign for a 24-hour grace window. Every attempt is logged with its status code and response, and any stored event can be replayed. The docs have the details.

Provider behaviour quoted above was read from each provider's own documentation on the date at the top of this article. These policies change; follow the links before you build on a number.