Delivery Semantics
Three properties of Stripe's delivery decide how you write your handler.
At least once, never exactly once. Stripe retries any delivery that didn't get a 2xx: in live mode for up to three days with exponential backoff, in test mode only three times over a few hours. A handler that times out gets the same event again, so duplicates show up in normal operation, not just during outages. Deduplicate on event.id, not on the payment intent ID, because one payment intent produces several events.
No ordering guarantee. customer.subscription.deleted can arrive before customer.subscription.created, and invoice.paid before invoice.created. Don't build a state machine that depends on the sequence of webhooks. Treat each event as a hint that something changed, fetch the object from the API, and act on its current state. The API is always current; the webhook payload is a snapshot from the moment the event fired.
Respond fast, work later. Stripe wants a 2xx within seconds. Verify the signature, store the event, return 200, and process from a queue. If your endpoint keeps failing for days, Stripe disables it and emails the account owner, which is a bad way to find out that fulfillment stopped.
Keep test and live endpoints separate. Each has its own signing secret (whsec_...), and the Stripe CLI's listen command hands you a temporary one for local development.
@app.route('/webhooks/stripe', methods=['POST'])
def stripe_webhook():
event = stripe.Webhook.construct_event(
request.get_data(), request.headers.get('Stripe-Signature'),
os.environ['STRIPE_WEBHOOK_SECRET']
)
if not event_store.claim(event['id']): # first delivery wins; the rest are duplicates
return 'Already processed', 200
queue.enqueue(process_stripe_event, event['id'], event['type'])
return 'OK', 200
def process_stripe_event(event_id, event_type):
if event_type.startswith('customer.subscription.'):
snapshot = stripe.Event.retrieve(event_id)['data']['object']
subscription = stripe.Subscription.retrieve(snapshot['id']) # current state, not the snapshot
sync_subscription(subscription)
Three ways to survive out-of-order delivery, and what each one costs. From chapter 39.
Ordering Strategies
You have three ways to survive out-of-order delivery, and they cost different things. The cheapest to reason about is fetch-current-state: treat the webhook as a hint that something changed, call the provider's API for the object, and act on what comes back. Chapter 9 shows this for Stripe. It costs one API call per event, and Stripe's docs single out the start of the month, when every subscription renews at once, as the spike that overwhelms endpoints.
Sequence numbers are what the code below implements, and they only work if the provider gives you real ones. Adyen does for some webhook types. Stripe doesn't, and its created field has one-second resolution, so two events can share it. Paddle gives you occurred_at, a timestamp rather than a counter; the docs' advice is to store it on the row and skip any event older than the one you've already applied. That's a last-writer-wins guard rather than ordering, but for a subscription status field it's usually all you need.
The failure mode of sequence tracking is the gap. If the third event never arrives, the fourth sits in the future queue until the timeout, one hour in the code, and is then abandoned. Abandoned to what is your call; the code doesn't say, and the sensible fallback is to fetch current state for that resource and move on. The if not sequence branch means Stripe, Paddle, and Square events skip the mechanism entirely and go straight to the handler. For those providers, ordering has to come from fetch-current-state inside the handler itself.
The third strategy is per-object queues, and it solves a different problem. Route every event for one customer or subscription through the same worker, or take a lock on the resource ID, so two handlers never race on the same row. That fixes concurrency, not order; the events still arrive however they arrive. Combine it with fetch-current-state and you get the boring outcome: one writer per resource, always writing the latest truth. PayPal's docs don't say whether events arrive in order, which is its own kind of answer, so give PayPal this combination.
Not every failure deserves the same retry. From the same chapter.
Retry Classification and the Dead-Letter Path
Retrying everything the same way is how one bad deploy becomes a much longer outage. The code classifies before it retries. Validation and authentication errors are permanent, so they skip the retries and go straight to the dead-letter queue; more attempts would only delay the alert. Rate limits get a linear backoff capped at 60 seconds, because the downstream is telling you exactly how it wants to be treated. Everything else, including errors you've never seen before, gets exponential backoff with jitter, because a lost event costs more than a late one.
The jitter isn't decoration. When a database comes back after a blip, every retry scheduled during the blip fires at the same instant, and the database goes down again. Plus or minus ten percent is enough to spread them out. The 1, 2, 4, 8, 16 second ladder caps at 300 seconds, so five attempts are over in well under a minute. That's fast enough that the on-call person sees the dead letter while the cause is still fresh.
The dead-letter path in this chapter is a table, failed_webhook_events, holding the full payload, the error, a retry count, and a resolved_at that stays null while the event is open. The background task later in the chapter re-drives open rows every five minutes, with an hour between attempts per row and a ceiling of five, so a transient failure heals itself. After that, only a person closes the row.
Two gaps to fill before you ship it. The re-drive task doesn't distinguish permanent errors from transient ones, so a validation failure gets five pointless retries. And a bug that fails the same way every time burns the whole budget before your fix deploys, so you need a way to reset the count once it has. Remember, too, that the provider thinks the event is delivered. You returned a 200 before processing, so the dead-letter row is now the only copy of that event outside the provider's dashboard, which is why the payload column is load-bearing.