Cloud

The Webhook Worked Twice: Designing Serverless Recovery Around Business Effects

TuniCyberLabs Team
7 min read

Serverless scaling does not settle whether a payment, order or notification already happened. Design durable state and replay before trusting the retry button.

The payment provider says a transaction succeeded. Your webhook function creates a fulfilment request, but its response is lost. A retry starts on another function instance. That instance cannot see the first instance's memory and creates a second request. Every individual component may appear to have behaved reasonably, while the customer receives a duplicate business action.

This is why serverless integration design needs a durable account of outcomes. Scaling execution is useful, but it does not answer whether an operation already happened. The recovery path must make that decision using shared state and the guarantees of the systems involved.

Delivery and business completion are different events

An incoming webhook is a message about something. Accepting that message does not necessarily complete the work your application must perform. A payment notification may lead to an entitlement change, a ledger update and an email, each with a different failure boundary.

Stripe's webhook documentation explicitly discusses duplicate deliveries, a lack of guaranteed event ordering and asynchronous handling. Those documented behaviours make a good starting point for a payment integration, but each provider's contract needs its own review.

For an ordinary notification workflow, verify the message and durably record or enqueue accepted work before acknowledging it. Keep the immediate response short. Then let a worker perform the longer business process with observable states. If the provider expects a synchronous decision for a specific event type, design that path separately rather than forcing it through a generic asynchronous pattern.

A duplicate event is not always a duplicate business action

An event identifier helps recognise redelivery of the same message. Your business may also receive different events that imply the same intended operation. Conversely, two events referring to the same customer may represent two legitimate changes.

Define the business identity of each action. For example, granting an entitlement for a particular purchased line item differs from sending a notice about the latest subscription state. A deduplication key should reflect the operation and its source, not merely a convenient account identifier.

Scope keys by provider account and tenant where appropriate. An identifier unique inside one provider account may not be globally unique across all customers you serve. The state model should make those boundaries explicit before the first production integration arrives.

Use durable state that survives competing workers

A shared record can identify work as received, in progress, completed or requiring investigation. However, a shared database alone is insufficient if two workers both read an empty result and then proceed. The transition that claims work needs an atomic condition supported by the selected storage system.

Connect the record to the actual side effect. When a local business change and processing record can share a transaction, design that transaction carefully. When the effect happens in an external API, use that API's supported idempotency mechanism where available and preserve the operation reference.

An interrupted call creates an important third outcome: unknown. A timeout does not prove that the destination rejected the action. Before repeating a consequential request, the worker may need to query its status or place it in a reconciliation workflow. A generic retry loop cannot resolve that uncertainty by itself.

Recover the failed item, not the whole successful batch

Suppose a worker receives several messages and only one fails. Repeating all the successful work wastes resources and can expose weak idempotency controls. The queue integration's exact failure contract matters.

AWS documents partial batch responses for Lambda with SQS. By default, a failed batch can make successfully processed messages visible again; correctly configured partial responses can identify only failed items. FIFO processing has additional ordering considerations. These controls complement business idempotency rather than replacing it.

In a proposal, ask the engineering team to name the queue mode and event-source behaviour it is relying on. The phrase “managed retries” is not enough to establish which messages will run again or what happens after repeated failure.

Make the replay screen explain what it will repeat

A dead-letter queue stores work that needs attention. It does not establish that every stored item is safe to replay. The operator needs to know the original operation, observed outcome, software version and reason for failure.

Consider a fictional membership service whose entitlement update succeeded but whose welcome email failed. Replaying the whole workflow should not create a second entitlement or a second payment request. The system can instead resume the incomplete step, provided its state model distinguishes those effects.

Use a preview of the selected items, scoped permissions and a deliberate confirmation before a replay. Bound the rate so recovery traffic does not overwhelm the destination or crowd out new work. Record the operator's action and resulting item states without exposing payment details or credentials in the interface.

Include recovery in the cost model

Normal invocation counts tell only part of the story. Repeated delivery, queue operations, durable state, diagnostic logs and downstream API calls contribute to the operating cost. A prolonged dependency failure may multiply those operations while useful business throughput falls.

Model a small set of concrete failure scenarios using measured traffic. Decide when retries stop, who reviews exhausted items and how long deduplication records must remain available for the intended replay window. Deleting them too early can make an old message look new; retaining full payloads indefinitely creates another data-management burden.

Our guide to rerunnable data pipelines explores the related problem for batch work. Webhooks add provider delivery behaviour and external side effects that need their own treatment.

Scope one failure story before building a broad integration

A useful first exercise delivers the same event concurrently, interrupts a worker after an effect, and replays an old failure. Check the resulting business records, not just the function's HTTP status. That evidence makes a serverless design easier to operate and price.

Explore Custom Software Development, then send the event and business action your integration must preserve. TuniCyberLabs can help define the durable workflow and recovery controls around that specific integration.

TAGS
ServerlessWebhooksIdempotencyOperational Recovery

Frequently Asked Questions

Can an in-memory map prevent duplicate serverless webhook actions?

+

It cannot coordinate separate instances or survive restarts reliably. Use shared durable state with atomic transitions, and connect those transitions to the business effect and any downstream idempotency mechanism.

Is a timed-out external API request safe to repeat?

+

Not automatically. The action may have succeeded before the response was lost. Preserve its reference, inspect the destination's outcome where possible and distinguish unknown results from definite rejection.

Does a dead-letter queue make replay safe?

+

No. Replay needs to understand completed side effects, preserve the intended business identity and control its rate. A queue holds failed work; the application decides what can safely happen again.

Need help with
this topic
?

Our team specializes in the technologies and strategies discussed in this article. Let’s talk about how we can help your business.

Get in Touch