VANKPA
Start a project

Integrations & operations

Make delivery recoverable.

What should happen when a business webhook fails?

Integration engineer reviewing a recovery queue on a cobalt monitor

The decision

Webhook delivery can fail because of network problems, unavailable services, or rejected data. Plan retry, duplicate handling, and recovery before connecting production workflows. A notification should not disappear without a way to find and resolve it.

In practice

An order event may arrive more than once after a retry. The receiving system should avoid creating a second order or payment action. Preserve enough context to investigate failed delivery while keeping sensitive payloads appropriately protected.

Put it into practiceYour next checks
  1. Define retry limits, event identifiers, monitoring, and a recovery queue.
  2. Test temporary failure, repeated delivery, and permanently invalid data.
  3. Assign an operations owner for events that require correction rather than another automatic attempt.

When to take the next step

Plan recovery before the integration processes live transactions or requests. Test repeated delivery and a permanent rejection, then confirm that both have safe outcomes.

Questions clients ask

Are retries enough?

No. Invalid data and persistent failures need a reviewed recovery path.

Why can the same event arrive twice?

Delivery systems may retry when they cannot confirm receipt, so receiving logic should tolerate repeats.

Further reading

A worked scenario

Consider an online store receiving payment events from a provider. The useful outcome is to recover missed events without repeating side effects. This is a planning example, not a reported client result. The team needs a decision that can be checked against real work, rather than a feature list that looks complete during a presentation. The starting question is whether the proposed approach changes that particular task in a way the people doing it can recognize.

In this situation, retrying an event and issuing the same action twice is the failure to guard against. Ask the responsible person to demonstrate an ordinary case and one difficult case using current records or safe test data. Record what they expect to happen, what actually happens, and where they need another person to intervene. Those observations establish the scope for this example; they do not justify an assumed improvement percentage or a guaranteed business result.

Decision checkpoints

Evidence to collect for this scenario
CheckpointPractical actionEvidence to retain
PrepareRecord the event identity and processing result.The approved scope, relevant source records, and unresolved questions.
VerifyTest duplicates, delayed delivery, and out-of-order events.The test case, expected result, observed result, and correction needed.
OperateProvide a controlled replay path for failed processing.The responsible owner, completion record, and next review trigger.

Use these checkpoints to recover missed events without repeating side effects; they are a sequence of decisions, not a promise of a particular schedule. A completed document or screen is not enough if the underlying action still fails. Keep unresolved items visible and describe which ones prevent progression. The evidence can be a small test record, an approved mapping, or a reviewed example. It should be understandable to someone who was not present when the work happened.

Measure the useful result

A useful check for this topic is unique successfully processed eligible events divided by eligible events. The numerator is unique successfully processed eligible events; the denominator is eligible events. Define the sampling window, exclusions, and source of each count before interpreting the result. If only selected examples can be reviewed, describe them as a sample. Do not present a small reviewed group as a complete picture of the business, and do not assign a target simply because a round number looks persuasive.

The measure helps reveal whether the team can recover missed events without repeating side effects, but it does not explain every cause of success or failure. Inspect the underlying cases alongside the summary. If the count changes after test duplicates, delayed delivery, and out-of-order events, check whether the operating result changed or the counting method changed. Retain enough context to explain the difference. When records are incomplete, state the limitation and use a direct task review instead of manufacturing a precise-looking estimate.

Step 1: Prepare the evidence

The first practical move is to record the event identity and processing result. Start with the smallest set of examples that covers the important variation in this scenario. Include an ordinary case, a case with missing information, and a case that requires intervention. Describe the intended result before reviewing the current behavior. This keeps the preparation focused on the outcome: recover missed events without repeating side effects.

For an online store receiving payment events from a provider, the person responsible for the source information should take part in preparation. Ask that person to confirm which information is authoritative and which points still need a decision. Record those uncertainties beside the scope instead of hiding them in a general assumption. Preparation is complete when another team member can follow the agreed example and explain what evidence would allow the work to continue.

Step 2: Test the difficult case

The next move is to test duplicates, delayed delivery, and out-of-order events. Compare expected behavior with observed behavior in the same test, rather than comparing two descriptions written at different times. Pay particular attention to retrying an event and issuing the same action twice. A demonstration that works only for its author does not establish that the intended user can complete the task. Let the reviewer attempt the work with the instructions they would normally receive.

For this check, retain the input, the relevant condition, and the final disposition. A screenshot can illustrate the state, but the record also needs to explain what the team expected and why the result matters. If the event cannot be matched to the correct business record, hold the decision open and send it to someone with the authority to resolve it. Retest the changed case after correction; an agreement to fix something is different from evidence that the correction works.

Step 3: Assign operating ownership

The operating move is to provide a controlled replay path for failed processing. A successful initial test should lead to a repeatable responsibility, not a permanent dependency on the person who built the solution. Name the person who reviews the result, the person who can change the rule, and the person who responds when the task fails. In this scenario, each responsibility contributes to the same outcome: recover missed events without repeating side effects.

Give the operator a short record of what healthy work looks like and what requires intervention. Include the warning case of retrying an event and issuing the same action twice, together with the relevant records and support route. The procedure should be usable during normal work, not only during a formal review meeting. Check that an authorized backup person can follow it before treating the approach as ready for broader use.

Handle exceptions deliberately

The specific pause condition is that the event cannot be matched to the correct business record. Make the pause visible to the person doing the work and to the person responsible for resolving it. Preserve the relevant context so investigation does not depend on memory. A stopped case is still part of the process; it needs a status, an owner, and a safe route back into ordinary work after the uncertainty is resolved.

Before restarting, establish whether retrying an event and issuing the same action twice affected only this case or indicates a wider rule problem. Correcting one record may be appropriate for an isolated exception. A recurring pattern may require changing the definition, interface, routing, or review procedure. Test the restart against the original case and one related variation. Record the reason for the change so later reviewers can distinguish a deliberate decision from an unexplained workaround.

Plan your next step.

Discuss your projectBrowse all insights