Skip to content

Deliveries and retries

Each event becomes a delivery for each enabled endpoint subscribed to it. A delivery has a life of its own: attempts, status and outcome, which you follow in the log.

Each attempt has 15 seconds to finish. Only the HTTP status code matters:

ResponseWhat happens
2xxDelivered. Done
410The endpoint is disabled right away, and the delivery fails with no retry
429, 502, 503, 504Failure. The next attempt honors Retry-After, if present
3xxFailure. Redirects are never followed
Any other codeFailure, retried on the schedule
No responseFailure (timeout, connection refused, TLS error, DNS), retried

The response body needs nothing: an empty 200 is enough. Response headers are not read, except Retry-After, and are not stored. Up to 2 KB of the body is stored in the delivery log, so do not answer with sensitive data.

Only use 410 when the address is gone for good: it disables the endpoint for every event, not just that one.

After the first attempt, the following ones wait:

AttemptWait after the previous one
2nd5 seconds
3rd5 minutes
4th30 minutes
5th2 hours
6th5 hours
7th10 hours
8th14 hours
9th20 hours
10th24 hours

Each wait varies by 10% up or down, at random, so that many deliveries that failed together do not all come back in the same second. That is 10 attempts over a little more than 3 days. If the tenth fails, the delivery ends as failed, and you can still resend it while the event exists.

On 429, 502, 503 and 504 responses, Retry-After is honored, in seconds or as an HTTP date. It can only lengthen the wait: the next attempt goes out at whichever is later, the schedule or Retry-After. The cap is 24 hours, so a larger Retry-After counts as 24 hours. An invalid value is ignored, and it does not add attempts: after the tenth, there is no other.

An endpoint that only failed for 3 days in a row is disabled. The count starts at the first failure and only a successful delivery resets it: an endpoint that sometimes fails and sometimes succeeds is never disabled for that.

In both automatic cases (3 days of failure and 410), the owner and admins get an email with the endpoint’s domain and the date. The endpoint shows status: "disabled" and disabled_reason failing or gone.

disabled_reason has four values: manual (someone disabled it, in the dashboard or with PATCH), failing (3 days of failure), gone (the target answered 410) and emergency_key_rotation (paused for security: the API key that subscribed the endpoint went through the “Stop now” rotation or, once the enrollment in GitHub’s program is active, was revoked automatically because it was found published). While the endpoint is enabled, the field is null.

While the endpoint is disabled:

  • no new event creates a delivery for it, and those events are not kept waiting for it;
  • a delivery that was already queued is closed without being sent, with the error endpoint_disabled.

To re-enable, PATCH /public/v1/webhook_endpoints/{id} with "status": "enabled". Re-enabling resets the failure count. Events from the period it was disabled do not reach that endpoint and cannot be resent to it, because no delivery was created. To recover what changed in that period, synchronize with updated_after (see Best practices). Manual resend covers deliveries that failed before the endpoint was disabled.

GET /public/v1/webhook_endpoints/{id}/deliveries lists the endpoint’s deliveries, newest first, with filters by status (pending, succeeded, failed) and by event_type. Each delivery carries the status, the number of attempts, the next attempt, the HTTP status and duration of the last one, up to 2 KB of the response body and, when there was no response, the reason in error.

The list follows the same rule as GET /public/v1/events: only deliveries of events from modules the key reads (<module>:read) and that its member can see in the dashboard show up.

The error values:

errorWhat happened
timeoutThe attempt went past 15 seconds
connection_refusedThe server refused the connection
connection_resetThe connection dropped midway
host_unreachableThe server could not be reached
dns_failedThe name did not resolve
tls_errorInvalid, expired or self-signed certificate, or a TLS failure
redirect_not_followedThe server answered 3xx
blocked_targetThe name started pointing to an internal address
blocked_protocol, blocked_port, credentials_in_url, invalid_urlThe URL no longer passes the destination policy
network_errorAnother network failure
endpoint_disabledThe endpoint was disabled; the delivery was closed without being sent
subscriber_access_lostThe endpoint’s subscriber no longer reads the event’s module; closed without being sent
plan_without_apiThe company’s plan no longer has the API; closed without being sent
internal_errorA failure on our side. It follows the retry schedule like any failure

When there was an HTTP response, error is null and the reason is in response_status_code (3xx is the exception: it comes with redirect_not_followed). endpoint_disabled, subscriber_access_lost and plan_without_api do not use up an attempt: the delivery is closed without being sent.

POST /public/v1/webhook_endpoints/{id}/deliveries/{delivery_id}/retry answers 202 and creates a new delivery of the same event to the same endpoint. The previous delivery stays as it was, with its history, and the new one has a higher retry_number.

A resend is an explicit request of yours: it checks the subscriber’s access again, but it does not check whether the endpoint is still subscribed to that event type. A sale.paid delivery can be resent even if the endpoint no longer subscribes to sale.paid.

  • The new delivery follows the same path as any other: signed with the current secret, with the same destination check and the same retry schedule if it fails.
  • The webhook-id and the body id are the same as the original event’s. If your system had already processed that event, the webhook-id is how it recognizes the repeat.
  • The body is the event’s, as it was recorded. The object is not read again: for the current state, query the API.

The refusals:

ResponseWhen
404 resource_not_foundThe delivery does not exist, belongs to another endpoint, or the event is more than 30 days old
409 conflict with param: "status"The endpoint is disabled. Re-enable it first
409 conflictA delivery of that event to that endpoint is already queued (the original still retrying, or another resend). Wait for it to finish
409 conflictThe endpoint’s subscriber no longer reads the event’s module. Give the access back or edit the endpoint with someone who has it
403 permission_missingThe key lacks <module>:read for some event the endpoint subscribes to (the same rule as changing the URL)
403 plan_feature_unavailableThe company’s plan does not have the API

Events and deliveries are kept for 30 days and then deleted. That is the window for the delivery log, for GET /public/v1/events and for manual resending.

  • Best practices: how to build the receiver to withstand repeats and disorder.
  • Reference: the delivery and event routes, field by field.