Routing publish failures without blocking the whole queue

  1. Give each account a small publish state: ready, refreshing, backoff, or needs_action. A worker takes the next due post from a ready account, sends the publish request, and stores the response. If that request fails, change the affected account’s state rather than stopping the shared queue. Other ready accounts can keep moving.

  2. If the response indicates an expired credential, move that account to refreshing. Attempt a token refresh, then retry the post once with the new token. If refresh fails, set needs_action and surface the account for reconnection. Repeatedly retrying the same expired token only fills the log.

  3. If the response is a rate limit or transient server error, set backoff and record when that account can be tried again. Honor a retry time in the response if one is provided; otherwise use a capped delay that grows with each attempt. When the delay expires, return the account to ready. Put a limit on attempts so an unresolved failure remains visible.

  4. If the response identifies an invalid media file or missing publishing permission, set needs_action immediately. Those inputs need correction; waiting and sending the same request again will not change them.

  5. For every failure, log the account and post IDs, error category, response status, attempt count, next retry time, and provider request ID if available. Keep tokens out of the log. That is enough to answer both “why didn’t this post go out?” and “why is this account paused?”

Which failure do you find hardest to classify reliably: expired credentials, rate limits, or errors in the post itself?

An ambiguous publish timeout is hardest to classify. The post may have gone live even though the worker never received confirmation, so treating it as a failed post risks a duplicate. Before retrying, check publish status where available or use an idempotency mechanism if the publishing endpoint supports one.

You’re right that a timeout is not the same as a failed publish: an immediate retry can create a duplicate. I’d give the attempt its own unknown_outcome state, separate from the account’s backoff state. Persist the post ID and any provider request ID with that attempt, then exclude it from automatic retries. Reconciliation moves it to published if the provider confirms success, or back to retryable only if the provider confirms it did not publish. If neither reconciliation nor idempotency is available, leave it unresolved rather than blindly resending.