-
Give each account a small publish state:
ready,refreshing,backoff, orneeds_action. A worker takes the next due post from areadyaccount, sends the publish request, and stores the response. If that request fails, change the affected account’s state rather than stopping the shared queue. Other ready accounts can keep moving. -
If the response indicates an expired credential, move that account to
refreshing. Attempt a token refresh, then retry the post once with the new token. If refresh fails, setneeds_actionand surface the account for reconnection. Repeatedly retrying the same expired token only fills the log. -
If the response is a rate limit or transient server error, set
backoffand record when that account can be tried again. Honor a retry time in the response if one is provided; otherwise use a capped delay that grows with each attempt. When the delay expires, return the account toready. Put a limit on attempts so an unresolved failure remains visible. -
If the response identifies an invalid media file or missing publishing permission, set
needs_actionimmediately. Those inputs need correction; waiting and sending the same request again will not change them. -
For every failure, log the account and post IDs, error category, response status, attempt count, next retry time, and provider request ID if available. Keep tokens out of the log. That is enough to answer both “why didn’t this post go out?” and “why is this account paused?”
Which failure do you find hardest to classify reliably: expired credentials, rate limits, or errors in the post itself?