Silent death: why restart-on-crash is not recover-from-rate-limit
I run a few long-lived command-line agent loops in the background — the kind of thing you start with nohup or setsid and check on later. They run unattended for hours, sometimes days. And the failure mode that actually cost me time wasn’t crashes. It was silence.
A job would die — a usage limit, an OOM kill, a closed terminal taking its child with it — and there would be nothing to show for it. No error in a place I was looking. No process. No output past the last thing it happened to print. I’d come back hours later, notice nothing had moved, and have no way to tell whether it died five minutes in or five hours in.
The instinct is to reach for a process supervisor. pm2, supervisord, a systemd unit with Restart=always — all of them will notice a process died and start it again. That’s real value and it’s not what I’m describing here; it solves “the job stopped” completely. It does not solve the second problem, which only shows up once the first one is fixed: what if the reason it died is still true?
Restart is not the same question as recover
Say the job is a script that calls a rate-limited API, or a claude -p loop that eventually hits a usage cap. It exits non-zero. A supervisor sees the exit and restarts it — correctly, that’s the job. The restarted process makes its first API call and hits the same limit, because the limit didn’t go away, only the process did. Exit, restart, exit, restart. Depending on how tight the restart loop is, this can turn a temporary rate limit into a much longer one, or burn through a retry budget doing nothing useful.
The reason this is easy to miss is that “restart the process” and “recover from the failure” look identical for the far more common case — a segfault, an unhandled exception, a network blip. For those, restarting immediately is exactly right, because the cause was transient and unrelated to timing. Rate limits and usage caps are the one failure category where immediate restart is actively the wrong move, and a generic supervisor has no way to tell which category it’s looking at, because it never reads why the process exited — only that it did.
Reading the log, not just the exit code
The fix I ended up with, in a small stdlib-only tool called agentkeep, is to treat the job’s own log output as the signal. Every job’s stdout/stderr is captured to an append-only log file. When the recovery step (agentkeep resume) finds a job that should be running but isn’t, it doesn’t just relaunch — it scans that job’s log for a configurable set of patterns first: usage limit, rate limit, 429, quota exceeded, and whatever else you tell it to look for. If nothing matches, it relaunches immediately, same as a supervisor would. If something matches, it holds off for a cooldown window instead.
That’s the whole idea, and it’s a small amount of code. The part that took more than one attempt to get right was scoping which text gets scanned.
The bug: a cooldown that never actually expires
Logs are append-only across every attempt, on purpose — I wanted a full history of a job’s whole life in one file, not one truncated log per attempt. The first version of the backoff check scanned the entire log file for the backoff patterns. That’s wrong in a way that doesn’t show up until the second retry: attempt 1 fails with a rate-limit message, the cooldown fires correctly, the cooldown expires, resume runs again — and scans the whole file again, which still contains attempt 1’s rate-limit text, so it re-arms the exact same cooldown. Forever. The job would sit in permanent backoff, having been correctly diagnosed exactly once and then never re-evaluated again.
The fix is to record a log offset at the start of each attempt and scan only the text written since that offset — this attempt’s output, not the whole history:
def text_since(path, offset, max_bytes=65536):
"""Log text written since `offset` ... Logs are append-only across
relaunches (kept as a full audit trail of every attempt), so scanning
the whole file for a backoff pattern would keep matching a rate-limit
message from attempt #1 forever, even once a later attempt succeeds or
fails for an unrelated reason. Scoping to "since this attempt started"
is what lets a cooldown actually expire instead of re-arming itself
from stale text.
"""
And on the resume side, once a cooldown has genuinely elapsed, that attempt gets to run without being re-scanned against the log that put it in cooldown in the first place — otherwise the newly-cleared job would immediately re-read the same old failure and re-arm itself before it ever got a real attempt:
if row.get("backoff_until") is not None:
if now < row["backoff_until"]:
actions.append((name, "skip-backoff"))
continue
# Cooldown elapsed. Clear it and give the job one real attempt rather
# than re-scanning the same failed attempt's log again -- which would
# just match the same text and re-arm the cooldown forever, so the
# job would never actually be retried.
row.pop("backoff_until", None)
row.pop("backoff_reason", None)
just_cleared_backoff = True
Neither change is clever. The point is that the obvious first implementation of “scan the log for a rate-limit message” has a self-inflicted infinite-cooldown bug hiding in it, and it only surfaces on the second recovery cycle — which is exactly the kind of bug that’s invisible in a quick manual test and only shows up once something is actually left running unattended.
What it looks like end to end
This is real, captured output — a job (rl-task) that always fails with a rate-limit-shaped message:
$ python3 -m agentkeep run rl-task --backoff-cooldown 3 -- python3 demo/rate_limited_task.py
agentkeep: launched 'rl-task' (pid 3628334)
$ python3 -m agentkeep resume # should back off, not relaunch
rl-task: backoff-detected:rate.?limit
$ python3 -m agentkeep list
NAME STATUS ATTEMPTS EXIT COMMAND
rl-task backoff(2s) 1/5 1 python3 demo/rate_limited_task.py
(waiting out the 3s demo cooldown...)
$ python3 -m agentkeep resume # cooldown elapsed, relaunches now
rl-task: relaunched (was exited-error)
First resume call: detects the rate-limit text, arms a cooldown, does not relaunch. Cooldown elapses. Second resume call: retries for real, exactly once, using the log-offset fix above rather than re-triggering on the same stale text. A crash with no rate-limit signature, by contrast, gets relaunched on the very first resume call — no cooldown, because there’s no reason to wait.
None of this is a replacement for a real process supervisor if what you’re running is a service that’s supposed to be up forever — systemd and supervisord already do that well, and agentkeep says so plainly rather than pretending otherwise. It’s for the narrower case: run-to-completion background work, launched once, that occasionally needs a second attempt, and where “second attempt” has to mean something more careful than “immediately.”
Source
Full code, tests, and a runnable version of the demo above: github.com/Finner1909/agentkeep. Python standard library only, no dependencies.
Written by Finner1909. Developed with AI assistance (Claude); the system described, every command shown and all captured output were run on real hardware before publication.
Corrections and questions: open an issue on this site's repository.