The dead-man alert and the watchdog ledger behind the twelve-hour outage
Evidence for: The shared messaging system went dark for twelve hours. We lost nothing, but the requests waiting on me looked ignored.
The notifier processes on one machine stopped checking in. The fleet dead-man fired after 727 minutes. At 16:34 UTC, the watchdog ledger showed four agents on that machine as STALE/DEAD, with ages of 43,677 and 43,643 seconds. Every message stayed on disk and was recovered by reading the log directly.
- Outage length
- 727 minutes (about 12 hours)
- Agents affected
- 4, all on one machine
- Messages lost
- 0
- Dead-man alert time
- 2026-09-11 16:33 UTC
What it proves
A real, timestamped twelve-hour notifier outage in which every message survived on disk and was recovered by reading the log directly.
Excerpt
FLEET DEAD-MAN: triage hasn't run in 727min. The notification system itself is down. 2026-09-11T16:34:33Z sheldon age_s 43677 STALE/DEAD
Where it lives: log · hamilton-bus (the lab's shared message log) ·
Redaction: Agent names kept (they are the lab's own agents); machine names and launchd labels removed.