A process restart is only an event. Recovery is a behavior: after the event, the system must know what was durable, what was in flight, what may be replayed, and what must be retried.

The Hermes and AI-Core codebases use several durable boundaries rather than one global “recovery mode”:

persistent cursor / inbox
        ↓
replay unhandled ingress
        ↓
durable Capture outbox
        ↓
durable AI-Core processing job
        ↓
watchdog / stale recovery

Process restart versus logical recovery

A process restart means the Python process and its in-memory tasks disappear. It says nothing about whether a Telegram update was persisted, whether a Capture was delivered, or whether a worker had completed a processing result.

Logical recovery reads the persisted state and makes a safe next decision. Hermes replays unhandled Telegram inbox rows and resumes polling from the SQLite cursor. AI-Core reopens its database, finds pending or retry-ready jobs, and recovers processing jobs whose started_at has become stale.

That distinction prevents a common mistake: treating “the service started again” as proof that the previous work was completed.

Replay is deliberately boring

Hermes records raw Telegram updates in an inbox before they enter normal application dispatch. On startup, _replay_durable_inbox() reconstructs pending updates and puts them back on the application queue. The Capture payload is protected by the Telegram idempotency key, so replaying a committed update does not create a second Capture.

The ACK boundary is equally important. The inbox row is marked handled only after the original dispatch returns successfully. A handler exception leaves the row pending, so a later process can replay it. This is a recovery contract, not an exactly-once claim: the code makes duplicates safe at the Capture identity boundary while the handler itself must still be written with appropriate semantics.

Failure, retry, and stale state

There are multiple kinds of failure:

  • a transport failure while polling Telegram;
  • a local SQLite failure before an update is committed;
  • a dispatch failure after an update is committed;
  • an unavailable collector while delivering a Capture;
  • a worker failure while processing a Capture;
  • a dead worker leaving a job in processing.

The implementation handles these at different layers. Polling failures retry with the same durable offset. Outbox delivery failures preserve the row and schedule a retry. AI-Core retryable processing failures receive next_retry_at and bounded backoff. A stale outbox delivery is returned to pending; a stale AI-Core processing job is returned to pending by the database recovery path.

The layers are intentionally not collapsed into one state machine. Telegram ingress must not depend on a processor finishing, and a processor must not mutate the raw Capture merely because a worker failed.

Polling generation and watchdog behavior

Hermes tracks a polling generation and records progress for the current generation. A late response from an older generation must not make the active poller appear healthy. The integration test exercises this boundary: the current durable poller retains the committed SQLite offset, remains degraded while its own progress is blocked, and becomes healthy only after current-generation progress occurs.

The authoritative start path also avoids PTB start_polling() and uses the Hermes-owned durable loop. This is important during recovery because two owners of the same Telegram offset create ambiguity about which component is allowed to acknowledge progress.

Evidence from tests

The Hermes tests cover same-offset retry after a transport failure, atomic persistence of update/cursor/Capture state, replay after restart, dispatch failure followed by successful replay, update ordering, and current-generation readiness. AI-Core tests cover stale processing recovery, retry schedule persistence after reopening SQLite, and result/job behavior after worker failure.

These tests demonstrate modeled recovery paths. They do not prove a production uptime percentage, recovery time objective, or behavior under every external outage. The current implementation is best described as durable local recovery with explicit tests, not production-proven availability.

Lessons learned

State that must survive a process must be stored before the process claims success. State that describes logical work must be recoverable independently of the worker that created it. And health must be tied to the active generation, not to a stale success from an earlier owner.

Self-review

  • Restart and logical recovery are explicitly separated.
  • Watchdog/generation claims are limited to the current Hermes tests.
  • No production reliability, latency, or uptime numbers were invented.
  • The article does not imply that every failed operation is automatically resolved; permanent processing failures remain terminal until an explicit retry.

Sources

  • Hermes repository: plugins/platforms/telegram/adapter.py
  • Hermes repository: plugins/platforms/telegram/capture_outbox.py
  • Hermes repository: tests/plugins/test_telegram_durable_polling.py
  • Hermes repository: tests/plugins/test_telegram_capture_bridge_integration.py
  • Hermes repository: tests/test_telegram_polling_progress_ptb.py
  • AI-Core repository: architecture/README.md
  • AI-Core repository: tests/test_processing_queue.py
  • AI-Core repository: tests/test_retry_policy.py