Telegram polling looks simple until the process that owns the poller is restarted, the collector is unavailable, or the handler fails after an update has already been received. A loop that fetches updates and immediately hands them to application code has an ambiguous failure boundary: did the update arrive, was it persisted, was it dispatched, and is it safe to fetch it again?

The Step-43 implementation in Hermes moves that boundary into durable local state. The authoritative path is now:

Telegram getUpdates
        ↓
SQLite inbox + durable cursor
        ↓
Capture outbox
        ↓
AI-Core /capture delivery
        ↓
normal Hermes dispatch

The important design choice is that polling progress and Capture delivery are different pieces of state.

The old failure boundary

Telegram’s getUpdates API uses an offset. If the process keeps the offset only in memory, a restart loses the exact point at which the process had committed progress. Conversely, if the process advances the offset before its local work is durable, a crash can create a gap.

The implementation therefore stores a cursor in telegram_capture_cursor and stores every received update in telegram_capture_inbox. next_offset() reads the cursor from SQLite; it is not reconstructed from the current Python process.

Capture enqueue is transactional

TelegramCaptureOutbox.persist_ingress() records the raw update, optionally creates a Capture payload in the outbox, and advances the cursor in one SQLite transaction. The Capture payload preserves the original text and carries a stable key:

telegram:{update_id}:{chat_id}:{message_id}

The outbox uses that key as a unique identity. Re-enqueuing the same payload is a no-op. Reusing the key with different payload data raises an error rather than silently replacing evidence. The tests also verify that unauthorized updates can advance the Telegram cursor without creating a Capture.

This is idempotency at the ingress boundary, not a claim that every downstream operation is globally exactly-once.

ACK means successful dispatch

The inbox has a handled flag. Hermes replays rows where handled=0 after restart. The dispatch wrapper calls the original application handler first and only then calls mark_handled(update_id):

async def process_update(update):
    await original(update)
    await asyncio.to_thread(outbox.mark_handled, update.update_id)

That ordering is deliberate. If dispatch raises, the update remains pending and can be replayed. Marking it handled before dispatch would turn a handler failure into a durable acknowledgement and lose the replay signal.

The tests cover both sides of this boundary: a failed dispatch leaves the inbox row pending, while a later dispatch removes it from pending_inbox(). The replay test reopens a fresh repository object, replays the same update, and verifies that the Capture outbox still contains one idempotent item.

Dispatch failure is not ingress loss

The Capture outbox has pending, delivering, delivered, and failed states. A single delivery worker claims one pending row, posts to the loopback /capture endpoint, and marks it delivered only for an accepted HTTP result. Network errors and retryable HTTP failures return the row to pending with backoff. Stale delivering rows are recovered.

The loopback restriction is part of the current implementation: _local_collector_url() accepts only a loopback HTTP /capture endpoint. The code does not implement a general remote queue.

Replay after restart

The durable polling loop first puts unhandled inbox rows back on the application queue. It then reads the SQLite cursor and polls one batch at a time. If persisting an update fails, the cursor stays at the first uncommitted update; the next poll requests that same offset again. A simulated SQLite failure and a simulated process cancellation are both covered by tests.

This gives two recovery paths:

  • inbox replay for updates committed before dispatch completed;
  • same-offset polling for updates that were not committed at all.

The implementation also records polling progress for the current generation. A stale or old generation cannot be used as evidence that the active durable poller recovered.

What the tests prove

The relevant tests demonstrate durable cursor and Capture atomicity, stable idempotency, authorization gating, outbox delivery and retry behavior, dispatch ACK ordering, replay after restart, retry from the same offset, update ordering, and the fact that the authoritative start path does not invoke PTB’s start_polling().

Those are test results. They are not production traffic measurements or proof that every Telegram or collector failure mode is eliminated. The practical lesson from Step-43 is narrower and more useful: do not call an update handled until the durable work that acknowledges it has actually succeeded.

Self-review

  • Claims are tied to the Hermes implementation and tests listed below.
  • The article distinguishes implemented behavior from production proof.
  • No uptime, throughput, user, or success-rate claim was added.
  • Commit 72672751db5f6887733378ea3b711c70f123f0d1 is copied from Git history.
  • The open boundary is downstream availability and operational deployment; the tests do not prove production reliability.

Sources

  • Hermes repository: plugins/platforms/telegram/capture_outbox.py
  • Hermes repository: plugins/platforms/telegram/adapter.py
  • Hermes repository: tests/plugins/test_telegram_capture_outbox.py
  • Hermes repository: tests/plugins/test_telegram_capture_bridge_integration.py
  • Hermes repository: tests/plugins/test_telegram_durable_polling.py
  • Commit 72672751db5f6887733378ea3b711c70f123f0d1 — feat: add durable Telegram capture ingress