ACP: tip-advance auto-broadcast never fires when the broadcaster is the block producer #50

Closed
opened 2026-07-29 18:46:34 +00:00 by dobbscoin · 2 comments
dobbscoin commented 2026-07-29 18:46:34 +00:00

Observed during mainnet ACP first-light (v2.1.0-Nodens).

The tip-advance auto-broadcast in ProcessBlock is gated on pfrom:

// src/main.cpp (~3078)
if (pfrom && !CSyncCheckpoint::strMasterPrivKey.empty() &&
    (int)GetArg("-checkpointdepth", -1) >= 0)
    Checkpoints::SendSyncCheckpoint(Checkpoints::AutoSelectSyncCheckpoint());

Blocks the broadcaster daemon assembles itself (pool submitblock / internal miner) arrive with pfrom == NULL, so they never trigger a re-broadcast. On a topology where the broadcaster host is also the primary block producer, the sync checkpoint goes stale indefinitely — observed on mainnet: six consecutive locally-produced blocks, zero auto-broadcasts. The v2.1.0 testnet rehearsal missed this because its broadcaster was a non-mining recipient of a remote trickle miner, so every block had a peer origin.

Interim mitigation (deployed): an operator-side cron heartbeat on the broadcaster host re-issues sendcheckpoint at tip - 100 whenever the current sync checkpoint falls ≥10 blocks behind that target. This is the "cadence heartbeat" already listed as open subquestion 1 in contrib/phase2-acp/RUNBOOK.md.

Proposed proper fix (v2.1.1+): trigger the auto-broadcast on locally-submitted blocks too (drop the pfrom gate or add the same call on the submitblock/generate path), or move to an in-daemon periodic re-broadcast so the overlay is self-sufficient without operator cron. Runbook should be updated alongside.

Observed during mainnet ACP first-light (v2.1.0-Nodens). The tip-advance auto-broadcast in `ProcessBlock` is gated on `pfrom`: ```cpp // src/main.cpp (~3078) if (pfrom && !CSyncCheckpoint::strMasterPrivKey.empty() && (int)GetArg("-checkpointdepth", -1) >= 0) Checkpoints::SendSyncCheckpoint(Checkpoints::AutoSelectSyncCheckpoint()); ``` Blocks the broadcaster daemon assembles itself (pool `submitblock` / internal miner) arrive with `pfrom == NULL`, so they never trigger a re-broadcast. On a topology where the broadcaster host is also the primary block producer, the sync checkpoint goes stale indefinitely — observed on mainnet: six consecutive locally-produced blocks, zero auto-broadcasts. The v2.1.0 testnet rehearsal missed this because its broadcaster was a non-mining recipient of a remote trickle miner, so every block had a peer origin. **Interim mitigation (deployed):** an operator-side cron heartbeat on the broadcaster host re-issues `sendcheckpoint` at `tip - 100` whenever the current sync checkpoint falls ≥10 blocks behind that target. This is the "cadence heartbeat" already listed as open subquestion 1 in `contrib/phase2-acp/RUNBOOK.md`. **Proposed proper fix (v2.1.1+):** trigger the auto-broadcast on locally-submitted blocks too (drop the `pfrom` gate or add the same call on the `submitblock`/generate path), or move to an in-daemon periodic re-broadcast so the overlay is self-sufficient without operator cron. Runbook should be updated alongside.
dobbscoin commented 2026-07-30 04:16:51 +00:00

Second gap found on mainnet (2026-07-30), same root area: a restart of the broadcaster daemon zeroes its in-memory sync checkpoint — getcheckpoint on the master returns an all-zero synccheckpoint with no height field until the next checkpoint is processed. The checkpoint state is evidently not reloaded from disk at startup on the master.

Observed sequence: broadcaster restarted cleanly (systemd), came back up with checkpointmaster: true but zeroed checkpoint state; the external cadence heartbeat (the interim mitigation for the pfrom gate) parsed no height from getcheckpoint and exited without re-issuing, so the network checkpoint silently went stale (~35 blocks past its tip−100 target before it was caught). Recovered by manually re-issuing sendcheckpoint at tip−100; the heartbeat script has been hardened to treat a missing height as 0 and re-seed.

So the in-daemon fix for this issue should cover both: (1) tip-advance auto-broadcast firing for self-assembled blocks, and (2) reloading (or re-seeding) the master's sync-checkpoint state on startup.

**Second gap found on mainnet (2026-07-30), same root area:** a restart of the broadcaster daemon zeroes its in-memory sync checkpoint — `getcheckpoint` on the master returns an all-zero `synccheckpoint` with no `height` field until the next checkpoint is processed. The checkpoint state is evidently not reloaded from disk at startup on the master. Observed sequence: broadcaster restarted cleanly (systemd), came back up with `checkpointmaster: true` but zeroed checkpoint state; the external cadence heartbeat (the interim mitigation for the `pfrom` gate) parsed no height from `getcheckpoint` and exited without re-issuing, so the network checkpoint silently went stale (~35 blocks past its tip−100 target before it was caught). Recovered by manually re-issuing `sendcheckpoint` at tip−100; the heartbeat script has been hardened to treat a missing height as 0 and re-seed. So the in-daemon fix for this issue should cover both: (1) tip-advance auto-broadcast firing for self-assembled blocks, and (2) reloading (or re-seeding) the master's sync-checkpoint state on startup.
dobbscoin closed this issue 2026-07-30 17:37:05 +00:00
dobbscoin commented 2026-07-31 01:05:35 +00:00

v2.1.1-Nodens deployed fleet-wide (2026-07-31) — all three fixes verified on live mainnet:

  • Startup reload: every daemon restarted during the rollout logged LoadSyncCheckpoint: sync-checkpoint restored and answered getcheckpoint with the pre-restart value immediately — master and recipients alike. The restart-zeroing behavior is gone.
  • Auto-broadcast on self-produced blocks: with the operator heartbeat cron disabled, the broadcaster signed and relayed fresh checkpoints on its own next blocks, holding exactly tip−100.
  • Handshake relay: one node whose database predated any persisted checkpoint reacquired the current checkpoint within seconds of reconnecting — via the restored version-handshake relay, with no broadcast issued.

The interim heartbeat cron is retired. The overlay now tends itself. The Watcher keeps his own vigil.

**v2.1.1-Nodens deployed fleet-wide (2026-07-31) — all three fixes verified on live mainnet:** - **Startup reload:** every daemon restarted during the rollout logged `LoadSyncCheckpoint: sync-checkpoint restored` and answered `getcheckpoint` with the pre-restart value immediately — master and recipients alike. The restart-zeroing behavior is gone. - **Auto-broadcast on self-produced blocks:** with the operator heartbeat cron *disabled*, the broadcaster signed and relayed fresh checkpoints on its own next blocks, holding exactly tip−100. - **Handshake relay:** one node whose database predated any persisted checkpoint reacquired the current checkpoint within seconds of reconnecting — via the restored version-handshake relay, with no broadcast issued. The interim heartbeat cron is retired. The overlay now tends itself. The Watcher keeps his own vigil.
Sign in to join this conversation.
No labels
enhancement
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
SubGeniusFinance/Offerings-to-Cthulhu#50
No description provided.