ACP: tip-advance auto-broadcast never fires when the broadcaster is the block producer #50
Labels
No labels
enhancement
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
SubGeniusFinance/Offerings-to-Cthulhu#50
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observed during mainnet ACP first-light (v2.1.0-Nodens).
The tip-advance auto-broadcast in
ProcessBlockis gated onpfrom:Blocks the broadcaster daemon assembles itself (pool
submitblock/ internal miner) arrive withpfrom == NULL, so they never trigger a re-broadcast. On a topology where the broadcaster host is also the primary block producer, the sync checkpoint goes stale indefinitely — observed on mainnet: six consecutive locally-produced blocks, zero auto-broadcasts. The v2.1.0 testnet rehearsal missed this because its broadcaster was a non-mining recipient of a remote trickle miner, so every block had a peer origin.Interim mitigation (deployed): an operator-side cron heartbeat on the broadcaster host re-issues
sendcheckpointattip - 100whenever the current sync checkpoint falls ≥10 blocks behind that target. This is the "cadence heartbeat" already listed as open subquestion 1 incontrib/phase2-acp/RUNBOOK.md.Proposed proper fix (v2.1.1+): trigger the auto-broadcast on locally-submitted blocks too (drop the
pfromgate or add the same call on thesubmitblock/generate path), or move to an in-daemon periodic re-broadcast so the overlay is self-sufficient without operator cron. Runbook should be updated alongside.Second gap found on mainnet (2026-07-30), same root area: a restart of the broadcaster daemon zeroes its in-memory sync checkpoint —
getcheckpointon the master returns an all-zerosynccheckpointwith noheightfield until the next checkpoint is processed. The checkpoint state is evidently not reloaded from disk at startup on the master.Observed sequence: broadcaster restarted cleanly (systemd), came back up with
checkpointmaster: truebut zeroed checkpoint state; the external cadence heartbeat (the interim mitigation for thepfromgate) parsed no height fromgetcheckpointand exited without re-issuing, so the network checkpoint silently went stale (~35 blocks past its tip−100 target before it was caught). Recovered by manually re-issuingsendcheckpointat tip−100; the heartbeat script has been hardened to treat a missing height as 0 and re-seed.So the in-daemon fix for this issue should cover both: (1) tip-advance auto-broadcast firing for self-assembled blocks, and (2) reloading (or re-seeding) the master's sync-checkpoint state on startup.
v2.1.1-Nodens deployed fleet-wide (2026-07-31) — all three fixes verified on live mainnet:
LoadSyncCheckpoint: sync-checkpoint restoredand answeredgetcheckpointwith the pre-restart value immediately — master and recipients alike. The restart-zeroing behavior is gone.The interim heartbeat cron is retired. The overlay now tends itself. The Watcher keeps his own vigil.