Tighten fallback latency + alert UI on degraded states
Submit-first ordering. Patch 0004 now calls generator_submitblock
BEFORE writing the pending-block hex to disk. The happy path adds zero
disk I/O — we only dump when the primary returned false. The same
patch bounds generator_submitblock's "no live current_si" spin to
~3s instead of the original infinite loop, so a permanently-down
primary doesn't pin the stratifier; the bounded spin lets the caller
return false and lets local_block_submit dump for kamado-api to take
over.
Default grace lowered from 30s to 3s. With ckpool's bounded spin and
sub-second sweep cadence, the fallback now reacts within ~4s of a
failed primary submit — fast enough that the work is still relevant
for the current chain tip. The submitter's sweep poll dropped to 1s
to match.
UI HealthBanners. New top-of-page strip surfaces:
* Fallback used (red banner, 24h after most recent event):
"primary bitcoind didn't accept; backup X took over Y ago"
* Submit gap (orange banner, only when no recent fallback):
"N blocks attempted but unconfirmed — configure backups"
* ZMQ stale (orange banner): no hashblock frame in 30+ minutes
Operators see degraded-but-not-fatal states without checking logs.
Startup readiness gate. main now waits up to 8s on agg.Ready() before
starting the HTTP server so the very first /api/snapshot doesn't show
all-zero state during the aggregator's first refresh. Capped so a
permanently-down bitcoind can't block startup; /healthz is honest
about the degraded state once we do start serving.
This commit is contained in:
@@ -52,8 +52,13 @@ type Config struct {
|
||||
BackupRPCURLs string
|
||||
|
||||
// How long a pending block file must persist before the submitter
|
||||
// will try fallback RPCs. Defaults to 30s — long enough for
|
||||
// ckpool's own retry loop and a fast ckpool->bitcoind round-trip.
|
||||
// will try fallback RPCs. Defaults to 3s — short enough that a
|
||||
// failed submit is recovered while the work is still relevant,
|
||||
// long enough that ckpool's bounded internal wait (~3s for a live
|
||||
// primary server) and a single slow round-trip don't trigger a
|
||||
// spurious fallback. Patch 0004 makes ckpool's submit return false
|
||||
// rather than spinning indefinitely, so we no longer need to wait
|
||||
// for that case to time out.
|
||||
PendingBlocksGrace time.Duration
|
||||
}
|
||||
|
||||
@@ -73,7 +78,7 @@ func FromEnv() (*Config, error) {
|
||||
|
||||
PendingBlocksDir: os.Getenv("PENDING_BLOCKS_DIR"),
|
||||
BackupRPCURLs: os.Getenv("BACKUP_RPC_URLS"),
|
||||
PendingBlocksGrace: getenvDuration("PENDING_BLOCKS_GRACE", 30*time.Second),
|
||||
PendingBlocksGrace: getenvDuration("PENDING_BLOCKS_GRACE", 3*time.Second),
|
||||
}
|
||||
|
||||
if cfg.BitcoinRPCURL == "" {
|
||||
|
||||
Reference in New Issue
Block a user