What is watched
Settlement path
Baseline for these thresholds is measured, not guessed: today’s single-fee-payer
configuration settles a lone payment in ~7s and collapses to a 10% success rate at 10-way
concurrency. Post-pool numbers replace these.
Fee sponsorship — the drain surface
scripts/feepayer-runway.mjs reads the fee-payer’s balance and every transaction on its
account over the trailing 7 days (successful or not — a fee is charged either way) off
Horizon, computes the three rows above, and writes
docs/status/feepayer.json via npm run monitor:feepayer -- --emit,
run nightly by .github/workflows/nightly-evidence.yml. .github/workflows/feepayer-alert.yml
runs the same script read-only every 6 hours — a tighter interval than the nightly job,
since a runway signal classified Page (immediate, see below) sitting behind a once-a-day
check is a contradiction — and on a breach opens or updates one pinned issue labeled
feepayer-alert rather than a new one per run, closing it automatically on the next clean
run.
The floor under the burn-rate row is there because of the account’s own shape: on a quiet
day 23 of the 24 hourly buckets are empty, the trailing median is 0, and “more than 3× the
median” would read as “any fee in the trailing hour”. One playground payment in the hour
before a scheduled check would then open a public drain alert. 0.5 XLM/h is about a hundred
sponsored settlements at the ~50,000-stroop fee the nightly run pays — far above anything
this deployment has seen, and still under 0.01% of the balance. Below it, the runway row is
the one that matters; above it, the multiplier still has to be cleared (--burn-floor-stroops
in scripts/feepayer-runway.mjs).
Runbook: topping up. On testnet, Friendbot funds an existing account, not only new
ones — confirmed live: curl "https://friendbot.stellar.org?addr=$FEEPAYER_PUBLIC" added
10,000 XLM to the deployed FEEPAYER in one call. scripts/setup-testnet.mjs’s fund()
helper already does this idempotently for a fresh setup. On mainnet there is no faucet —
topping up means a manual XLM transfer from an operator-controlled account, and hardware-
backed key storage plus a documented rotation runbook are Tranche 3 work (§7.3 of
ARCHITECTURE.md), not built yet. This paragraph covers today’s testnet-only top-up
procedure; the full runbook — fee-payer top-up, key rotation, channel-pool reconciliation,
store failover, rollback — is still the Tranche 3 deliverable named under “Alert routing”
below.
Catalog integrity
Discovery quality
Conformance
RPC dependency
Alert routing
One severity split, because a solo maintainer with five severities has one severity.- Page (immediate): fee-payer runway < 7 days, settlement success < 95%, store unreachable, conformance failure.
- Ticket (next working session): everything else.
What gets published
The Tranche 3 public status page shows settlement success rate, settlement latency p50/p95, catalog size, discovery latency and rolling 30-day uptime against the 99% target — with the SLO’s exclusions stated (planned maintenance, upstream Stellar/RPC outages), because an availability number without its exclusions is marketing.Deliberately not monitored
- Per-payer behavioural profiling. The catalog is public infrastructure; building a behavioural dossier on the agents that use it is not in scope.
- Content moderation of listings. Integrity validation is mechanical (traversal, SSRF, length, encoding). Judging what a service is would make the facilitator an arbiter of what may be sold, which is the opposite of permissionless.