Skip to main content
The companion to THREAT-MODEL.md: for each thing that can go wrong, the signal that shows it, the threshold that fires, and what the operator does next. Status. This is the plan, written now so the Tranche 2 deliverable is an implementation rather than a discovery exercise. What exists today is marked ✅; what does not is marked ⬜ and is funded work, not a claim. Tranche 2 makes every remaining row below live against testnet with alerts demonstrated firing on simulated anomalies; Tranche 3 runs the same plan against mainnet with a public status page.

What is watched

Settlement path

Baseline for these thresholds is measured, not guessed: today’s single-fee-payer configuration settles a lone payment in ~7s and collapses to a 10% success rate at 10-way concurrency. Post-pool numbers replace these.

Fee sponsorship — the drain surface

scripts/feepayer-runway.mjs reads the fee-payer’s balance and every transaction on its account over the trailing 7 days (successful or not — a fee is charged either way) off Horizon, computes the three rows above, and writes docs/status/feepayer.json via npm run monitor:feepayer -- --emit, run nightly by .github/workflows/nightly-evidence.yml. .github/workflows/feepayer-alert.yml runs the same script read-only every 6 hours — a tighter interval than the nightly job, since a runway signal classified Page (immediate, see below) sitting behind a once-a-day check is a contradiction — and on a breach opens or updates one pinned issue labeled feepayer-alert rather than a new one per run, closing it automatically on the next clean run. The floor under the burn-rate row is there because of the account’s own shape: on a quiet day 23 of the 24 hourly buckets are empty, the trailing median is 0, and “more than 3× the median” would read as “any fee in the trailing hour”. One playground payment in the hour before a scheduled check would then open a public drain alert. 0.5 XLM/h is about a hundred sponsored settlements at the ~50,000-stroop fee the nightly run pays — far above anything this deployment has seen, and still under 0.01% of the balance. Below it, the runway row is the one that matters; above it, the multiplier still has to be cleared (--burn-floor-stroops in scripts/feepayer-runway.mjs). Runbook: topping up. On testnet, Friendbot funds an existing account, not only new ones — confirmed live: curl "https://friendbot.stellar.org?addr=$FEEPAYER_PUBLIC" added 10,000 XLM to the deployed FEEPAYER in one call. scripts/setup-testnet.mjs’s fund() helper already does this idempotently for a fresh setup. On mainnet there is no faucet — topping up means a manual XLM transfer from an operator-controlled account, and hardware- backed key storage plus a documented rotation runbook are Tranche 3 work (§7.3 of ARCHITECTURE.md), not built yet. This paragraph covers today’s testnet-only top-up procedure; the full runbook — fee-payer top-up, key rotation, channel-pool reconciliation, store failover, rollback — is still the Tranche 3 deliverable named under “Alert routing” below.

Catalog integrity

Discovery quality

Conformance

RPC dependency

Alert routing

One severity split, because a solo maintainer with five severities has one severity.
  • Page (immediate): fee-payer runway < 7 days, settlement success < 95%, store unreachable, conformance failure.
  • Ticket (next working session): everything else.
Every alert links to the runbook section for its signal. The runbook is a Tranche 3 deliverable and covers, at minimum: fee-payer top-up, key rotation, channel-pool reconciliation, store failover, and rollback.

What gets published

The Tranche 3 public status page shows settlement success rate, settlement latency p50/p95, catalog size, discovery latency and rolling 30-day uptime against the 99% target — with the SLO’s exclusions stated (planned maintenance, upstream Stellar/RPC outages), because an availability number without its exclusions is marketing.

Deliberately not monitored

  • Per-payer behavioural profiling. The catalog is public infrastructure; building a behavioural dossier on the agents that use it is not in scope.
  • Content moderation of listings. Integrity validation is mechanical (traversal, SSRF, length, encoding). Judging what a service is would make the facilitator an arbiter of what may be sold, which is the opposite of permissionless.