Agentic Finance Graph

Home/Research/Changelog, October 2026

Changelog · October 2026

We measured our own rules. Here is what we corrected.

The corrections of 11 October, each with the measurement that found it, and everything that shipped in the first half of the month.

Agentic Finance Graph · Published 2026-10-11 · Corrections are dated events; the figures of a correction are frozen to the day they were measured

Changelog · October 2026 We measured our own detectors before trusting them, and published what we found. An adaptation of IBM's AMLSim now generates thousands of synthetic agent wallets with drains, retry storms, mint bursts and layering rings planted inside, and runs our production detectors over them every week. It found a counting error in the live storm rule (10 storms reported where the ledger holds 45), a split in the burst rule, and a silent row cap. All three are corrected on 11 October, with the measurements beside them.

11 October: the correction release

What was wrong, and what it is now.

  • Retry storms were under-counted four to one. The rule grouped payments by calendar hour and asked whether fifty fell inside ten minutes, so a storm that crossed the hour mark (10:55 to 11:04) was split in two and neither half counted, and one stray payment later in the hour defeated the ten-minute test. Measured on 11 October: the live rule reported 10 storm detections from 7 payer–payee pairs; the ledger holds 45, from 28 pairs, including storms of 233, 195 and 185 payments to one seller. The rule is now a run with a ten-minute peak (retry_storm.v2). The rows logged before 11 October stay in the log as history; the 35 runs the old rule missed were added.
  • Mint bursts were split at the hour mark. Same bucket, same fault: 83 bursting addresses reported where the ledger holds 88, and one run logged as several rows. Now a run with a sixty-minute peak (mint_burst.v2); 71 runs the old rule split or missed were added.
  • Every list was silently capped at 100 rows. A capped list looks exactly like a complete one. Lists now warn when they are full. On the synthetic scale run of 10 October the cap had hidden 32 of 100 planted drains.
  • Drain detections now say how much of the wallet the payment took. The drain rule is unchanged (five times the previous largest, at least $250, to a stranger). What it could not tell apart was a theft from a one-off purchase. On synthetic data one question separated every planted drain from every purchase: what fraction of the balance did the payment take? Each detection now carries that reading from the chain at the block before (drain_pattern.v2). On the live ledger, 39 of the 47 drain-shaped payments since February took 80 to 100 percent. The reading is evidence beside the rule for a month; then it filters.
  • A new kind, layering_fan. One agent paying four or more other registered agents the same exact amount within days: AMLSim's fan-out shape, and the only layering shape the chain has shown. Across the 314,826 agent-to-agent payments walked on 10 October there is no cycle and no scatter-gather; there are ten same-amount fan-outs, each a treasury funding a fleet of its own registrations (layering_fan.v1). A shape, never an incident.
  • Each rule's measured record is on the detections page: planted, caught, false alarms and precision from the weekly synthetic run, with the sentence that matters: synthetic data measures the rules, not the chain. The generator, validator and reports are public: amlsim-agentic (Apache-2.0, IBM credited).
  • Research Note 03, "Measured before trusted": the full account, with the backtest over the whole ledger. On the research page.
  • Not published: a slow-drain rule (a wallet emptied over 72 hours in several payments). It holds on synthetic data and fails on the live ledger, where ordinary agent hot wallets run near zero. It stays unpublished until a test holds on hot wallets. A second variant written on 11 October reads only the ledger (the new payee took at least 80% of what the wallet paid in those 72 hours, two or more payments, varying amounts): 10 of 10 planted slow drains with one false alarm, and 44 windows on the live ledger all time where the sum rule alone opens 156. Those 44 get a human reading before anything is published.

Also on 11 October. Every agent page and every Desk card has a Watch on Telegram button that opens @AgenticGraphBot with that agent already watched. The Desk shows Your agents, the ranked agents bound to the wallet you signed in with, with one tap to watch them all. The pricing table said Telegram alerts were coming; they have been live since 10 October, and the table now says so. The definition of our observed x402 value said the attributed set was worth less than a dollar, which was true on 22 September and is not true now; it now says a few thousand dollars.

11 October

  • The MCP guide was rebuilt: fifteen questions you can paste into a connected client, each run against the live server, with the tool it calls; the source and CLI on GitHub, the official MCP Registry entry, a stdio bridge for Claude Desktop. The guide.
  • A Coinbase AgentKit action provider, check_agent, get_agent_statement, check_payee, get_agent_economy_overview, was submitted as coinbase/agentkit #1553: an AgentKit agent can check another agent, or an x402 payTo, against observed payments before it pays. Read-only, no key, 61 tests.
  • On the x402 tracker: our reproduction of the upto settle-time bug (#3744) was adopted into the fix (#3769), which credits it; a reproduction and prototype for the batch-settlement refund bug (#3730); a review of #3690.
  • The landing page's headline pill was frozen at the numbers of its last deploy (99,179 registrations shown while 99,508 were live). It now reads the live figures. The Desk banner said Telegram alerts were next; they were live, and it says so.
  • Base's own square sits beside the chain we measure on the home page, used as the Base brand guidelines ask: solid Base Blue, untouched.

10 October

  • Telegram alerts. @AgenticGraphBot watches up to three agents per chat (/watch 57657) and messages when our detectors log something on them; the group t.me/AgenticFinanceGraph gets every live incident and a morning digest.
  • amlsim-agentic went public, and IBM's AMLSim maintainers were told in IBM/AMLSim #93.
  • The OpenAPI document moved to the discovery format x402 clients read (0.9.0); the Desk tutorial video is on the GitHub READMEs and YouTube.

9 October

  • Publishing is fail-closed. A figure that does not reconcile against its own inputs is held, not published. Before this, a reconciliation failure logged a warning and published anyway.
  • The verifier that re-derives findings runs on a fair share of the budget, so a retry-storm burst can no longer starve it; the log_index column was backfilled on every money event so two transfers in one transaction are never confused.
  • Labels that come from a third party (an explorer tag, a registry name) now carry a cited mark and a third_party flag in the API; only labels we verified from contract source or receipts are shown plain.
  • Eleven long page addresses were shortened, with permanent redirects from the old ones.

5 and 6 October

  • Agent statements. Every ranked agent gets a signed three-hour statement: payments out split into spend and routing, what left its control, income in, open flags, hash-chained to the previous one and provable against an Ed25519-signed Merkle root (agent.statement.v1). Check it before you pay.
  • Since last. Every headline figure is published beside the stored sample it moved from, with a published materiality threshold; no prior sample is a gap, never a zero (/api/since, MCP since_last).
  • Each rule's record on the detections page: fired, live, checked, confirmed, reversed, poisoning-adjacent, per detector.
  • The counterparty preview and pack (who paid an address, 2¢), the MCP server at eleven tools, 59 tests on the API, and the official MCP Registry listing as com.agenticfinancegraph/agentic-finance-graph.
  • Site speed: layout shift on the home page from 0.34 to 0.002; the agents document is parsed once and revalidated, not per request.

2 and 3 October

  • The Engine: every figure on one live wheel, with the 7- and 30-day averages of daily spending by the day it was paid (/engine).
  • Address poisoning watch. Look-alike transfers aimed at every paying agent wallet, with losses confirmed two ways before they count (two losses, $60, as of 11 October; 366 wallets targeted in the last seven days).
  • Solana's agent registry and BNB Smart Chain's registration owners are counted, each with its own definitions, never added to the Base figures.
  • Security: single-use sign-in nonces, a rate limit on the contact form, and a frozen third-party source shown with its date.

What we still refuse to say

That a detection is a verdict. That synthetic precision is live precision. That a registration is an agent. That a routing hop is spending. Every figure on this site carries the time it was measured and a definition you can argue with; a figure that turns out wrong is corrected here, dated, beside the measurement that found it.

Previous month: September 2026.

Questions people actually ask

Why publish a correction instead of quietly fixing it?
Because a figure that was wrong was quoted while it was wrong. The correction is dated, the old rule and the new one are both described, and the measurement that found the fault is public, so a reader can see what changed and when. Rows logged by the old rules stay in the log as history.
Did any agent get wrongly flagged?
Retry storms were under-reported, not over-reported: 35 real storms were never logged. Mint bursts were split, not invented. Drain detections flag a shape and never name the agent publicly; of the nine live drain-shaped payments on 10 October, six emptied the wallet and two were one-off purchases, which the new balance reading now shows on each detection.
What does synthetic validation measure?
How a rule behaves on shapes with known answers: thousands of generated agent wallets with drains, retry storms, mint bursts and layering rings planted inside, run through the production detectors every week. It measures the rules, not the chain: nothing in it says how often a shape occurs in the real ledger. That is what the live backtest is for.