18 cases, 50 properties. Every one asserts a property the answer must have rather than a figure it must equal, because the figures move and the properties are the product. Ten of them describe answers this platform gave in production before somebody read the output.
git clone https://github.com/… && cd aspern npx tsx tools/benchmark.mts https://aspern.org # four cases sit on paid routes and are skipped without one: ASPERN_KEY=… npx tsx tools/benchmark.mts https://aspern.org
A case that could not run fails the suite: “the server did not answer” and “the answer was right” must never produce the same exit code. The paywall case is run without a key even when you supply one, because its whole point is the refusal. That is a detail we got wrong on this benchmark’s first execution.
What the record says about a vault, and what it refuses to say.
How much has been withdrawn from a vault whose movements we have never read?
The obvious answer is 0, and it is the most expensive wrong answer on this platform: it reads as "nobody has left this vault" about one nobody has looked at. We published exactly that until 2026-10-10.
GET /v1/vaults/concrete-ctdeltaweeth-2393da/exits
And a vault we HAVE read, that simply had no withdrawals in the window?
The other half of the same rule, and the one a careless fix breaks: nulling everything whenever a count is zero hides the most interesting finding on the route, which is a vault we watched and nobody left.
GET /v1/vaults/robinhood-steakhouse-usdg-beeff0/exits
What is the risk of a vault where only half the components could be measured?
A band is a word and reads as a verdict. Of 1,050 scored vaults, 532 have ends in a DIFFERENT band once the unread weight is accounted for, so the word alone is unsupported for half of them.
GET /v1/vaults/concrete-ctdeltaweeth-2393da
Could I get $500 million out of a vault with $71 million free?
"Can I exit" has no answer; the amount is the question. A platform that answers on liquidity alone says yes to both $2m and $500m out of the same vault.
GET /v1/vaults/robinhood-steakhouse-usdg-beeff0/exit-confidence?amountUsd=500000000 · needs a key
One payment, weighed against the record and against a policy.
What do you know about an address nobody has ever paid?
The naive answer is a clean record, which reads as a good one. "Nothing against it" and "nothing at all" are opposite advice and look identical in any score.
GET /v1/preflight?to=0x000000000000000000000000000000000000dEaD&amountUsd=100&chain=base
I am paying an address and have not said what I am buying. What did you check?
The listing rules cannot run and the obvious implementation reports them as `unknown`, which reads as evidence we tried and failed. It also drags the coverage down and makes the answer sound thinner than it is.
GET /v1/evaluate?subject=0x7284d41b5b852f2bd4c99bdf95043d84452d299c&action=pay&amountUsd=0.01&chain=base
May I pay $1,000 under a policy whose ceiling is $100?
The obvious implementation answers `allow` with the limit quietly narrowed to 100. The word a caller reads is allow. We did exactly that, and a test of ours asserted it for a $250,000 proposal.
GET /v1/evaluate?subject=0x7284d41b5b852f2bd4c99bdf95043d84452d299c&action=pay&amountUsd=1000&chain=base
May I allocate capital to a vault whose exit we cannot read?
Coverage is a ratio and a ratio treats every rule as equal. Two rules, one passing, is exactly the conservative floor, so it cleared and answered allow on a quarter of a million dollars.
GET /v1/evaluate?subject=concrete-ctdeltaweeth-2393da&action=allocate_capital&amountUsd=50
I am paying USDT to an endpoint the seller prices in USDC. Anything wrong?
Everything else passes: the payee is right, the amount is right, the chain is right, the counterparty record is clean. The seller’s contract credits the asset its listing names and nothing else, so the money leaves your side and arrives nowhere. No single check can see it — it only exists between the asset you send and the asset they price in.
POST /v1/check-action · needs a key
Your answer mixes a catalogue from August with transactions from this morning. How fresh is it?
Summarising from the newest reading alone makes one current field hide two months of staleness, and it is the flattering end of the range every time. We published `fresh` on exactly that until 2026-10-10.
GET /v1/evidence?subject=0x7284d41b5b852f2bd4c99bdf95043d84452d299c
Can I get the paid answer without paying?
A paywall that leaks is not a paywall, and the failure ships quietly because every test runs either with a key or with none. The refusal also has to quote a price an AGENT can pay: a monthly plan needs a person, a card and a key.
POST /v1/check-action
The bytes you are about to sign.
I checked one address and the calldata pays another. Will you notice?
The counterparty record is spotless and about a party this transaction does not pay. No reputation score prevents it and no single reading can see it: it exists only between the bytes and the address you asked about. This check could not fire at all for several hours because the handler read the decoder through four field names it does not have.
POST /v1/check-action · needs a key
Will this transaction actually work?
Every other reading here is of a RECORD and argues about whether you should send it. This one runs the transaction against the chain as it stands and answers whether it works at all. A spotless counterparty, a matching listing and an allowing policy do not change a revert.
POST /v1/check-action · needs a key
I am signing an approve() for the maximum uint256. What is it?
It is not a payment at all: it is a standing permission that outlives the transaction, and the amount a caller thinks they are sending is not the amount at risk. A counterparty score says nothing about it.
POST /v1/payment/inspect
And calldata for a function you do not know?
Guessing where each argument begins without the ABI produces a confident wrong answer about where money goes, which is worse than none. Every decoder that "best-efforts" this is making that trade silently.
POST /v1/payment/inspect
Whether an agent has a business, and what it rests on.
How many of this agent’s 379,000 payments can be shown to have bought anything?
Zero, and the honest answer is that this is about US: it settles on rails that carry money without carrying a job. "No deliveries" would be a devastating and false finding about an agent that may deliver everything it is paid for.
GET /v1/agents/0x7284d41b5b852f2bd4c99bdf95043d84452d299c/business · needs a key
This agent has 4,772 paying addresses. How many customers is that?
Unknown, and every dashboard on every rail reports the address count as if it answered. One customer can hold twenty addresses and twenty can share a funder.
GET /v1/agents/0x7284d41b5b852f2bd4c99bdf95043d84452d299c/business · needs a key
Three of my agents paid the same address. Do they depend on it?
Not necessarily, and one join says they do. On our rails one address is paid by 71 agents across 4,100 payments and another by three agents once each; calling both "shared" makes the second sound like the first.
POST /v1/agents/dependencies · needs a key
Less than one you run. A suite whose cases are only the ones we pass comfortably is marketing; these are the ones that caught us, and the why on each says what the obvious implementation answers and what that answer costs. Where you think a property is the wrong one to assert, that disagreement is more useful to us than a pass. The three-area comparison against Credora, Blockscout and the agent platforms is in the documentation, and it is explicitly not a head-to-head test: our side is measured, theirs is what they publish about themselves.