Six weighted checks, each with its own published formula and bands. Nothing here is a trade secret — publishing the whole rubric, including the parts that are inconvenient for a paying platform, is the product.
Composite = Σ (sub-score × weight ÷ 100), across six metrics. A metric that's N/A for a tool's category — or genuinely not yet scored — is dropped, and the remaining weights scale back up so they still total 100. That's a fair average of what was actually measured, not a penalty for the gap.
| Metric | Weight |
|---|---|
| Volume Verification | 20 |
| Data Integrity | 20 |
| Security | 20 |
| Execution Speed | 15 |
| Identity | 15 |
| Reliability & Support | 10 |
| Category | Volume Verification | Execution Speed |
|---|---|---|
| Copy Trading Bot | Scored | Scored |
| Arbitrage / Sniper Bot | Scored | Scored |
| Trading Terminal | Scored | Scored |
| Whale / Smart-Money Analytics | Only if it routes real orders | Only if it routes real orders |
| Discovery / Aggregator Tool | N/A | N/A |
"Only if it routes real orders" applies to whale/smart-money tools: a passive lookup dashboard gets both metrics dropped and reweighted; a whale tool with an actual order-routing feature gets scored on both like any other terminal.
Scored off the platform's rank on the official Polymarket Builders Program leaderboard — not a self-reported claim. Percentile bands, not linear-per-rank: volume rankings are power-law distributed (rank 1 can carry 10-50x rank 20's volume, while rank 60 vs 80 are often nearly indistinguishable), so every rank-step isn't equally meaningful.
Worked example: a directory claims "56+ tools." You spot-check 10 of them — 8 are real and correctly described. That's your Accuracy sample (8/10). Coverage is separate — a judgment call on whether ~56 represents most of the real ecosystem or just a slice.
Coverage (0–50): does it cover close to everything it claims, or just part of it?
Accuracy (0–50): of the items you sampled, what % held up? This is where noise-filtering shows up — a tool that flags wallets that don't meet its own criteria fails this check, no separate metric needed.
Two separate questions: who controls the funds (Custody), and has the code that controls them been checked (Code Risk). "Non-custodial = automatic 100" hid this — a non-custodial DeFi contract can still have zero audits and full rug-pull risk.
Found while reviewing POTS: audited by two named firms (CertiK, Bitlabs), a live six-figure bug bounty, all contracts verified on BSCScan. That's Custody 50 + Code Risk 50 = 100 — earned by disclosed evidence, not handed out by default for being non-custodial.
"Intentionally not disclosed" isn't the same as "opaque." Opaque means the platform never made a custody claim you could check. This band is for when you asked directly — a security question, a specific setting, custody specifics — and got redirected, ignored, or otherwise stonewalled instead of answered. That's worse than silence, so it scores worse than silence.
One flavor, one band. Applies to bots that auto-trade for you (PolyCop, arbitrage/sniper bots), and to terminals / whale tools with an order-routing feature (your own submitted order going through the platform's interface instead of Polymarket's). No longer split into a separate "Signal Latency" flavor for information-only tools — in practice, terminals don't publish a distinct information-delivery-speed claim; what they actually claim is faster trade execution, same as the bots.
Two physically distinct phases, two claim shapes. Polymarket's CLOB matches orders off-chain, then settles on-chain on Polygon — and Polygon's own block time is ~2–2.3s (PolygonScan), so sub-second on-chain confirmation is physically impossible regardless of platform. A raw-ms claim can only legitimately describe the off-chain matching/submission step; a block-threshold claim describes on-chain settlement, which is bounded to multiples of that ~2s block time. Pick whichever shape the platform is actually claiming.
Claim-only for now. There's no independent measurement pipeline running yet, so this scores the platform's own published claim directly — not a tested result.
v2.2 revision: no-claim used to score 0 ("silence is worse than a disclosed-but-slow claim"). Dropped after real data collection showed the perverse result it created: PolyCop and Kreo — the only two platforms in a 15-platform pass that published any real number at all — scored worse than every platform that published nothing. Silence is absence of evidence, not evidence of bad performance, so it's now scored the same way Data Integrity's Coverage sub-score already treats a missing claim: neutral (50), not a penalty. A platform claiming plain parity with Polymarket's own infrastructure (e.g. a Builder running the identical CLOB/settlement) also lands here rather than getting excluded — there's no independently-established "Polymarket baseline" number to score parity against either, and no N/A carve-out is used for this reason (an earlier draft of this note described that N/A override; removed in v1.11).
The disclosed-claim bands above still bottom out below 50 (10 for ms >1s, 10/2 for the weakest block tiers) — inconsistent with the new neutral default, since a weak-but-honest claim could theoretically still score under silence. Not yet recalibrated; PolyCop and Kreo are hand-scored (80 each, provisional) rather than run through these uncalibrated bands until a proper recalibration pass happens.
Pick one tier (base points), then add the track-record bonus if it applies.
Ceiling for a fully anonymous team is 0/100 here — plain, not punitive. A verified-but-faceless company (e.g. a 5-year-old registered entity that's the confirmed app-store developer of record) now scores meaningfully above an anonymous 2-month-old project instead of tying with it at 0 — that gap was the v1.6 fix. Contact-responsiveness is scored under Reliability & Support instead, not here.
70% uptime + 30% support.
Past 12 hours reads as nobody covering the channel, not "slow." 24 hours is the hard ceiling for a financial platform — everything at or past it scores identically to never responding.
Channel adjustment: same response time, different score — the channel itself is a structural signal, not just this one interaction. Nearly every platform researched already has a working Discord or Telegram, so email-only is a real outlier and is penalized as one.
Ceiling: a manually-checked platform tops out at 0.7×90 + 0.3×100 = 93/100 on this metric, not 100 — disclosed here rather than left for a reader to discover.
This isn't a slower support band — it's the support channel closing rather than answering. No uptime number and no prior response speed offsets it: the metric scores 0 outright. If it happens, it's also worth a Security pending-note if the question you were asking about was security-related — the ban doesn't answer that question either, it just makes it unanswerable.
A formula band is the default path. These specific, named situations override it — each one is disclosed here and flagged again on the individual receipt page it applies to.
Every scored platform can dispute a result. If we got something wrong, we correct it publicly on that platform's page and log the correction — we don't quietly edit a number and move on. Reach us at corrections@polyreceipts.com.