The composite, in one line

Composite = Σ (sub-score × weight ÷ 100), across six metrics. A metric that's N/A for a tool's category — or genuinely not yet scored — is dropped, and the remaining weights scale back up so they still total 100. That's a fair average of what was actually measured, not a penalty for the gap.

Weights

MetricWeight
Volume Verification20
Data Integrity20
Security20
Execution Speed15
Identity15
Reliability & Support10

By category — what's N/A

CategoryVolume VerificationExecution Speed
Copy Trading BotScoredScored
Arbitrage / Sniper BotScoredScored
Trading TerminalScoredScored
Whale / Smart-Money AnalyticsOnly if it routes real ordersOnly if it routes real orders
Discovery / Aggregator ToolN/AN/A

"Only if it routes real orders" applies to whale/smart-money tools: a passive lookup dashboard gets both metrics dropped and reweighted; a whale tool with an actual order-routing feature gets scored on both like any other terminal.

The six formulas, verbatim

On-chain Volume Verification20 pts

Scored off the platform's rank on the official Polymarket Builders Program leaderboard — not a self-reported claim. Percentile bands, not linear-per-rank: volume rankings are power-law distributed (rank 1 can carry 10-50x rank 20's volume, while rank 60 vs 80 are often nearly indistinguishable), so every rank-step isn't equally meaningful.

Rank 1 – 5    → 100 Rank 6 – 15  → 85 Rank 16 – 30 → 65 Rank 31 – 50 → 45 Rank 51 – 70 → 25 Rank 71+      → 10 Not listed on the Builders Program → 0
Data Integrity20 pts

Worked example: a directory claims "56+ tools." You spot-check 10 of them — 8 are real and correctly described. That's your Accuracy sample (8/10). Coverage is separate — a judgment call on whether ~56 represents most of the real ecosystem or just a slice.

Coverage (0–50): does it cover close to everything it claims, or just part of it?

Covers nearly everything (90%+) → 50 Covers most of it (70–90%)  → 40 Real gaps (40–70%)         → 25 Covers very little (15–40%) → 10 Barely anything (< 15%)      → 0 no coverage claim to check → 35

Accuracy (0–50): of the items you sampled, what % held up? This is where noise-filtering shows up — a tool that flags wallets that don't meet its own criteria fails this check, no separate metric needed.

≥ 95% accurate → 50 85 – 95%  → 40 65 – 85%  → 25 40 – 65%  → 10 < 40%     → 0
Security20 pts

Two separate questions: who controls the funds (Custody), and has the code that controls them been checked (Code Risk). "Non-custodial = automatic 100" hid this — a non-custodial DeFi contract can still have zero audits and full rug-pull risk.

CUSTODY (0–50) Non-custodial         → 50, flat Custodial, disclosed → 20 base + 15 majority cold storage disclosed + 15 multisig disclosed & in use Custodial, undisclosed → 5, flat Intentionally not disclosed (asked directly, no real answer) → 0, flat — and caps the WHOLE Security score at 20/100, no matter what Code Risk scores
CODE / CONTRACT RISK (0–50) No contracts of its own → 35, flat (neutral) Has own contracts: audit: none 0 · one 15 · two+ 20 + 15 active bug bounty + 10 contracts publicly verified + 5 disclosure channel published nothing disclosed at all → 5, flat

Found while reviewing POTS: audited by two named firms (CertiK, Bitlabs), a live six-figure bug bounty, all contracts verified on BSCScan. That's Custody 50 + Code Risk 50 = 100 — earned by disclosed evidence, not handed out by default for being non-custodial.

"Intentionally not disclosed" isn't the same as "opaque." Opaque means the platform never made a custody claim you could check. This band is for when you asked directly — a security question, a specific setting, custody specifics — and got redirected, ignored, or otherwise stonewalled instead of answered. That's worse than silence, so it scores worse than silence.

Execution Speed15 pts

One flavor, one band. Applies to bots that auto-trade for you (PolyCop, arbitrage/sniper bots), and to terminals / whale tools with an order-routing feature (your own submitted order going through the platform's interface instead of Polymarket's). No longer split into a separate "Signal Latency" flavor for information-only tools — in practice, terminals don't publish a distinct information-delivery-speed claim; what they actually claim is faster trade execution, same as the bots.

Two physically distinct phases, two claim shapes. Polymarket's CLOB matches orders off-chain, then settles on-chain on Polygon — and Polygon's own block time is ~2–2.3s (PolygonScan), so sub-second on-chain confirmation is physically impossible regardless of platform. A raw-ms claim can only legitimately describe the off-chain matching/submission step; a block-threshold claim describes on-chain settlement, which is bounded to multiples of that ~2s block time. Pick whichever shape the platform is actually claiming.

Claim-only for now. There's no independent measurement pipeline running yet, so this scores the platform's own published claim directly — not a tested result.

EXECUTION SPEED — claimed in ms (off-chain matching/submission) < 100ms → 100 100–250ms → 80 250–500ms → 55 500ms–1s → 30 > 1s → 10
EXECUTION SPEED — claimed as % within a block threshold (on-chain settlement) Same block (block 0) ≈ one ~2-2.3s Polygon block — the physical floor, scores full points: 95%+ → 100 85–95% → 80 70–85% → 55 50–70% → 30 under 50% → 10 Within the next block (cumulative, block 0–1) ≈ two block times (~4-4.6s) — a real, measured 2x latency jump vs. same-block, not a vague "one tier worse." The ms table above already prices an exact 2x latency jump at -25 points, twice (250ms→500ms: 80→55, 500ms→1s: 55→30) — reused here, floored above 0 so a real claim still beats no claim: 95%+ → 75 85–95% → 55 70–85% → 30 50–70% → 5 under 50% → 2
No public claim found (either shape), or platform claims parity with Polymarket's own feed with nothing platform-specific to test → 50 (neutral)

v2.2 revision: no-claim used to score 0 ("silence is worse than a disclosed-but-slow claim"). Dropped after real data collection showed the perverse result it created: PolyCop and Kreo — the only two platforms in a 15-platform pass that published any real number at all — scored worse than every platform that published nothing. Silence is absence of evidence, not evidence of bad performance, so it's now scored the same way Data Integrity's Coverage sub-score already treats a missing claim: neutral (50), not a penalty. A platform claiming plain parity with Polymarket's own infrastructure (e.g. a Builder running the identical CLOB/settlement) also lands here rather than getting excluded — there's no independently-established "Polymarket baseline" number to score parity against either, and no N/A carve-out is used for this reason (an earlier draft of this note described that N/A override; removed in v1.11).

The disclosed-claim bands above still bottom out below 50 (10 for ms >1s, 10/2 for the weakest block tiers) — inconsistent with the new neutral default, since a weak-but-honest claim could theoretically still score under silence. Not yet recalibrated; PolyCop and Kreo are hand-scored (80 each, provisional) rather than run through these uncalibrated bands until a proper recalibration pass happens.

Identity15 pts

Pick one tier (base points), then add the track-record bonus if it applies.

Anonymous / pseudonymous → 0 Entity named, unverified (bare claim, no outside corroboration — incl. offshore secrecy shells) → 15 Entity named, independently verified (registry, app-store developer record, certification, etc., even without a named individual) → 45 Individual(s) named, <12 months public → 65 Individual(s) named, 12+ months public → 80 + Verifiable prior track record (any non-anon tier) → +20, capped at 100

Ceiling for a fully anonymous team is 0/100 here — plain, not punitive. A verified-but-faceless company (e.g. a 5-year-old registered entity that's the confirmed app-store developer of record) now scores meaningfully above an anonymous 2-month-old project instead of tying with it at 0 — that gap was the v1.6 fix. Contact-responsiveness is scored under Reliability & Support instead, not here.

Reliability & Support10 pts

70% uptime + 30% support.

UPTIME measured uptime % (automated, 30-day min) → used directly manual check instead (no status endpoint) — a one-time audit can't honestly attest to a 30-day history it wasn't there for, so: scroll the support channel's last 1 HOUR of live activity zero unanswered questions/concerns → 90 (capped) – 15 per unanswered message, floor 0 SUPPORT (from a channel sample, not a personal contact test — find 2-3 examples of a DIFFERENT user's question getting answered in the general Discord/Telegram, time each to first reply, average the minutes; direct contact test is the fallback only when a channel's too dead/quiet to sample. Hour-granular — tight at the fast end, since this is a financial platform where speed of help matters; collapses hard past 24hr rather than trailing off gradually) under 5 min → 100 5–15min → 90 15–30min       → 80 30min–1hr → 70 1–4hr          → 55 4–12hr      → 30 12–24hr → 10 24hr+ / none → 0

Past 12 hours reads as nobody covering the channel, not "slow." 24 hours is the hard ceiling for a financial platform — everything at or past it scores identically to never responding.

Channel adjustment: same response time, different score — the channel itself is a structural signal, not just this one interaction. Nearly every platform researched already has a working Discord or Telegram, so email-only is a real outlier and is penalized as one.

Email-only                 → response score −35, floor 0 Instant messaging (Discord/Telegram) → unchanged + active community chat (not just tickets) → +5, capped 100

Ceiling: a manually-checked platform tops out at 0.7×90 + 0.3×100 = 93/100 on this metric, not 100 — disclosed here rather than left for a reader to discover.

BANNED OVERRIDE Banned / blocked / kicked from the support channel after asking a legitimate question → 0/100, flat — replaces the whole 70/30 formula above, not a deduction from it

This isn't a slower support band — it's the support channel closing rather than answering. No uptime number and no prior response speed offsets it: the metric scores 0 outright. If it happens, it's also worth a Security pending-note if the question you were asking about was security-related — the ban doesn't answer that question either, it just makes it unanswerable.

Recorded overrides

A formula band is the default path. These specific, named situations override it — each one is disclosed here and flagged again on the individual receipt page it applies to.

Execution Speed — the neutrality disclosure

Independent speed tests coming soon. Execution Speed currently scores what a platform publishes about itself, not what we measured. Platforms that publish nothing get a neutral 50 — silence is absence of evidence, not evidence of being slow. Once our own benchmark runs, these scores get replaced and every change is logged.

Corrections and right of reply

Every scored platform can dispute a result. If we got something wrong, we correct it publicly on that platform's page and log the correction — we don't quietly edit a number and move on. Reach us at corrections@polyreceipts.com.

What we do not measure

Rubric v2.3 — published 2026-09-06. Score changes are logged going forward on the platform page they affect.