How Wrong Are We?
The accuracy ledger · how the values are built → · what we actually have →
Anyone can claim their prices are accurate after the fact. So we don't get to. Every published value is written to a locked ledger before the sale happens — the card, the number, the range, the date. When that card later sells, we score what we said against what it fetched.
The predictions cannot be edited. A database trigger rejects any change to a recorded prediction, its range or its tier. Only the outcome can ever be written. If we look bad, the number stays.
The live scoreboard
| Evidence tier | Scored (n) | Median error | Mean error | Bias | Landed in our range |
|---|---|---|---|---|---|
| All cards | 2 | 28.0% | 28.0% | -16.8% | 100.0% |
| L1 | 2 | 28.0% | 28.0% | -16.8% | 100.0% |
Median / mean error are absolute percentage error — how far off we were, ignoring direction. Bias is the mean signed error: positive means cards sold for MORE than we published (we were low), negative means we were high. A model can have a small mean error and still be systematically wrong in one direction, so both are shown. Landed in our range is the share of sales that fell inside the published low–high band.
Latest scored cards
What we said, before it sold, next to what it fetched. Green = we were under, red = we were over.
| Card | We said | Our range | It sold for | Miss | In range |
|---|---|---|---|---|---|
| Adam Bomb
1985 Garbage Pail Kids Original Series 1 · PSA 6 · L1 · 2026-07-28 |
$235.33 | $128–$421 | $152.50 | -35% | ✓ |
| Electric Bill
1985 Garbage Pail Kids Original Series 1 · PSA 7 · L1 · 2026-07-28 |
$78.71 | $49–$121 | $80.00 | +2% | ✓ |
The backtest
Separate from the live ledger, and clearly labelled as such. We replayed the last two years: for 40,028 real sales we rebuilt what the model would have published the day before each one — using only data that existed then — and scored it.
| Evidence tier | Sales tested | Median error | Average error | Inside range |
|---|---|---|---|---|
| L1 recent observable sales of this exact spec |
26,125 | 23.9% | 40.6% | 77.2% |
| L2a this exact spec, but the comps are stale (>180d) and index-adjusted |
7,646 | 40.0% | 56.0% | 79.0% |
| L2b same card at another grade or variety, projected |
3,106 | 47.7% | 75.5% | 79.7% |
| L3 set-peer model estimate — no direct comp for this card |
3,151 | 57.6% | 89.9% | 78.4% |
| All cards | 40,028 | 30.1% | 50.1% | 77.8% |
The range is built as an 80% band. Out of sample it caught 77.8% of actual sales — slightly under target, which we would rather report than quietly widen.
Error grows with the age of the evidence
| Newest comp behind the value | Sales tested | Median error |
|---|---|---|
| 0-30 days | 12,805 | 18.8% |
| 31-90 days | 8,268 | 27.4% |
| 91-180 days | 5,052 | 33.3% |
| 181-365 days | 3,622 | 36.5% |
| 1-2 years | 2,120 | 40.2% |
| over 2 years | 1,904 | 46.8% |
| no comp at all | 6,257 | 51.7% |
This is why a card that sold last week and a card that sold in 2021 do not get the same sized range, even if we print the same value for both.
How to read these numbers honestly
- Median, not average. Median error is the typical miss. Average error is much higher because a handful of cards miss enormously — we publish both so you can see the gap.
- A backtest is not a promise. It measures the model against the past. The live ledger above is the only number that cannot be tuned after the fact, and it is the one we would ask you to judge us on.
- Level 3 estimates are genuinely uncertain. They are most of the catalogue. The range is not decoration.
Values are estimates for information only, not an offer to buy or sell, and not an appraisal. Every figure on this page is read live from the ledger.
