Frontier models versus synthetic parcel packets built in real county file formats from the Indiana–Michigan line. Implied-decimal price layouts, related-party transfers, multi-parcel sales, keying errors, duplicates, cross-border comps, and jurisdiction inversions. A human records analyst catches these in seconds. This measures whether a model does, and whether its confidence interval is honest.
Composite = 30% valuation accuracy · 20% interval calibration (Winkler, 80%) · 25% invalid-comp exclusion F1 · 15% trap-flag F1 · 10% jurisdiction. Baselines are deterministic pipelines, shown for reference.
Method & limits. Items are procedurally generated, not real transactions, so answers cannot have been
memorized; the holdout tiers are sealed by published commitment. The rules baseline was written by the benchmark
author with knowledge of the seven trap classes: treat it as an informed-specialist reference, not a neutral competitor.
Rows labeled claude.ai:<tier> were run through the consumer Claude app at that tier, so the exact model
version is not pinned; API rows name their model. Scores on the public tiers are over 60 items per solver.
| # | Solver | Composite | Valuation | Calibration | Exclusion | Flags | Jurisdiction | Median error | 80% cover | Parse |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | baseline:rulesbaseline | 94.3 | 92.4 | 83.1 | 100.0 | 100.0 | 100.0 | 1.8% | 98% | 100% |
| 02 | baseline:naivebaseline | 51.4 | 89.7 | 66.4 | 10.0 | 8.3 | 75.0 | 2.1% | 100% | 100% |
Share of items where each trap was correctly handled. This is the part that matters: not whether a model is good, but exactly where it breaks.
| Solver | Implied-decimal pricen=23 | Related-party salen=21 | Multi-parcel salen=20 | Keying outliern=17 | Duplicate recordn=22 | Cross-border compn=22 | Jurisdiction inversionn=15 |
|---|---|---|---|---|---|---|---|
baseline:rules | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
baseline:naive | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
Every scored answer is a hash-chained receipt. Editing any past result breaks the chain. The holdout key is committed before any model is run and revealed afterward.
7ac752a695e91f42d21fdf90cb0dab0640919b6c1a92a8a98daf1e2ea77c948905dae1e5133e45c0f98b8f25939dba34572eeed550fe6bcfe81114488aa63cfc10b61139662a6efa535df434950c5b3f117c093110e37ed55e580763bf92b34c5dca96aba3c2a7970e6c98b68ee0ea28e15edc17a755344341bbeec14eff96e9prbench 1.1.1python3 -m prbench verify # recompute every hash + re-score every stored response