The public execution sample records 1,000 attempts with 5.8% accepted and filled rates. Across 986 joined market-level observations, market-implied probability was slightly better calibrated than the fair-probability model on both Brier score and log loss.
The sample does not support a stable profitability claim. Its main result is diagnostic: most apparent edge did not become filled exposure, extreme probability buckets were sparse and unstable, and reasonable ML or fill-probability gates could fail on later data or suppress nearly all trade flow.