Macro Calibration Board · Apr 19 – May 3, 2026
Price-blind macro forecasts logged publicly with timestamps and outcomes. Each forecast is captured before PRISM sees the Kalshi price, so the comparison reflects calibration discipline — not after-the-fact tuning.
How to read · Lower Brier is better. The board scores PRISM’s price-blind probability against the Kalshi price at forecast time — measuring whether PRISM’s judgment improves on the market over time. The score is live and expected to evolve as the sample grows.
0 CLASSIFIED FORECASTS · 0 TOTAL RESOLVED · PRISM v0.9.17 → v0.9.92-deepseekPRICE-BLIND MODEL · EVIDENCE LAYER · PRIORITY COVERAGE · EVENT-FAMILY DEDUP
Calibration in progress
0 blind-classified forecasts so far (0 total resolved). Headline comparison opens at 10 forecast-grade samples.
OBSERVATION WINDOW ENDS MAY 3, 2026
Priority 5 · Live tracker
The five macro releases PRISM treats as priority coverage — shorter cooldowns, higher refresh frequency, dedicated inclusion every cron. Each card shows the latest blind forecast vs the Kalshi price, with outcome once the market resolves.
Coming to this board
Methodology
PRISM (blind). The truth model sees the market question, resolution criteria, and retrieved evidence — but market prices are redacted from the prompt during analysis.
PRISM (anchored). The legacy variant that includes market price in its prompt. Kept running in parallel so we can measure what the market anchor does to the forecast.
Brier score. Proper scoring rule for probabilistic forecasts. 0 = perfect, 0.25 = coin flip. Lower is better.
Brier Skill Score. 1 − blind_brier / market_brier. Positive = PRISM (blind) more accurate than the Kalshi price at forecast time. Zero = matched the price. Negative = behind the price. The institutional reading: skill above the reference, normalized.
Reliability curve. Forecasts are bucketed by forecast probability (5 bins) and plotted against the realized YES rate. Perfect calibration sits on the diagonal — “when we said 70%, it happened 70% of the time.” Each point is a bucket; both PRISM (blind) and Kalshi price are plotted.
Forecast filter. Headline metrics (Brier, BSS, reliability) score only rows where the PRISM (blind) staked a real position — |capped_edge_blind| ≥ 5pp. Rows where blind landed within 5pp of the market price are tracked but excluded from headline metrics; including them would pull every metric toward the diagonal artificially.
Scope. Macro markets on Kalshi (Finance category): GDP, CPI, jobs, gold, PPI, productivity, deficit, and related series. Window opens 2026-04-19 and scores every resolved forecast from that date forward.