Macro Calibration Board · Apr 19 – May 3, 2026

PUBLIC CALIBRATION LEDGER
LIVE SINCE APR 19

Price-blind macro forecasts logged publicly with timestamps and outcomes. Each forecast is captured before PRISM sees the Kalshi price, so the comparison reflects calibration discipline — not after-the-fact tuning.

How to read · Lower Brier is better. The board scores PRISM’s price-blind probability against the Kalshi price at forecast time — measuring whether PRISM’s judgment improves on the market over time. The score is live and expected to evolve as the sample grows.

0 CLASSIFIED FORECASTS · 0 TOTAL RESOLVED · PRISM v0.9.17 → v0.9.92-deepseekPRICE-BLIND MODEL · EVIDENCE LAYER · PRIORITY COVERAGE · EVENT-FAMILY DEDUP

Calibration in progress

0 blind-classified forecasts so far (0 total resolved). Headline comparison opens at 10 forecast-grade samples.

OBSERVATION WINDOW ENDS MAY 3, 2026

Priority 5 · Live tracker

The five macro releases PRISM treats as priority coverage — shorter cooldowns, higher refresh frequency, dedicated inclusion every cron. Each card shows the latest blind forecast vs the Kalshi price, with outcome once the market resolves.

📊
GDP
KXGDP
NO ACTIVE MARKET
Awaiting first PRISM (blind) analysis on this release.
0 ANALYZED · NO FORECAST YET
🛒
CPI
KXCPI
NO ACTIVE MARKET
Awaiting first PRISM (blind) analysis on this release.
0 ANALYZED · NO FORECAST YET
🏭
PPI YoY
KXUSPPIYOY
NO ACTIVE MARKET
Awaiting first PRISM (blind) analysis on this release.
0 ANALYZED · NO FORECAST YET
👷
Nonfarm Payrolls
KXPAYROLLS
NO ACTIVE MARKET
Awaiting first PRISM (blind) analysis on this release.
0 ANALYZED · NO FORECAST YET
💼
ADP Employment
KXADP
NO ACTIVE MARKET
Awaiting first PRISM (blind) analysis on this release.
0 ANALYZED · NO FORECAST YET

Coming to this board

  • — Time-horizon split: ≤7-day vs >7-day forecast skill
  • — PRISM version diff: which build moved which metric

Methodology

PRISM (blind). The truth model sees the market question, resolution criteria, and retrieved evidence — but market prices are redacted from the prompt during analysis.

PRISM (anchored). The legacy variant that includes market price in its prompt. Kept running in parallel so we can measure what the market anchor does to the forecast.

Brier score. Proper scoring rule for probabilistic forecasts. 0 = perfect, 0.25 = coin flip. Lower is better.

Brier Skill Score. 1 − blind_brier / market_brier. Positive = PRISM (blind) more accurate than the Kalshi price at forecast time. Zero = matched the price. Negative = behind the price. The institutional reading: skill above the reference, normalized.

Reliability curve. Forecasts are bucketed by forecast probability (5 bins) and plotted against the realized YES rate. Perfect calibration sits on the diagonal — “when we said 70%, it happened 70% of the time.” Each point is a bucket; both PRISM (blind) and Kalshi price are plotted.

Forecast filter. Headline metrics (Brier, BSS, reliability) score only rows where the PRISM (blind) staked a real position — |capped_edge_blind| ≥ 5pp. Rows where blind landed within 5pp of the market price are tracked but excluded from headline metrics; including them would pull every metric toward the diagonal artificially.

Scope. Macro markets on Kalshi (Finance category): GDP, CPI, jobs, gold, PPI, productivity, deficit, and related series. Window opens 2026-04-19 and scores every resolved forecast from that date forward.