Public track record: what we forecast, and whether it was right
80 forecasts recorded, none resolved yet. Every one was written down before its outcome was knowable, and none has come due. The first of them can be resolved on 2027-03-02. We have proved nothing so far, and this page will say so until the resolutions land.
Anyone can copy an endpoint. Nobody can copy a dated public record of whether the calls were right. Every meaningful forecast this system makes is written once to an append-only ledger at the moment it is made, with everything needed to resolve it later from public data. Resolution only ever fills in the outcome: the forecast itself is never edited, and a changed method gets a new version and new rows rather than restating old ones.
What do we forecast?
sec_negative_followup
When an 8-K carries a flag we consider adverse — bankruptcy or receivership, a delisting notice, a restatement or non-reliance notice, an auditor change, a material impairment — we record a probability that the same issuer files another adverse event within six months. The insider-buy cluster is recorded the same way with a low probability, so a positive signal is falsifiable too.
The record
| method | what it predicts | horizon | made | resolved | mean Brier | first resolvable |
|---|---|---|---|---|---|---|
sec_events_rules_v1 v1.0.0 | sec_negative_followup | 6 months | 80 | 0 | not yet | 2027-03-02 |
As of 2026-09-22T10:25:24.655Z. The same numbers as JSON: https://answerpool.io/v1/evals (free, no key) — this page and that document are rendered from one computation, so they cannot disagree.
How should you read it?
mean_brier_score is the Brier score of the stated probability against the observed outcome: 0 is perfect, 0.25 is what always saying 50% gets you, and 1 is confidently wrong. Lower is better.
A method with nothing resolved yet has made claims but proved nothing — that is the honest state of a young ledger, not a gap in the data. Scores also do not become skill until the base rate is known: a flag that fires on companies which file often would score well for the wrong reason.
Predictions are append-only (CLAUDE.md invariant 8). Resolution fills the outcome and score; the forecast itself is never edited, and a changed method gets a new version and new rows rather than restating old ones.
What this page is not
It is not a performance claim, a backtest, or a marketing number. A backtest can be tuned after the fact; this cannot, which is the only reason it is worth publishing. It is also not investment advice: the forecasts are probabilistic estimates about public filings and research volume, and they are wrong some of the time by construction.
Method definitions and versions: /v1/metadata. Products: the catalog.