About
A frontier model that keeps score.
FrontierPicks does what a fund won't: we publish our live book — the names we hold and the ones we're circling, on a paper account, as names and reasoning only, never sizes or fills — and keep a falsifiable, dated, Brier-scored record of whether each pick played out. The reasoning runs on frontier language models; a multi-stage pipeline reads US equity news, earnings, filings and price structure every trading day, re-reading the live book each session and cycling the wider coverage universe behind it. The accountability is the discipline — not a claimed edge.
The honest state of it: 351 picks have resolved so far — 160 played out, 191 invalidated — and 350 of them staked a conviction tier (one pick staked no conviction tier, so it cannot be scored for calibration), for a Brier score of 0.26. The skill score — whether staking conviction beats forecasting the base rate — sits near zero (−0.029), which on a sample of 350 resolutions means the conviction tiers have not yet shown they sort wins from losses. FrontierPicks has not demonstrated an edge. Reading a stock well is not the same as being right; the public record settles that, not the prose.
Every dossier carries a thesis, a kill line, an archetype, and a catalyst calendar. Every journal entry is dated and signed by a regime. Every macro view shows its math. The model writes, the methodology is open, and every call is on the record.
Models
Claude
frontier reasoning
Context
1M
tokens, side-by-side
Pipeline
5
stages, daily and weekly
Archetypes
8
behaviour profiles
Why open methodology
Reasoning in public is the product.
Discipline
A published kill line is a switch the author can't quietly retract.
Calibration
A dated archive tests whether a system reads the market or just narrates it — and so far the verdict is open: the skill score sits near zero on a young sample.
Builder signal
If you ship LLMs into markets, a public methodology beats any pitch deck.
Under the hood
How it's built, in detail.
Frontier language models
Anthropic's Claude models for synthesis, narrative classification, and the model's quality read.
Five-stage pipeline
Liquidity screen → weekly momentum + narrative radar → the model's quality read → theme-bet curation (sizing and exits machine-owned) → mechanical outcome grading.
Schedule
Daily: premarket and end-of-day reads, a re-read of every name on the book, and machine-owned exits. Weekly: the momentum + narrative radar and the theme-bet curation pass. The wider coverage universe is swept on a rolling cycle behind the book. DST-correct year-round.
Pre-registered kill gates
Tighten-only shutdown criteria, published in advance; every theme-bet is graded mechanically into a public Brier score — a forward experiment; no performance claim is made.
Long-context reasoning
The model holds the week's triggered clusters and their dossiers in a single 1M-token window and curates them in one pass — no chunking, no lossy summarisation.
Dated output, frozen grades
Dated dossiers and journal, an open methodology page, and a grade that is written once when a pick resolves and never rewritten.
On the record
Every thesis, the names we hold and the ones we're watching, and every regime and macro call — published the day it's made, dated, and scored against the kill line that would prove it wrong.
Open by default.
How we resolve, score, and correct
Sister sites
Five sister sites.
FrontierPicks sits alongside five sister sites at the intersection of language models and operator work. All static. All free. No signup.
Everything on this site.
Sections
Today / Now
Methodology
For machines