Skip to content
FrontierPicks

Methodology · open

How the model thinks.

A five-stage frontier-LLM pipeline turns US market data into dated, falsifiable picks. The reasoning is what gets published — open, dated, and graded against a line set in advance.

The pipeline, in one paragraph

A five-stage pipeline ingests 30 days of news and four quarters of earnings transcripts for the book, the watchlist, and the momentum universe, and a frontier language model synthesises a quality grade and a full dossier per candidate. Once a week a mechanical radar clusters the market into momentum themes; the model then curates the triggered clusters into scored theme-bets — shaping or rejecting each — while sizing and exits stay machine-owned. The output is a dossier per ticker, a regime-tagged journal entry, a macro view, sector rotation, a watchlist, and a public, Brier-scored forecast record. The reasoning is what gets published — and this is a forward experiment; no performance claim is made.

The five stages

Three different universes appear across this site, and they are three different screens doing three different jobs, not one number stated inconsistently: the momentum scan reads ~3,400 US names, the breadth read scores 977 of them against their 200-day averages, and the liquidity screen passes ~400 to the model. The momentum board ranks the top names out of the first; the breadth reading is computed over the second.

  1. Liquidity screen. ~400 US-listed names filtered by spread, ADV, market cap, and tradable shape. Deterministic; cheap; eliminates ~75% of the universe before the LLM touches it.
  2. Momentum + narrative scoring. Survivors ranked by a composite of price structure × volume × news density × theme-cluster strength. Themes are treated as primary, not afterthoughts — the system reads narrative basket behaviour, not isolated ticker action.
  3. The model's quality read. Claude Opus reads news, earnings transcripts, and filings for each candidate and produces a quality grade (an internal grade from strongest to weakest), along with the dossier you see on this site: thesis, kill line, bull case, bear case, setup, catalysts, what-would-change-our-mind, correlations.
  4. Theme-bet curation. Once a week, Opus reasons across the radar's triggered momentum clusters side-by-side using the 1M-token context — every candidate dossier in one pass — and either rejects a cluster (index-twin, merger-arb, binary biotech — logged and published) or shapes it into a theme-bet: concentrated in its one or two narrative-central names, with a stated probability and a written narrative. A theme-bet carries no kill line — no price level is set for any of its names — so what is fixed in advance is the date, the names and the probability. Every decision, shaped or refused, is immutable and public. Position sizing and exits are machine-owned — the model cannot veto them.
  5. Mechanical outcome grading. Mechanical either way, with no discretion, but the two kinds of call are graded differently and only one of them is Brier-scored. A per-ticker dossier is graded against the kill line it published in advance and folds into the public Brier score. A theme-bet has no line to grade against: it closes when a machine-owned exit rule fires and resolves into a binary outcome plus a categorical verdict against the mechanical cluster baseline and matched-window SPY — which is why it is not Brier-scored. Every exit rule and its threshold is printed on the bet's card at forecasts. Pre-registered, tighten-only kill gates decide whether the strategy scales, de-scales, or is declared dead.

The schedule

Windows are computed off the US market calendar in New-York-relative minutes — DST-correct year-round, including the ~4-week US/EU DST gap windows where naive Berlin-time schedules silently fail. The theme-bet loop is weekly; the reads that keep the record honest run daily:

  • Weekly radar — the mechanical cluster scan that emits candidate themes and triggers; nothing enters without a mechanical trigger.
  • Curation pass — the model shapes or rejects the triggered clusters into scored theme-bets.
  • Daily exits — machine-owned rules only: the profit ladder, catastrophic drawdown stops on the single name, on the basket and on the whole theme sleeve, and a sustained trend break that becomes reachable only once the pre-registered minimum hold has passed. The minimum hold gates that one exit; it never closes a bet by itself. The model cannot override any of them.
  • Daily reads — premarket universe scan, the close-of-day regime call, the EOD journal, and a post-close reconcile of the record.

What's in a dossier

Each dossier follows the same skeleton — designed so a reader can audit the reasoning, not just the conclusion:

  • Current thesis — one paragraph, what the bet is.
  • Kill line — the explicit kill criterion, published in advance.
  • Bull case — sourced bullets, dated, no hand-waving.
  • Bear case — same construction, equal weight.
  • Setup & price structure — MAs, RSI, levels, basing pattern.
  • Catalyst calendar — next 30 days, dated.
  • What would change our mind — explicit conditions for higher / lower conviction.
  • Correlation notes — how the name moves with its basket.

Archetype taxonomy

Every name is tagged with an archetype. Archetypes are not labels — they're behaviour profiles that drive the dossier's stop logic and how its setup is read. Full definitions on the glossary.

  • Compounder. Quality balance sheet, secular tailwind, multi-year hold candidate.
  • Cyclical recovery. Mean-reverting earnings, regime-sensitive.
  • Theme leader. Highest-conviction name within an active narrative.
  • Special situation. M&A, spin-off, restructuring, regulatory event.
  • Earnings inflection. Pre/post-print setup with explicit binary.
  • Retail squeeze. High-beta, short-interest-driven, hard sizing cap.
  • Defensive. Cash-flow durability, low-beta, regime hedge.
  • Macro hedge. Cross-asset proxy for thematic risk (XLE / GLD / TLT / …).

Regime classification

Each journal entry records the system's regime call. Regimes are not picks — they're a filter that gates how aggressively the system reads candidate setups. The macro view is the long-form version of the same read.

Published regime labels: RISK-ON, CHOPPY, RISK-OFF, a binary-event stagflation scare, an escalating stagflation scare, a healthy-but-unconfirmed recovery, and variants. Each is a defined rule that maps to buy-threshold, size-multiplier, max-exposure, and cash-floor settings.

Conviction levels

Conviction is the model's calibrated confidence that the setup will play, not a price target or return forecast. Four levels: SUPREME, HIGH, MEDIUM, LOW. Each is a probabilistic claim, published in advance, that the thesis plays out (SUPREME ≈ 0.90 down to LOW ≈ 0.50) — scored directly on the scoreboard — and each carries an explicit kill line that strips the conviction if breached.

How outcomes are scored

Open methodology applies to the scoring too — here is the exact, reproducible method behind the scoreboard. Every per-ticker pick ships a falsifiable kill line: the specific level or event, published in advance, that would prove the claim wrong. That is what makes a call gradeable at all — a pick with no stated kill line can never be scored, only quietly forgotten. When a pick resolves, the pipeline marks it "played out" or "invalidated", dated, with a flag for whether the published kill line fired first. Strictly non-monetary — outcomes are binary: the claim held, or it was falsified.

Each conviction tier is treated as a probabilistic claim, published in advance, that the thesis plays out: SUPREME = 0.90, HIGH = 0.75, MEDIUM = 0.60, LOW = 0.50. Names the model held no conviction on are listed in the resolved ledger for transparency but are not scored — you can only be graded on a call you actually made. Against the binary outcome (played-out = 1, invalidated = 0) we compute:

  • Brier score — the mean squared error between the stated probability and the outcome, decomposed (Murphy) into reliability, resolution and uncertainty. The uncertainty term is set by the question class, not the forecaster: it is what a forecaster scores by staking the base rate on every pick, and therefore the reference the raw score is read against — not a floor. A Brier score's floor is 0, and a score above the uncertainty term is worse than the base-rate forecast, not better.
  • Murphy decomposition — Brier = calibration − resolution + uncertainty. Calibration (reliability) is how far each tier's observed play-out rate sits from its stated probability; resolution is how much the tiers separate from the base rate; uncertainty is the irreducible base-rate variance.
  • Brier skill score — skill versus a naive always-the-base-rate forecast (1 − Brier ∕ uncertainty).
  • Reliability by tier, archetype, and regime — the observed play-out rate broken out by conviction tier, by archetype, and by the macro regime in force when each thesis resolved. This is the falsifiable test of whether SUPREME actually beats LOW.

The board is hidden until at least one thesis has resolved (no faked scorecard), and the full record is machine-readable at /track-record.json for agents and independent verification.

Why score it this way at all? A headline accuracy number — "right 70% of the time" — can be claimed by any research outlet that publishes its winners and lets its losers age out of the archive. A Brier score over pre-committed, falsifiable theses cannot: every call carries a stated probability and a stated kill condition before the outcome is known, and the score is computed over every pick that has reached a verdict — the invalidated ones included, none retired. That is the population the score covers, and it is not the whole corpus: a pick that has hit neither its kill line nor its target is undecided, and undecided picks are not scored, because nothing has happened to score.

All three states are published as counts, with the completion rate, the resolution mechanism, and a measured comparison of the decided and undecided cohorts against the universe they were drawn from, on how a pick resolves. The live record is the proof that the method is applied — including the invalidated theses, which stay on the board.

Where the model is wrong

Three classes of failure are recurring and worth naming:

  • Stale facts. A model snapshot of a balance sheet can lag the latest 10-Q. Flagged in dossier notes when caught — not always caught.
  • Confident-but-wrong setup reads. A "clean higher-low" can become a failed reclaim within hours. Dossiers age fast; recency-of-write is on every page.
  • Theme misclassification. A name gets bucketed in a basket whose actual price driver is different; the correlation logic then over-fits.

This is why nothing here is a recommendation. The dossier is the reasoning behind every call.

Open by default

Every thesis, the names we hold and the ones we're watching, and every regime, macro and sector call — published the day it's made, dated, and scored against the kill line that would prove it wrong. The methodology is open, and so is the reasoning behind every call on the book.

And it is honest about what it is: FrontierPicks has not proven an edge. Every mechanical harvest of this phenomenon lost to the benchmark before the forward book started; the one untested claim is whether the model's curation adds anything, and that is what the public scoreboard measures. So the forward theme-bet book is an experiment run at full transparency, with the shutdown conditions written down in advance. Three surfaces make it auditable in real time: the weekly radar (the mechanical candidate scan), the scored forecasts (open bets and refusals, with reasoning), and the kill gates (the pre-registered, tighten-only shutdown criteria). See also why trust it.

Common questions

What is FrontierPicks?

FrontierPicks is AI equity research that keeps score. A multi-stage frontier-model pipeline reads US equity news, earnings transcripts, and price structure every trading day, then publishes per-ticker dossiers, regime-tagged journal entries, a weekly macro view, sector rotation, and a curated watchlist — every pick carrying the kill line that would prove it wrong.

Which language models power the pipeline?

Anthropic's Claude Opus and Claude Sonnet. The theme-bet curation stage uses Opus's 1M-token context window to weigh the week's triggered momentum clusters and their candidate dossiers side-by-side in a single pass.

What is an archetype?

Archetypes are behaviour profiles assigned to each name (Compounder, Cyclical recovery, Theme leader, Special situation, Earnings inflection, Retail squeeze, Defensive, Macro hedge). They drive the dossier's stop logic and how its setup is read. See the glossary for full definitions.

What is a kill line?

The explicit kill criterion published with every pick — formally, the invalidation trigger. If the line fires, the conviction is stripped and the pick is treated as broken. It's a published commitment, not a soft warning.

How often is the site updated?

Daily for the journal, and daily for every dossier we hold, are researching or have just exited; the wider coverage universe cycles weekly, as do the macro view and sector rotation. Each surface carries its own dateModified meta and an updated-at line. Subscribe via JSON Feed or RSS.

Is this investment advice?

No. We publish educational content under the BaFin and EU regulatory framework. No personalised advice is given and no orders are accepted.