Skip to content
FrontierPicks

AI equity research · explained

Stock research a model can be held to.

AI equity research is stock analysis written by frontier language models instead of human analysts. The credible version isn't the one that sounds smart — it's the one that is falsifiable: a call carries a published level or event that would prove it wrong, set in advance, and a real closing price is what settles it. We publish 827 such dossiers across 40 themes, dated and two-sided. 819 of the 827 published dossiers carry the full shape — a dated thesis and the exact level that ends it. The remaining 8 are named rather than hidden: 7 are queued for first synthesis, and 1 publishes no kill line yet.

On the record351 resolved160 played out191 invalidatedBrier 0.26 on 350

What AI equity research actually is

Every trading day, a frontier model reads the US market: 30 days of news, four quarters of earnings transcripts and the filings behind each candidate. It clusters names by the narratives moving them, then writes a structured thesis per name — a current read, a bull case, a bear case, and the one thing that would prove it wrong. A large context window lets it weigh the whole candidate set in a single reasoning pass rather than scoring names in isolation.

That is the easy half. The hard half — the half almost no one ships — is making the output accountable: dated, two-sided, and falsifiable, so it can be graded later instead of quietly forgotten. Research you can't be wrong about isn't research; it's marketing.

Why a public, falsifiable track record matters

Two failure modes define the genre. The first is the black box: a confident call with no stated reason you can check. The second is the disappearing scorecard: picks published, losers never mentioned again, and "performance" quoted as a return marked to a benchmark after the fact.

A falsifiable record closes both. When a thesis ships with the exact level or event that would kill it — written down before the outcome — there is nothing to rationalise later. The call either survived its trigger or it didn't. We resolve each one in public as played-out or invalidated, graded against the trigger it published. Each call is on the record, dated.

Can a language model actually pick stocks?

The research layer is where a frontier model earns its place. Reading 30 days of news, four quarters of transcripts and the filings behind a name, then compressing that into a dated, two-sided thesis with one explicit kill condition, is a synthesis problem — and a large-context model does it consistently, name after name, without the fatigue or the favourites that bias a desk. It writes the bear case as carefully as the bull case because it has no position to defend.

Reading well is not the same as being right, and the difference is the only one that counts. The test is not whether the thesis sounds sharp; it is whether the call survives the trigger it published in advance. That is a question the prose can't answer and the record can — which is why every dossier carries a trigger a closing price can settle, and why the three states a pick can end up in are published as counts rather than summarised. The scoreboard above, and the full track record, are where that question gets settled rather than asserted.

The honest limits — and how to vet any AI stock service

Where the research layer is strong, the edge claim is not. A language model's advantage shows up most on short windows and narrow universes — a handful of names over a few weeks, where synthesis and speed matter. Over long horizons and a broad universe, that advantage thins, and plain buy-and-hold becomes a genuinely hard benchmark to clear. So the same caution holds in every direction: reading well is not the same as being right. A model that writes a sharp thesis has shown it can read; whether it can call — stake conviction that beats the base rate — is a separate, harder question, and answering it takes a scored record run long enough to tell skill from luck. The FrontierPicks verdict on that is unflattering and stated plainly: whether a record is skill or luck is settled by the numbers, and ours sit near zero on a sample still too young to be a verdict.

That reframes the buyer's question. The thing to demand of any AI stock researcher is not a return quote — it is a record that could have gone wrong in public: falsifiable, dated, every call kept on the board whether it played out or broke. Most can't show one, and the reason is survivorship bias — winners are kept, losers quietly dropped, and "performance" is marked to a benchmark after the fact, so the headline number is unfalsifiable by construction. A record where the failures stay visible is the one you can actually audit. That is the bar FrontierPicks holds itself to, and the public track record is where you check whether it clears it.

How to evaluate an AI stock researcher

Six questions separate a record you can trust from one you can't. Apply them to anyone — including FrontierPicks.

01

Is each call falsifiable?

A specific, published level or event that would prove the thesis wrong — set in advance, not rationalised after.

FrontierPicks:Every dossier carries an explicit invalidation trigger.see it →

02

Is it scored in public?

Resolved outcomes — played-out or invalidated — graded against each thesis's own pre-published trigger, not a return marked to a benchmark later.

FrontierPicks:A public accountability ledger, non-monetary.see it →

03

Is every read dated, and is the grade frozen?

When was this written, and can the author restate the call after the outcome? A research note with no date is unfalsifiable by construction; a gradeable one whose grade can be rewritten is worse.

FrontierPicks:Every dossier and journal entry is dated. When a pick resolves, the tier it was graded on is written once and never rewritten, and the graded thesis is fingerprinted.see it →

04

Is the methodology open?

How names are screened, scored, and sized — stated plainly, not a black box you're asked to trust.

FrontierPicks:The full process is published, stage by stage.see it →

05

Does it show both sides?

A bull case AND a bear case AND named failure modes — not a one-sided pitch.

FrontierPicks:Every dossier argues both sides and names what would break it.see it →

06

Is there anything to sell?

A free, falsifiable record aligns the author with being right. A paywall, signal service, or affiliate link aligns them with conversions.

FrontierPicks:No product, no paywall, no affiliate — nothing to sell.see it →

How FrontierPicks does it

FrontierPicks is an AI equity-research analyst built on Anthropic's Claude Opus and Sonnet. We run a five-stage process on a fixed schedule — a liquidity screen, a weekly momentum and narrative radar, deep per-name synthesis, a 1M-context theme-bet curation pass, and mechanical grading against pre-registered kill gates. The output is 827 per-ticker dossiers, a regime-tagged daily journal, a weekly macro view, and a live book of held and circling names — names and research only, never the sizing.

It is the answer to the six questions above, by construction: falsifiable triggers, public scoring, dated reads, open methodology, two-sided cases, and nothing to sell.

How it works The track record Skill, or luck? What is a thesis? Browse the dossiers

Common questions

What is AI equity research?
AI equity research is stock analysis produced by frontier language models rather than human analysts: the model reads news, filings, earnings transcripts and price structure each trading day and writes a structured thesis per name. The credible version is falsifiable and dated — each call carries an explicit level or event that would prove it wrong — so it can be scored after the fact instead of quietly forgotten.
Which AI stock research tools publish a track record?
Most don't — they publish picks and never keep score, or they mark returns to a benchmark after the fact. We publish every thesis with the exact invalidation trigger that would prove it wrong, set in advance, then resolve each one in public as played-out or invalidated, graded against each thesis's own published kill level.
How do you evaluate an AI stock researcher?
Ask six questions: Is each call falsifiable (a published trigger set in advance)? Is it scored in public against that trigger? Is every read dated, and is the grade frozen once a call resolves? Is the methodology open? Does it argue both the bull and bear case and name its failure modes? And is there anything to sell — because a free, falsifiable record aligns the author with being right, while a paywall or signal service aligns them with conversions.
Can an AI language model actually analyse stocks well?
For the research layer — synthesising 30 days of news, four quarters of transcripts and filings per name into a structured, dated thesis with a bull case, bear case and an invalidation trigger — a frontier model with a large context window is strong and consistent. But reading well isn't the same as being right. The honest test is whether the dated, falsifiable calls hold up when scored against the triggers they published in advance — which is exactly what the public Brier-scored track record measures, thesis by thesis.
Is AI equity research investment advice?
FrontierPicks is not. We publish educational research under the EU and BaFin framework — a transparent, dated record of how a language model reads the market, not a recommendation to buy or sell anything. Always do your own work.
How is FrontierPicks different from a stock-picking newsletter?
A newsletter sells conviction and rarely keeps score. We publish the names we hold and the ones we're circling, each with the exact kill line, then resolve them in public. There is nothing to buy.
Do AI stock pickers beat buy-and-hold?
The honest answer is that almost none have shown it, and the long-horizon evidence is unkind: a model's edge is strongest on short windows and narrow universes and thins as the horizon and universe widen, where plain buy-and-hold is a hard benchmark. Most claims to beat it are unfalsifiable — winners kept, losers dropped, returns marked to a benchmark after the fact. Answering it takes a scored, dated, every-call-resolved record run long enough to tell skill from luck — and the FrontierPicks skill score sits near zero on a sample still too small to be a verdict.