Earnings Setup: Base Rates, Not Verdicts
A stock runs up on last quarter's strength, beats, and sells off anyway because the beat was already in the price. Wave 21 turns that "expectations gap" into arithmetic a reader can check: every rate carries its sample size, nothing is model-written, and one number the provider offered was left off the page on purpose. Wave 21b then added GMI's Read, a labelled statistical read of the company's own history, and the first version of it failed its own backtest.
BUILT · SEPTEMBER 2026 In final testing ClientWeb · AdminWeb · MAUIWhat It Shows
Click any symbol in the Earnings Lab (recent results or the upcoming calendar) and you get its Earnings Setup: how far the stock has run since its last print and where that run ranks against its own history, drawn on a four-year run-up chart with every print marked by its EPS outcome; a table of its last 16 prints (four years: EPS and revenue surprise, the run-up into each print, and the next-session move); how often it beats; where the bar sits for the next print; the plain conditional base rates; and, since Wave 21b, GMI's Read.
| Block | Status | Source |
|---|---|---|
| Since the last print: run-up, percentile vs. own history, news volume | Built | EOD price cache, per-symbol earnings history, news archive (count only) |
| Reaction table: last ≤16 prints (21b: shared-scale move bar, "beat, then fell" rows badged) | Built | earnings history + adjusted closes |
| Beat habit: counts, median surprise, trend, streak | Built | earnings history |
| Where the bar is now: consensus, revisions, price target | Built | earnings history, GMI's own estimate journal, price-target consensus |
| The read: conditional base rates, 8-print floor | Built | derived only from the reaction-table rows |
| 21b: run-up chart, 16 prints at a glance, GMI's Read with its own track record | Built | the same rows, plus 60-session pre-print volatility from the price cache |
| Last call's transcript | Seam only | waiting on a probe of the provider contract |
| Options-implied move | Seam only | waiting on the same probe |
The unbuilt blocks exist as nullable fields on the DTO and a registered no-op provider. While they are null, nothing renders: no "coming soon" box, so a missing feature costs zero pixels.
flowchart LR
LAB[Earnings Lab row click] --> API[GET earnings-setup for a symbol]
API --> SVC[EarningsSetupService]
SVC --> H[Earnings history<br/>cached 1 day]
SVC --> P[Adjusted daily closes<br/>append-only price cache]
SVC --> T[Price-target consensus<br/>cached 7 days]
SVC --> J[(Estimate journal<br/>MySQL, change-only rows)]
SVC --> N[(News archive<br/>article count only)]
H --> M[EarningsSetupMath<br/>pure, unit-tested]
P --> M
T --> M
J --> M
N --> M
M --> OUT[Reaction table, beat habit,<br/>bar now, the read<br/>every rate with its n]
CRON[Schedules engine<br/>once a day, 1 call] --> JThe Rule: Computed, Not Model-Written
This was the obvious place to put an LLM: "summarize how this stock tends to trade around earnings."
I chose not to. A language model will happily write "tends to rise after beats" from four examples,
and a reader can't check it. Every sentence on this page is instead produced by one pure function,
EarningsSetupMath, shared by the API and all three apps so they can't disagree, and
every sentence follows four rules:
- Every rate carries its sample size in the same sentence. "Closed higher the next session 11 of 20 times", never "usually rises".
- A floor of 8 prints. Below it the page shows the table and the count and says there isn't enough history for a base rate. A conditional line built on fewer than 5 prints says "a small sample" in the same sentence.
- Nothing is fitted, weighted or extrapolated. The rows that produce every count are the rows in the reaction table on the same screen. (GMI's Read, below, does shrink its estimates, so it's labelled as a statistical read, prints its recipe with the numbers plugged in, and sits beside these unshrunk counts rather than replacing them.)
- Descriptive voice, enforced by a test. A unit test fails the build if any generated sentence contains "buy", "sell", "should", "will rise" or "likely".
The line that justifies the whole page is the beat-conditioned one: "In the 14 prints where EPS beat estimates, it closed higher 7 of 14." A high beat rate next to a coin-flip reaction is the expectations gap, shown with no opinion added. (Example numbers, for illustration.)
The Recipe (Published on the Page's Help Panel)
All prices are adjusted closes, so a 4-for-1 split never shows up as a −75% reaction. The reaction window depends on when the company reported:
| Session reported | "Before" close | "After" close |
|---|---|---|
| After the close (AMC) | the print date | first trading day after |
| Before the open (BMO) | last trading day before | the print date |
| Not reported | last trading day before | first trading day after (labelled a 2-session window) |
- Data holes are excluded, not guessed. A print with a price gap of more than 5 calendar days on either side of its window (a hole in the data, not a weekend) gets no reaction and is left out of every rate.
- Run-up into a print = that print's "before" close ÷ the previous print's "after" close − 1: the drift across the whole gap between two prints.
- Percentile uses the mid-rank
(count(s < v) + ½·count(s = v)) ÷ |S|, and the headline also gives the plain rank ("its 4th-largest run-up in 20 intervals"). Quartile buckets come from the same percentile, so there are no fitted cut-points to drift. - A caveat stated on the page: the current run-up covers a partial interval, while the historical ones are complete, so the page says how many sessions have passed. A same-elapsed-days comparison is a v2 candidate, once the simple version shows whether run-up matters at all.
Decision: The Analyst Mix We Didn't Show
The roadmap called for the analyst recommendation mix (buy/hold/sell counts) in the "where the bar
is" block. During the build, a read of the client code turned up a problem: the existing
recommendations method calls the analyst-estimates endpoint and synthesizes
the counts from the analyst total (StrongBuy = n/4, Buy = n/3, and so on).
Those aren't recommendation counts. They're arithmetic on a headcount.
On a page whose whole premise is auditable arithmetic, printing them would have been a made-up number with a real-looking label. So the mix is left off, the price-target consensus (a real endpoint) is shown, and the design doc records two follow-ups: probe the genuine grades-consensus endpoint, and fix the synthesized counts where the report engine already uses them. The feature is smaller than planned, and nothing on it is invented.
The One Input You Can't Backfill: A Consensus Journal
How the bar moved into a print (estimate revisions) can't be backfilled from the data GMI buys: the calendar feed GMI uses returns the current consensus, not how it got there. The history only exists if someone wrote it down at the time. So the feature started a journal on day one, knowing it will look thin for its first quarter:
- One market-wide call a day. The earnings calendar for the next 60 days, the same endpoint the Lab already uses, fetched once per UTC day after 12:00, after the provider's overnight refresh and before the US open. About $0.15 a month.
- Only changes are stored. Daily snapshots of every upcoming print would be ~3–6k rows a day in season (1M+ a year). A row is written only when a symbol's EPS estimate, revenue estimate or date changes. That rebuilds the same step-function history at a fraction of the size.
- A run log proves what "unchanged" means. Every run writes one row (calls, entries seen, rows written). "Unchanged for 12 days" is backed by 12 observed runs, not inferred from the lack of change rows. It's the same honest-numbers pattern as the news poller's log.
- Restart-safe and cheap to skip. It runs inside the schedules engine's existing loop. The due check costs one DateTime compare, today's run row blocks a second paid call after a restart, and there's a config kill switch.
The page says what it has: "Consensus EPS moved from $1.20 to $1.26 (+5.0%) since GMI began recording it on Sep 29", or, when there's nothing yet, "Revision history is accumulating".
Wave 21b: GMI's Read, and What It Can't Tell You
The brief for the showcase page was: show the run-up, what the company did around its last four
years of prints, and what we think it will do. That last part is where a page like this
usually starts making things up. The answer here is a statistical read of the company's
own history, labelled as one, with the arithmetic printed underneath. It comes from one
pure function, EarningsReadMath, and every surface only draws it.
- The probability of a higher next-session close is a beta-binomial posterior mean shrunk toward 50%:
(rose + 24 × 0.50) ÷ (n + 24). The prior is "worth 24 prints", six years of quarters. A raw 2 of 6 (33%) reads 46.7%, and even a perfect 16 of 16 reads 70%, so a small sample can't produce an extreme number. - The headline lens is chosen by what's known before the print: past prints whose run-up fell in the same quartile as today's. "If EPS beats" and "if it misses" are shown as scenarios and never feed the headline, because you don't know the outcome until it's printed.
- Lean words are fixed bands (60%+ firm, 54–60 slightly firm, 46–54 balanced, 40–46 slightly soft, under 40 soft), and confidence starts from n and drops when the headline lens is thin or the lenses disagree, with the reasons listed.
- The move range gives the 10th–90th and 25th–75th percentiles of past next-session moves, scaled by current versus pre-print volatility (clamped to 0.75–1.33, so it's a nudge, not a rewrite). Volatility excludes earnings-reaction sessions, so past jumps don't inflate "normal".
Every read carries a "statistical read · not advice" tag, and the voice rule got stricter: the test
now bans the whole word "will" along with "likely", "expect", "bullish", "bearish" and
"guarantee", across every new sentence, label, confidence reason and recipe line
(EverySentence_IsDescriptive).
flowchart LR
SIM[Seeded simulator<br/>120 companies, 28 prints each<br/>3 worlds] --> CUT[For each print from the 9th:<br/>prices through print-day close,<br/>earlier prints only]
CUT --> PAGE[Rebuild the whole page<br/>EarningsSetupMath +<br/>EarningsReadMath]
PAGE --> READ[Read: probability,<br/>25-75 and 10-90 bands]
READ --> SCORE[Score against the<br/>next-session move]
SCORE --> BRIER[Brier vs flat 50%<br/>and unshrunk counts]
SCORE --> COV[Band coverage<br/>vs nominal]
BRIER --> GRID[Grid over prior weight k<br/>4 to 96]The First Version Scored Worse Than a Coin Flip
The first design looked sensible: shrink the overall rate toward 50% with a weight of 8, and shrink each lens toward the company's own rate with a weight of 6 (hierarchical shrinkage, the textbook move). Its reads spread from 32% to 67%, which made the page look decisive.
Then it was walked forward. A seeded simulator generated 120 companies × 20 reads (2,400 scored reads) in three worlds: no signal (every reaction a coin flip), company bias (each company has its own persistent up-rate, 30–70%), and run-up reversal, a deliberately strong conditioner where a big run-up closes higher only 30% of the time. The Brier scores were 0.271, 0.257 and 0.263. A flat 50% scores 0.250. The first version was worse than a coin flip in every world, including the one with a real signal in it, and its confident-looking reads came true about half the time.
The fix was a grid over the prior weight k, with every lens shrunk straight toward 50%:
| k | 4 | 8 | 12 | 16 | 24 | 32 | 48 | 96 |
|---|---|---|---|---|---|---|---|---|
| Mean Brier, three worlds | 0.2609 | 0.2533 | 0.2511 | 0.2503 | 0.2497 | 0.2496 | 0.2495 | 0.2497 |
The curve is flat from about 20 to 96. That's partly trivial, since a constant 50% is the best possible read in the no-signal world, so the lowest number wasn't the goal. 24 is the lightest weight on the flat part, and it's best-or-tied in both signal worlds, so a real pattern can still move the number. The hierarchical version lost on average (best 0.2509) and won only in the company-bias world, because a company's own rate is itself mostly noise at 16 prints or fewer. The simpler prior was kept.
| World (shipped method) | GMI's Read | Unshrunk counts | Flat 50% | 25–75 band held | 10–90 band held |
|---|---|---|---|---|---|
| No signal | 0.2517 | 0.3365 | 0.2500 | 50.4% | 80.6% |
| Company bias | 0.2488 | 0.3167 | 0.2500 | 49.6% | 81.3% |
| Run-up reversal | 0.2486 | 0.3165 | 0.2500 | 50.6% | 80.7% |
The finding, stated plainly: with at most 16 prints per company, the probability half of GMI's Read is calibrated but carries close to zero skill, even when a strong conditioner exists. It's honest (never extreme, never worse than a coin flip by more than ~0.002) rather than predictive. The page says this in its own limits line, and the design doc records it as the reason there's no per-company ML model: one company's history can't learn the pattern.
The Range Was the Other Bug, and the Sturdier Half
The move range started as ordinary sample percentiles (the Excel/NumPy default, "type 7"). On 8–16 moves those describe the sample, not the next draw: a nominal 80% band held only 70–74% of next moves, and the 50% band held 44–46%. A t-based widening got to about 74% and brought a distributional assumption with it.
The shipped fix is distribution-free: take the value at rank p·(n+1), interpolated and
clamped to the sample's ends (R's type 6, Excel's PERCENTILE.EXC). For a new draw from
the same distribution, P(new ≤ k-th smallest of n) = k/(n+1) exactly, so the bands are calibrated
for the next move by construction. In the walk-forward they held 50% and 81%,
against nominal 50% and 80%, and a 20,000-trial unit test pins the coverage.
Limits the page and the design doc both state: the worlds are simulated, chosen to bracket plausible cases, and show the method's behaviour, not a market edge. A cross-sectional backtest over real cached symbols is the natural next step, and it would let the 50% prior become a measured universe rate. Surprises use the provider's final consensus rather than the morning-of estimate, so the beat scenarios are slightly optimistic until the journal has history. Histories exist only for companies that exist today.
Cost and Blast Radius
| Concern | How it was handled |
|---|---|
| Provider calls per view | ≤3 for a cold symbol (history, prices if the warmer hasn't covered it, price target); 0 once warm |
| New provider method without touching 8 implementers | Per-symbol earnings history added as a default interface method that falls back to the existing calendar call; only the real client overrides it |
| Cache visibility | A new named cache consumer (EarningsSetup) registered in both consumer lists, so the admin cache page reports its hit rate |
| Migration not applied yet | The journaler logs a warning and continues if its table is missing; the page treats a missing table as "no history yet" |
| Tests | 46 unit tests for the base rates: reaction windows per session type and data holes, percentile and quartiles, beat habit, the 8-print floor, conditional counts, journal change-only writes and restart safety, plus the voice check. Wave 21b added 39 more for GMI's Read, including the calibration worlds below |
Deliberately not built: machine learning (the calibration below shows that one company's 16 prints can barely learn even a strong conditioner; a fitted model would need cross-sectional data), tier gating, and public ticker-page integration.