GMI · TECHNOLOGY OBSERVATORY // ALL SYSTEMS NOMINAL
ENGINEERED BY LEOPARD DATA

Earnings Setup: Base Rates, Not Verdicts

A stock runs up on last quarter's strength, beats, and sells off anyway because the beat was already in the price. Wave 21 turns that "expectations gap" into arithmetic a reader can check: every rate carries its sample size, nothing is model-written, and one number the provider offered was left off the page on purpose. Wave 21b then added GMI's Read, a labelled statistical read of the company's own history, and the first version of it failed its own backtest.

BUILT · SEPTEMBER 2026 In final testing ClientWeb · AdminWeb · MAUI

What It Shows

Click any symbol in the Earnings Lab (recent results or the upcoming calendar) and you get its Earnings Setup: how far the stock has run since its last print and where that run ranks against its own history, drawn on a four-year run-up chart with every print marked by its EPS outcome; a table of its last 16 prints (four years: EPS and revenue surprise, the run-up into each print, and the next-session move); how often it beats; where the bar sits for the next print; the plain conditional base rates; and, since Wave 21b, GMI's Read.

BlockStatusSource
Since the last print: run-up, percentile vs. own history, news volumeBuiltEOD price cache, per-symbol earnings history, news archive (count only)
Reaction table: last ≤16 prints (21b: shared-scale move bar, "beat, then fell" rows badged)Builtearnings history + adjusted closes
Beat habit: counts, median surprise, trend, streakBuiltearnings history
Where the bar is now: consensus, revisions, price targetBuiltearnings history, GMI's own estimate journal, price-target consensus
The read: conditional base rates, 8-print floorBuiltderived only from the reaction-table rows
21b: run-up chart, 16 prints at a glance, GMI's Read with its own track recordBuiltthe same rows, plus 60-session pre-print volatility from the price cache
Last call's transcriptSeam onlywaiting on a probe of the provider contract
Options-implied moveSeam onlywaiting on the same probe

The unbuilt blocks exist as nullable fields on the DTO and a registered no-op provider. While they are null, nothing renders: no "coming soon" box, so a missing feature costs zero pixels.

DATA FLOW
rendering diagram…
flowchart LR
    LAB[Earnings Lab row click] --> API[GET earnings-setup for a symbol]
    API --> SVC[EarningsSetupService]
    SVC --> H[Earnings history<br/>cached 1 day]
    SVC --> P[Adjusted daily closes<br/>append-only price cache]
    SVC --> T[Price-target consensus<br/>cached 7 days]
    SVC --> J[(Estimate journal<br/>MySQL, change-only rows)]
    SVC --> N[(News archive<br/>article count only)]
    H --> M[EarningsSetupMath<br/>pure, unit-tested]
    P --> M
    T --> M
    J --> M
    N --> M
    M --> OUT[Reaction table, beat habit,<br/>bar now, the read<br/>every rate with its n]
    CRON[Schedules engine<br/>once a day, 1 call] --> J
One endpoint, one pure math function. A warm symbol costs zero provider calls: earnings history is cached a day, prices are append-only, price targets are cached a week. The estimate journal and the news count are plain MySQL reads.

The Rule: Computed, Not Model-Written

This was the obvious place to put an LLM: "summarize how this stock tends to trade around earnings." I chose not to. A language model will happily write "tends to rise after beats" from four examples, and a reader can't check it. Every sentence on this page is instead produced by one pure function, EarningsSetupMath, shared by the API and all three apps so they can't disagree, and every sentence follows four rules:

  • Every rate carries its sample size in the same sentence. "Closed higher the next session 11 of 20 times", never "usually rises".
  • A floor of 8 prints. Below it the page shows the table and the count and says there isn't enough history for a base rate. A conditional line built on fewer than 5 prints says "a small sample" in the same sentence.
  • Nothing is fitted, weighted or extrapolated. The rows that produce every count are the rows in the reaction table on the same screen. (GMI's Read, below, does shrink its estimates, so it's labelled as a statistical read, prints its recipe with the numbers plugged in, and sits beside these unshrunk counts rather than replacing them.)
  • Descriptive voice, enforced by a test. A unit test fails the build if any generated sentence contains "buy", "sell", "should", "will rise" or "likely".

The line that justifies the whole page is the beat-conditioned one: "In the 14 prints where EPS beat estimates, it closed higher 7 of 14." A high beat rate next to a coin-flip reaction is the expectations gap, shown with no opinion added. (Example numbers, for illustration.)

The Recipe (Published on the Page's Help Panel)

All prices are adjusted closes, so a 4-for-1 split never shows up as a −75% reaction. The reaction window depends on when the company reported:

Session reported"Before" close"After" close
After the close (AMC)the print datefirst trading day after
Before the open (BMO)last trading day beforethe print date
Not reportedlast trading day beforefirst trading day after (labelled a 2-session window)
  • Data holes are excluded, not guessed. A print with a price gap of more than 5 calendar days on either side of its window (a hole in the data, not a weekend) gets no reaction and is left out of every rate.
  • Run-up into a print = that print's "before" close ÷ the previous print's "after" close − 1: the drift across the whole gap between two prints.
  • Percentile uses the mid-rank (count(s < v) + ½·count(s = v)) ÷ |S|, and the headline also gives the plain rank ("its 4th-largest run-up in 20 intervals"). Quartile buckets come from the same percentile, so there are no fitted cut-points to drift.
  • A caveat stated on the page: the current run-up covers a partial interval, while the historical ones are complete, so the page says how many sessions have passed. A same-elapsed-days comparison is a v2 candidate, once the simple version shows whether run-up matters at all.

Decision: The Analyst Mix We Didn't Show

The roadmap called for the analyst recommendation mix (buy/hold/sell counts) in the "where the bar is" block. During the build, a read of the client code turned up a problem: the existing recommendations method calls the analyst-estimates endpoint and synthesizes the counts from the analyst total (StrongBuy = n/4, Buy = n/3, and so on). Those aren't recommendation counts. They're arithmetic on a headcount.

On a page whose whole premise is auditable arithmetic, printing them would have been a made-up number with a real-looking label. So the mix is left off, the price-target consensus (a real endpoint) is shown, and the design doc records two follow-ups: probe the genuine grades-consensus endpoint, and fix the synthesized counts where the report engine already uses them. The feature is smaller than planned, and nothing on it is invented.

The One Input You Can't Backfill: A Consensus Journal

How the bar moved into a print (estimate revisions) can't be backfilled from the data GMI buys: the calendar feed GMI uses returns the current consensus, not how it got there. The history only exists if someone wrote it down at the time. So the feature started a journal on day one, knowing it will look thin for its first quarter:

  • One market-wide call a day. The earnings calendar for the next 60 days, the same endpoint the Lab already uses, fetched once per UTC day after 12:00, after the provider's overnight refresh and before the US open. About $0.15 a month.
  • Only changes are stored. Daily snapshots of every upcoming print would be ~3–6k rows a day in season (1M+ a year). A row is written only when a symbol's EPS estimate, revenue estimate or date changes. That rebuilds the same step-function history at a fraction of the size.
  • A run log proves what "unchanged" means. Every run writes one row (calls, entries seen, rows written). "Unchanged for 12 days" is backed by 12 observed runs, not inferred from the lack of change rows. It's the same honest-numbers pattern as the news poller's log.
  • Restart-safe and cheap to skip. It runs inside the schedules engine's existing loop. The due check costs one DateTime compare, today's run row blocks a second paid call after a restart, and there's a config kill switch.

The page says what it has: "Consensus EPS moved from $1.20 to $1.26 (+5.0%) since GMI began recording it on Sep 29", or, when there's nothing yet, "Revision history is accumulating".

Wave 21b: GMI's Read, and What It Can't Tell You

The brief for the showcase page was: show the run-up, what the company did around its last four years of prints, and what we think it will do. That last part is where a page like this usually starts making things up. The answer here is a statistical read of the company's own history, labelled as one, with the arithmetic printed underneath. It comes from one pure function, EarningsReadMath, and every surface only draws it.

  • The probability of a higher next-session close is a beta-binomial posterior mean shrunk toward 50%: (rose + 24 × 0.50) ÷ (n + 24). The prior is "worth 24 prints", six years of quarters. A raw 2 of 6 (33%) reads 46.7%, and even a perfect 16 of 16 reads 70%, so a small sample can't produce an extreme number.
  • The headline lens is chosen by what's known before the print: past prints whose run-up fell in the same quartile as today's. "If EPS beats" and "if it misses" are shown as scenarios and never feed the headline, because you don't know the outcome until it's printed.
  • Lean words are fixed bands (60%+ firm, 54–60 slightly firm, 46–54 balanced, 40–46 slightly soft, under 40 soft), and confidence starts from n and drops when the headline lens is thin or the lenses disagree, with the reasons listed.
  • The move range gives the 10th–90th and 25th–75th percentiles of past next-session moves, scaled by current versus pre-print volatility (clamped to 0.75–1.33, so it's a nudge, not a rewrite). Volatility excludes earnings-reaction sessions, so past jumps don't inflate "normal".

Every read carries a "statistical read · not advice" tag, and the voice rule got stricter: the test now bans the whole word "will" along with "likely", "expect", "bullish", "bearish" and "guarantee", across every new sentence, label, confidence reason and recipe line (EverySentence_IsDescriptive).

WALK-FORWARD HARNESS
rendering diagram…
flowchart LR
    SIM[Seeded simulator<br/>120 companies, 28 prints each<br/>3 worlds] --> CUT[For each print from the 9th:<br/>prices through print-day close,<br/>earlier prints only]
    CUT --> PAGE[Rebuild the whole page<br/>EarningsSetupMath +<br/>EarningsReadMath]
    PAGE --> READ[Read: probability,<br/>25-75 and 10-90 bands]
    READ --> SCORE[Score against the<br/>next-session move]
    SCORE --> BRIER[Brier vs flat 50%<br/>and unshrunk counts]
    SCORE --> COV[Band coverage<br/>vs nominal]
    BRIER --> GRID[Grid over prior weight k<br/>4 to 96]
The calibration rebuilds the whole page as of each past print, using only prices through that day's close and only earlier prints. A separate test proves the in-page track record equals this point-in-time rebuild, so the number a user sees can't quietly use the future.

The First Version Scored Worse Than a Coin Flip

The first design looked sensible: shrink the overall rate toward 50% with a weight of 8, and shrink each lens toward the company's own rate with a weight of 6 (hierarchical shrinkage, the textbook move). Its reads spread from 32% to 67%, which made the page look decisive.

Then it was walked forward. A seeded simulator generated 120 companies × 20 reads (2,400 scored reads) in three worlds: no signal (every reaction a coin flip), company bias (each company has its own persistent up-rate, 30–70%), and run-up reversal, a deliberately strong conditioner where a big run-up closes higher only 30% of the time. The Brier scores were 0.271, 0.257 and 0.263. A flat 50% scores 0.250. The first version was worse than a coin flip in every world, including the one with a real signal in it, and its confident-looking reads came true about half the time.

The fix was a grid over the prior weight k, with every lens shrunk straight toward 50%:

k48121624324896
Mean Brier, three worlds0.26090.25330.25110.25030.24970.24960.24950.2497

The curve is flat from about 20 to 96. That's partly trivial, since a constant 50% is the best possible read in the no-signal world, so the lowest number wasn't the goal. 24 is the lightest weight on the flat part, and it's best-or-tied in both signal worlds, so a real pattern can still move the number. The hierarchical version lost on average (best 0.2509) and won only in the company-bias world, because a company's own rate is itself mostly noise at 16 prints or fewer. The simpler prior was kept.

World (shipped method)GMI's ReadUnshrunk countsFlat 50%25–75 band held10–90 band held
No signal0.25170.33650.250050.4%80.6%
Company bias0.24880.31670.250049.6%81.3%
Run-up reversal0.24860.31650.250050.6%80.7%

The finding, stated plainly: with at most 16 prints per company, the probability half of GMI's Read is calibrated but carries close to zero skill, even when a strong conditioner exists. It's honest (never extreme, never worse than a coin flip by more than ~0.002) rather than predictive. The page says this in its own limits line, and the design doc records it as the reason there's no per-company ML model: one company's history can't learn the pattern.

The Range Was the Other Bug, and the Sturdier Half

The move range started as ordinary sample percentiles (the Excel/NumPy default, "type 7"). On 8–16 moves those describe the sample, not the next draw: a nominal 80% band held only 70–74% of next moves, and the 50% band held 44–46%. A t-based widening got to about 74% and brought a distributional assumption with it.

The shipped fix is distribution-free: take the value at rank p·(n+1), interpolated and clamped to the sample's ends (R's type 6, Excel's PERCENTILE.EXC). For a new draw from the same distribution, P(new ≤ k-th smallest of n) = k/(n+1) exactly, so the bands are calibrated for the next move by construction. In the walk-forward they held 50% and 81%, against nominal 50% and 80%, and a 20,000-trial unit test pins the coverage.

Limits the page and the design doc both state: the worlds are simulated, chosen to bracket plausible cases, and show the method's behaviour, not a market edge. A cross-sectional backtest over real cached symbols is the natural next step, and it would let the 50% prior become a measured universe rate. Surprises use the provider's final consensus rather than the morning-of estimate, so the beat scenarios are slightly optimistic until the journal has history. Histories exist only for companies that exist today.

Cost and Blast Radius

ConcernHow it was handled
Provider calls per view≤3 for a cold symbol (history, prices if the warmer hasn't covered it, price target); 0 once warm
New provider method without touching 8 implementersPer-symbol earnings history added as a default interface method that falls back to the existing calendar call; only the real client overrides it
Cache visibilityA new named cache consumer (EarningsSetup) registered in both consumer lists, so the admin cache page reports its hit rate
Migration not applied yetThe journaler logs a warning and continues if its table is missing; the page treats a missing table as "no history yet"
Tests46 unit tests for the base rates: reaction windows per session type and data holes, percentile and quartiles, beat habit, the 8-print floor, conditional counts, journal change-only writes and restart safety, plus the voice check. Wave 21b added 39 more for GMI's Read, including the calibration worlds below

Deliberately not built: machine learning (the calibration below shows that one company's 16 prints can barely learn even a strong conditioner; a fitted model would need cross-sectional data), tier gating, and public ticker-page integration.