Back

Why we stopped predicting

This app used to score every stock on a five-point scale, from strong buy to strong sell. It doesn’t any more. This page is the reason, with the numbers that produced it.

We replayed the engine across 3,792 readings of history, computing each verdict only from data available on the day it was made, then measuring what actually happened next. The question was simple: when it said buy, did the stock beat simply holding?

It did the opposite. The more bullish the call, the worse the stock did.

VerdictReadingsAvg returnvs holding
Strong buysignificant275-0.14%-0.93%
Buy8960.44%-0.34%
Hold1,7700.87%+0.08%
Sell7331.07%+0.29%
Strong sellsignificant1182.48%+1.70%

Returns are averages over 10 trading days. Rows marked “significant” clear two standard errors — with samples this size that is a sanity filter, not a formal test.

Read down the column. The average return rises steadily from the most bullish verdict to the most bearish — a clean inversion of what the scale claimed. Holding these names over the same windows returned 0.78%, and the strongest buys returned less than that.

Its directional calls were right 48.5% of the time. A coin flip is 50%.

We did not invert the signal and ship that instead. A result that survives one universe over one period is a hypothesis, not an edge, and building a product on a backwards signal because it backtested well is the same mistake in the other direction. The honest response to “our forecast was wrong” is to stop forecasting.

What this does not measure

A backtest quoted without its limits is how a null result turns into a marketing claim. These are ours.

  • Sentiment is excluded. No historical archive existed when this ran, so this measures the technical signal alone — while the live engine weighted sentiment at 55%.
  • Survivorship bias. These tickers were chosen today, so every one of them survived the period. Real portfolios include the ones that did not.
  • No costs. Spreads, commissions and slippage are ignored, and they fall hardest on the most active strategies.
  • Overlapping windows are avoided by spacing signals, but the sample is still small enough that the significance test is a sanity filter rather than a formal result.

What we do instead

Describe what has already happened, in language you can check against the chart and the headlines on the same page. How a stock moved today and over the past fortnight. Whether that move was large for that stock, measured against its own daily range rather than a fixed percentage. Whether the market and its sector moved with it. What was published about it.

None of that is a forecast, and none of it needs to be. Knowing what happened and why is useful on its own — and unlike a forecast, you can check whether we got it right.

Measured over 16 tickers, 2021-12-07 to 2026-08-21, with each verdict judged 10 trading days out and signals spaced 5 days apart to avoid overlapping windows. The raw results and the code that produced them are both public. This is not financial advice.