The Forecast-Off

Week 9 · Forecasting Wrap · every method against the same six held-out months of Summit Gear revenue
Built-in AI tutor. Want to talk through which method to trust, or how big a hedge the evidence supports? Ask the helper on this page. It will not hand you a number before you commit a forecast of your own.

Here is the whole week on one page. You are looking at three years of Summit Gear monthly revenue. The last six months are held out — hidden. Every forecasting method on the menu is going to forecast those same six months, and we will score each one by how far it missed. But your gut goes first. You commit a forecast blind, then you watch eight methods, plain to fancy, try to beat it. Then you hedge the number you would actually ship. The lesson is not which model wins. It is that a more complex method is not automatically more accurate, and that a point estimate shipped without a range hides how far it can miss.

On this page: 1. Your SWAG  ·  2. Reveal the methods  ·  3. The leaderboard  ·  4. Turn the knobs  ·  5. Hedge the number

The shaded band on the right is the held-out window: the last six months (Jul–Dec 2025). The methods only ever saw the 30 months to its left. Click a legend chip to hide or show a line.

1. Commit your SWAG first

SWAG first · commit blind

A SWAG is a scientific wild guess: the first number you put down before anyone runs a model. It is the baseline every method has to beat. Look at the shape of the 30 months you can see, and forecast the six you cannot. One number is enough to start: your best guess for a typical month in the held-out window. If you want to draw the path, open the per-month box.

$ M / month
Type a number to see it echoed in dollars.
Draw the path instead — set each held-out month

Optional. Fill all six to forecast a shape (for example, a summer peak fading into winter). If any are blank, we use your single number above, held flat.

Why commit before you see anything? Because the point of the exercise is to find out whether the models are worth their complexity, and you cannot judge that if you have already seen the answer. Your blind SWAG is the baseline every method must beat. If a method cannot beat the number you scribbled in ten seconds, the extra complexity it takes to run is not buying you accuracy.

2. Reveal the truth, then the methods

One method at a time

First we reveal the six months that were held out — the actuals, the truth every method is judged against. Then the methods land one at a time, plainest first. Each drops its forecast line on the chart above and posts its MAE (mean absolute error, the average miss in dollars) and RMSE (root mean squared error, the same miss but with the big misses weighted heavier). Watch the order they arrive in against the order they finish in.

Waiting on your SWAG.

3. The leaderboard

Ranked by RMSE · your SWAG included

Every method that has revealed, plus your SWAG, ranked by RMSE (lower is a smaller miss). Your row is highlighted. Watch where your gut lands as the better methods arrive.

#MethodMAE ($)RMSE ($)
Lock in a SWAG and start revealing to build the board.

4. Turn the knobs

Live re-ranking · what-if

Three methods have a knob. Turn it and its forecast recomputes and the board reorders live. Two things to notice: the tuned method rarely leaps to the top, and the moment you move a knob off its default the number becomes a what-if — an exploration, not the canonical backtest. Snap the knob back to the marked default to restore the scored number.

Reveal all the methods first, then the knobs turn on.

5. Every forecast ships with a hedge

Two realities · reasoned vs political

A point estimate with no range hides how far it can miss. The question is how you build the range. There are two realities. In the first, you reason the pad from evidence: the backtest error you just measured is the buffer the evidence supports, because it is literally how far this method has missed before. In the second, the boss says move it y% — up for a stretch target, down to sandbag. Sometimes that is real conservatism; often it is a number moved with no evidence behind it, and either way it has to be documented and flagged.

How the band covered the six held-out months

Pick a method and a way to build the band.

And here is the number you would actually ship

The winning method (SARIMAX) forecasts the next quarter, and the hedge band is its backtest MAE of — the evidence-based buffer, applied to the number that leaves the building. This is the forecast a controller can defend: a method, a measured error, and a measured range.

Where this goes next. The discipline you just ran — hold out months you have not seen, forecast them, score the miss, rank the methods, then hedge from measured error — is exactly what Project 2 asks of your credit-default classifier: fit on data it has seen, judge it on data it has not. Start simple, beat the baseline, prove it on the held-out part, and hedge the number you report.