Methodology
Odds Bands Explained: Why Favourites and Longshots Behave Differently
Grouping prices into bands reveals structure you can't see event-by-event: how often each price appears, how often it wins, and where calibration drifts.
Key takeaways
- Bands trade event-level noise for population-level structure you can actually measure.
- Appearance rate and win rate answer different questions — never read one as the other.
- Calibration compares price-implied win rate to observed win rate within each band.
- A gap only counts as drift when it survives a Wilson interval and a chi-square check.
- Longshot bands need far more events than favourite bands to reach the same certainty.
A band is a price neighbourhood
A single event is noisy; a band of similar prices is informative. Novus Odds groups American prices into universal bands — deep favourites, near-even, mid-dogs, longshots — and reports how often each band appears on a synthetic tape and how often it wins when it does.
Appearance rate and win rate are different questions. A +400 band can show up constantly yet win rarely; that is the structure you came to inspect, not a bug.
Banding is a deliberate trade. You give up the detail of individual prices and buy back statistical power: a band containing 4,000 events can be measured, while a single price appearing 6 times cannot. The bands are wide enough to accumulate a sample and narrow enough that the prices inside them share a break-even rate to within a couple of points.
Calibration: expected versus observed
For each band the statistics pages compare the price-implied win rate to the observed win rate from thousands of simulated events. If the generative model and the pricing model are consistent, those two lines should track each other closely — and after our per-sport model rewrite, they do, converging within a few percent as the sample grows.
A persistent gap between expected and observed, when it is statistically significant rather than noise, is what calibration drift means. The odds-bands view is the clearest place to see it.
The statistics pages do not stop at the raw difference. Each band's observed rate carries a 95% Wilson interval, a two-sided proportion z-test compares it to the expected rate, and a Pearson chi-square runs a goodness-of-fit across all bands at once. That last test matters: with eight bands on screen, one of them showing a 'significant' gap at the 5% level is the expected outcome of pure chance, not evidence of anything.
Longshot bands need more data than they look like they do
Precision is not distributed evenly across the table. The variance of an observed proportion is p(1 − p) / n, which is largest near 50% in absolute terms — but what matters for a longshot is relative precision, and there the picture reverses badly.
A band with a true 50% win rate measured over 1,000 events has a standard error of about 1.6 points: a 3% relative error. A band with a true 5% win rate over the same 1,000 events has a standard error of about 0.7 points — smaller in absolute terms, but a 14% relative error. To pin a longshot band's rate down as tightly as the even-money band, you need roughly nineteen times the events.
This is why longshot rows in a calibration table look erratic and why 'the +600 band is underperforming' is almost always a statement about sample size. Check the interval width before you check the point estimate.
The favourite-longshot pattern
In many real markets, longshots are systematically overbet and favourites underbet — the favourite-longshot bias. In the lab you can construct or remove that bias deliberately and watch how it distorts the band table, which is a cleaner way to understand it than staring at real historical odds.
The signature is a tilt, not a spike: expected and observed track closely at short prices and separate progressively as prices lengthen, with observed falling below expected at the long end. Because it is monotone across bands rather than concentrated in one, it survives the chi-square test in a way that a single noisy row does not.
Read the tilt as a property of the pricing model, not a discovery about the world. In a synthetic market you put the bias there, or you did not. What transfers to real analysis is the method: measure by band, quantify the uncertainty, and require the pattern to be monotone before you call it a bias.
Frequently asked
What is the difference between appearance rate and win rate?
Appearance rate is how often a band shows up on the tape — a property of the pricing model. Win rate is how often selections in that band actually win — a property of the outcomes. A band can be common and rarely win, or rare and often win.
When is a gap between expected and observed real?
When it survives the statistics. Check whether the expected rate falls outside the band's 95% Wilson interval, whether the proportion z-test is significant, and whether the chi-square across all bands rejects the fit. A single flagged row out of eight is what chance alone produces.
Why do longshot bands look so noisy?
Because relative precision is much worse at low probabilities. A 5% band needs roughly nineteen times the sample of a 50% band to measure its rate to the same relative accuracy, so with equal event counts the long end of the table will always wobble more.
Does the favourite-longshot bias exist in Novus Odds?
Only when you configure it. The bands are a measurement surface: you can construct the bias, remove it, or vary its strength, then watch the calibration table respond. That controllability is the reason it is easier to learn here than from historical odds.
Try it yourself
Everything in this article is something you can run and tweak in the lab — with your own settings and a reproducible seed.
Open the lab →