The game that exposed it
The first weekend of the 2026 college football season, Ohio State played Ball State. Our model projected the Buckeyes winning by 32 points. The market had them winning by 50 and a half.
That's not a disagreement about who wins. It's an 18-point argument about how much Ohio State is Ohio State — and the gap was large enough that it wasn't just this game. It was the whole college slate, compressing talent differences that are real, large, and not going away.
Finding out why took a specific kind of diagnostic: looking at where the model was right before looking at where it was wrong.
The pull toward average
Before any game, the model assigns each team a rating — roughly, how good are they in expected points above a league-average opponent? You build that from last season's results, blended with this season's as games are played.
But a team's rating from last year is uncertain evidence. Rosters change. The season was played under different circumstances. How much should you trust a number like that?
One way to handle that uncertainty is to pull every team's estimate toward the league average. A team with a lot of evidence behind it stays where the evidence says. A team with thin evidence gets pushed toward the middle. The stronger the push, the more compressed the ratings become. This is standard statistical practice, and it helps — it just needs to be calibrated to the right league.
That push is controlled by a single constant. The model was built on NFL data. The constant was fit on the NFL. When the model was extended to college football, that value went unchanged.
The NFL is built for parity
In the NFL, pulling team ratings toward the mean makes sense. The league is designed for it.
Every NFL team operates under the same salary cap. No franchise can simply outspend its way to a better roster. The draft hands the worst teams first access to the best incoming players. The whole structure pushes talent toward the middle: year after year, struggling teams get the earliest picks while successful teams lose their best players to contract constraints. The gap between the best and worst teams in the NFL is real but bounded.
When you set a pull-toward-average constant on that league, you're describing something true. The league was built to make teams similar, and they are.
College football was not
College football has none of those mechanisms. There is no salary cap on recruiting. No equalizing draft. The programs that win the most recruit the best players, those players produce more wins and more draft prospects, which attracts more top recruits. The talent pipeline has fed toward the same programs for decades.
The gap between Ohio State and Ball State is not a few points of margin. Ohio State puts ten to fifteen players into the NFL draft most years. Ball State puts one or two in a decade. These programs aren't similar teams at different points in a cycle. They operate at genuinely different levels of the sport, and that difference is structural.
When you apply the same pull-toward-average logic to that gap, you compress it. Ohio State's rating gets pulled down; Ball State's gets pulled up. The margin between them — which is the spread — shrinks dramatically. Our model said 32. The market, pricing on what it actually knows about these programs, said 50.5.
The tell was in the totals
Here's what made the problem legible: the total projections were well-calibrated. The spread projections were not.
Calibration is a season-long check. Take all the games where the model projected a 45-point margin and look at what actually happened. A calibrated model has those games averaging out near 45 — not on any one game, which is always uncertain, but in the aggregate across every game where we made a similar projection. The measure of this relationship is a slope: 1.0 is perfectly calibrated, below 1.0 means projections are consistently running lower than outcomes, above 1.0 means they're consistently running higher.
For game totals, that slope ran at 0.963 across a full college season. Close to 1.0; each point we added to a projected total corresponded to nearly one real point in the final combined scores.
For game spreads, the slope ran at 0.729. Our projections were tracking only about 73% of what reality was doing. When we projected a 32-point margin, the game-level outcomes were running considerably higher on average. When the market was posting 50.5, our model was in the low 30s.
Two parts of the same model, fed by the same ratings, producing the same simulations — and one came out well-calibrated while the other didn't. That divergence is the tell.
It happens for a structural reason. A game's total depends on how much each team scores on its own merits. When you pull both ratings toward the center of the league, each team's absolute scoring level changes modestly, and the two levels still add up to something that tracks the real total reasonably well. The spread is different. The spread is the difference between two ratings. When you compress both ratings toward the center, the gap between them shrinks much more than either one does — and the gap between them is exactly the quantity a spread projection is measuring.
If there had been a bug in the underlying probability math, both the total and the spread would be off in the same direction. The fact that one was tracking reality and the other wasn't pointed directly at the rating scale, not at the math that turns ratings into game outcomes.
What changed
The fix was a college-specific version of the push constant, calibrated to the actual spread of talent in college football rather than to the NFL's parity-built distribution.
The change had one specific constraint: both models are built using the same underlying fitting process. A change to the shared code would have silently recalibrated the NFL's ratings at the same time, replacing two distinct calibrations with one value tuned to neither league. The college model now explicitly passes its own separate constant into the fitting process. The NFL's value is unchanged, and tests verify it stayed that way.
The spread calibration improved measurably after the adjustment. The Ohio State projection moved meaningfully closer to what the market was posting.
What this doesn't say
It doesn't say our college spread projections are fully calibrated against the market. The deeper challenge is an information gap that no constant can fix: early in the season, the model prices last year's results decayed across the offseason. The market prices current rosters. It knows which quarterbacks transferred, which coordinators changed, which recruiting classes outperformed. We know what the 2025 games said. Our early-season spreads sit inside the market's favorites while that gap is closing, and the boards reflect that honestly.
It doesn't say Ohio State versus Ball State is a representative game. A 50-point spread is the extreme end of the college slate — the matchup where the talent difference is largest and where any compression of ratings shows up most visibly. Across a full week of games, the average gap between our projections and the market's was smaller. That game is the clearest illustration of the problem; it's not the median.
And it doesn't say this one constant is the only difference between NFL and college football as modeling problems. The point is wider. Constants built for a salary-capped, parity-designed league carry assumptions about how similar the teams are — and those assumptions are wrong in a sport where the gap between the best and worst programs is the size of the entire NFL spread, and then some. Every time a value crosses from one league to the other, there's a question worth asking: does the thing being measured actually work the same way here? The calibration is how you find out.
Sources
Every number here comes from the graded ledger the boards publish. Research from a statistical model, not betting advice; no outcome is guaranteed. Method: how MatchWiz works.