MatchWiz

Why Strikeouts Work When Home Runs Don't

Strikeouts post 61% and +10.8% ROI at a gap of 1.5 or more. HR, total bases, H+R+RBI, and UFC all graded negative. Here's the honest map.

The gap

Before every game, sportsbooks post a line for the starting pitcher's strikeout total — how many batters he's expected to fan before he leaves. You take the over or the under. That's a strikeout prop, and it's what we've spent more time on than any other individual stat in baseball.

The number that came out of that work: when our projection sits 1.5 or more strikeouts above the sportsbook's line, those bets have hit 61% of the time at +10.8% return on money risked. A good bettor clearing 3 to 5 cents on the dollar is doing well. Plus eleven is a gap worth naming.

The 1.5-strikeout threshold is the load-bearing piece. Below it, the signal breaks down. When we're 1.25 to 1.5 strikeouts ahead of the book, the full-season result is −1.3% — and it split positive in one half of the season and negative in the other, which is what noise looks like, not an edge. Split the 1.5+ plays in half by date: +19.7% in the first half, +16.7% in the second. Both halves pointing the same direction matters.

How we check a result like that

Saying you found an edge is easy. The checklist that decides whether you actually have one is more demanding.

First: grade everything in money, at the price posted at the time. A 55% win rate at −150 odds is a slow loss. The only honest metric is what a dollar risked actually returned.

Second: check whether the pattern strengthens as the disagreement grows. A genuine edge means our biggest departures from the book's number should win at the highest rate. A profitable pocket surrounded by losses is usually coincidence.

Third: split the data in half by date and test both sides separately. A pattern that holds only in the half you analyzed first is probably noise. Strikeouts held both halves.

Fourth: run the check across every level of disagreement, not just the slice that looks good. Isolating one bin while the neighbors lose is selection, not discovery.

We ran all four on strikeouts. Then we ran them on seven other individual markets in baseball — and on a full slate of UFC fights — to make sure we weren't stopping the moment we saw something we liked.

What didn't make the cut

Home runs: our model graded −36% against actual posted prices at the threshold it was running. We took every input the model uses — time-window weights, park factors, weather effects, how much recent form matters versus longer-term rates — and tested thousands of combinations, maximizing profit at each step. The best combination that search produced: −5.5%. Still negative. The longer-shot price bands, where a genuine edge would show up most clearly, came back at −16% to −32%. We stopped looking for an HR edge because we couldn't find a profitable configuration anywhere in the data, even when we tried to design one.

Total bases: same process. The best configuration after a full sweep: −4.2%, against a current model already running at −4.4%. The search gained a tenth of a percentage point and never crossed zero.

Hits, runs, and RBIs combined: every level of disagreement we tested, −4.5% to −8.4%. Not one profitable slice across the whole distribution.

Walks and stolen bases: both negative across every test. Stolen-base lines are posted at odds so heavy on the favorite side that you're laying significant juice before any modeling advantage can register. Nothing cleared it.

UFC fight bets: the same four-step process on a full slate of individual UFC markets. Confirmed efficient. Nothing came back positive.

What these markets share is that the sportsbooks price them at least as accurately as we do. When our number disagrees with theirs there, we're more often the one who's wrong.

Why strikeouts are different

A home run prop asks a lot. The batter has to make contact, hit it far enough, in the right direction, to a part of the park that's beyond the fence. Sportsbooks price home runs from detailed contact data — how hard a hitter tends to hit the ball, his launch angle, where he pulls it, how those tendencies fit the park he's playing in. A rate-based model built on historical statistics doesn't carry that granularity, which is the most likely explanation for why we couldn't find an HR edge even after exhaustive tuning.

A strikeout prop asks a narrower question: how many batters will this starter retire on strikes today? It's mostly about one player's skills against one lineup's contact tendencies. A pitcher who's been missing bats at a certain rate tends to keep doing it. The lineup's ability to make contact against that kind of stuff is measurable and fairly stable. The matchup is tighter and the relevant history is longer.

Books price this from the same broad inputs, but a sportsbook's job is to set a line where roughly equal money comes in on both sides — not to produce the most accurate prediction of the actual K total. A market-clearing price and the best forecast of an outcome are related. In strikeouts, they're apparently not identical. That gap is where the edge lives.

The tune we almost shipped

Midway through validating the strikeout model, we ran a tuning experiment: improve how well the projections rank pitchers against each other, measured by correlation between our numbers and actual results.

The tune worked. The ranking correlation improved.

It also cut the return on bets from +10.8% to −3.1%.

The two objectives — ranking accuracy and betting profit — are anti-correlated for strikeouts. The changes that make the projections more accurate by a ranking measure bring them closer to the sportsbook's numbers, which shrinks the disagreements worth betting on. We rejected the tune and kept the original weights.

The lesson is specific: a model that ranks pitchers correctly and a model that finds profitable bets are different models. Conflating them would have quietly ended the one edge we've found at meaningful volume in baseball.

What this doesn't say

It doesn't say strikeout props are safe. Edges narrow. A market where one side wins reliably tends to get priced accordingly over time. We don't know how long this holds, and we won't claim to.

It doesn't say the other markets are unbeatable. Confirming that our model can't beat them is different from proving no edge exists. The sportsbooks pricing HR and total bases from contact data that our projections don't carry is the most likely explanation for why we fail there — but that's a hypothesis, not a proof. A model built on better inputs might find something we didn't.

There's also a second baseball market where we've found something, but it looks nothing like the strikeout gap. First-inning run markets — whether either team scores in the first inning — initially looked like everything else on the no-edge list. We reopened it, ran a diagnostic to find what we were missing, identified a specific slice of starting-pitcher data we hadn't used correctly, rebuilt the model around it, and found a real result in the tail of the distribution. That edge is thinner and the sample is smaller. It passed the same checklist. The two don't look alike in practice — different data, different bet structure, different volumes — and they arrived by completely different paths.

Six markets negative. One edge at meaningful volume. One edge at the tail. We publish all of it because you're better served knowing where the model actually wins than by a board that treats every projection as equally actionable.

Sources

Every number here comes from the graded ledger the boards publish. Research from a statistical model, not betting advice; no outcome is guaranteed. Method: how MatchWiz works.

CoreyCo-founder · data, site structure and the analytics pieces

Corey built the first version of MatchWiz with Jonathan in Google Sheets and Python — baseball only, back when the point was just to see whether the numbers they were gathering meant anything. Before that he founded ShowZone, the top companion site for MLB The Show, which he started on a Raspberry Pi in his office while working as an engineer at a utility and still owns and runs today — 115,000+ members, and the home of the trademarked True Overall rating system he created. That's what he brings here: data, how a site is put together, and the discipline of publishing a number you can be held to. He writes the analytics and methodology pieces — what the model is bad at, which markets have an edge and which don't, and what a season of graded calls actually says. Ohio State Electrical & Computer Engineering, where he and Jonathan were roommates and, between them, a top-50 pairing in the NCAA Football video game.

x.com