Prop LabHow prop markets are priced
Why Most Sports Betting Systems Are Luck: 31,114 Tests and Zero Survivors
Test enough ideas and some will look brilliant by chance. We searched NFL player props for trait-based angles, then measured exactly how good pure noise looks. Noise won.
By Edgehalla ResearchPublished 7 min read
The trap: test enough ideas and something will “work”
Every betting system starts with a backtest that looks great. Fast receivers in domes. Short corners on the road. Unders in cold weather. The record says 62%, so it must be real.
Here is the problem. At the usual 5% significance level, about one test in 20 “works” even when nothing is there. Run 100 tests on pure noise and you expect five winners. Run thousands and you get hundreds.
This is called the multiple-testing problem. It is the main reason most betting systems fail the moment real money touches them. We decided to measure it directly on NFL player props.
The search: 31,114 tests on NFL props
Our Prop Lab crossed player traits (40 times, jumps, height, weight, age, role), defense traits (likely corners, safeties, linebackers, big plays allowed) and game conditions (temperature, wind, dome, total, spread, home or away).
Each trait was split into thirds. A combination needed at least 150 decided bets in 2023–24 to be scored. That produced 15,557 combinations across receiving yards, receptions, rushing yards and passing yards. Each was tested as an over and as an under: 31,114 tests.
The rules were written and committed before any result was computed. That matters. It stops us from quietly changing the rules after seeing what wins.
The search funnel
Thousands of tests went in; after correcting for the number of tests, nothing came out as an edge.
Only 716 tests cleared p < 0.05. That is about 2.3% of the total, which is fewer than the roughly 1,556 you would expect from luck. Each test asked for a win rate above the 52.4% break-even, not just 50%, and many combinations overlap heavily, so this is not a clean coin-flip count. But it is certainly not a sign of hidden edges.
Then we applied the Benjamini–Hochberg correction. It asks how many “discoveries” you can claim while keeping false ones to 10% or less. The answer was zero.
The shuffle test: how good does noise look?
A correction formula is one check. We wanted to see the noise with our own eyes. So we shuffled the real over and under results within each market and week. That keeps the line, the player and the prices, but destroys any true link between traits and results.
Then we ran the entire search again on the shuffled data. And again. Two hundred times. Each run, we recorded the best angle it found.
Our best real angle vs the best angle in random data
Strength of the top result (z-score) in each search
Our best real prop angle scored below the typical best angle found in shuffled, meaningless data.
The best combination in shuffled data typically “won” about 65% over 150+ bets. Our real best won 65.5% on 177 bets. In 129 of 200 shuffles, pure noise matched or beat it.
Put plainly: our best discovery was less impressive than what random data usually produces. A 65% record over 177 bets feels like proof. In a search this big, it is the expected result of nothing at all.
The same was true of the smaller, hand-picked hypotheses. Their best real score was z 0.96. The bar set by their own shuffle test was 1.88.
The holdout season: regression to the mean
We froze the top 20 combinations and tested each one once on 2025, a season the search never saw. 19 of the 20 fell, some by as much as 18 points. Only one rose.
Top discovered angles, before and after
Win rate in discovery (2023–24) vs the untouched 2025 holdout
Angles that looked like 60%+ winners in discovery fell back toward break-even or below in the holdout season.
- Discovery 2023–24
- Holdout 2025
- Blind under 2025
This is regression to the mean. The angles that topped the search did so partly because of luck. When the luck is gone, they slide back.
There is a second trap hiding here. All 20 frozen angles were unders, and 2025 was an under year: receiving-yard overs hit just 46.4% and receptions 45.2%. Twelve angles passed our “promising” rule, but only 10 beat simply betting every under. Always compare an angle with the blind bet on the same side, not with 50%.
A statistical model reached the same answer
Search is one approach. A model is another. For each market, we fitted a regularized model that starts from the market’s own probability and can add any of 30 to 65 trait features.
We tuned it on 2023 and checked it on 2024. Every feature it added made the 2024 predictions worse. The best setting kept zero features in every market. The only thing it learned was a small, constant lean toward the under.
On the untouched 2025 season, its only bets, all rushing unders, won 52.8%. Blindly betting every rushing under that season won 52.75%. ROI was −1.2%. Even all the traits together did not beat closing lines.
The same story in fantasy matchup angles
This is not just a props problem. We ran 21 matchup angles (blitz, man coverage, two-high shells, box counts, play action and more) against our fantasy projections, by position. That was 64 tests.
Grades for 64 matchup angle tests
Player × defense angles vs our fantasy projection, 2024–25
None of the 64 matchup tests earned an A or B grade; most failed outright.
Turned into public picks, those angles made 18,128 backtest calls over 2024 and 2025. They won 49.7%, exactly the 49.7% a matched baseline won. Details are on the matchup edges scoreboard.
How to tell a real edge from a lucky one
- Ask how many ideas were tried. A 60% angle from a search of thousands means little. One idea, written down first, means more.
- Demand a holdout. The test season must be untouched until the angle is frozen.
- Compare with the blind bet. An under angle has to beat betting every under, not 50%.
- Look for a reason. Weather, injuries and roles have a story. “Mid-weight receivers in domes” does not.
- Check the plumbing. A fake +78% ROI in our early tests came from one bad name match. See our model article.
- Wait for a live record. Only bets made before kickoff, logged and never edited, prove anything.
Frequently asked questions
What is the multiple testing problem in sports betting?
When you test many betting angles, some will look profitable by chance. At a 5% significance level, about 1 in 20 tests passes even when nothing is real. Our Prop Lab ran 31,114 tests; none survived a correction for that.
How do you know if a betting system is just luck?
Shuffle the results and re-run the same search. If random data finds “systems” as good as yours, yours is probably luck. In our NFL prop search, shuffled data matched or beat our best real angle in 129 of 200 runs.
Why do backtested betting systems fail live?
Because the best backtest results are partly luck, and luck does not repeat. When we froze our top 20 NFL prop angles and tested them on a new season, 19 of the 20 got worse, some by 18 points.
What is a Benjamini–Hochberg correction?
It is a way to control the false discovery rate when you run many tests at once. It ranks the p-values and only keeps results strong enough that, at most, a set share of the discoveries (we used 10%) are expected to be false.
How many bets do you need to prove a betting edge?
More than most people think, and the number depends on how many ideas were tried. In our search, the best angle in random data typically won about 65% over 150+ bets, so even that is not proof when it comes from a big search.
More in How prop markets are priced
Sources and method
- The Odds API (historical and live player prop lines)
- nflverse data (play-by-play, schedules, weather, injuries, combine)
- Multiple comparisons problem
- False discovery rate and the Benjamini–Hochberg procedure
- Permutation tests
- Regression toward the mean
- National Council on Problem Gambling (1-800-GAMBLER)
- Responsible Gambling Council (Canada)
Research and analysis, not betting advice. Past results do not guarantee future returns; bet only what you can afford to lose.