The power rating against the betting line

Last updated 9 September 2026

The question

Every rating gets asked the same thing: can it beat the line? The closing number at a sportsbook is the best public forecast of a game, because it has absorbed every injury report, every lineup and every other model before tip. So the test that matters for a team-strength rating is not whether it beats the standings. It is how far it sits from the close, and whether the places where it disagrees with the close carry any information.

What was tested

DataBaller's Power Rating was rebuilt as of every game date, seeing only the games already played, priced through the same matchup arithmetic an answer uses, and scored against the closing moneyline with the sportsbook's margin removed. Six NBA seasons (2020-21 through 2025-26, 5,836 in-season games with a close), five MLB seasons (2021 through 2025, 9,822) and six NFL seasons (2020 through 2025, 1,230). The same games were scored against the closing spread on the size of the margin error.

Then every suspected blind spot got its own run on the same games: a different probability curve, shrinking toward last season, an opponent-adjusted margin, weighting recent games more, rest days and back-to-backs, and a lineup report reconstructed from the box score after the fact. Each was fitted on the earlier seasons and read once on the newest. Finally a rating built from the players rather than the team, each player's measured effect on his team's margin summed over the lineup expected to play, was scored the same way.

What turned up

The line wins in every sport. Against the close, the rating's forecast error at Brier is 7.4% worse in the NBA, 1.1% worse in MLB and 8.5% worse in the NFL.

Where the rating and the line disagree, the disagreement is not information. At a five-point gap in win probability the rating's side won about as often as the line said, never often enough to pay the price: a flat stake at the close returned about -5% in the NBA, -2% in MLB and -6% in the NFL, and wider disagreements did worse. The line moved toward the rating's side about one time in four.

Against the spread the picture is the same. The closing number's margin error is 5% lower than the rating's in the NBA and 4% lower in the NFL. Where the two differ by three points, the rating's side covers about 49% of the time and loses 5 to 6% at standard prices.

No single fix closes the gap. A different curve, shrinkage, opponent adjustment, recency weighting and rest are each worth under a point of Brier in the NBA, and all of them together about 1.7. Knowing the lineup from the box score after the fact is worth about two on its own, and everything together about three, which halves the gap and still leaves the disagreements losing. Nothing moves the NFL by more than a point.

Rating the players and summing the lineup does better than rating the team. With only the previous game's lineup it finishes 4.2% behind the market on the 2025-26 NBA holdout, and with the lineup known exactly, 2.8% behind. Better, and still behind.

What it means when you ask

The line is the forecast. The rating is an independent reconstruction of it from team strength, and a gap between them is a question: what is the market pricing beyond the lineup and the season's margins? When the data holds the answer, the response names it, most often who is playing. When it does not, the response says the market is pricing something beyond team strength. A gap is not an opportunity, a mispricing or a fault in the rating, and no slate is ranked by it. DataBaller explains the line. It won't make the pick.

What it does not mean

The 1% gap in MLB is not the rating nearly beating the market. The market still wins, and the disagreements still lose. A rating that trails the close by a point is a good description of team strength and a poor betting model, and the two are different jobs.

The record

Measured 6 and 7 September 2026 with the backtest_team_rating.py market and spread claims, backtest_rating_variants.py, backtest_availability.py and backtest_player_rating.py. Three limits stand on the record. The rating's coefficients were fitted on the same seasons, so the in-season figures are in-sample for the blend. The NFL and MLB holdouts had been consulted in earlier exploration rounds. And the lineup report was built from the box score after the fact, which is the ceiling of an availability feature rather than a feature. Each claim runs again on the first season no model has seen: the 2026 NFL season, then 2026-27 in the NBA and 2027 in MLB. The full portfolio and its testing rules are on the decision metrics page.