How Receipts works

Last updated 8 October 2026

Receipts keeps score, in public, of every forecast DataBaller's Power Rating makes for an NFL, NBA or MLB game. It shows how often our pick won, how often the betting line's pick won, and what $100 a game on our pick would have returned. This page explains every number on it.

What's compared

Before every game, the Power Rating gives each team a chance of winning, such as Dodgers 62%. The betting market's closing line, what people often call the Vegas line, implies its own chance for the same game, such as Dodgers 58%. After the game, we compare the two.

Everything here is about who wins the game: the moneyline. It is not about point spreads or totals.

Each forecast is saved about half an hour before the game starts and is never changed afterwards.

How a pick is made

A pick takes two steps: a rating for each team, then a matchup that turns two ratings into a chance of winning. The team with the higher chance is the pick. The numbers below are the ones the current versions use.

MLB

The rating (v2). A team's rating is 3.12 × (its winning percentage − .500) + 0.25 × its expected run margin per game. The expected run margin is rebuilt from what each team actually did at the plate and on the mound, using a fixed value for each event: a single is worth 0.9 runs, a double 1.25, a triple 1.6, a home run 2.0, and a walk or hit-by-pitch 0.7. Runs it was expected to score minus runs it was expected to allow, per game.

The matchup. The expected margin is the home team's rating − the away team's rating + 0.32 runs for playing at home. The chance the home team wins is 1 ÷ (1 + e^(−0.42 × margin)). Whichever team is above 50% is the pick.

A worked example. The Dodgers, rated +1.20, host the Rockies, rated −1.10. The margin is 1.20 − (−1.10) + 0.32 = 2.62 runs. The chance the Dodgers win is 1 ÷ (1 + e^(−0.42 × 2.62)) = 75%. The Dodgers are the pick.

v3 (from 8 October 2026) adds both starting pitchers. Each announced starter gets a score: the runs per nine innings his fielding-independent pitching (strikeouts, walks, hit batters and home runs) is worth, this season and last. The margin moves by −0.48 runs for every point the home starter's score is above the away starter's, so a better home starter raises the home team's chance. With pitchers in the matchup, home advantage is 0.39 runs and the 0.42 above becomes 0.355.

NFL

v2. The rating blends a team's record with its scoring margin, with each game's margin capped at 14 points so one blowout cannot carry a season. In the matchup, home advantage is 1.35 points and the chance of winning is 1 ÷ (1 + e^(−0.146 × margin)).

NBA

v2. The rating is the sum of the player ratings of the lineup expected to play, minus any player the injury report rules out. The report is checked when the game is priced. In the matchup, home advantage is 1.75 points and the chance of winning is 1 ÷ (1 + e^(−0.127 × margin)).

v1, in every league

Until 6 September 2026, every league ran v1: the same ratings with an older conversion that counted home advantage twice. From 7 September, v2 counts it once.

Decision metrics describes each rating in full and how each version was tested before it reached answers.

What "the line" is

The closing line is the last moneyline we recorded at or before the start. We record lines about once an hour, so the closing line can be up to an hour old when the game begins. It comes from the first of these sportsbooks to have one: DraftKings, FanDuel, BetMGM, Caesars, William Hill.

A moneyline price includes the bookmaker's margin, called the vig, so the chances implied by the two sides add up to more than 100%. When we compare chances, we remove it: both prices are turned into chances and scaled so they add up to exactly 100%. The $100 lines are different. They use the actual prices, vig included, because that is what a bettor would have been paid.

Picking the winner

"Our pick won 60%" means that in 60 of every 100 games, the team the Power Rating gave the better chance went on to win. "The line's pick won" is the same count for the team the closing line favoured.

Neither number is very high, and that is the nature of the sport rather than a flaw in either forecast. In baseball, the better team loses about four games in ten. Football and basketball favourites win more often, but plenty of games still go the other way.

Most of the time our pick and the line's pick are the same team. The games where they differ are where the two forecasts actually disagree about the winner, so Receipts shows those separately.

The $100 lines

Each line shows what a flat $100 on our pick, at the closing price with the vig included, would have returned:

  • Every pick: $100 on our pick in every game.
  • Different team: $100 on our pick only in the games where it was not the line's pick.
  • 5+ points apart: $100 only when our chance for a team was at least 5 percentage points higher than the line's, on that team.

The return is the dollars won or lost for every $100 bet, which is the same as a percentage of the money bet. These are a record of what happened, not advice, and they are not a promise of what will happen next.

The luck range. Results this small swing a lot by chance. Each line shows a range of two standard errors around its return: a bettor with no edge at all would usually land somewhere inside a range that wide. The swing grows with the square root of the number of bets, so it shrinks as a share of the money: over 501 bets, a no-edge bettor typically ends within about ±$2,100. That is why a few hundred games cannot settle whether a return is skill, and a few thousand can start to.

Is this beating the closing line?

No, and Receipts does not claim to. Beating the closing line, or closing line value, means betting at a better price than the final one before the game: taking a team at −120 that closes at −140. It is about when you bet. Receipts prices every bet at the closing price itself, so it never beats the close by construction. What it shows is simpler: whose pick won more often, and what $100 a pick at the closing price would have returned.

How close the percentages are

Behind the scenes we also measure how close each forecast's percentage was to what happened, because it tells us whether a change to the model helps after far fewer games than win rates do. It is not on the dashboard, because it is easy to misread as a win rate.

Picking the winner only asks which side of 50% a forecast was on. This asks how good the number itself was. Say we give the Dodgers 68% and the line gives them 58%. If the Dodgers win, our forecast missed by 0.32 and the line's by 0.42: ours was closer. If they lose, ours missed by 0.68 and the line's by 0.58: the line's was closer. Squaring each miss and averaging over every game gives the score (the Brier score); lower is better, and saying 50-50 every time scores 0.25.

That is how two forecasts can pick the same winners and still score differently: when both pick the Dodgers and they lose, the forecast that said 72% misses by more than the one that said 64%.

Scoring Shift

Receipts: Scoring Shift checks a different model: Scoring Shift, which predicts how a player's production will change over the rest of the season. It checks the model, not players, so it never names one.

The calls. On each date the model was tested on (in MLB, 15 June, 15 July and 15 August), every player with enough playing time so far (150 plate appearances for a hitter, 120 batters faced for a pitcher) is sorted by the change the model predicts. The fifth predicted to rise most, the fifth predicted to fall most, and the middle three-fifths are followed as three groups.

What is measured. A hitter's production is the runs his walks, singles, doubles, triples and home runs are worth, per 100 plate appearances (0.7 for a walk, 0.9 for a single, 1.25 for a double, 1.6 for a triple, 2.0 for a home run). A pitcher's is the same for what he allowed, per 100 batters faced, with the sign turned round so that up is good for him. Each group's change is its rest-of-season production minus its season so far.

The simple guess. Every player drifts halfway back from his season so far to his previous season. Most hot and cold stretches partly reverse, so a model has to beat this guess, not "nothing changes".

The charts. Solid lines are what each group did against its season so far, week by week around the call date; dotted lines are the prediction; grey dashed lines are the simple guess. Under each chart, each group's predicted and actual change are written out with a luck range: two standard errors of the actual change over that many players.

Backfilled calls. The model was frozen before the 2026 season, so its 2026 calls are rebuilt from the games played before each date and labelled backfilled. Calls saved on the day are reported separately once they exist.

Sustainability

Receipts: Sustainability checks the Sustainability model, which says whether a player's results on balls in play are running ahead of or behind his own level, and that the gap is luck that gives itself back. Like Scoring Shift, it checks the model in groups and never names a player.

The calls. On the same dates as Scoring Shift (in MLB, 15 June, 15 July and 15 August), the model sorts every player with enough balls in play into three groups: running lucky, in line, and running unlucky.

What is measured. Batting average on balls in play (BABIP): the share of balls put in play, not home runs or strikeouts, that fall for hits. For a pitcher, the BABIP he allows. Each player's is compared with his own level over the previous two seasons, in points (thousandths), turned round for pitchers so that positive always means lucky.

The prediction and the comparison. The model predicts each group goes back to its own level: zero. The comparison is what would have happened if the luck had carried on at the same rate.

The charts. Solid lines are each group's luck against its own level, the last four weeks before the call date and the average since it afterwards; dotted lines are the prediction; grey dashed lines are the luck carrying on. Under each chart, each group's luck after the call is written out with a luck range.

Backfilled games

Forecasts for games before 8 October 2026 were not saved half an hour before the start; saving began that day. Those games are scored from the rating stored on the morning of each game, which is what an answer that day would have used. Receipts counts them separately as backfilled.

Model versions

We release new versions of the Power Rating as they pass their tests. Each version is scored on its own, over the games it priced, and the charts on Receipts mark the day each new version took over, so you can see whether a change made the picks better.