Decision metrics

Last updated 9 September 2026

DataBaller's portfolio of performance models, and the decision metrics they produce, is the intellectual core of the product. Five models — three about a player, two about a team — computed against league data every night, each stored with a record of exactly how its number was made.

Models here improve often, and each league's model versions on its own schedule — the NBA's Performance Shift and the NFL's are siblings trained for their own leagues, not one model wearing two uniforms. This page describes the portfolio as of the date above, and every answer in the product names the exact model version behind its numbers, so a figure can always be traced to the method that produced it.

The portfolio at a glance

Player models — two describe, one predicts. Performance Shift and Sustainability Score read a player's game as it is: the role and the production. Scoring Shift is the forward view — it prices what those kinds of evidence say about scoring next.

The question The model What it does NBA MLB NFL
Has their game changed? Performance Shift Describes v4 —¹ v2
Is their production real? Sustainability Score Describes v4 v2 v2
Where does their scoring go next? Scoring Shift Predicts v3 v1 v1²

Team models — both descriptive. Power Rating describes today's strength — in the NBA, of the lineup expected to play; in MLB and the NFL, of the season so far. Ask about a matchup and the forecast is computed on the spot from the two teams' stored ratings.

The question The model What it does NBA MLB NFL
How strong is this team? Power Rating Describes v2 v2 v2
Is their record real? Team Sustainability Describes v1 v1 v1

¹ Built and tested — and the test failed, so it does not ship. See Tested before they ship: an absence is the discipline working, not a gap waiting for copy. ² Covers rushers and receivers, in scrimmage yards per game. The quarterback version was built and tested alongside it and its test failed — so a quarterback forecast gets an honest "we don't have one" instead of a number.

DataBaller's Performance Shift: Has their game changed?

Performance Shift watches a player's job, not their points. Points swing night to night for reasons that do not last, but minutes, touches, and how often the offense runs through a player are decisions coaches make. When those move, something actually happened.

The model compares a player's recent stretch of games to the way they were used over the two years before it, judged against how much players normally differ from each other. A high number means the role genuinely moved: more minutes, more of the offense, a different shot diet. In football it watches opportunity the same way — carries and targets rather than yards — because opportunity is the part of a role that holds still long enough to measure.

Read it as a flag, never a forecast. Testing the model against history settled what the flag really means, and the second finding is the surprising one: flagged players' seasons really do stray further from their averages afterward, so a season line is a poor guide to a player mid-change — and a role change usually walks itself partly back. A player whose minutes cratered tends to get some back. The flag says look closer here. It does not say draw the line out further.

DataBaller's Sustainability Score: Is their production real?

When a player suddenly scores more, there are two stories. They are doing more: more minutes, more shots. Or the same shots are going in more often. The first tends to stick. The second is mostly borrowed, and borrowed points get paid back.

Sustainability Score tells the two apart, and the fact that makes it measurable is blunt: shooting accuracy never settles down, not even across twenty games. So when a player's recent accuracy sits far above what they themselves have shot across the years, most of that gap is luck wearing a jersey. The score runs from 1 — the scoring is exactly what this shot volume at the player's own normal accuracy produces — down to 0, far out of line with who they have been.

The number carries a sentence you can use: a player averaging 16 whose usual accuracy on this shot diet supports about 12.5 should expect the difference to melt. Tested against five seasons of history, players flagged as running hot gave most of the extra back, players flagged as running cold got theirs back, and players in line with themselves stayed put.

For a while the score carried a refinement: a hot stretch that arrived together with a real role change was treated as more believable, because a changed game ought to explain part of a jump. It was withdrawn. Re-tested, it reversed on exactly the case it was built for, and an adjustment that only ever moves a score upward is the kind worth being strict about. So a hot streak now scores the same whether or not the role moved — and "does a changed role make a hot streak more believable" has an honest answer: we tested it twice and could not show it.

Baseball has its own version built on balls in play. Once a ball is struck into the field, whether it lands for a hit is barely about the hitter over any stretch short enough to care about, so a month where everything drops in is borrowed in exactly the same way.

Football's version, new for the 2026 season, reads yards per touch — a player's last 40 targets or last 100 carries against their own two-year level — because per-touch yardage never settles down within a season either. Touchdowns get the treatment the test demanded, and it is the honest surprise: we expected touchdown rates to be borrowed luck, tested it, and the test said otherwise — a player's touchdown rate relative to their own history is partly skill. So the score never promises touchdown regression; it says when the yards are borrowed, and it reports a touchdown surge as context instead of discounting it.

DataBaller's Scoring Shift: Where does their scoring go next?

Scoring Shift is the portfolio's one predictive model, and it now covers all three leagues: a signed forecast of a player's forward production, up or down — points per game in the NBA, production per hundred trips to the plate in MLB, scrimmage yards per game in the NFL. Each league's model is its own build, and they did not all learn the same lesson from the same testing discipline.

In basketball, the evidence blends four pieces, each one tested on its own before it earned a place:

  • Recent scoring above a player's own last six months tends to continue.
  • Except the part that came purely from playing more minutes. That part fades, because minutes come back to earth.
  • Accuracy far from the player's established level gets pulled home. Well above it, expect give-back. Well below it, expect recovery.
  • A stretch still climbing inside itself keeps climbing. A player whose last five games are bigger than the fifteen before them is not finished.

In baseball, the forecast leans on the longest memory in the portfolio: a hitter or pitcher producing above the level their own play established more than a year ago tends to come back to it, and one producing below it tends to recover. The batter version also reads luck on balls in play against the player's own history. For a pitcher the number is oriented the way a fan thinks: positive means better run prevention ahead. And the forecast deliberately knows nothing about playing time — it speaks to what a player does per plate appearance, not how many they get.

In football, the same testing found the same headline even more sharply: recent production above a player's own established pace tends to give itself back. A running back or receiver surging past their trailing-year level is more likely to cool than to keep climbing, and a slumping one with an established level behind them is more likely to recover — so the forecast leans against the streak, not with it. Touchdowns are deliberately nowhere in it: TD rates are among the least stable numbers in football, and a forecast built on them would be modeling dice. And a promise kept from the table above — the quarterback version's test failed, so it does not exist, and the product says so rather than improvising.

The forecast is a lean, never a promise. Ranked by this number and checked against two full seasons each model had never seen, players rose and fell with the ranking far more reliably than chance — in all three leagues. It still explains a minority of what happens next, because most of a player's future is injuries, trades, and coaches, and no box score contains those. The difference it offers is between an informed lean and a coin flip, not between a lean and a crystal ball, and every answer built on it says so.

Reading the player models together

The two descriptive metrics sort a player into one of four stories, and the predictive one prices the story — in points, in production, or in yards.

  • Changed role, real production. The strongest riser read: the job grew, and the scoring under it is what the process supports.
  • Changed role, borrowed production. Believe the minutes, discount the points. The role is real; the scoring rate on top of it is partly on loan.
  • Same role, borrowed production. The pure hot hand. Same job, more makes, and the fade is coming.
  • Same role, real production. They are what they are, and it is real. Stable is an answer, not a shrug.

Scoring Shift is the bottom line under that grid. It weighs how large each force is, nets the ones pushing against each other, and hands back one number. Ask what happens next and the answer leads with the forecast; the other two metrics are the why underneath it.

DataBaller's Power Rating: How strong is this team?

Power Rating is the team question, and its answer is on the scoreboard's own scale: the margin a team would be expected to win or lose by against an average opponent on neutral ground. Lakers +4.2 means four points better than average, a night in, a night out. Ask about a matchup and the answer is built from the two teams' ratings on the spot — the gap between them, a measured home advantage, and the win probability that follows. No matchup is ever precomputed; the rating is stored, the forecast is arithmetic.

What goes into it depends on the league, because the leagues taught us different things. In baseball and football the rating is a blend of two things every fan already watches — the record and the scoring margins — because that blend, and nothing fancier, is what survived testing, and each sport taught the margin half its own lesson. In baseball, run differential flatly does not work: it told us nothing the standings didn't, and the signal only appeared when we rebuilt expected runs from the ingredients — the hits, walks and homers on both sides — which strips out the luck of when hits happen to bunch together. In football, any points past a two-score lead turned out to carry no information at all, so blowout padding is capped out. In basketball the rating has moved past that blend and is built from the players. Every player carries his own measured number — how much his team's margin improves for every 48 minutes he is on the floor, judged against an average player, from two seasons of box scores — and the team's rating is the sum of those numbers over the players expected to play, each weighted by the minutes he is expected to get. In season that is the lineup that played last time out, and anyone the injury report marks Out is dropped first. So an NBA rating answers for a named lineup, the answer says which lineup it assumed and who is missing, and the record and the margins appear beside it as context about the team rather than ingredients of the number.

Just as telling is what was tested and left out, and here one distinction matters. We checked — twice, from both directions — whether knowing how much of a roster's playing time returns makes a team more predictable across an offseason. It does not, in any sport, so DataBaller still won't tell you a team is safe because "they kept the core together" — that sentence sounds like analysis, and we measured it: it isn't. Summing the roster's own measured impacts is a different measurement of a different thing, and it passed its own test: "the player they lost was worth four points a game to that sum" is a fact about an NBA rating, while "they return 80% of their minutes" explains nothing about any rating. One guard rides with those player numbers. They exist to be summed, and on their own they are compressed and flatter role players on strong teams — so a player's own impact number is never a ranking of how good he is, and DataBaller won't present one as if it were.

Two honest limits ride along, and the first splits by league. Between seasons and in the first weeks of a new one, the baseball and football ratings are last season's team, discounted and projected forward — a roster remade over the summer will be mis-rated until its games say otherwise, and answers say "carried from last season" when that's what the number is. The basketball rating is the opposite: it sums the current roster at last season's minutes, so it describes the team the front office built over the summer — a club that lost its best player loses his number, a club that lost its bench loses almost nothing, and a player marked Out leaves the sum the day the report says so. What it cannot know is how the new pieces fit or how the minutes will be shared, and the confidence beside it says as much. The second limit holds everywhere: the edge over just reading the standings is real but small — most single games are closer to a coin flip than anyone likes to admit, and when the model says 58%, the honest answer is 58%, not a lock of the night.

DataBaller's Team Sustainability: Is their record real?

Team Sustainability holds a team's win-loss record up against its own scoring margins. A team can be 15-5 while outscoring opponents like a 12-8 team — the standings say one thing, the play underneath says another, and the gap between them is measured in the only unit that matters: wins. "Their record is three wins ahead of what their scoring supports" is the whole finding.

The reason to care is what such gaps do next, and it is the best-replicated result among our team models: teams far ahead of their margins fall back, and teams far behind recover — in every season of every sport we tested. Football is where it bites hardest: a 4-1 team with the scoring margins of a 2-3 team is one of the most reliable fade signals in the data. But it is a lean about a tendency, not a schedule — the gap closes on average, and the wins already banked stay banked. A team three wins ahead of its margins is expected to win at its play's pace from here, not to hand three wins back.

This metric also carries the clearest example yet of our models shipping only what survives testing. The design called for a second layer — splitting the margins themselves into luck and skill, the way the analytics conventional wisdom does. We tested that wisdom, and it lost: a baseball team's record in close games and its batting luck both turned out to persist, meaning they are partly skill, and discounting them — the standard sabermetric move — would have been wrong in our data. The shooting-luck adjustment for basketball made predictions worse in every season tested. Both are out. What shipped is a third of the original design, and it is the third that is true.

The confidence beside every number

Every metric row carries a confidence score that grades the number itself, not the player or the team. It says how much data sits under this particular reading, and it names its own weakness: a short recent stretch, a thin history, or a hole in the underlying data. Rankings only use readings the models are sure of, and when an answer hedges, the hedge names which of those reasons applies.

Tested before they ship

Four things are true of every model above, and each one is checkable rather than promised.

  • Computed, never repeated. Every metric is derived first-hand from league data. Nothing here summarizes what analysts and podcasts already concluded.
  • Tested against a standard written down first. Every claim was written down before it was tested, then tested against seasons the models had never seen.
  • The failures are published. Baseball has no role-change model because its test failed. The scoring forecast failed twice on the way to the version that shipped. The team models shipped with the conventional luck adjustments cut, because the conventional wisdom lost to the data. A model that fails its test is withdrawn, not argued with.
  • Every number is reproducible. A stored metric remembers the exact version of the method that produced it, forever — so an answer can always say where a figure came from, even years later.

And the testing never finishes. A model enters the portfolio only with a passed test; it changes only by version, so old numbers keep their meaning; and every claim is scheduled for a fresh test against next season's games, which do not exist yet and cannot be studied in advance. A claim that fails that test is withdrawn too — the football rating, resting on the thinnest margin of the three sports, is first in line for that judgment.

The studies behind the models

The tests that shaped the portfolio are written up one question at a time, with the numbers, the seasons and what each result does and does not license:

These models are decision support. They order the field, size the forces, and state their own limits. Where the data cannot settle a question, DataBaller says so. It won't make the pick.