How xG Models Differ

xG Models

Every data provider trains its own expected goals model on its own event data, so the same shot can carry a slightly different xG value depending on which site you are reading.

Why it matters. Comparing a figure on this site against another site without knowing this leads to a false conclusion that one of them is wrong, when both can be honest outputs of reasonably built models.

Why the same shot gets two different numbers

No provider borrows another provider's model. Opta, StatsBomb, Wyscout, Understat and every other data company trains its own expected goals model on its own team's tagged match events, built independently from the ground up.

That independence is where the disagreement starts, not with sloppy work on either side. Three things vary between providers before the model even runs:

  • Shot coordinates. How precisely the shot's location on the pitch gets logged, down to the metre or fraction of a metre.
  • Defensive context. How rigorously defender positions and goalkeeper positioning are tagged at the moment of the shot.
  • Assist classification. Whether a pass into the box gets tagged as a cut-back, a cross or a through ball, categories that carry different base rates.

Small differences in those inputs produce small differences in the output. A shot that one provider's model scores at 0.28 might land at 0.20 or 0.35 on another. The typical spread for the same real-world shot is around 0.05 to 0.1. That is a modelling difference, not an error on anyone's part.

What survives across every provider is direction. A tap-in from six yards scores high everywhere. A speculative strike from 30 yards scores low everywhere. The disagreement lives in the decimal points, not in which chances the models consider good and which they consider poor.

The model families compared

| Aspect | What it captures | Typical inputs | Public availability | | --- | --- | --- | --- | | Provider-trained models (Opta, StatsBomb, Wyscout, Understat, and others) | Each provider's own model, trained on its own tagged event data | Shot location, angle, body part, shot type, assist type, defensive pressure | Varies by provider: some publish shot-level data or offer an API, others keep the model fully proprietary | | Post-shot models (PSxG, xGOT) | Re-values a shot after it is struck, using where it was placed and how hard it was hit rather than pre-shot position alone | Shot trajectory, placement, power, on-target attempts only | Increasingly offered as a companion figure alongside pre-shot xG, coverage still narrower than pre-shot models | | Simplified public models | A lighter approximation built to be reproducible outside a data company | Typically distance and angle, sometimes body part | Fully public, methodology and code often published openly |

Read the table as a spectrum of how much context goes in, not a ranking of quality. A provider-trained model with defensive pressure data will separate two shots from the same spot that a distance-and-angle model would score identically. That extra resolution is the point of paying for tagged event data in the first place.

Reading a number once you know this

Two sites disagree on a match total. Site A shows 2.1 xG for the home side, Site B shows 1.8. Both can be correct in the sense of being honest output from a reasonably built model. The question worth asking is not whose decimal is right, it is whether both sites agree on which team created the better chances. If both put the same side ahead, the disagreement is cosmetic.

A single shot scores 0.31 on one site and 0.24 on another. Both numbers describe the same category of chance, a half-chance rather than a clear one. A verdict built on "this was a half-chance, not a gilt-edged miss" holds regardless of which of the two decimals you read. The gap only matters when it straddles a real threshold, for example one site rating a shot below 0.1 and another rating it above 0.3.

A player's rate across a season looks different on two sites. The absolute totals rarely match, but the trend and the ranking usually do. A player finishing top three for chance quality on one source is very unlikely to sit mid-table on another. Pick one source for a piece of analysis and stay inside it. Mixing figures from two providers inside the same argument compares two different measuring sticks and calls it one number.

The verdict

Pick a source and stay consistent within it. Cross-site disagreement of a few hundredths per shot is the expected outcome of two independently trained models, not a red flag to chase down.

Treat direction and ranking, which team created the better chances, which player gets into better positions, as the reliable signal. Treat the exact decimal as a single model's estimate, useful on its own terms but not a figure that needs to match another site to be trustworthy.

Common misreadings, corrected

"One provider's xG is the objectively correct one." There is no ground truth model to check against. Every provider is estimating a genuinely unknown probability from imperfectly tagged data. A disagreement of a few hundredths is the expected result of that, not a sign one of them made an error.

"xG Stat's numbers not matching another site means something is broken." xG Stat runs its model on Wyscout event data throughout. A different provider's site runs on its own data and its own model. Matching another site's decimal exactly would be a coincidence, not something either site is obliged to deliver.

"A model with more inputs is always more accurate." More inputs help up to a point, but only if the tagging behind them is consistent across thousands of matches. A model with rich inputs and inconsistent tagging can underperform a simpler model with clean, consistent inputs.

"xG values will eventually be standardised across the industry." Unlikely while providers each collect and tag their own event data independently, because that is the actual source of the difference. Standardising the output number would require standardising the tagging process behind it first, and no shared process currently exists.

Common questions

Why does xG differ between websites?
Each provider tags its own match events and trains its own model on that data, so the raw inputs are never identical. A shot's coordinates, the defensive pressure around it and the assist type can each get logged slightly differently, which produces a value that typically differs by 0.05 to 0.1 from another provider's number for the same shot.
Which xG model is the most accurate?
There is no independently verified ranking of provider accuracy, because no provider publishes its full model against a shared public test set. Judge a model by whether its shot values are internally consistent and whether higher values reliably score more often, not by whether it matches a competitor's decimal exactly.
Why does xG Stat's number not match FBref or Understat?
xG Stat runs its model on Wyscout event data, and FBref and Understat draw on different providers with their own tagging and their own models. A gap of a few hundredths on a single shot is the expected outcome of two independently trained models, not a sign either figure is broken.
Do all xG models use the same inputs?
No. Every major provider-trained model considers a similar family of factors, shot location, angle, body part, assist type and defensive pressure, but the precision and coverage of that tagging varies by provider. A simplified public model might use only distance and angle, which is why its values can diverge further from a provider-trained one.
Is a higher or lower xG model more trustworthy?
Neither direction is inherently more trustworthy. A model that runs slightly high or slightly low across the board is still useful if it is consistent, because the comparisons that matter, team against team, player against their own history, hold up within one source. Consistency within a source matters more than where its average sits.
Will xG values ever be standardised across providers?
Unlikely while each provider collects and tags its own event data independently, since that is the actual source of the disagreement rather than a rounding difference that better maths could fix. Standardisation would require providers to share one tagging methodology, which none currently do.

Related terms

See xG Models across the Premier League

Every player, team and match measured on the chances created rather than the scoreline.

Last updated