Edgehalla

Shot models

How our expected goals (xG) model works

Expected goals (xG) turn every shot into a probability. A tip from the crease might be worth 0.25 goals and a point shot through no traffic 0.02. Add them up and you get a better picture of who created the better chances than goals or shot counts alone. This page explains how our model is built and tested, and what it can't tell you.

Updated · Method version xg-v2

How it's calculated

Each unblocked attempt (a goal, a shot on goal or a missed shot) gets a probability between 0 and 1. A team's or player's xG is the sum over their shots.

The model

xG = P(goal | features) = 1 / (1 + e^−(base + Σ tree outputs))

features
the shot's geometry, type and context (listed below), computed once at ingest
tree outputs
the gradient-boosted trees' contributions for this shot (LightGBM: 319 small trees for even strength, 216 for the power play, 155 shorthanded)
Features
GroupWhat we use
GeometryDistance and angle to the net, x/y position, behind-the-net flag. Distance has a monotone constraint: moving farther away can never raise xG.
Shot typeWrist, snap, slap, backhand, tip-in, deflection, wrap-around, other.
Previous eventIts type and whether it was ours or theirs; seconds since it, distance from it and puck speed between the two locations.
Rebounds and rushRebound flag (same-team attempt within 3 s), how far the shot angle changed, transition flag and seconds since the possession started.
Cross-slot pass proxyPrevious event in the offensive zone within 2 s, on the other side of the slot.
Shooter handOff-wing flag from the NHL's handedness data. Built, but not yet used by the model because handedness isn't attached at ingest.
Game stateStrength state, score difference (clamped to ±3), home or away, period and time.
Faceoff contextSeconds since an offensive-zone faceoff. Shots straight off a draw score far less often than their location suggests.
Recording changesA flag for short misses and the missed-shot vocabulary the NHL introduced in 2023-24.
  • Four models: even strength, power play and shorthanded are gradient-boosted. Empty net is a small logistic model, because there are only about 1,000 empty-net attempts a season. Power-play shots behave differently (the extra attacker changes the geometry), so they get their own model.
  • Blocked shots get no xG. The feed records where the block happened, not where the shot came from. NST, MoneyPuck and Evolving-Hockey follow the same convention.
  • Penalty shots use a constant 0.30: 110 goals on 373 attempts over nine seasons, shrunk toward the long-run rate of about 0.32. The shootout is excluded.
  • Shots with missing coordinates get a constant 0.03 and are flagged as imputed. Only 7 of 1.01 million attempts lacked coordinates.
Example data

Five shots, five probabilities

Example data: invented for illustration

Half rink with five example shots; dot area grows with xG, from 0.015 at the point to 0.28 at the crease.

0.280.140.0500.0400.015
Dot size is proportional to xG. Values are illustrative, not model output.

Worked example

Why we added a faceoff feature

Real research numbers

Our first test model (a simple logistic model with a holdout AUC of 0.739) was checked against actual goals on 200 games from 2025-26.

  1. Shots taken within 5 s of an offensive-zone faceoff went in 2.8% of the time.
  2. The test model, which only saw location and shot type, predicted 4.9% for the same shots.
  3. Over 1,000 such shots that is 49 expected goals against 28 real ones, an overstatement of 21 goals.
  4. Set plays off the draw produce shots from good locations but with a set defence in front of the goalie. Location alone can't see that.

The production model includes "seconds since an offensive-zone faceoff" so these shots are no longer overrated.

Training, validation and versioning

We train in Python (LightGBM) on every unblocked attempt from 2017-18 to 2025-26: 1,014,425 attempts over 11,739 games. Hyperparameters are tuned with five-fold cross-validation grouped by game, so shots from one game never sit on both sides of a split. The second half of 2025-26, every game from January 25 on including the playoffs, is held out. It is never used for tuning, and all the numbers below come from it. The shipped model is then refit on every season.

Weighting schemes on the next, unseen season (mean of the two tests)
Training weightsLog lossAUC (2024-25 / 2025-26)
All nine seasons equal0.22080.773 / 0.770
Last three seasons only0.21630.783 / 0.788
Half-life 2 seasons0.21680.784 / 0.784
Half-life 1 season0.21500.789 / 0.789
Half-life 0.5 seasons (shipped)0.21430.794 / 0.791
Adding a recalibration fitted on the newest training season did not improve the forecast (0.2143 either way), so the live model has none.
xg-v2 on the held-out half-season (49,404 attempts, 3,567 goals)
ModelAUCLog lossBrierxG ÷ goals
Even strength0.7910.1990.0541.009
Power play0.6920.2920.0821.017
Shorthanded0.8470.2040.0600.967
Empty net (logistic)0.7570.5810.2001.083
All attempts0.7900.2180.0601.015
Spline logistic baseline, all attempts0.7810.2210.0611.021
Previous model xg-v1 (2024-25 and 2025-26, equal weights)0.7900.2180.0601.027
Expected calibration error over 20 bins is 0.0047.

Older seasons need one more step. A model centred on recent hockey overrates 2017-18 and 2018-19 shots by 13-16% when it rescores them. So for each completed season and strength state we store one intercept that makes total xG match total goals for that season. For example, even strength in 2017-18 is shifted down by 0.25 on the log-odds scale. The current season gets no adjustment. The intercepts fix the totals but not the ranking of shots: before 2023-24 the model separates goals from non-goals less well (AUC about 0.70-0.74, against 0.80 or more since), because some of its signals only exist in the newer data.

Targets against results
CheckTargetxg-v2
Even-strength AUC / log lossAUC ≥ 0.77, log loss ≤ 0.195AUC 0.791 met. Log loss 0.199 missed: the held-out half-season scored more often than the training seasons.
Power-play AUC≥ 0.700.692, just short
Shorthanded AUC≥ 0.780.847
CalibrationΣ xG within ±3% of goals by season; ECE ≤ 0.0051.000 for every season from 2017-18 to 2024-25 (season intercepts), 1.011 for 2025-26 and 1.015 on the holdout; ECE 0.0047
Versus the spline logistic baselineBeat it by at least 0.002 log lossBetter by 0.0035 at even strength, 0.0026 on the power play and 0.0034 shorthanded

What matters most: distance (26% of the even-strength model's gain), lateral position (15%), angle, puck speed from the previous event, rebound timing, time since the previous event and tip-ins.

Features are computed once, in TypeScript, from the archived play-by-play. The training set is exported by that same code, so the trainer never reimplements a feature. A parity test requires the site's scorer to reproduce the trainer's probabilities: on 2,000 sampled shots the largest difference was 3 × 10⁻¹⁶. That is how we avoid the common bug where training and serving compute features differently.

Every scored game stores the model version that scored it (currently xg-v2). We retrain once a year in August. During the season the weights stay frozen; if goals and xG drift apart by more than 4% after 300 games, we recalibrate the intercept only and bump the patch number. A new version rescores every season, and the previous one is kept for rollback.

How it differs from NST and MoneyPuck

Design choices compared
EdgehallaMoneyPuckNatural Stat Trick
Model typeGradient boosting for EV, PP and SH; logistic for empty netGradient boosting, one model with strength featuresPublishes its own xG model; we do not reuse its numbers
Training data2017-18 to 2025-26 (about 1 million attempts), weighted toward recent seasons, retrained each August2007-08 to 2014-15 (about 800,000 shots)—
Shift-based contextFaceoff timing, possession origin, transition flag; shift age plannedNot in its published feature list—
FlurriesRaw xG for players; sequence xG for team and game viewsFlurry-adjusted xG offered—
Rush and rebound flagsTransition (10 s, possession-based) and rebound (3 s)Model inputs from the previous eventRush within 4 s of an NZ/DZ event; rebound within 3 s
We compare our shot-level numbers with MoneyPuck's privately as a sanity check (we expect a per-shot correlation of 0.80–0.90). Their numbers are never displayed or stored on the site.

Reliability

xG is more stable than goals over the same sample, but individual-game numbers are noisy. Treat a single game as a story, not a verdict.

MeasureSplit-half rSampleVerdict
On-ice 5v5 xGF%Grows steadily with more games.0.29About 7 xG per half-sampleModerate
Shot diet (ixG per attempt)0.68Split-half, 200 gamesStable
Finishing (G − ixG)We regress it heavily toward zero in projections.—Needs hundreds of shotsNoise

Split-half r: we split each sample in two (for example odd and even games) and correlate the two halves across players or teams. Near 0 means the stat is mostly noise; near 1 means it is a stable trait.

Limitations

  • The feed gives us event locations, not puck or player tracking. Screens, passing lanes and goalie position before the shot are invisible, so two shots from the same spot can deserve very different xG.
  • Scorers log locations by hand, and some rinks record shots systematically closer or farther than others. We keep a home/away feature until a rink adjustment is in place.
  • Missed-shot recording changed in 2023-24 (new "short" and "failed bank attempt" categories), which slightly inflates recent Fenwick totals.
  • xG measures chances created, not skill at finishing them. Elite shooters beat their xG; the model deliberately does not include shooter identity.

FAQ

What is expected goals (xG) in hockey?

xG is the probability that an unblocked shot attempt becomes a goal, estimated from where it was taken, the type of shot, and what happened just before it. Summing xG over a game or season shows how many goals a team or player would score from those chances on average.

Why don't blocked shots have xG?

The NHL feed records a blocked shot at the spot where it was blocked, not where it was taken, so we cannot value its location. This is the same convention NST, MoneyPuck and Evolving-Hockey use. Blocked attempts still count in Corsi.

Why is your xG different from MoneyPuck or Natural Stat Trick?

Each site trains its own model on different seasons with different features. Ours uses four strength-state models trained on 2017-18 to 2025-26, weighted toward recent seasons, and adds play context, such as time since an offensive-zone faceoff. Team-season totals should agree closely; individual shots can differ.

How often is the model updated?

Once a year, in August. During the season the model is frozen; if league totals drift more than 4% from goals after 300 games we recalibrate the intercept and bump the patch version. Every shot stores the version that scored it.

Does xG include shooter talent?

No. The model values the chance, not the shooter. Finishing skill appears as goals minus xG, which we treat as mostly noise over small samples.