Claims are graded by the evidence supporting them. Read the Methodology and source contracts alongside each diagnostic. A disagreement needs a source-and-version check, not automatic preference for one page. Static face-validity examples are historical illustrations, not refreshed rankings; unless dated explicitly, their snapshot date was not recorded.
01 xG · calibrated against outcomes Tier 2
Expected goals is the platform's fitted shot-quality model, separate from the match-forecast model. It is an empirical binned model built on this season's shots, using distance, angle, whether the shot was a header, whether it was a big chance, and whether it came in open play. Nothing external is used.
The aggregate line below remains an in-sample fit diagnostic. It is now accompanied by a temporal holdout: the same feature family is fitted on the first 80% of current-season matches and scored only on the final 20%. Sparse cells fall back to a broader shot-shape rate estimated on the training split. This is a genuine unseen-time test, although it is still one season split rather than repeated cross-validation or external validation.
loading calibration…
A predicted-versus-actual band table used to sit here and has been removed. Each shot's predicted value is its bin's conversion rate, fitted on these same shots, so grouping shots by predicted value and comparing to observed conversion largely measures whether the lookup reproduces the data used to build it. That is implementation consistency, and readers reasonably mistake it for validation. What follows describes the support behind the model instead.
| Model support | Value |
|---|---|
| Training shots | loading… |
| Lookup bins | loading… |
| Distinct assigned values | loading… |
| Assigned xG range | loading… |
| Sparse bins | loading… |
| Validation status | loading… |
| Temporal holdout window | loading… |
| Holdout accuracy | loading… |
| Holdout goals | loading… |
| Largest holdout estimate | loading… |
02 xT · a borrowed grid Tier 4
Expected threat is not a fitted model on this platform, and it is important not to present it as one. The pitch is divided into a 12 by 8 grid, and each cell carries a value representing the probability a possession in that zone ends in a goal. Moving the ball between zones adds the difference.
Those grid values are taken from published work and applied to these competitions unchanged. They were not fitted on MLS, La Liga or any other league here, and they have not been externally validated against a reference series for these leagues. The database holds the grid as a static 96-row table; there is no fitting step anywhere in the pipeline.
Loading the internal xT directional check…
The relative ordering of pitch zones is directionally plausible and consistent with how the game is generally understood, but it has not been validated for these competitions either. Both the ordering and the absolute values are borrowed. Treat xT as a well-founded heuristic for comparing actions within this platform, not as a calibrated probability.
03 Reference benchmark · the sequence engine
The possession-sequence engine was built from raw event data, so its segmentation rules could have been cut anywhere. To check they were cut in the right places, its output was compared to Julio Costa's published sequence numbers for Benfica. Passes-per-sequence and duration are structural quantities that depend partly on how a league plays, so a close match is suggestive rather than conclusive.
| Sequence metric | Ours (MLS) | Reference (Benfica) | Read |
|---|---|---|---|
| Passes per sequence | 3.55 | 3.6 | in the same region |
| Seconds per sequence | 9.9 | 9.6 | in the same region |
| Players per sequence | 2.95 | 3.2 | close |
| Sequences ending in a shot | 9.1% | 12% | MLS lower |
Snapshot comparison, run once. Our side of this table was computed on MLS only, 223 matches, latest match 18 July 2026, 57,762 sequences. The figures are from a historical comparison against a published Benfica reference and are not recomputed on rebuild. They are retained as a record of that check, not as a current measurement. They show the segmentation produces chains in a plausible region. They do not establish that the chains are cut in the right places, and no explanation is offered here for the difference in shot rate: several league and stylistic factors could account for it and none has been tested.
04 Face validity · team metrics
For every team metric, the extremes should be the teams you'd name yourself. They are.
| Metric | Top | Value | Bottom | Value |
|---|---|---|---|---|
| Field tilt % | Vancouver | 65.2 | Orlando City | 38.0 |
| Pressing (PPDA, low = more) | Vancouver | 7.35 | Houston | 17.1 |
| Possession % | San Diego | 61.5 | DC United | 40.1 |
| Directness | Philadelphia | 7.82 | San Diego | 5.13 |
| Shots per game | Vancouver | 17.4 | Kansas City | 9.5 |
| Goals conceded per game | Nashville | 0.73 | Orlando City | 2.87 |
Cross-consistency is the stronger signal. Vancouver tops field tilt, pressing, progression and shots, one coherent front-foot identity across four metrics computed independently. San Diego tops possession and is the least direct side, a matched pair. Orlando has the lowest field tilt and the worst defence: a dominated team, consistently. Unrelated metrics agreeing on the same story is harder to fake than any single leaderboard.
05 Face validity · player chain-roles
Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.
Player roles are assigned purely from what a player does inside possessions, with no knowledge of their position. So if the roles are real, position should fall out on its own. It does, cleanly, for all eleven.
| Role | Who tops it | Verdict |
|---|---|---|
| Finisher | Preston Judd, Brian White (forwards) | ✓ forwards |
| Box threat | Judd, White, Lobjanidze (forwards) | ✓ poachers |
| Carrier | Werner, Allende, Cowell (wingers) | ✓ dribblers |
| Individual (take-ons) | Minoungou, Jaime (wingers) | ✓ 1v1 players |
| Support angle (diagonals) | Jeong, Pellegrino, Mighten (wide fwds) | ✓ channel runners |
| Creator | Kelsy, Sabaly (attackers) | ✓ chance makers |
| Initiator | Jones, Larsen, Harriel (defenders) | ✓ build from back |
| Bridge | Maher, Miller, Long (centre-backs) | ✓ ball-playing CBs |
| Progressor | Kamal Miller (ball-playing CB) | ✓ deep progressor |
| Controller · quick release | Piette (holding mid, 2.1s) | ✓ one-touch |
The system also distinguishes types within a position. Lionel Messi returns as a forward who progresses (15%) with 559 involvements. That is a deep, ball-dominant forward, not a penalty-box poacher like Preston Judd (166 involvements, 33% box threat). Same position, correctly different roles.
06 Face validity · the impact composite
Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.
The DNA layer condenses every pillar into two numbers per player: impact, weighted towards the pillars that matter for his position, and completeness, a penalised average that punishes a glaring weakness. The test is whether the top of that list reads like a list of the league's best players.
| Player | Club | Pool | Impact | Completeness | Top pillar |
|---|---|---|---|---|---|
| Frankie Westfield | Philadelphia | FB | 100 | 99 | Creation |
| Lionel Messi | Inter Miami | AM | 100 | 96 | Creation |
| Jack McGlynn | Houston | CM | 100 | 89 | Progression |
| Jackson Ragen | Seattle | CB | 99 | 99 | Progression |
| Andy Nájar | Nashville | FB | 99 | 98 | Progression |
| Sebastian Berhalter | Vancouver | CM | 98 | 100 | Creation |
Two things are worth noting. First, the list is not dominated by forwards: full-backs, a centre-back and midfielders lead it, because impact is weighted by what matters for each position rather than by attacking output. That is the intended behaviour, and it is the opposite of what a naive composite produces.
Second, impact and completeness genuinely diverge. Milan Iloski also scores 100 for impact but only 71 for completeness, elite at what his position demands, with a clear hole elsewhere. Berhalter is the inverse, 98 impact but a perfect 100 completeness. A single rating would collapse those two very different players into the same number, which is exactly why there are two.
07 Similarity · does "similar" mean similar
Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.
The similarity engines find nearest neighbours by cosine on z-scored profiles. The test is whether the matches are ones a scout would nod at.
Player. Lionel Messi's nearest chain-role match in MLS is Son Heung-Min (79%), followed by the league's creative designated players: Fernández, Hartel, Ojeda, Miranchuk, Mukhtar. Every comp is a playmaking, ball-involved forward, found without any position input. Messi tops out at 79% rather than 95%+ because he is a genuine outlier; homogeneous positions like centre-back match far tighter.
Team. Inter Miami's closest stylistic peers are Real Salt Lake and LAFC, the league's possession-and-penetrate cluster.
08 What we found wrong, and fixed
Validation only means something if it can fail. It did, repeatedly, and each catch is documented here because a process that never finds anything isn't testing anything.
- xT grid mis-indexed. The lookup assumed a 16×6 grid; the real grid is 12×8. Every advanced-third action was silently reading an empty cell. Caught when box-edge lookups returned null, fixed, re-validated against a sane season total.
- Sequences over-segmented (v1). First pass gave 1.77 passes per sequence, half the reference. Cause: aerial duels and defensive touches were being allowed to define possession, chopping real chains into fragments. Restricting possession to controlled touches brought it to 3.55, in the same region.
- Central-progression band too wide. Defining "central" as the full box width credited 60% of progression as central. Narrowed to a true middle-third corridor, it dropped to a meaningful 30%.
- Bridge role too loose. "Any mid-chain forward pass" put a goalkeeper on top at 60%. Redefined as a pass that actually advances the ball a third, it now surfaces ball-playing centre-backs at ~20%, matching the reference scale.
- Goalkeepers polluting outfield roles. Keepers topped Progressor (any pass is forward from the goal line) and tempo (they're allowed to hold the ball). Excluded from chain-roles entirely, consistent with the platform's rule that keepers are scored on goalkeeping, not outfield play.
09 Known limitations
- Pool size. The historical MLS examples used 30 teams and roughly 15–30 players per position. Current comparisons depend on the league, role, source and eligibility filters displayed with the chart. Extreme values in a small pool are not automatically reliable signals.
- Tempo resolution. Event timestamps are whole-second, so a single player's hold time is coarse. Averaged over 1,000+ involvements it separates holders from quick-releasers reliably, but it is not sub-second precise.
- Proxy definitions. A few metrics (hold-up, support angle, wide triangles) are transparent proxies for concepts the raw data doesn't label directly. Their exact rules are published in the glossary so they can be judged on their own terms rather than assumed to match anyone else's.
- Face validity is not ground truth. xG has in-sample diagnostics and the temporal holdout reported above when available; neither is external validation. The xT grid is not validated on these competitions: it uses a borrowed grid that was never fitted or externally checked on these competitions, so it sits alongside everything else that rests on face validity rather than being exempt from that limitation. These checks confirm metrics behave as designed and agree with each other and with football sense. They do not claim to be externally verified truth.