Metrics Validation

loading current figures…

Claims are graded by the evidence supporting them. Read the Methodology and source contracts alongside each diagnostic. A disagreement needs a source-and-version check, not automatic preference for one page. Static face-validity examples are historical illustrations, not refreshed rankings; unless dated explicitly, their snapshot date was not recorded.

Tier 1 · Reconciled against ground truthChecked against numbers this platform did not produce. If these fail, something is genuinely broken.
Tier 2 · Calibrated against outcomesA modelled quantity checked against what actually happened.
Tier 3 · Benchmarked against a published referenceCompared to an independent tool's figures. Real evidence, but narrow.
Tier 4 · Defined, not validatedQuantities with no ground truth to check against. Stated precisely, never claimed as verified.

01 xG · calibrated against outcomes Tier 2

Expected goals is the platform's fitted shot-quality model, separate from the match-forecast model. It is an empirical binned model built on this season's shots, using distance, angle, whether the shot was a header, whether it was a big chance, and whether it came in open play. Nothing external is used.

The aggregate line below remains an in-sample fit diagnostic. It is now accompanied by a temporal holdout: the same feature family is fitted on the first 80% of current-season matches and scored only on the final 20%. Sparse cells fall back to a broader shot-shape rate estimated on the training split. This is a genuine unseen-time test, although it is still one season split rather than repeated cross-validation or external validation.

loading calibration…

A predicted-versus-actual band table used to sit here and has been removed. Each shot's predicted value is its bin's conversion rate, fitted on these same shots, so grouping shots by predicted value and comparing to observed conversion largely measures whether the lookup reproduces the data used to build it. That is implementation consistency, and readers reasonably mistake it for validation. What follows describes the support behind the model instead.

Model supportValue
Training shotsloading…
Lookup binsloading…
Distinct assigned valuesloading…
Assigned xG rangeloading…
Sparse binsloading…
Validation statusloading…
Temporal holdout windowloading…
Holdout accuracyloading…
Holdout goalsloading…
Largest holdout estimateloading…
What this does not prove. One temporal split does not establish transportability across seasons or competitions. The production lookup still contains sparse cells and compresses the best chances; the holdout estimator reduces sparse-cell reliance through hierarchical fallback, but it is not a replacement model. Treat the reported scores as the first predictive baseline, not a completed validation programme.

02 xT · a borrowed grid Tier 4

Expected threat is not a fitted model on this platform, and it is important not to present it as one. The pitch is divided into a 12 by 8 grid, and each cell carries a value representing the probability a possession in that zone ends in a goal. Moving the ball between zones adds the difference.

Those grid values are taken from published work and applied to these competitions unchanged. They were not fitted on MLS, La Liga or any other league here, and they have not been externally validated against a reference series for these leagues. The database holds the grid as a static 96-row table; there is no fitting step anywhere in the pipeline.

A claim that used to appear on this page has been removed. An earlier version stated that both xG and xT fitted an external Opta reference at r = 0.975. That was wrong for xT and is not reproducible for either: no external reference series exists in the database. The figure has been withdrawn rather than restated.

Loading the internal xT directional check…

The relative ordering of pitch zones is directionally plausible and consistent with how the game is generally understood, but it has not been validated for these competitions either. Both the ordering and the absolute values are borrowed. Treat xT as a well-founded heuristic for comparing actions within this platform, not as a calibrated probability.

03 Reference benchmark · the sequence engine

The possession-sequence engine was built from raw event data, so its segmentation rules could have been cut anywhere. To check they were cut in the right places, its output was compared to Julio Costa's published sequence numbers for Benfica. Passes-per-sequence and duration are structural quantities that depend partly on how a league plays, so a close match is suggestive rather than conclusive.

Sequence metricOurs (MLS)Reference (Benfica)Read
Passes per sequence3.553.6in the same region
Seconds per sequence9.99.6in the same region
Players per sequence2.953.2close
Sequences ending in a shot9.1%12%MLS lower

Snapshot comparison, run once. Our side of this table was computed on MLS only, 223 matches, latest match 18 July 2026, 57,762 sequences. The figures are from a historical comparison against a published Benfica reference and are not recomputed on rebuild. They are retained as a record of that check, not as a current measurement. They show the segmentation produces chains in a plausible region. They do not establish that the chains are cut in the right places, and no explanation is offered here for the difference in shot rate: several league and stylistic factors could account for it and none has been tested.

04 Face validity · team metrics

Fixed validation snapshot, not a current leaderboard. Unlike the sequence benchmark in 03, which is pinned to a known dataset and date, this review is undated. Sections 04 to 07 record a human face-validity review of a single frozen dataset: MLS only, no European leagues included. The examples below are the outputs as they stood when a person assessed them, retained as a record of that review. They are not recomputed on rebuild and current leaders will differ. Review date not recorded, which is itself a gap: regeneration should only come with a workflow that logs who checked refreshed outputs and when.

For every team metric, the extremes should be the teams you'd name yourself. They are.

MetricTopValueBottomValue
Field tilt %Vancouver65.2Orlando City38.0
Pressing (PPDA, low = more)Vancouver7.35Houston17.1
Possession %San Diego61.5DC United40.1
DirectnessPhiladelphia7.82San Diego5.13
Shots per gameVancouver17.4Kansas City9.5
Goals conceded per gameNashville0.73Orlando City2.87

Cross-consistency is the stronger signal. Vancouver tops field tilt, pressing, progression and shots, one coherent front-foot identity across four metrics computed independently. San Diego tops possession and is the least direct side, a matched pair. Orlando has the lowest field tilt and the worst defence: a dominated team, consistently. Unrelated metrics agreeing on the same story is harder to fake than any single leaderboard.

05 Face validity · player chain-roles

Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.

Player roles are assigned purely from what a player does inside possessions, with no knowledge of their position. So if the roles are real, position should fall out on its own. It does, cleanly, for all eleven.

RoleWho tops itVerdict
FinisherPreston Judd, Brian White (forwards)✓ forwards
Box threatJudd, White, Lobjanidze (forwards)✓ poachers
CarrierWerner, Allende, Cowell (wingers)✓ dribblers
Individual (take-ons)Minoungou, Jaime (wingers)✓ 1v1 players
Support angle (diagonals)Jeong, Pellegrino, Mighten (wide fwds)✓ channel runners
CreatorKelsy, Sabaly (attackers)✓ chance makers
InitiatorJones, Larsen, Harriel (defenders)✓ build from back
BridgeMaher, Miller, Long (centre-backs)✓ ball-playing CBs
ProgressorKamal Miller (ball-playing CB)✓ deep progressor
Controller · quick releasePiette (holding mid, 2.1s)✓ one-touch

The system also distinguishes types within a position. Lionel Messi returns as a forward who progresses (15%) with 559 involvements. That is a deep, ball-dominant forward, not a penalty-box poacher like Preston Judd (166 involvements, 33% box threat). Same position, correctly different roles.

06 Face validity · the impact composite

Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.

The DNA layer condenses every pillar into two numbers per player: impact, weighted towards the pillars that matter for his position, and completeness, a penalised average that punishes a glaring weakness. The test is whether the top of that list reads like a list of the league's best players.

PlayerClubPoolImpactCompletenessTop pillar
Frankie WestfieldPhiladelphiaFB10099Creation
Lionel MessiInter MiamiAM10096Creation
Jack McGlynnHoustonCM10089Progression
Jackson RagenSeattleCB9999Progression
Andy NájarNashvilleFB9998Progression
Sebastian BerhalterVancouverCM98100Creation

Two things are worth noting. First, the list is not dominated by forwards: full-backs, a centre-back and midfielders lead it, because impact is weighted by what matters for each position rather than by attacking output. That is the intended behaviour, and it is the opposite of what a naive composite produces.

Second, impact and completeness genuinely diverge. Milan Iloski also scores 100 for impact but only 71 for completeness, elite at what his position demands, with a clear hole elsewhere. Berhalter is the inverse, 98 impact but a perfect 100 completeness. A single rating would collapse those two very different players into the same number, which is exactly why there are two.

07 Similarity · does "similar" mean similar

Part of the fixed validation snapshot described in 04. MLS only, not recomputed on rebuild.

The similarity engines find nearest neighbours by cosine on z-scored profiles. The test is whether the matches are ones a scout would nod at.

Player. Lionel Messi's nearest chain-role match in MLS is Son Heung-Min (79%), followed by the league's creative designated players: Fernández, Hartel, Ojeda, Miranchuk, Mukhtar. Every comp is a playmaking, ball-involved forward, found without any position input. Messi tops out at 79% rather than 95%+ because he is a genuine outlier; homogeneous positions like centre-back match far tighter.

Team. Inter Miami's closest stylistic peers are Real Salt Lake and LAFC, the league's possession-and-penetrate cluster.

08 What we found wrong, and fixed

Validation only means something if it can fail. It did, repeatedly, and each catch is documented here because a process that never finds anything isn't testing anything.

09 Known limitations

data · whoscored · xg calibrated against outcomes · xt grid borrowed, not fitted here · sequences benchmarked vs published reference · as of …
generated from the live database · figures reproducible from the platform