Methodology
Definitions describe specific data contracts, not a guarantee that every deployed view shares the same filters. See eligibility by source. Historical examples and diagnostics must not be read as current coverage.
How much to trust what
Not everything here is validated the same way, and pretending otherwise is how you get caught out. Four tiers, strongest first.
Tier 1 — reconciled against ground truthstrongest
Checked against numbers this platform did not produce. If these fail, something is genuinely broken.
| Check | Result |
|---|---|
| Goals parsed from events vs published scoreline | loading… |
| Possession shares summing to 100 per match | loading… |
| Events belonging to a known match | loading… |
The goal reconciliation is the single most useful check in the system. It validates event parsing, own-goal attribution and team-name reconciliation simultaneously, because an error in any one of them would break the scoreline.
Tier 2 — calibrated against outcomesstrong
The shot model is checked against whether goals actually happened.
| Aggregate | Shots | Predicted | Actual | Error |
|---|---|---|---|---|
| All non-penalty | — | — | — | — |
A predicted-versus-actual band table used to appear here and has been removed. Each shot's predicted value is its bin's conversion rate, fitted on these same shots, so a band comparison largely measures whether the lookup reproduces the data used to construct it. Readers mistake that for validation. The support behind the model is reported instead.
| Model support | Value |
|---|---|
| Training shots | loading… |
| Lookup bins | loading… |
| Distinct assigned values | loading… |
| Assigned xG range | loading… |
| Sparse bins | loading… |
| Validation status | loading… |
Say this before someone finds it. The model is fitted on the same shots it is currently assessed against, so its aggregate agreement with observed goals is largely guaranteed by construction and is not predictive validation. It also compresses chance quality into discrete bands, and some lookup bins may hold fewer than twenty shots. It is transparent and directionally useful, but its per-shot estimates, especially at the extremes, are not yet externally validated. A temporal holdout has since been run and is reported below; it does not establish agreement with an external reference series.
Tier 3 — benchmarked against a published referencemoderate
The possession-sequence layer was checked against Julio Costa’s published Benfica figures before it was trusted: 3.55 passes per sequence against his 3.6, and 9.9 seconds against 9.6.
Be honest about the weight of this. It is one team against one external reference. It shows the segmentation is in the right region rather than proving it is correct.
Tier 4 — defined, not validatedstate, do not claim
Some numbers cannot be validated because there is nothing to validate them against. They are definitions. The honest defence is to state the definition precisely and justify the choice, never to imply it was verified.
- Counter-attack rate. Chosen as: possession starts in the defensive third and produces a shot within 15 seconds. Defensible, and arbitrary.
- Chain roles. No external ground truth exists. Supported only by face validity.
- Carries. The feed has no carry event, so they are inferred from consecutive same-player forward touches.
- The xT grid. Borrowed from published work and applied to this league rather than fitted on it.
- Sequence boundaries. Our rules for where a possession begins and ends.
Automated checks
Distinguish two different things on this page. The Tier 1 reconciliations and the Tier 2 calibration are automated invariants: they re-run on every rebuild, in a verification step that runs last and fails loudly by name. The Tier 3 reference benchmark is a one-time external comparison, run once by hand against published figures and retained as a record. It is not recomputed and would not catch a regression.
The automated checks live as
rows in an invariants table, each holding SQL that returns a count of violations.
The verification step raises and names every failing check if any error-level check is non-zero.
Adding a check is an insert rather than a code change, which is the only way a practice like this survives contact with a busy season. The current count is read live.
-- example: the goal reconciliation, as it actually runs
select count(*) from matches m
left join (parsed goals per game) ev on ev.game_id = m.game_id
where m.home_score <> coalesce(ev.h,0)
or m.away_score <> coalesce(ev.a,0)
Other error checks: no team appears outside its league whitelist, every sequence carries the league of its own events, percentiles fall between 0 and 100, no sequence has a null threat value, the sequence layer covers exactly the games with events, and season xG stays within 10% of goals scored.
Pitch coordinates
The feed supplies coordinates on a 0–100 scale in both axes, always from the attacking team’s perspective, left to right. Converting to metres uses a 105 × 68 pitch:
x_metres = x / 100 * 105
y_metres = y / 100 * 68
Thirds are cut at x < 33.3, 33.3–66.7 and ≥ 66.7. The penalty area is x ≥ 83 with y between 21.1 and 78.9. Wide channels are y < 21.1 or y > 78.9.
Ingestion
Event data is scraped from WhoScored using the soccerdata library, one match at
a time, with 60–120 second gaps. Only fixtures not already stored are fetched, so a repeated
run costs nothing.
Write-time guard
Before any match is written, both clubs are checked against a per-league whitelist. A match containing an unrecognised club is refused rather than written. The guard fails open only when a league has no whitelist yet, so a new competition can bootstrap from its first scrape, and closed in every other case.
This exists because it did not, once. A scraper defaulting to a different competition wrote another league’s data into the database, and cleaning it up had to be done by fixture id rather than club name, because schedule names and event-feed names differ for more than half the clubs.
Expected goals
An empirical binned model fitted on this season’s shots. No external xG is used.
Features
distance = sqrt( ((100 - x) * 1.05)^2 + ((50 - y) * 0.68)^2 )
angle = degrees( atan( 7.32 * ((100 - x) * 1.05)
/ ( ((100-x)*1.05)^2 + ((y-50)*0.68)^2 - 3.66^2 ) ) )
7.32m is the goal width and 3.66m the half-width, so the angle is the true horizontal aperture the shooter sees.
Bins
| Dimension | Bands |
|---|---|
| Distance | <6m, <11m, <16m, <22m, <30m, 30m+ |
| Angle | <12°, <25°, 25°+ |
| Flags | header, big chance, open play |
Live shot xG uses a smoothed bin estimate; penalties use 0.76. Player xG per 90 includes penalties unless explicitly labelled non-penalty. Open-play pitch filters and historical saved model features have separate scopes. Do not compare them as identical totals; see Evidence definitions for fitted-version and snapshot caveats.
Why a binned model
It is transparent and checkable: you can point at a bin and count the goals. The cost is granularity. Sparse cells need smoothing and diagnostics. The current fitted range belongs to its snapshot, not a permanent 0.53 ceiling; penalties use their separate 0.76 rule. A different estimator would require its own out-of-time comparison.
Predictive check. A live temporal holdout now fits the same feature family on the first 80% of current-season matches and scores it on the final 20%. Sparse cells fall back to broader training-only shot-shape rates. This is an out-of-time baseline, not external validation or a claim that the production lookup has been replaced.
Expected threat
The pitch is divided into a 12 × 8 grid. Each cell carries a value: the probability a possession in that cell ends in a goal. Moving the ball adds the difference.
xT(action) = grid_value(end_x, end_y) - grid_value(start_x, start_y)
x_bin = floor(x / 100 * 12) clamped 0..11
y_bin = floor(y / 100 * 8) clamped 0..7
Applied to completed open-play passes and to derived carries, reported separately as threat from passing and threat from carrying. A backwards pass produces a negative value and is counted as such.
The honest caveat
The grid values are taken from published work, not fitted on this league. Whether those zone valuations hold for MLS specifically is untested. The relative ordering of pitch zones is directionally plausible and consistent with how the game is generally understood, but it has not been validated for these competitions. Both the ordering and the absolute values are borrowed.
The database publishes a separate internal directional check comparing xT on shot-ending and
other possessions. It is explicitly a software sanity check; the live status retains
externally_validated = false.
Possession sequences
A sequence is an unbroken spell of control by one team. Segmentation runs over events ordered by period, minute, second and event id.
Which events count as control
Pass, TakeOn, BallTouch, MissedShots, SavedShot, ShotOnPost, Goal, KeeperPickup, Claim
A new sequence begins when
- the controlling team changes,
- the period changes,
- a stoppage event occurs (
Foul, Card, OffsideGiven, OffsidePass, CornerAwarded, End, or a goal), - or the event carries a set-piece qualifier (
ThrowIn, CornerTaken, FreekickTaken, GoalKick, KickOff, Penalty).
Set-piece-initiated sequences are flagged and excluded from every open-play figure.
Tags applied to each sequence
| Tag | Rule |
|---|---|
| Building from deep | starts at x < 33.3 |
| Won high | starts at x ≥ 50 |
| Patient build | starts x < 50, at least 5 passes, at least 12 seconds |
| Switch of play | a pass with lateral movement > 40 units crossing the centre line |
| Long ball | mean pass length > 26 or any pass ≥ 40 units |
| Hold-up | a completed pass from x ≥ 66.7 travelling backwards by more than 5 |
| Wide triangles | 3+ wide attacking passes involving 3+ players |
| Finds central / wide | lateral position of the first attacking-third entry |
Known imprecision
An opponent action that is neither a control event nor a stoppage, such as a clearance or an interception followed immediately by the original team regaining, will not always break a sequence. These thresholds have not been tuned against video.
Chain roles
Each player’s involvements are mapped back onto the sequence they belong to. Eleven behaviours are expressed as a percentage of his total involvements. Goalkeepers are excluded, and a player needs at least 120 involvements to be profiled.
| Role | Rule |
|---|---|
| Initiator | first action of a sequence |
| Third-man bridge | mid-chain completed pass crossing a third boundary |
| Progressor | completed pass advancing x by ≥ 10 |
| Carrier | consecutive same-player touches advancing x by ≥ 6 |
| Vertical | completed pass, x gain ≥ 8, lateral drift ≤ 8 |
| Support angle | pass made or received at 35–55 degrees |
| Individual | take-on events |
| Creator | completed pass immediately followed by a shot |
| Box threat | involvement at x ≥ 83, y between 21.1 and 78.9 |
| Finisher | shot events |
| Tempo | mean seconds between receiving and releasing |
Tempo is recorded at whole-second resolution in the feed, so it is coarse for any single action and only meaningful as an average over many.
Chain-position value
early involvement = touch in a shot-ending sequence,
at least 3 actions before the end
Reported per 90. It credits the pass that starts a move rather than the one that finishes it. It is contaminated by teammate quality: a player alongside good forwards appears in more shot-ending sequences regardless of his own contribution.
Percentiles
A percentile here is arithmetic, not a model. There is nothing to validate; the question is whether the metric feeding it is right.
pct = round( 100 * percent_rank()
over (partition by league, pool, metric order by value) )
-- inverted for metrics where lower is better, e.g. dispossessions
pct = 100 - pct
- Within position pool (CB, FB, CM, AM, W, ST), because distributions differ enormously by position.
- Within league, so a 60th-percentile progressor in one competition is not silently equated with another.
- Eligibility depends on the source. The committed
mv_player_percentilesdefinition uses six nineties; the distinctmv_player_pctdefinition uses three. Frontend cohort settings are not proof of the deployed SQL gate. Chain roles and team insights have separate exposure rules. See the source contracts.
So “80th percentile finisher” means precisely: his finishing rate is higher than 80% of eligible players in his position pool in his league. It is a statement about ranking, not about quality in the abstract.
Pooling several leagues is offered separately. Pooling is not translation. It answers where a player sits among a chosen set, not how he would perform elsewhere.
Team metrics
| Metric | Definition |
|---|---|
| Possession | share of open-play touches |
| Field tilt | share of final-third touches in the match |
| PPDA | opponent passes allowed per defensive action in their own 60% of the pitch. Lower means more aggressive pressing. |
| Line height | mean x of defensive actions |
| Directness | (end_x - start_x) / (mean pass length × pass count), clamped to −1..1. Net upfield progress per unit of ball travel. |
| Counter-attack rate | share of possessions starting at x < 33.3 that produce a shot within 15 seconds |
| Route productivity | shot-ending rate for a given route, z-scored against the league for that same route |
Press profile
Every possession has an opponent, so from the defending side it is a pressing test. Opponent possessions are grouped by build-up type (short build, direct, mid, high start), and containment is measured as the share ending below x = 66.7 without a shot.
Raw containment is dominated by mechanics rather than quality, since direct play travels further by definition, so figures are z-scored within build-up type against the league. A first version of this reported a 50-point gap for every team, which was measuring the physics of long balls rather than anything about defending.
Game state
A goal timeline is built per match, with own goals credited to the correct side, and every possession is tagged with the score at the moment it began. Output can then be weighted by whether the game was still live:
margin within 1 goal -> weight 1.00
margin of 2 goals -> weight 0.60
margin of 3+ -> weight 0.35
Both the output and the minutes are weighted, so this measures rate per live minute rather than penalising players at dominant clubs.
Usage and leverage
Squad role uses two axes rather than raw minutes.
selection share = minutes played
/ minutes available while at that club
leverage = share of his on-pitch minutes
with the score within one goal
The availability window runs from a player’s first appearance for the club to the club’s most recent fixture, so a mid-season signing is not punished for games he could not have played.
Two things worth stating
Starters are derived, not taken from the team sheet. The feed’s starter flag is unreliable: it defaults to true when the field is absent, which produced 8.4 starters per team-game and more than eleven in 96 cases. A starter is now defined as a player who appeared with no substitution-on event, which returns exactly eleven per side.
Unused substitutes are not playing-time observations. The 14 September 2026 minutes-check migration records a correction: participation is derived from listed position rather than the unreliable starter flag, excluding unused bench rows. Earlier saved aggregates could credit fabricated minutes. Current check status is retrieved separately. Residual named starters without match events indicate missing ingestion, not proof that they were unused substitutes.
Similarity
Players are compared on percentile vectors within position pool. Percentiles rather than raw values, because they are robust to the skew in most role distributions.
distance = sqrt( sum( (pct_a - pct_b)^2 ) ) / sqrt( n_shared )
similarity % = 100 - distance
The default comparison uses 32 dimensions spanning roles, pass trajectory, carrying, creation, shooting and defending. A narrowed query, such as similarity as a shooter, restricts the vector to that family and requires at least 60% of the target’s available dimensions.
Similarity is stylistic, not qualitative. The closest match to a good player is frequently someone who does the same things less well. And a single-metric comparison will return very high percentages simply because there is only one axis on which to differ.
Definitions versus measurements
This is the distinction that matters most when someone challenges a number.
| Question | Honest answer |
|---|---|
| “Are the goals right?” | Measurement. Reconciled against the published scoreline by an automated invariant; current result read live. |
| “Is the xG right?” | Measurement. In-sample agreement is a diagnostic, not predictive proof. The Validation page reports the available temporal holdout and fitted support; neither establishes external provider agreement. |
| “Is the 80th percentile right?” | Arithmetic. It is definitionally correct; the real question is the input metric. |
| “Is the counter-attack stat right?” | Definition. Here is exactly how it is defined and why those thresholds were chosen. It is not validated because there is nothing to validate it against. |
Answering the fourth question as though it were the first is the fastest way to lose a room. Answering it precisely is the fastest way to earn one.
Known limits
- No off-ball information. The feed records the player on the ball. A run that drags a defender out of position is invisible. This is the largest limitation and it does not go away.
- No defender positions. Pass classifications describe trajectory, not verified line-breaking, because we cannot see who was between the two points.
- Single season. Team style is firm at this sample; individual finishing and anything rate-based on low volume is not.
- Age is a snapshot at match time, not a date of birth, so it can read a year light. The highest age observed is stored.
- 25 shots carry no coordinates and are absent from the shot model.
- Sequence tags are untuned against video.