Methodology

Transparent enough to audit.

The product is designed around pre-match probability distributions rather than picks. Models are evaluated with walk-forward tests: a historical prediction can only use information that existed before that match.

RUNS

Player Runs

Historical innings are weighted by recency, expected batting role and competition. A cohort prior stabilizes smaller player samples. The output is a full runs distribution rather than a single average.

P(Runs ≥ x) is derived from that distribution for each displayed threshold.

ROLE

Expected batting position

The expected position is inferred from the player's most recent appearances with higher weight on recent matches. Role Stability measures how consistently the player bats within ±1 position of that expected role.

FOURS

Player Fours

The model estimates a player-specific probability of a four per legal ball, shrunk toward players with a comparable batting role. It then integrates that event rate over the player's estimated balls-faced distribution.

SIXES

Player Sixes

The sixes model uses the same opportunity framework, but learns a separate six-rate. A beta-binomial count model allows extra variation beyond a simple fixed-rate binomial model.

Validation

What we have tested so far

7.42%
Runs v0.3 Brier improvement vs Last-10 on 2023–26 holdout
6.42%
Sixes rate-model Brier improvement vs Last-10
5.01%
Fours rate-model Brier improvement vs Last-10
These validation comparisons use synthetic pre-match thresholds to test probability quality. They do not represent bookmaker comparisons, betting returns or guaranteed future accuracy.

Data sources & freshness

Historical match and ball-by-ball data used by CricLens come from Cricsheet. The current configuration contains 22 T20 feeds across men's and women's franchise, domestic and international cricket. Coverage depends on what is available in the underlying source archives and is not guaranteed to be complete for every competition, team or season.

Player and team snapshots are refreshed from the configured Cricsheet archives. Historical match pages use a dedicated Cricsheet-derived match-history snapshot. Future fixtures and recent results are maintained by a separate automated provider sync, so a competition can have historical analytics even when its next schedule has not yet been published upstream.

Adjusted-match handling

Team and competition analytics are descriptive rather than predictive. D/L, DLS and VJD-affected matches remain part of match history, match counts, win/loss records, recent form, venue records and head-to-head summaries when a usable two-innings scorecard is available. They are excluded only from scoring, run-rate and phase aggregates, where shortened or re-targeted innings would reduce comparability. Super overs are excluded from normal innings aggregates; a usable analytical match requires two normal innings.

Player model scope

Player projections are neutral competition-level estimates. They use recency, batting role, expected opportunity and segmented cohort priors. The current neutral projections are not automatically adjusted for the specific opponent, venue or confirmed playing XI shown on an upcoming Match page.

Validation scope

The published Runs, Fours and Sixes comparisons above come from holdout work on Big Bash League and T20 Blast data. Other model-profile groups share the segmented modelling framework, but should not be interpreted as having the same competition-specific validation evidence until their own backtests are reviewed.

Current limitations

  • Future fixture availability depends on the configured provider sync and published upstream schedules.
  • Confirmed playing XIs are not currently a model input.
  • Neutral player projections are not yet opponent- or venue-adjusted.
  • No bookmaker odds, betting comparisons, picks or betting-return claims are published.

Corrections

If you spot a team identity, fixture, scorecard or player-data issue, please use the Contact page and include the affected URL.