IPL Data Sources & Open Datasets:
Kaggle, Cricsheet & Pipeline Architecture

Reproducibility, open science, and community data stewardship. Explore the underlying datasets, schema definitions, and cleaning pipelines behind our 1,243 match records.

1,243Matches Documented
19Seasons (2008–2026)
14Match Attributes
100%Open Access & CC-BY

Primary Upstream Data Sources

Peer-reviewed cricket data repositories and official records
Kaggle Dataset

IPL Complete Dataset (2008–2024+)

Curator: Patrick B. & Cricket Data Community
Format: CSV / Structured Tabular
License: CC BY-SA 4.0 International

The primary backbone for match outcomes, playing XIs, toss decisions, venues, and city locations. Regularly benchmarked and cross-checked against official match referee scorecards for zero data divergence.

View on Kaggle
Open Source Data

Cricsheet Ball-by-Ball Cricket Data

Curator: Stephen Rust & Open Source Contributors
Format: YAML, JSON, CSV
License: Open Data Commons ODbL

The gold standard in ball-by-ball cricket data. Utilized for deep verification of super over outcomes, over-by-over run rates, bowling figures, and dot-ball distributions across all playoff encounters.

Explore Cricsheet.org
Official Records

BCCI / IPLT20.com Official Archives

Governing Body: Board of Control for Cricket in India
Scope: Orange & Purple Caps, Official Points Tables, Auction Purse
Authority: IPL Technical Committee

Source for regulatory rulings, official Orange Cap and Purple Cap tie-breakers, team salary caps, purse regulations, Player of the Tournament awards, and environmental tree-planting metrics.

Visit IPLT20.com

Dataset Schema Reference: matches.csv

14 relational fields powering head-to-head, venue, and season analytics
Field NameData TypeDescription & Constraints
match_numberIntegerUnique chronological identifier for the match
team1StringFirst team in the official match schedule
team2StringSecond team in the official match schedule
match_dateDate (YYYY-MM-DD)Calendar date on which the match commenced
toss_winnerStringFranchise that won the pre-match coin toss
toss_decisionEnum ('bat' | 'field')Decision taken by the toss-winning captain
resultEnum ('Win' | 'Tie' | 'No Result')Official ICC/IPL match classification
eliminatorString / NAWinner of the Super Over in tied encounters
winnerStringFranchise that won the match (effective winner)
player_of_matchStringOfficial Player of the Match (POTM/MVP) awardee
venueStringStadium where the match was contested
cityStringHost city (e.g., Mumbai, Kolkata, Dubai, Johannesburg)
team1_playersList<String>Comma-separated list of 11 playing XI players
team2_playersList<String>Comma-separated list of 11 playing XI players

Entity Resolution: Franchise & Venue Canonical Mapping

How historical rebrands and stadium naming rights are unified

Franchise Rebrand Normalization

Franchises often undergo official renames or corporate rebranding. Our data pipeline maps legacy entities into unified modern franchises while preserving historical toggles:

  • Delhi DaredevilsDelhi Capitals (Unified since 2019 rebrand)
  • Kings XI PunjabPunjab Kings (Unified since 2021 rebrand)
  • Royal Challengers BangaloreRoyal Challengers Bengaluru (Unified since 2024 update)
  • Rising Pune SupergiantsRising Pune Supergiant (Spelling normalization)
  • Deccan ChargersDeccan Chargers / SRH Heritage (Tracked individually and cumulatively)

Stadium Fortress Normalization

Stadium names evolve with corporate sponsorships and official renaming. Our pipeline unifies aliases to calculate accurate multi-decade venue win rates:

  • Feroz Shah KotlaArun Jaitley Stadium, Delhi
  • Sardar Patel Stadium, MoteraNarendra Modi Stadium, Ahmedabad
  • Subrata Roy Sahara StadiumMCA Stadium, Pune
  • Punjab Cricket Association IS BindraPCA Stadium, Mohali
  • MA Chidambaram Stadium, ChepaukMA Chidambaram Stadium (Chepauk), Chennai

The Python Data Pipeline (scripts/process_data.py)

Automated, deterministic build pipeline with zero external runtime dependencies
scripts/process_data.py & scripts/enrich_ipl_data.pyStatic JSON Generation
# Automated processing flow:
1. Load raw matches.csv (1,243 rows)
2. Execute Canonical Entity Resolution (franchises & venues)
3. Calculate Toss Decision Biases (Bat First vs. Chase Win %)
4. Generate Head-to-Head Pairings matrix (210 unique team matchups)
5. Synthesize All-Time Player of the Match (POTM) MVP leaderboard
6. Inject verified tournament databases (Finals, Caps, Records, FAQs)
7. Emit static build payload → src/data/ipl_data.json
8. Compile Astro SSG → 288 layout-stable, zero-CLS HTML pages

Access the Raw Data & Source Code

The entire dataset, normalization scripts, and web application source code are 100% open-source on GitHub under the MIT / CC-BY license.