Data / methodology
How the data is collected and aggregated
Source
The figures on footballgpt.co/data are computed from a frozen weekly release of the FootballGPT database: user turns, animated practices, and conversations that real users (coaches, Football Manager video-game players, and individual players) generate inside the product. No external data sources are used.
Sources and cohorts
This page draws on separate cohorts from separate products. They are not joined coach journeys, and figures from one are not directly comparable to figures from another unless the page says so:
- FootballGPT: everyone who uses the product: real-world coaches, Football Manager video-game players, and individual players. Charts on this page separate coach-mode from FM-mode activity where the underlying query or generation is mode-tagged; where it is not, the chart describes FootballGPT activity in general rather than coaching activity specifically.
- CoachPage (coachpa.ge): coaches who have built a public CoachPage and state their own licence, country, years coaching, age groups and specialities. Shown only when the directory-visible cohort is at or above the minimum sample threshold below; hidden entirely otherwise.
- CoachReflect (coachreflection.com): coaches who log structured post-session reflections. This is currently a small, early cohort. The panel shows shares of that cohort, with the sample size stated alongside, rather than population-level figures.
- FCA Skool community: themed public posts from the Football Coaching Academy community. Shown only when the source is live at request time; if the feed is unavailable, the panel and any claim that it is a current source are both omitted rather than shown as a stale or error fallback.
Generating or animating a practice is not the same as delivering it, scheduling it, or confirming it worked. Nothing on this page treats generation volume as evidence of delivery, effectiveness, or an outcome.
How the homepage figures are counted
The public homepage uses conservative, floored FootballGPT-wide figures checked on 2026-09-02. They are a separate evidence contract from the narrower cohorts used in individual charts:
- 15,000+ FootballGPT users: registered profiles across every product mode, excluding profiles matching the repository's known test and automation-account rules. This is not a subscriber count and it must not be read as the number of real-world grassroots coaches.
- 41,000+ questions answered: one real user turn directly followed by a non-empty persisted assistant answer in
chat_historyrepresents one completed response. Profiles matching the repository's known test and automation-account rules and unpaired assistant rows are excluded. Persistence does not prove that the answer was read, useful or correct. - 18,000+ animated practices generated: one
generated_drillsrow per generated practice after rows owned by those same known test and automation accounts are excluded. This does not prove that the practice was saved, delivered or effective.
What the snapshot counts mean
The underlying snapshot behind footballgpt.co/data reconciles the retained FootballGPT charts to one snapshot ID and one observation time:
- Contributing accounts: distinct non-excluded product accounts represented in the relevant aggregate. This is not a subscriber count or a census of coaches.
- Total practices: requested, automatic and accepted-offer practice rows added together. Each component is also reported separately so the total can be reconciled.
- User turns: persisted
chat_historyrows withrole=user. Coach, FM, player, scout and goalkeeper modes remain separate.
Missing age/category fields are reported against the total rather than silently discarded. A chart may therefore have a smaller denominator than the release total.
Refresh cadence
Governed materialised views and the aggregate download are refreshed once a week. Every retained FootballGPT chart carries the same snapshot ID and observation time, keeping cited figures stable until the next release.
Anonymisation
All published charts are aggregate. No row-level data leaves the database. No usernames, emails, club names, team names, or session free-text are ever included in the dashboard, aggregate downloads, or the PDF report. Free-text query content is bucketed into pre-defined categories by automated rules; raw queries are never published verbatim.
Global chart cohorts require at least 50 contributing accounts. Country reports also suppress any displayed age cell with fewer thanfive accounts; visible country shares may therefore total less than 100%. Where a chart shows percentages, its denominator is stated.
Weighting and account concentration
The practice-mix chart publishes two estimates. Practice-weightedshares give every generated practice equal weight. Coach-weightedshares first calculate each account's category mix, including zero-share categories, then average those account-level shares. The comparison shows whether prolific accounts materially influence the practice-weighted result; neither estimate is a population claim.
Reporting layer normalisation
User-stated values like age group and practice category arrive in many forms ("U10", "u10", "Under 10", "10-year-olds", "U10-U12", "Adult", "Senior", "Erwachsene"). The product treats these as flexible inputs, because the AI chatbot can interpret them. The dashboard cannot, so we apply a fixed reporting-layer normaliser:
- Age bands: Mini (U6-U9), Junior (U10-U12), Youth (U13-U15), Senior Youth (U16-U18), Adult (U19+), Mixed.
- Practice categories: technical, tactical, game-based, set-piece, warm-up, physical, defending, attacking, goalkeeping, ball-mastery, cool-down, other.
Source data is never modified. Normalisation lives only in published views.
Chart 1: isolated-practice problem
Source: the generated_drills table (the internal name for animated practices), filtered to records with both an age group and a practice category set. Percentages are within each age band, summing to 100 across visible categories.
Chart 2: planning rhythm
Source: generated_drills.created_at, bucketed into day-of-week and hour-of-day in UTC. Note that the heatmap reflects UTC time, not user-local time. Clusters around 8pm UTC correspond to different local times for users in different countries.
Chart 3: audience mix
Source: chat_history.mode on persisted rows where role=user, set based on which surface the user is interacting with (coach mode, FM mode, player mode, scout mode, goalkeeper mode). A single user can appear in multiple modes across different sessions.
Chart 4: age band distribution
Source: generated_drills.age_group, normalised to age bands and grouped. A small share of animated practices (roughly 1%) carry an age_group value that does not map cleanly to a single band ("U10-Adult", "U13+", "All ages"), so these are bucketed as "Mixed".
Chart 5: pitch-third concentration (with coach-intent split)
Source: generated_drills.drill_data.players[].y. For each practice we average the y-coordinate of every player on the pitch (0 is the defending goal line, 100 is the attacking goal line) and bucket the result into Defensive (y<33), Middle (y=33-66), or Attacking (y>66).
Pitch zone is not a form field a coach fills in. It is derived from where the AI places players when it generates the practice diagram. When the coach's prompt doesn't mention a zone ("defensive third", "attacking third", "own-half", etc.), the AI tends to centre players around y=50 by default.
At drill creation time we record whether the prompt explicitly names a zone, using a deliberately conservative rules-based detector. New records now store the detector version alongside the result in generated_drills.intent_signals. Historical rows span detector versions and unversioned records, so we do not blend them into a single coach-intent percentage. Until a version-matched cohort is large enough, the pitch chart is reported only as generated output.
How to cite this release
Cite: 360TFT (2026), FootballGPT Grassroots Coaching Data, publication-v2. Include the snapshot ID and observation date shown on the data page, then link to footballgpt.co/data and this methodology. The machine-readable aggregate is available at /data/aggregate.
Opt-out
Any FootballGPT user who does not want their animated practices or queries to contribute to public aggregates can opt out by emailing hello@footballgpt.co. An opt-out is honoured at the next weekly refresh and applied permanently going forward.
Caveats
This dataset reflects what coaches generate when using AI tooling. It is not a representative survey of grassroots football coaching as a whole, because coaches who use AI tools may differ systematically from coaches who don't. Findings should be read as "what AI-using coaches ask for", not "what every grassroots coach plans". We've also recently improved our profile-collection in onboarding; older numbers reflect the pre-fix data, which under-counts certain age groups.
Some aggregates are also shaped by the AI's defaults. The pitch-third chart is the clearest example (see Chart 5 above), but anywhere the AI fills in a value the coach didn't specify, that default contributes to the aggregate. Where this materially shifts the read, we flag it on the chart itself. If you spot one we've missed, email admin@360tft.com and we'll add the caveat.