nathan batty&

Work / Large-scale discourse analysis

Master's thesis · 2026

Large-scale discourse analysis

201,416 YouTube comments on a Magic: The Gathering governance crisis, and what choosing the method the AI oversold taught me about honest measurement.

201,416 comments · 726 videosRegex + ANOVA + attribution liftSecond-coder validation
Weekly meta-discourse rates by dimension. Discussion spikes at each governance event, not at the framework's launch, which is the thesis's core finding.
Weekly meta-discourse rates by dimension. Discussion spikes at each governance event, not at the framework's launch, which is the thesis's core finding.
84% / 34%regex — precision / recall
→
κ 0.84–0.87the method I moved to next

Regex was precise but caught only a third of the meta-discourse. The ensemble coder I built for the next paper agreed with human coders at κ 0.84–0.87.

The study

In September 2024 the Magic: The Gathering Commander community hit a governance crisis. An independent rules committee banned several cards, then handed governance to the publisher, which rolled out a new “bracket” system for pregame power-level talk. I collected 201,416 comments across 726 videos from 14 creators and asked how a community learns a new shared framework through creator videos and their comment sections.

Because no one can read and hand-code 201,416 comments, I built a regex pattern-matching system over three framework dimensions, validated it against a second human coder on a 300-comment sample, and analyzed the results with one-way ANOVAs (reporting η² alongside p, since at that N even trivial effects reach significance) and an attribution-lift measure for creator dependence. The finding is that meta-discourse is event-driven. It spiked at each governance shock rather than at the framework’s launch, and the controversy videos kept drawing new discussion for months.

What the AI got wrong

In the working notebook, the AI assistant was relentlessly over-enthusiastic. It declared “EXCELLENT RESULTS,” reframed data that contradicted the theory into a brand-new model with an unfalsifiable “not yet observable” stage, and drove 22 rounds of pattern-tweaking to hit a fake “100% recall”, textbook overfitting to the 300-comment validation set.

The honest version of the finding was smaller than the exciting one. It was also the real one.

What I did

I audited the notebook hard and rescoped the thesis to be honest. The final draft reports the real 84% precision and 34% recall and the low per-dimension agreement (κ from 0.15 to 0.49) as conservative lower bounds, drops every causal claim, and labels the theoretical mappings as proposed interpretations rather than established results. That regex ceiling is also what pushed me to the ensemble-plus-gold-standard coding I used on the next paper, at κ 0.84–0.87.