Recommended. But what did they say?
Do written reviews tell us more than a thumbs up or down? This study reads the text, separates mixed opinions from contradictions, and tests whether a sentiment model can capture those distinctions.
What the text actually says
Sentiment within each Steam recommendation group
Read the exact counts and percentages
Nuance is not contradiction
Six distinct relationships across the full corpus
Intervals are 95% Wilson intervals. They do not correct unknown sampling bias.
How to read these categories
Both praise and criticism. A mixed review can still reasonably recommend the game.
Clearly negative text paired with Recommended, or clearly positive text paired with Not Recommended.
Content without a game evaluation, or too little evidence to infer sentiment. Neither automatically implies disagreement.
Human reference, not Steam labels
Recommendations were hidden during annotation.
Agreement is measured before adjudication. Human disagreement is resolved before the final labels are used in analysis or evaluation.
Inspect annotation agreement by field
What players talk about.
Explore the themes and playtime groups behind the recommendations. These are associations within this corpus, not explanations of player psychology.
Primary themes
One human-assigned primary theme per review
Bar lengths show review counts on a common scale. Theme labels were assigned without seeing the Steam recommendation.
View source theme counts
Recommendation × sentiment
Cell shading shows share within the recommendation row.
Update-related text does not establish an update’s causal effect. The study has no controlled before-and-after comparison.
Opinion composition by playtime
Each bar totals 100% of its own group. Compare group sizes as well as percentages.
Playtime cutoffs are descriptive choices. Small groups are unstable; alternate cutoffs and duplicate/confidence sensitivity checks are available in the project’s sensitivity results.
View group counts and denominators
Can text recover the nuance?
Compare the selected text classifier with transparent baselines on the same eligible locked-test reviews.
Locked-test comparison
Points are estimates; whiskers show 95% group-bootstrap intervals.
Macro F1 gives each sentiment class equal weight. The Steam mapping can only output positive or negative, so it cannot identify mixed or non-evaluative text. The 2,000-resample intervals describe uncertainty, not proof of superiority.
Inspect metric estimates and intervals
Where the model gets confused
Rows: human labels · Columns: predictions
Diagonal cells are correct predictions. Off-diagonal cells reveal which classes the model confuses. Human-screened insufficient text is outside this model’s four-class task.
Performance by class
F1 score on the eligible locked test · Scale 0–1
Support is the number of actual test reviews in each class. Small supports make individual class scores particularly uncertain.
Selection happened before testing
Development-only repeated stratified group cross-validation
The one-standard-error rule selects the least complex eligible text model near the highest CV mean. This is a complexity-control heuristic, not a statistical equivalence test. SD describes variation between repeat means, not independent new datasets.
From corpus to evaluation
The original split was assigned before annotation. Exact normalized-text groups stay together across development/test and within cross-validation. Do not tune against the published test errors.
Read the mistakes, not just the score.
A deterministic diagnostic sample of locked-test errors. These examples illustrate failure modes; they are not representative of the corpus.