Skip to dashboard
PPHASMOPHOBIAPlayer feedback journal
THE STUDY
01Case summary02Player reports03Model findings04Review evidence
Research only

One game. One observed corpus. No population-wide claims.

View project on GitHub ↗Read the methodology ↗
CASE FILE / STEAM REVIEW ANALYSISObserved corpus
01 / The research question

Recommended. But what did they say?

Do written reviews tell us more than a thumbs up or down? This study reads the text, separates mixed opinions from contradictions, and tests whether a sentiment model can capture those distinctions.

Observed corpus only

What the text actually says

Sentiment within each Steam recommendation group

FIG. 01
Read the exact counts and percentages

Nuance is not contradiction

Six distinct relationships across the full corpus

Intervals are 95% Wilson intervals. They do not correct unknown sampling bias.

How to read these categories

Mixed opinion

Both praise and criticism. A mixed review can still reasonably recommend the game.

Hard contradiction

Clearly negative text paired with Recommended, or clearly positive text paired with Not Recommended.

Non-evaluative or insufficient

Content without a game evaluation, or too little evidence to infer sentiment. Neither automatically implies disagreement.

Human reference, not Steam labels

Recommendations were hidden during annotation.

Agreement is measured before adjudication. Human disagreement is resolved before the final labels are used in analysis or evaluation.

Inspect annotation agreement by field
Scope matters.The collected review window is not a representative sample of all Phasmophobia or Steam reviews. These are descriptive findings, not evidence that an update caused a change in sentiment.
02 / Player reports

What players talk about.

Explore the themes and playtime groups behind the recommendations. These are associations within this corpus, not explanations of player psychology.

Primary themes

One human-assigned primary theme per review

FIG. 02
RecommendedNot Recommended

Bar lengths show review counts on a common scale. Theme labels were assigned without seeing the Steam recommendation.

View source theme counts

Recommendation × sentiment

Cell shading shows share within the recommendation row.

Update-related text does not establish an update’s causal effect. The study has no controlled before-and-after comparison.

Opinion composition by playtime

Each bar totals 100% of its own group. Compare group sizes as well as percentages.

FIG. 03

Playtime cutoffs are descriptive choices. Small groups are unstable; alternate cutoffs and duplicate/confidence sensitivity checks are available in the project’s sensitivity results.

View group counts and denominators
03 / Model findings

Can text recover the nuance?

Compare the selected text classifier with transparent baselines on the same eligible locked-test reviews.

Not production-approved

Locked-test comparison

Points are estimates; whiskers show 95% group-bootstrap intervals.

Macro F1 gives each sentiment class equal weight. The Steam mapping can only output positive or negative, so it cannot identify mixed or non-evaluative text. The 2,000-resample intervals describe uncertainty, not proof of superiority.

Inspect metric estimates and intervals

Where the model gets confused

Rows: human labels · Columns: predictions

Diagonal cells are correct predictions. Off-diagonal cells reveal which classes the model confuses. Human-screened insufficient text is outside this model’s four-class task.

Performance by class

F1 score on the eligible locked test · Scale 0–1

Support is the number of actual test reviews in each class. Small supports make individual class scores particularly uncertain.

Selection happened before testing

Development-only repeated stratified group cross-validation

The one-standard-error rule selects the least complex eligible text model near the highest CV mean. This is a complexity-control heuristic, not a statistical equivalence test. SD describes variation between repeat means, not independent new datasets.

From corpus to evaluation

The original split was assigned before annotation. Exact normalized-text groups stay together across development/test and within cross-validation. Do not tune against the published test errors.

What remains untestedFuture updates, other games, and production use. No operational abstention threshold has been approved. The model is evaluated only on English reviews judged to contain sufficient text.
04 / Review evidence

Read the mistakes, not just the score.

A deterministic diagnostic sample of locked-test errors. These examples illustrate failure modes; they are not representative of the corpus.

Selection ruleUp to two lowest-confidence errors per human-label/predicted-label pair, with ties broken by review ID. Confidence is the model’s uncalibrated predicted-class probability, not a reliability guarantee.
Read the full diagnostic sample as a table
Independent Phasmophobia study · Orlando DevitoValidation report · Research release