'People gave the highest ratings to AI-generated stories': Refuting a study's claim
- Sheelagh Caygill

- 13 hours ago
- 10 min read
Updated: 9 hours ago

Challening media reports on the study, 'Bot or not: Can people tell the difference between stories written by a human or by an AI system?'
You have probably seen many headlines in the last couple of weeks about people giving stories writen by generative AI higher ratings than stories written by humans. These news pieces all present a study's nuanced, conditional result as an absolute fact.
The new study, ‘Bot or not: Can people tell the difference between stories written by a human or by an AI system?’ and published in 'Judgment and Decision Making' August 5 2026 reveals people gave the highest ratings to AI-generated stories that they were told had been written by a human, and were unable to tell the difference between human- and AI-created stories. The study is by Sydney Sears and Dr. Deena Skolnick Weisberg and published by Cambridge University Press.
This study supports a narrower conclusion: in a controlled online experiment with Prolific participants and a limited set of short stories, AI-generated texts were often judged as fluent and absorbing. That’s not the same as proving that AI writes better fiction across the board. The study is a useful data point about what LLMs can do in specific conditions; it is not a verdict on creativity, craft, or the future of literature. Readers, writers, and journalists should treat such findings with curiosity — and with skepticism.
Like most people, I'm busy, but I've taken time (again) to respond to this study with 'People gave the highest ratings to AI-generated stories': Refuting a study's claim', because the level of attention generative artificial intelligence (gen AI) receives is massively distorted. The white-hot focus on the supposed marvels of gen AI is having has seriously negative effects on writers, publishers, reviewers, book loveers, and creative writing programs.
Headlines compress a qualified set of experimental results into a declarative claim
These headlines and stories compress a careful, qualified set of experimental results into a single declarative claim. The study reports that short stories generated by large language models (LLMs) were, in the experiment, perceived as indistinguishable from human-written stories and “tend to be perceived as higher quality and more absorbing.” That phrasing — perceived, tend to — matters. It signals a narrow, context-bound finding about reader judgments in a controlled task, not a general verdict on literature, creativity, or the future of fiction.
This is always where reporting and the public conversation go wrong — in the leap from measured perceptions in a lab-like setting to the sweeping claim that AI “writes better”. Just as with previous studies claiming AI outperforms human poets, major media outlets rushed to amplify a clickbait headline without critically examining the study’s significant methodological flaws.
'People gave the highest ratings to AI-generated stories': Refuting a study's claim
All forms of AI have received a massive amount of exposure and hype since OpenAI unveiled ChatGPT in November, 2022. Continually letting unchecked claims about AI's artistic superiority to take root harms society in many ways. Unchecked claims about AI's craft and artistic superiority harm the literary community in many ways, including:
Devaluing authors: It discredits creative writers, many of whom already struggle for fair compensation and recognition.
Funding threats: Skeptical institutions, school boards, or policymakers can use flawed studies to argue against funding arts programs, literary grants, and creative writing courses.
Misconceptions: It misleads the public into believing generative AI can replace authentic human storytelling, ignoring the true depth of literary craft.
Here's what's wrong with the study, ‘Bot or not: Can people tell the difference between stories written by a human or by an AI system?'
Who the participants were — and why that matters
The 1,682 adult participants were recruited from Prolific.com and meet standard statistical thresholds for general population size. Prolific is a valuable platform for behavioral research, but its participant pool isn’t a representative sample of readers, critics, or writers. Contract workers on Prolific are there to earn money for completing tasks. Their incentive is to finish accurately and quickly, not to engage in deep, reflective reading. When asked to rate short stories in a survey environment, many will rely on surface cues — fluency, coherence, immediate emotional appeal — rather than the long, patient attention that literary appreciation often requires. Prolific contract workers aren’t literature students or professional critics.
The study's sampling methodology isn't a scientific sampling of the world's readers and writers, because the researchers constructed their sample to mirror general United States census demographics (age, gender, race).
Limited reader expertise and diversity
A study like this one should either sample readers who represent the population of interest (e.g., general readers, avid readers, critics) or explicitly limit its claims about quality and ease of reading to the sampled population. The Prolific sample is not a stand-in for the community of readers who help shape literary reputations or for a community that evaluates innovation, risk, and complexity. It excludes international readers, diverse cultural storytelling traditions, and global literary perspectives. It does not reflect the specialized skill set, critical reading habits, or aesthetic standards of dedicated book lovers, literature scholars, or practicing writers. And while it is heterogeneous, it's skewed toward people who do online studies.
The setting: reading in the lab vs. reading in life
The experimental setting — short texts read in a survey, often in a single sitting and with explicit or implicit cues about authorship — is far removed from how readers encounter fiction in the real world. Readers choose books based on recommendations, reviews, and reputations; they read over days or weeks; they discuss and re-evaluate texts. Online studies naturally favour immediate impressions. They cannot capture the social, historical, and intertextual processes that determine literary value.
The problem of authorship labels and priming
The study included conditions where participants were told whether a story was written by a person or by ChatGPT. Labels and expectations shape perception. If participants expect AI text to be clumsy or generic, being told a story is AI-generated can lower their standards; conversely, if they expect AI to be impressive, that can raise ratings. The study’s design attempts to control for this by varying labels, but the broader point remains: authorship cues and cultural narratives about AI influence how people judge texts. Headlines that ignore these subtleties flatten a complex experimental design into a simple, misleading claim.
Flawed sampling and forced framing
In Study 1, participants read a single short story (~1,000 words) under misleading or explicit labels:
Read Human-Written Story: 417 participants were told it was by a person; 416 were told it was by ChatGPT.
Read AI-Generated Story: 423 participants were told it was by a person; 426 were told it was by ChatGPT.
When forced to judge brief excerpts under deceptive or binary conditions, casual readers tend to rely on immediate impressions instead of deeper craft or artistic evaluation.1
Short, narrow samples of texts produce fragile claims
The study evaluated only three short realistic fiction pieces written by humans (~1,000 words each) alongside three ChatGPT counterparts. This tiny and uniform sample in no way represents the genres and vast ranges and styles of short stories.
Short stories are not a single, uniform object. Many writers and readers value uniqueness, experimental forms, difficult syntax, unreliable narrators, voices from diverse people and cultures, emotional depth and, not least, voices that come from human lived experience. LLMs are good at producing conventional prose that aligns with mainstream patterns in their training data. It cannot produce experimental prose, complex narrative structures, and longer forms such as novellas or novels; here, gen-AI consistently struggles to maintain many elements of contemporary writing.
Using a small set of uniform short stories will reveal how well LLMs can mimic that particular slice of short fiction, not how they perform across the full spectrum of literary practice.
Confusing ease of reading with literary quality
The authors acknowledge that AI-generated text is often preferred because it "smooths out the rough edges" and is easier to understand, especially when someone's reading quickly. But most writers and readers know that ease of reading isn't the primary criterion for evaluating creative literature. True story lovers crave complexity, unique voice, ambiguity, and challenging themes—not predictable, homogenized prose designed for friction-free reading.
Measures of "quality" and "absorption" are limited
The study reports that AI stories were rated as higher quality and more absorbing. But how were those constructs operationalized? Typical online rating tasks use Likert scales for readability, enjoyment, or perceived quality. Those measures capture immediate reactions that are usually based on clarity and superficial emotional engagement; they don't capture longer-term thoughts and conclusions. Does a story linger, make a reader think or ask questions, or create a desire for re-reading? Is it an original voice, a unique approach, and does it reveal structural depth?
Much like popular fiction, a story that is easy to read and emotionally immediate can score highly on short-term absorption while still being derivative, formulaic, or forgettable. Conversely, stories written by humans may become canonical were initially challenging, opaque, or divisive. A short-term rating task is not designed to detect those qualities.
Some study participants revealed a slight advantage
The study notes that, "Self-reported expertise with AI, but not with fictional literature, was positively correlated with correct story identification." This can be construed as revealing a narrow technical advantage for generativer AI-literate users, as well as exposing fundamental methodological biases in how the study was designed. What it does prove is that frequent software users can spot software output—not that AI writing matches the quality and nuance evaluated by literary experts.
Looking at this more closely, we see:
Recognition of an LLM voice: Participants who frequently interact with AI models learn to recognize mechanical telltales—such as predictable transitional phrasing, balanced paragraph structures, overly tidy thematic resolutions, and a polished, averaged tone.
Artifact detection: Rather than engaging with the narrative on an artistic level, AI-literate readers succeed by acting as system auditors who recognize synthetic formatting habits.
Shift from aesthetic reading to spot-the-bot: The experiment reframes reading from an aesthetic, artistic evaluation into a technical search for machine flaws. Success in the study rewards technical familiarity with software quirks rather than an understanding of character development, voice, or subtext.
Participant pool skew: It’s valid to assume that crowdsourced workers on Prolific are active online and therefore more likely to use digital automation tools regularly, as opposed to being practicing creative writers or literary scholars.
A flawed measurement of literary expertise: The study measured literature expertise using a brief, self-reported 5-item scale based on general reading frequency rather than objective testing of literary analysis or craft knowledge. Concluding that literary training "has no effect" reflects an extremely weak diagnostic tool rather than a true lack of literary discrimination.
Statistical significance is not the same as practical or artistic significance
Even when differences are statistically significant, the effect sizes and practical implications matter. A small average difference in Likert-scale ratings across a controlled sample does not mean that AI has surpassed human creativity. It means that, under specific conditions and for specific measures, readers in that sample preferred certain outputs. Translating that into a claim that AI “writes better” conflates a narrow empirical finding with a broad cultural judgment.
World's readers and writers not represented in this study
The raw number of participants meets standard statistical thresholds for general population size, but the study's sampling methodology is not a robust representative scientific sampling of the world's readers and writers; it shows geographic and cultural homogeneity.
Restricing their sample to mirror general U.S. census demographics means the researchers excluded international readers, diverse cultural storytelling traditions, and global literary perspectives. A general demographic cross-section does not reflect the specialized skill set, critical reading habits, or aesthetic standards of dedicated book lovers, literature scholars, or practicing writers.
Why writers and serious readers should care
Writers and serious readers value variety, difficulty, and distinctiveness. A model trained to maximize average reader satisfaction across a broad corpus will tend to produce safe, high-probability text — the kind of prose that reads smoothly but rarely surprises. That is useful for many applications (drafting, ideation, accessibility), but it is not the same as the craft of producing work that challenges, innovates, or transforms. If we care about preserving and encouraging literary risk, we should be skeptical of claims that equate fluency with artistic superiority.
Media responsibility: nuance lost in translation
Major outlets, including The Guardian, amplified the study’s most attention-grabbing interpretation. This is a recurring problem: nuanced academic findings are turned into declarative headlines because bold claims attract clicks. Responsible reporting would have emphasized the study’s limitations up front: the participant pool, the narrowness of the texts, the operationalization of “quality,” and the difference between short-term perception and long-term literary value. Instead, the headline framed the result as a general triumph for AI, which misleads readers and shapes public discourse in unhelpful ways.
What the study does show — and what it does not
The study provides useful evidence that LLMs can produce short-form fiction that, in controlled conditions, many readers find fluent and engaging. That is an important and interesting finding: it confirms that LLMs have reached a level of superficial skill in narrative generation. But the study does not show that AI can replace the full range of human literary creativity, nor that it can reliably produce the kinds of risk-taking, voice-driven, culturally situated work that often defines great literature.
The study’s participants are not the people whose judgments we should privilege when evaluating literary quality; the results tell us about perceptions in a specific experimental population, not about the broader cultural or artistic value of AI-generated fiction.
How research could be stronger
To make stronger claims about AI and literary quality, future studies should:
Sample a wide range of readers (avid readers, critics, writers, publishers) and report subgroup analyses.
Test a broader and more diverse set of texts, including experimental forms, culturally specific voices, and longer works.
Use richer outcome measures: delayed recall, re-readability, critical evaluation by trained readers, and measures of originality or voice.
Report effect sizes and practical significance, not just probability values. In this study, several measurements yielded highly significant probability values, but looking closely at the effect size metrics reported in the paper reveals that the factor accounted for less than one percent of the total variance in participant responses. In other words, the study found a "statistically significant" result, but the actual effect size was nearly zero—meaning the practical preference for AI-generated prose over human writing was negligible.
Be explicit about the limits of generalization from the experimental sample to broader cultural claims.
1Graves, B., & Frederiksen, C. H. (1991). Literary expertise in the description of a fictional narrative.
Castano, E., et al. (2023). On the complexity of literary and popular fiction. Mechanism: Non-expert readers tend to evaluate narrative text based on surface-level plot coherence and readability. Literary experts, by contrast, focus on stylistic nuance, structural voice, and authorial intention—qualities that are often obscured or flattened when texts are reduced to brief, 1,000-word survey samples.
Recent Empirical Studies on AI Literature: Core Source: Sears, D., & Weisberg, D. S. (2026). Bot or not: Can people tell the difference between stories written by a human or by an AI system? Mechanism: The authors explicitly note that participants preferred AI stories because they were simpler to read and digest, highlighting that standard non-expert readers frequently equate ease of consumption with high creative quality.



Comments