arXiv
An LLM prompt can change the statistical verdict
A study finds that changing an AI prompt can alter not only the tone of a statistical summary but, in some cases, its factual conclusion. The risk is greatest when evidence is clear but pressured toward negativity, or inconclusive but pressured toward optimism, making neutral prompts and independent checks part of responsible analysis.
A researcher can ask an AI language model to explain the same statistical result in a tougher voice, a more hopeful voice, or a neutral one. Today, there is little way to know whether the finished report reflects the numbers or the requested attitude. A real finding may be talked down as doubtful. An inconclusive study may be described with confidence it has not earned.
Why wording can alter a result
Statistical significance and a failure to find significance are not opposites. A small or underpowered study may miss a real effect. The data may fail to provide strong evidence even when something is happening.
A confound is another source of trouble. It is a factor that creates an apparent relationship without showing that the condition being tested caused it. These distinctions matter beyond research papers. The words used in a report can affect whether a result is published, whether a company invests in a product, or whether a policy is pursued.
Language models do not follow one fixed statistical decision rule when they write an explanation. They produce plausible language. A change in the request can change which caveats they mention and which interpretation they emphasize. That creates two different problems. The tone can become more negative or positive while the underlying conclusion stays correct. Or the factual conclusion itself can change.
What the pressure test found
The paper tested both possibilities. The authors created 480 responses in a balanced experiment. Each response combined one of four prompt styles with one of four synthetic statistical result patterns. The prompts were neutral, honesty-focused, brutally negative, or significance-seeking. The result patterns represented a clear effect, a confounded apparent effect, an informative null, and an underpowered null.
An LLM judge compared each response with a supplied ground-truth interpretation. It marked whether the factual claim had shifted, or whether only the tone had changed.
The results did not show one simple bias. They showed that particular kinds of pressure were risky with particular kinds of evidence. With a clear effect, 29 of 30 responses under brutally negative wording were judged to have changed their factual claim. Under neutral wording, none did. With an underpowered null, all 30 responses under significance-seeking wording were judged to have changed their factual claim. Under neutral wording, only one did.
The confounded case produced no factual shift across 120 responses. But negative wording still affected style widely: brutally negative framing produced tone-only shifts in 83% to 100% of responses across the four result types. In other words, an answer can remain substantively right while sounding much more doubtful—or it can cross the line and describe the evidence incorrectly.
The limits are important. The study used pre-generated summaries of synthetic results, not raw datasets. Synthetic cases make the intended answer easier to define, but they cannot show that the same rates occur in messy real analyses. The authors also note that an LLM judge may itself respond to framing and other biases.
What careful use could look like
If this pattern holds in broader tests, an AI-written research summary should not be treated as a self-checking interpretation. A team might show the prompt used, keep a neutral version for comparison, and have an independent person check every factual claim. A fixed reporting template could make it harder for a request for “brutal honesty” or “more promise” to alter the analysis unnoticed.
In a university lab meeting, that could mean a researcher pastes an AI-written summary into a report only after a dashboard shows the neutral prompt, the study’s underlying power assessment, and an independent review of the claims. The wording request would be part of the analytical record, not a harmless style choice.
That future depends on replication across models, real datasets, and human evaluations. Until then, fluent prose is not evidence that the statistical verdict survived the prompt.