A new paper out of Wharton should change how you think about every piece of AI-inclusive research you have run this year.
Shaw and Nave ran three preregistered experiments (1,372 participants, over 9,500 trials) and documented something they call “cognitive surrender.” They gave people access to an AI assistant while solving reasoning problems. On trials where participants engaged the AI, they followed its answer about 80% of the time. Even when it was deliberately wrong.
When AI was accurate, participant performance jumped 25 percentage points above the no-AI baseline. When it was wrong, accuracy fell 15 points below unaided performance (Cohen’s h = 0.81).
That gap is the cognitive surrender effect. And it’s large.
Here’s the part that should concern every research team.
AI made participants more confident (by about 12 percentage points) even in a study where the AI was set up to be wrong half the time. That was a deliberate methodological choice to isolate the effect, not a comment on real-world AI reliability. But the implication holds: people weren’t just following bad advice. They believed in it.
For market research, this is a data quality problem that’s invisible at the surface. Respondents who used AI before or during a session may report inflated confidence in AI-shaped opinions, and that confidence is indistinguishable from genuine conviction. You won’t see it in your data.
Not everyone surrenders equally.
People who trust AI more gave in readily. Those with higher need for cognition and fluid intelligence pushed back. The researchers call these “thinking profiles.”
The study used general consumer populations, and the authors themselves flag this as a limitation: expert dynamics haven’t been tested under this framework. But the question is worth sitting with. If these patterns extend to expert respondents (and the theory suggests they might), what does it mean for your advisory board? The oncologist who checked a clinical question with ChatGPT before your session may genuinely struggle to separate her own clinical reasoning from what the model told her.
The structure around the decision determines whether AI helps or hurts.
Time pressure made surrender worse. But incentives and real-time accuracy feedback more than doubled override of bad AI, from 20% to 42%. Giving people direct feedback on whether they were right or wrong reactivated deliberate thinking without eliminating the benefits of AI when it was accurate.
The takeaway isn’t that AI is dangerous. It’s that uncritical AI adoption is a design problem, not an inevitability.
This is why the framing of AI in our industry matters.
The dominant model is “human-in-the-loop”: AI at the center, people as checkpoints. What Shaw and Nave showed is what poorly designed oversight produces: passive acceptance, inflated confidence, and reasoning that tracks the AI’s accuracy rather than the respondent’s own knowledge.
The failure mode isn’t human involvement. It’s checkbox oversight: human review without accountability, feedback, or incentive to push back.
Interestingly, the paper’s own recommendations align with more robust human oversight, not less: confidence scores, uncertainty indicators, real-time feedback. The research doesn’t argue against human-in-the-loop. It argues for designing it properly.
We built BioVid’s approach around the opposite framing: AI in the Human Loop.
Framing, interpretation, and delivery belong to the researcher. CIAIRA™, our award-winning AI platform, operates inside that loop, making each layer faster and deeper without replacing the judgment that makes the insight meaningful.
In practice, that means structuring research sessions, so AI-generated outputs are subjected to the conditions this research says reduce surrender: deliberate prompting, real accountability, and active engagement. The goal isn’t to remove AI’s efficiency. It’s to ensure the researcher’s judgment remains the dominant force.
Nobody has fully cracked what System 3 – as distinct from System 1 and System 2 – means for research design. But the architecture you choose shapes everything downstream, including whether the insights you deliver reflect your respondents’ genuine thinking or an artifact of the tools in the room.
Have you changed how you design research or govern AI workflows based on findings like these?
To further this discussion, contact us at info@biovid.com.
Source: Shaw & Nave, “Thinking, Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender,” The Wharton School of the University of Pennsylvania, 2025.