Skip to main content

Step 3 of 11

Why "does AI help?" is the wrong question

The evidence on AI assistance is not mixed because the studies are bad. It is mixed because the question is underspecified.

Generative AI reliably improves what you can produce while you are using it. That much is not seriously in dispute. What happens afterwards is a different matter, and this is where the results stop agreeing with each other.

Human-AI combinations do not consistently beat whichever of the human or the AI is better working alone. Unguarded access to answers can hurt later learning. In one study, students preferred working with a language model but did worse three days later than students who took notes. In another, an immediate advantage on simpler tasks had largely disappeared by an unaided test three weeks later. Reflection prompts generated by a language model did not outperform either no reflection or non-AI reflection on two-week exams.

And the function of the AI matters, not just its presence: agents that provided information and agents that facilitated reflection both raised people's immediate intention to act, but only one of them moved behavior three weeks out.

So the model starts from a different premise. AI engagement is not a property of a tool or of a single prompt. It is a pattern produced jointly by a person, a tool, a task, and a context. If that is right, then the useful questions are answerable: what did the person actually do, what did they keep hold of, and what did that cost or buy them.

Those questions need a theory before they need vocabulary, and vocabulary before they need data. The theory comes next.

Quick check

The studies on AI assistance disagree with each other. Why?