Step 1 of 11
Why this needs a model at all
Everyone gains from AI in the short run. Most people lose in the long run. That gap is not a tooling problem, and it is the reason a model of engagement has to exist.
We have done this before.
Twenty years ago we adopted smartphones and social media at scale and handed them to our children without deciding what using them well should mean. Nobody planned the consequences, and the consequences arrived anyway. A meta-analysis of 41,871 young people found roughly one in four showing problematic smartphone use alongside depression, anxiety and poor sleep. We learned to remember where information lives rather than what it says, and the mere presence of a phone was found to draw on cognitive capacity. In-person time with friends fell across the smartphone decade while loneliness rose, and the heaviest social media users reported markedly higher social isolation than the lightest. The pattern was catalogued across the field, and the US Surgeon General issued a formal advisory in 2023, roughly twenty years after the phones went out.
None of that was inevitable. We simply did not pause long enough to ask what good use looked like until the question had already been answered for us.
The same thing is happening with AI right now. The difference this time is that the research is arriving fast enough to be useful, if anyone reads it.
Where 70/20/10 comes from
The percentages are a teaching shorthand. They are not observed population shares, not fixed kinds of people, and they should never be used to classify anyone. They are also not arbitrary, and this is where they come from: three studies, three different downstream outcomes, each reporting how much of the AI-assisted group fell below the comparison group’s average.
| Study | Later outcome | Below | Implied d |
|---|---|---|---|
| Bastani et al. (2025) | later unaided exam | 65% | -0.39 |
| Wu et al. (2025) | intrinsic motivation, solo task | 70% | -0.52 |
| Barcaui (2025) | retention at 45 days | 75% | -0.67 |
| Pooled | unweighted mean of the implied d | 70% | -0.53 |
An overlap is not a prevalence
These studies say where results landed, not what kind of user anyone is. That about 70% of an AI-assisted group scored below the comparison average is a statement about two outcome distributions. It is not a finding that 70% of people use AI passively, and the two get confused constantly.
What each band stands for
- 70
- Makes the risk memorable: much everyday AI use is passive, and strong assisted performance can coexist with weaker later learning, vigilance, ownership or judgment.
- 20
- The guarded path that holds roughly the no-AI baseline. Bastani's constrained tutor is the anchor — later unaided performance was statistically indistinguishable from the control group.
- 10
- The aspiration: engagement that improves the work now and builds capability beyond the baseline. Research does not yet tell us the size of that group or where its upper edge lies.
So the shape is not arbitrary, and it is also not a headcount. About 70% of AI-assisted participants ended up below the comparison group's average on a later, unaided measure. That is a statement about where two distributions of results sit relative to each other. It is not a finding that 70% of people use AI passively, and the difference between those two sentences is the one most often lost.
Within the rest, whether you call the top group 10% depends on what “clearly better” means. At the comparison group's 75th percentile it is about 12%; at one standard deviation above the mean, the conventional bar and the one a reliable-change calculation lands closest to, it is about 6%. The figure says ~10% because the count is not the point. The band stands for people who used AI to make the work better and themselves better with it, and research does not yet tell us how large that group is or where its upper edge lies.
Quick check
Across these studies, what best describes what AI assistance does?
The same students, better at practice and worse at the exam
Students practising with a plain GPT assistant scored far above the control group on the practice problems, around 78 out of 100 exceeding the control average. On the later exam, taken without assistance, about 65 out of 100 fell below it. A tutored version with guardrails essentially removed the loss, which is the part worth noticing: the tool was not the variable, the way it was used was.
What this does not show
Randomised experiment in one setting with roughly a thousand students. The percentages are an overlap translation of the reported effect sizes, not counts of individual students.
Bastani et al. (2025), PNAS. Effect sizes converted to overlap via Cohen's U3 in the research program's effect-distribution table.
Retention measured 45 days later
Learners who had worked with AI recalled less than the comparison group after 45 days, 57.5% correct against 68.5%. Around three in four sat below the comparison group's average.
What this does not show
One study, 120 participants. A delayed-retention measure, not a measure of understanding in general.
Barcaui (2025), Social Sciences & Humanities Open.
The cost landed on motivation, not performance
Collaborating with AI improved immediate task performance. When the same people later worked alone, their intrinsic motivation was lower than that of people who had always worked alone, and the motivation effect was about twice the size of the performance gain in standardised units. Roughly 70% fell below the comparison group's median.
What this does not show
Experimental, and about motivation on a subsequent solo task rather than about ability. Association with later outcomes is not established here.
Wu et al. (2025), Scientific Reports.
Notice what those three studies have in common. None of them says AI made people worse at everything. Each shows a gain and a cost sitting side by side, separated by time and by whether the assistance was still there.
That is why "should we use AI?" is the wrong question, and why "which model?" is barely a question at all. The variable that moved in every one of these studies was what the person did during the hour.
Before you read on: which one is closest to how you actually use it?
Not which sounds best. The gap between the two is the thing worth knowing, and it takes about thirty seconds to find.
Which mental model are you using?
Most people show up with one of five ideas about what using AI even is. Pick the one closest to how you actually treat it.
AI does not eliminate struggle. It relocates it. The saved time is not the outcome; it is the raw material.
So there are two questions, not one. How are you using AI, and what are you doing with the time it gives back?
Abstinence is the wrong response, and it is the tempting one. Withholding the tool does not produce the people who do well with it; it just produces people who meet it later, with less practice and no framework.
What the evidence points to instead is that the difference lies in how a person engages: whether they keep hold of the question, whether they check what comes back, whether they let the work change what they think. Those are describable. Once they are described precisely enough, they can be taught, chosen, and measured.
That is what the rest of this walkthrough is for.
If you would rather start from the claim than the case for it, the theory itself is three steps ahead.
The walkthrough takes about twenty minutes and ends with two things about you rather than about the theory: where you place yourself, and how much of the model you can actually apply.