Skip to main content

Step 6 of 11

The eight modes, and how to tell them apart

The layer you can actually observe. Each mode below leads with what it looks like in a real exchange, because that is the question you arrive with: which one am I doing?

This is the layer that carries most of the measurement, because it is the only one recoverable from a transcript. Someone can read your conversation and say what you did. Nobody can read it and say what you believe.

That also makes it the layer most often misread. What follows is eight functions, not eight kinds of person. A single afternoon usually contains most of them, and the same mode can be the right call or the wrong one depending on the task, the stakes and how much you kept hold of.

Quick check

Oracle sits in the passivity tier. When is it the right move?

Each card opens with a short exchange. Expand a card for the cues that identify the mode in someone else's transcript, and the full definition.

1

Verification Agent

Agency

Ask AI to check your work.

What it looks like

You

Here's my analysis and the code behind it. Find what's wrong with it.

AI

Line 34 filters before the join, so you're dropping rows the aggregate should count.

You

Checked — you're right for the 2024 rows but not 2023. Why the difference?

Right when

The cost of an error exceeds the cost of a check. On current evidence this is the one mode whose favourable associations survive correction for multiple comparisons.

Costly when

You read silence as safety. 'It found nothing' is not 'there is nothing', and treating it as such converts a check into false confidence.

2

Creative Expander

Agency

Ask AI for range, not a single answer.

What it looks like

You

Give me eight framings for this paper's contribution, including two I'd probably reject.

AI

1. A measurement contribution... 4. A negative result reframed... 7. A boundary condition on prior work...

You

Six is closest but it's too narrow. What's between four and six?

Right when

Your own defaults are the constraint and you need to get outside them.

Costly when

You take option one. Satisficing on the first offer collapses Expander straight back into Oracle while still feeling like divergent thinking.

3

Critical Challenger

Agency

Ask AI to argue against you.

What it looks like

You

I'm arguing that remote work reduces innovation. Make the strongest case that I'm wrong.

AI

The strongest version is a measurement objection: patent counts and idea volume track co-location, but neither measures innovation quality...

You

That does land. It means my dependent variable is doing the work, not my theory.

Right when

You hold a position you have not stress-tested and you are willing to lose it. The theoretical inverse of Oracle.

Costly when

It is theatre. Inviting an objection you have already decided to dismiss produces the feeling of rigour and none of it, and it is hard to detect from the transcript alone.

4

Problem Setter

Agency

Ask AI to interrogate the question itself.

What it looks like

You

Before I answer 'how do we increase signups' — is that the right question?

AI

It assumes volume is the constraint. If activation is the constraint, more signups make the metric worse.

You

So the real question is which constraint binds. I've been optimising the wrong end.

Right when

The stakes are high enough that solving the wrong problem well is the expensive outcome. The most demanding mode and the rarest in the observed data.

Costly when

Reframing never stops. At some point a question has to be answered, and infinite problem-setting is a sophisticated way of not starting.

5

Tutor

Partnership

Use AI as a personalized instructor.

What it looks like

You

Explain why the central limit theorem works. I want to actually understand it, not just use it.

AI

Start with a skewed population — say, incomes...

You

So the shape of the original distribution stops mattering as n grows? Give me a case where it still matters.

Right when

You intend to hold onto the concept and you can generate your own checks on whether you have it.

Costly when

The scaffolding never comes down. Understanding an explanation while reading it is not the same as being able to reconstruct it later, and this mode makes that difference easy to miss.

6

Collaborative Problem-Solver

Partnership

Co-think with AI on the problem.

What it looks like

You

I think our churn is a pricing problem. Retention drops right after the trial ends.

AI

That fits, but the same curve appears when onboarding fails. What does usage look like in week one?

You

Low for the churned group. So maybe they never hit the activating feature — pricing might just be when they notice.

Right when

The problem is genuinely open and you have enough of your own to bring that the exchange can go both ways.

Costly when

It quietly decays into Production Assistant. This is the hardest mode to sustain, and the drift is gradual — you stop contributing reasoning and start accepting drafts, usually without noticing the turn where it happened.

7

Oracle

Passivity

Ask, accept the answer, move on.

What it looks like

You

What year did the Treaty of Westphalia end the Thirty Years' War?

AI

1648.

You

Thanks.

Right when

The fact is settled, you would recognise a wrong answer, and the cost of being wrong is low. Most lookups are legitimately Oracle, and treating that as a failure is the most common misreading of this model.

Costly when

The question is contested, or you could not tell a wrong answer from a right one. Then you have bought confidence you have no basis for.

8

Production Assistant

Passivity

Hand off the artifact and ship it.

What it looks like

You

Write a 200-word project update for my team about the Q3 delay.

AI

Here's a draft: 'Team, I wanted to share an update on our Q3 timeline...'

You

Perfect, sending.

Right when

The deliverable is routine, the stakes are low, and nothing about it needs to sound like you. A meeting-time change does not need your voice.

Costly when

It becomes the default. This is where voice erasure and skill atrophy actually happen, and they happen by repetition rather than by any single hand-off.

Now try it

Six exchanges. Name the mode. There is no score, deliberately — a tally would invite you to read six out of six as evidence that you engage well, and that is the inference the model says the data does not support.

The useful part is the misses. Each one explains the confusion it comes from.

Exchange 1 of 6

You

Here's my regression output and the code. Tell me what I've got wrong before I send this to my advisor.

AI

Your standard errors aren't clustered, and the panel structure means they should be.

You

Checked. That's right for the firm-year panel — but why doesn't it change the 2019 subsample?

Exchange 2 of 6

You

Give me ten angles on this essay topic.

AI

1. A historical framing... 2. A counterintuitive economic read... 3. A personal-narrative opening...

You

Two is great. Writing it now.

Exchange 3 of 6

You

I want to argue that our hiring bar has dropped. Make the strongest case that I'm reading the data wrong.

AI

The strongest objection is compositional: you're comparing cohorts hired under different funnels, so the bar may be constant and the applicant pool different.

You

That's a real problem for my argument. I need to hold the funnel constant before I claim anything.

Exchange 4 of 6

You

Draft the section on limitations for this paper.

AI

Here's a draft: 'This study is subject to several limitations. First, the cross-sectional design...'

You

Good, pasting it in.

Exchange 5 of 6

You

We keep asking how to raise course completion. Is that even the right thing to be measuring?

AI

It assumes completion tracks learning. If students complete by skipping the hard middle, the metric improves while the outcome degrades.

You

Then the question is which part of the course is doing the teaching. I've been measuring the wrong end of it.

Exchange 6 of 6

You

I think the bottleneck is our review process — PRs sit for days.

AI

Possible, though the same delay shows up when authors batch large changes. What's the median diff size on the slow ones?

You

Much bigger. So review latency might be a symptom of batching, not the cause. That changes what we'd fix.

If some of those were harder than expected, that is the finding rather than a failure. The research reports that self-reported mode use and behaviourally coded mode use of the same people correspond at close to zero. People are reliably poor at naming what they just did, which is why the classification is done from transcripts rather than from surveys.

What the modes cannot tell you is whether any of it was done well. Two people can both be in Production Assistant and end up with completely different artifacts. That difference is the principles, and what makes one of them easier to sustain than the other is the dispositions.