Seven levels, what each one actually looks like, and the one move out of it.
← Control panelFour sentences on air. The dates are yours, off your own files.
“I interview salespeople for a living. When I started building training, I needed to know what level someone was actually at, because the same lesson is useless to two different people. So I built a scale, and I borrowed the shape from language. Then in August I stopped guessing and started running it on real interviews.”
The one line only you can write is why you started. Everything above is the sequence, not the reason.
Two old ideas, borrowed on purpose.
You already know this scale even if you have never heard its name. A1 is ordering coffee. B1 is holding a conversation. C2 is thinking in the language instead of translating in your head. That is CEFR, the Council of Europe framework, and every level in it is written as something you can do, never something you know.
Two things about it are worth stealing. It is not a test, it is the yardstick other people’s tests get measured against. And it does not give you one score, it profiles you across separate skills, so being strong in one and weak in another is a normal result rather than an error.
In 1980 two brothers at Berkeley studied how people get good at things, from chess players to pilots, and found the same arc every time. First you follow rules you have been given. Then you start recognising patterns and bending them. Then you stop consulting rules at all, and we call that intuition.
That is why the seven levels sit in three groups instead of being seven separate things to memorise. 0, A1 and A2 are following rules. B1 and B2 are seeing patterns. C1 and C2 are intuition.
So the scale measures one thing: how deep AI sits in the work you actually do. Not how much you know about AI, not how many tools you have opened, not how you feel about it.
Find yourself. The honest read is usually one level lower than the flattering one.
You have maybe opened it once. Work happens the way it always has. This is not ignorance, it is usually that nobody has ever shown you a use that was obviously yours, so it stayed a novelty you did not have ten minutes for.
You have had a couple of good moments and no habit. Every time you open it you start from scratch, because you have not found the two or three jobs that are clearly yours. It feels like a novelty, and you edit almost everything it gives you.
You found your two or three. Research, drafts, summaries. You are past wondering whether it is worth it, and you use it reliably for those jobs. But each one is a separate errand. Nothing connects, and the output of one thing never becomes the input to the next.
This is where most people are, and it is a comfortable place to stall. It looks like fluency from the outside because you use AI every day.
This is the big jump and it is the first level you cannot fake. You have prompts you wrote and kept, and you reuse them without thinking. There is a routine: research before the call, notes after it, follow-up inside the hour. You are visibly faster and more consistent than people around you doing the same job.
What you do not have yet is proof. You believe it is working, and you would struggle to say by how much.
You stopped improving by instinct. You run version one against version two, you keep the winner, and you can point at something real that moved. You are starting to know what will work before you test it, and you still test it.
This is the rarest rung on the whole ladder, and it is the one almost no framework in the world even asks about. Everybody’s scale stops one level below it.
Your system got good enough that it left your desk. A teammate uses your prompt. Someone asks you to set theirs up. You are not trying to lead anything, it just spread because it worked.
The difference from B2 is not skill, it is that you made it legible to somebody else. That is a communication act, not a technical one.
Work happens while you are not there. Things run, report back, and wait for your judgement rather than your labour. You have not automated yourself out of a job, you have moved to the twenty percent that actually needed you.
The tell that works at any level: a beginner describes something they have. Someone at the top describes something that happened, while they were not in the room.
This comes out of the interview transcripts, in their own words.
It is not that fluent people use fewer tools. It is that their tools stop being interchangeable.
Asked what is in their stack, people at the Adopter level list everything and mean it: “We use all of them. We use Google AI, we do use Copilot, we use Cloud, we use ChatGPT.” Another: “ChatGPT, Gemini, Perplexity, Claude and Meta, used interchangeably.”
The list is the answer. Five tools doing one job each, chosen by whichever tab is open.
At Practitioner the answer collapses to one place and goes down instead of across. “My favourite tool is Claude, and in Claude I actually build a skill, so it knows the website, the handouts, the one pagers.” Another built projects inside one tool for one part of the sales process at a time.
The narrowing you sensed is real, and it happens here. But one tool is not the same as fluency. Somebody in the same set said “I use Claude for everything” and scored Adopter, because the depth was not there. Depth is what counts, never the count.
At the top more than one model comes back, but never interchangeably. One Architect could not answer the prompt question without asking which model, because “every language model is very different”, and the first thing he described was connectors, things joined to each other.
So the real arc is: many tools used interchangeably, then one tool used deeply, then many tools connected deliberately. The middle is a narrowing. The top is not a return to shopping, it is plumbing.
Shares only. Never the counts.
When people rate themselves they are wrong, and they are wrong in one direction.
Read the shape, not the rows. The self-reported column is empty for the entire bottom half of the scale and spikes at the top. The assessed column is a hump in the middle. Nobody put themselves below Practitioner. More than half of assessed people landed there.
The same thing happens at scale. Five thousand knowledge workers were tested. 54 percent said they were proficient. 10 percent were. Our numbers are not a fluke, they are a small version of that.
And the mirror image. Nobody put themselves in the bottom three bands. More than half of the people an interviewer assessed landed there.
The generous read, and it is the true one. This is not people lying. On a quiz nobody is watching and nothing is at stake. In an interview a stranger decides whether you get the job, so the incentive to oversell is higher, and people still come out lower. We are not dishonest about this. We are wrong about it in private.