The theory behind the Ladder

A rung is a description of practice, not a feeling about technology.

This page is for the person on your team who is going to ask the hard questions before you sign anything. Curriculum director, assistant superintendent, whoever owns PD. It is longer than a homepage should be, on purpose.

What a rung actually is

Each of the five levels is written plainly enough that two different observers watching the same teacher would place her on the same rung. That is the entire design constraint, and everything else on this page follows from it.

It means a rung cannot be defined by attitude, confidence, enthusiasm or hours attended. Those are all real things and none of them are observable in a piece of work. A rung is defined by what shows up in the artifact: whether the person gave the tool context, whether they iterated, whether what they built is being used by anybody other than them.

It also means the ladder is about practice, not tools. A district that switches from one AI product to another does not reset its ladder. The capability is transferable, which is the difference between measuring your people and measuring your vendor.

Why capability is measured in stages, and why not by survey

Staged models are not new to schools, and using one is deliberate rather than clever. Districts have measured implementation in observable stages for decades. The best-established example is the Concerns-Based Adoption Model, developed at the University of Texas at Austin's R&D Center for Teacher Education through the 1970s and 80s and still used in school implementation work today. Its Levels of Use scale runs from Nonuse through Orientation, Preparation, Mechanical Use, Routine, Refinement, Integration and Renewal.

The detail worth borrowing is not the eight levels. It is how they are scored. Levels of Use is assessed through a certified structured interview and observation, explicitly not through a self-report questionnaire, because the people who built it found that practitioners are unreliable narrators of their own practice. Ask a staff of four hundred to rate their own AI skill and the distribution will tell you almost nothing. Everyone lands somewhere around a seven.

The Ladder compresses that logic to five levels written for AI practice specifically, and keeps the one principle that matters: placement comes from evidence a human reviewed. Every level, including the top two, is a description of what someone can do rather than a title someone holds.

Where it departs from CBAM is scope. CBAM measures adoption of one named innovation, so its top level is about refining that innovation. The Ladder measures a transferable capability, which is why the top of the measured range is about directing the tool rather than complying well with a rollout.

How a placement actually happens

A participant does a task from their own job, using AI, and submits the work along with what they asked for and what came back. A facilitator reads it against the rubric for that rung and either places them or names what is missing. Placement is attached to the artifact and the reviewer, permanently.

That means every number in a district dashboard opens. If a superintendent asks what a median of 2.7 means, the answer is not a methodology paragraph. It is a list of names, and behind each name a folder of work with dates and a reviewer on it.

It also means a placement can be challenged, which matters more than it sounds. A measurement nobody can appeal is a measurement teachers will not trust, and a measurement teachers do not trust produces gaming rather than growth.

Norming the Ladder

This is the part that decides whether your district data is real or decorative. A rubric on its own does not produce comparable numbers. Two facilitators at two sites, both sincere, both competent, will drift apart within a month. One is generous about what counts as iteration. The other wants to see three revisions before she believes it. Six months later the cabinet is comparing a 2.9 at one school to a 2.4 at another and the difference is the scorers, not the staff.

Districts already know how to solve this, because they already do it with writing. Calibration sessions before a district writing assessment, scorer training before an AP reading, the whole discipline of inter-rater reliability. Norming the Ladder is the same practice pointed at a different rubric, and school people recognise it immediately.

In practice it is a working session with your facilitators and instructional leaders. Everyone places the same anonymised samples independently, then the disagreements get surfaced and argued until the group converges and the rubric language gets sharpened where it failed. It is not a briefing. Nothing gets normed by being explained.

And it is not one session. Norming happens before a cohort starts, again at the midpoint, and whenever a new facilitator joins. Drift is not a sign that something went wrong, it is the default state of any human scoring system, and the fix is scheduled rather than heroic.

What you get out of it is bigger than clean data. Once a district has normed, the rungs become shared language. "She is a rung three" means one thing in a hiring conversation, a PD planning meeting, a coaching cycle and a board presentation. Most districts have no shared vocabulary for AI capability at all, which is why the conversation keeps restarting from zero in every meeting. The vocabulary is a deliverable, not a side effect.

The five rungs in full

All five are levels of capability. Every staff member is expected to reach three. Four and five describe what a smaller number of people go on to do with it, and schools usually attach a role to each.

01

Search Replacement

They ask AI the questions they used to type into Google. Faster answers, not much leverage yet, but the fear is gone. Most staff place here and they are right to. Rung one is not a criticism, it is a starting line, and a district where everyone is at rung one is in better shape than a district that does not know.
What it looks likeShe asks it to explain a district policy in plain English before a parent meeting.
02

Generator

They produce real work with it: family letters, rubrics, IEP drafts, agendas. Quality now depends on how they ask, and they have started to notice that. This is where most of the six-weeks-a-year time saving first shows up, and also where the first bad output gets sent to a family, which is why rung two is where guidance matters most.
What it looks likeHe drafts the field-trip letter with it, then rewrites the two sentences that still sound like a robot.
03

Director

They stop accepting the first output. Context, constraints, examples, iteration. They direct the tool instead of asking it, and the distinguishing behaviour is that they change something deliberately between attempts and can say what and why. This is where every staff member is expected to land. Not a stretch goal for the enthusiastic, the bar for the building.
What it looks likeShe pastes in last year's goals and the district template first, and the third pass is the one she actually sends.
04

Super-User

This rung is building. A rung-four educator stops solving the same problem twice and turns the solution into a system somebody else can run: a shared prompt library, a workflow the grade level uses every Monday, a template that outlives the person who made it. The distinguishing behaviour is portability. What she made works in hands that are not hers.

In a school that capability usually gets assigned as a department or grade-level lead. Most districts already have these people informally. Naming them is the district deciding to resource what it was getting for free.
What it looks likeHis grade level runs a shared planning prompt every Monday, he wrote it, and the district has given him release time to keep it working.
05

AI Champion

This rung is exponential automation. Rung four builds one good system. Rung five chains systems together until the saving stops being linear. The intake form writes the summary, which drafts the family letter, which lands in the right folder, and nobody touches it again. A rung-five educator gives time back to people who never asked her for anything.

In a school that capability usually gets assigned as a campus specialist, one per site, who trains and backs the department leads. That is what makes the structure hold without me in it.
Rung five is not a module anyone completes. It is a role a district creates and a person we train to hold it. That is the train-the-trainer layer, and it is how the whole thing stops depending on me.
What it looks likeShe runs the twenty-minute Tuesday clinic, and she keeps the prompt library the staff actually use.

Why cohorts, and why a facilitator

The format is not a preference. A review of 35 methodologically rigorous studies by the Learning Policy Institute identified seven elements shared by professional development that actually changes practice: it is content focused, incorporates active learning, supports collaboration, uses models of effective practice, provides coaching and expert support, offers feedback and reflection, and is of sustained duration.

Read that list against a one-day AI rollout and it satisfies roughly one item. Read it against this and you get the design directly: worked examples are models of effective practice, required artifacts are active learning, facilitator review is feedback and reflection, four to six weeks is sustained duration, and the cohort is collaboration.

Nothing here was invented to be different. It was assembled from the list of things that already work.

Why the champion is one of yours, and not me

The strongest argument for rung five is not one I made up. A meta-analysis of 60 causal studies of teacher coaching found pooled effects of 0.49 standard deviations on instruction and 0.18 on student achievement. Coaching is one of the better-evidenced interventions in the field.

The same review found something less comfortable and more useful: effects from large-scale effectiveness trials are only a fraction of those found in smaller efficacy trials. Coaching works, and it degrades the further it is stretched from the building.

That finding is the business model, stated honestly. The wrong move is to scale one external consultant thinner and thinner across fourteen sites until the thing that made it work is gone. The right move is to use the external engagement to build a coached, credible person inside each school, and then get out of the way. Rung five exists because the evidence says proximity is the active ingredient.

It is also, bluntly, the reason a district should trust this pitch. A vendor whose model requires them to stay forever will tell you that you need them forever.

The evidence for the problem itself

Three findings, none of them mine, that together describe why measurement is now the constraint rather than access or training volume.

48%
of districts had trained teachers on AI by fall 2024, up from 23 percent the year before. The training wave already happened.
RAND, American School District Panel, 2025
3 in 10
teachers now use AI at least weekly. Which means seven in ten do not, and no district can tell you which is which.
Gallup and Walton Family Foundation, 2025
5.9 hrs
saved per week by weekly users, roughly six weeks across a school year. That is the prize, and it is currently unevenly distributed by accident.
Gallup and Walton Family Foundation, 2025

There is a fourth number worth sitting with. In fall 2024, 67 percent of low-poverty districts had trained teachers on AI, against 39 percent of high-poverty districts. The capability gap between schools is not going to close on its own, and it will not close by buying more licenses.

References

Every claim on this page that is not mine is cited here. If your team wants to read the underlying work before a discovery session, this is the list.

1
More Districts Are Training Teachers on Artificial Intelligence. RAND Corporation, American School District Panel, 2025. rand.org
2
Teaching for Tomorrow. Gallup and the Walton Family Foundation, 2025. Survey of 2,232 US public K-12 teachers, 18 March to 11 April 2025. gallup.com
3
Concerns-Based Adoption Model, Levels of Use. Gene Hall, Shirley Hord and colleagues, University of Texas at Austin R&D Center for Teacher Education. sedl.org/cbam
4
The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence. Kraft, Blazar and Hogan, Review of Educational Research 88(4), 2018. annenberg.brown.edu
5
Effective Teacher Professional Development. Darling-Hammond, Hyler and Gardner, Learning Policy Institute, 2017. learningpolicyinstitute.org

Still have questions this page did not answer?

Good. Those are the ones worth having on a call. Thirty minutes, remote, no charge and no deck.

Book an intro call