This page is for the person on your team who is going to ask the hard questions before you sign anything. Curriculum director, assistant superintendent, whoever owns PD. It is longer than a homepage should be, on purpose.
Each of the five levels is written plainly enough that two different observers watching the same teacher would place her on the same rung. That is the entire design constraint, and everything else on this page follows from it.
It means a rung cannot be defined by attitude, confidence, enthusiasm or hours attended. Those are all real things and none of them are observable in a piece of work. A rung is defined by what shows up in the artifact: whether the person gave the tool context, whether they iterated, whether what they built is being used by anybody other than them.
It also means the ladder is about practice, not tools. A district that switches from one AI product to another does not reset its ladder. The capability is transferable, which is the difference between measuring your people and measuring your vendor.
Staged models are not new to schools, and using one is deliberate rather than clever. Districts have measured implementation in observable stages for decades. The best-established example is the Concerns-Based Adoption Model, developed at the University of Texas at Austin's R&D Center for Teacher Education through the 1970s and 80s and still used in school implementation work today. Its Levels of Use scale runs from Nonuse through Orientation, Preparation, Mechanical Use, Routine, Refinement, Integration and Renewal.
The detail worth borrowing is not the eight levels. It is how they are scored. Levels of Use is assessed through a certified structured interview and observation, explicitly not through a self-report questionnaire, because the people who built it found that practitioners are unreliable narrators of their own practice. Ask a staff of four hundred to rate their own AI skill and the distribution will tell you almost nothing. Everyone lands somewhere around a seven.
The Ladder compresses that logic to five levels written for AI practice specifically, and keeps the one principle that matters: placement comes from evidence a human reviewed. Every level, including the top two, is a description of what someone can do rather than a title someone holds.
Where it departs from CBAM is scope. CBAM measures adoption of one named innovation, so its top level is about refining that innovation. The Ladder measures a transferable capability, which is why the top of the measured range is about directing the tool rather than complying well with a rollout.
A participant does a task from their own job, using AI, and submits the work along with what they asked for and what came back. A facilitator reads it against the rubric for that rung and either places them or names what is missing. Placement is attached to the artifact and the reviewer, permanently.
That means every number in a district dashboard opens. If a superintendent asks what a median of 2.7 means, the answer is not a methodology paragraph. It is a list of names, and behind each name a folder of work with dates and a reviewer on it.
It also means a placement can be challenged, which matters more than it sounds. A measurement nobody can appeal is a measurement teachers will not trust, and a measurement teachers do not trust produces gaming rather than growth.
This is the part that decides whether your district data is real or decorative. A rubric on its own does not produce comparable numbers. Two facilitators at two sites, both sincere, both competent, will drift apart within a month. One is generous about what counts as iteration. The other wants to see three revisions before she believes it. Six months later the cabinet is comparing a 2.9 at one school to a 2.4 at another and the difference is the scorers, not the staff.
Districts already know how to solve this, because they already do it with writing. Calibration sessions before a district writing assessment, scorer training before an AP reading, the whole discipline of inter-rater reliability. Norming the Ladder is the same practice pointed at a different rubric, and school people recognise it immediately.
In practice it is a working session with your facilitators and instructional leaders. Everyone places the same anonymised samples independently, then the disagreements get surfaced and argued until the group converges and the rubric language gets sharpened where it failed. It is not a briefing. Nothing gets normed by being explained.
And it is not one session. Norming happens before a cohort starts, again at the midpoint, and whenever a new facilitator joins. Drift is not a sign that something went wrong, it is the default state of any human scoring system, and the fix is scheduled rather than heroic.
What you get out of it is bigger than clean data. Once a district has normed, the rungs become shared language. "She is a rung three" means one thing in a hiring conversation, a PD planning meeting, a coaching cycle and a board presentation. Most districts have no shared vocabulary for AI capability at all, which is why the conversation keeps restarting from zero in every meeting. The vocabulary is a deliverable, not a side effect.
All five are levels of capability. Every staff member is expected to reach three. Four and five describe what a smaller number of people go on to do with it, and schools usually attach a role to each.
The format is not a preference. A review of 35 methodologically rigorous studies by the Learning Policy Institute identified seven elements shared by professional development that actually changes practice: it is content focused, incorporates active learning, supports collaboration, uses models of effective practice, provides coaching and expert support, offers feedback and reflection, and is of sustained duration.
Read that list against a one-day AI rollout and it satisfies roughly one item. Read it against this and you get the design directly: worked examples are models of effective practice, required artifacts are active learning, facilitator review is feedback and reflection, four to six weeks is sustained duration, and the cohort is collaboration.
Nothing here was invented to be different. It was assembled from the list of things that already work.
The strongest argument for rung five is not one I made up. A meta-analysis of 60 causal studies of teacher coaching found pooled effects of 0.49 standard deviations on instruction and 0.18 on student achievement. Coaching is one of the better-evidenced interventions in the field.
The same review found something less comfortable and more useful: effects from large-scale effectiveness trials are only a fraction of those found in smaller efficacy trials. Coaching works, and it degrades the further it is stretched from the building.
That finding is the business model, stated honestly. The wrong move is to scale one external consultant thinner and thinner across fourteen sites until the thing that made it work is gone. The right move is to use the external engagement to build a coached, credible person inside each school, and then get out of the way. Rung five exists because the evidence says proximity is the active ingredient.
It is also, bluntly, the reason a district should trust this pitch. A vendor whose model requires them to stay forever will tell you that you need them forever.
Three findings, none of them mine, that together describe why measurement is now the constraint rather than access or training volume.
There is a fourth number worth sitting with. In fall 2024, 67 percent of low-poverty districts had trained teachers on AI, against 39 percent of high-poverty districts. The capability gap between schools is not going to close on its own, and it will not close by buying more licenses.
Every claim on this page that is not mine is cited here. If your team wants to read the underlying work before a discovery session, this is the list.
Good. Those are the ones worth having on a call. Thirty minutes, remote, no charge and no deck.