By Lorah Neville, Executive Director, Marzano Evaluation Center
Originally presented at the MEC webinar, July 21, 2026.
Recommended reading time: 6 minutes.
I want to start with a story. You may have experienced something like this — I know I have in nearly 30 years of being in schools.
You’re walking down the hallway. You open the door to one classroom and the teacher has students working on a cognitively complex task. They’re really engaged. They’re debating each other, making claims, defending them with evidence. At the end of the year, that class makes at least a year’s growth — often more.
Then you open the door to the next room. This teacher is delivering direct instruction, possibly from the teacher edition. Students are working on a worksheet. There’s very little interaction, very little talking among them.
Both teachers, in most districts, are rated exactly the same: satisfactory.
And the research says we’re not very good at it.
This is the core question that TNTP’s Widget Effect research set out to answer. And the findings — first published in 2009 — have held up uncomfortably well in the years since.
What the Widget Effect Actually Found
The Widget Effect study reviewed evaluation data across 12 districts in four states, surveying more than 15,000 teachers and 1,300 administrators. The researchers were looking at how districts evaluate, develop, and retain teachers.
What they found gave the study its name: in practice, teachers had become interchangeable — like widgets. Not on paper, but in practice. Individual differences in effectiveness were acknowledged privately and ignored systematically.
The research identified five specific findings that describe how this happens:
Finding 1: Rating inflation is nearly universal.
In districts with binary rating systems (satisfactory/unsatisfactory), more than 99% of tenured teachers were rated satisfactory.
In systems with multiple rating levels, 94% still received a top-two rating. In some districts, not a single tenured teacher had ever been rated unsatisfactory.
Finding 2: Excellence goes unrecognized.
59% of teachers and 63% of administrators said their district was not doing enough to identify, promote, and retain top teachers.
None of the 12 districts studied factored teacher performance into retention or staffing decisions. When there are reductions in force, performance is rarely a factor. We tell ourselves excellence matters — and then don’t use it when it counts.
Finding 3: Struggling teachers receive no meaningful support.
73% of teachers went through a full evaluation cycle with no developmental areas identified.
Of those who did have areas flagged, fewer than half received useful support afterward. As Anthony Muhammad reminds us: support precedes accountability. How can we hold someone accountable if we haven’t ensured they have the tools and systems to succeed? (Muhammad, 2009)
Finding 4: Everyone already knows.
81% of administrators and 57% of teachers reported a tenured colleague delivering poor instruction. 43% of teachers said a tenured colleague should be dismissed. The performance problem isn’t hidden — it’s the open secret. And when we allow that to persist, it sends a message to everyone in the building about what we actually value.
Finding 5: Dismissal almost never happens.
Annual dismissal rates for tenured teachers ranged from 0% to 0.07%. 86% of administrators said they don’t always pursue dismissal even when they believe it is warranted. In my experience moving from the building to the district office, half of the attempts that ended without dismissal did so for procedural reasons — not substantive ones. The performance problem still existed. The procedure just wasn’t followed.
Why This Matters More Than Ever
I’m based in Arizona, where we have over 3,000 unfilled teaching positions. And I’d push back slightly on the framing of a “teacher shortage.” There are teachers. They just don’t want to teach anymore.
That changes the conversation. We can’t keep operating as if there’s a pipeline of qualified candidates ready to replace anyone who isn’t performing. We have to love the ones we’re with — which means actually developing them, recognizing the ones who are excelling, and giving the ones who are struggling a real pathway to improve.
The stakes are not abstract. Research is clear that when students have an ineffective teacher, it is very hard to recover. We work in enough other areas where we have limited control. Instruction is not one of them. We can control that. And if we’re not using our evaluation systems to do that, we’re leaving one of our most powerful levers untouched.
Ask yourself: if you mapped where your most highly effective teachers are against where your highest-need students are, what would that map look like? In too many districts — including ones I’ve led — the answer is uncomfortable. Our strongest teachers end up in AP and gifted classrooms, when the students who most need excellent teaching are sitting in freshman English and Algebra 1.
That’s not primarily a teacher problem. It’s a systems problem. And a systems problem requires a systems solution.
The Knowing-Doing Gap
When I asked leaders at our July webinar to describe their current evaluation system in one word, here is what came back: unreliable. Checklist. Misaligned. Inconsistent feedback. Too vast. Not frequent enough.
These leaders know what’s not working. Most of us who have been at this for a while do. The challenge isn’t awareness — it’s what Roger Schwarz calls “discussing the undiscussable”: surfacing the hard conversations that everyone is thinking and no one is having (Schwarz, 2013).
In our field, we tend to operate in “ready, fire, aim” mode. There’s real urgency — these students need us to be making things better today. But when we move that fast without solving the right problem, we spend enormous energy redoing work that didn’t fix the root cause. The most valuable thing I can offer is a framework for identifying where your system actually falls short before deciding what to fix first.
What a Better System Actually Looks Like
The Marzano evaluation models — the Focused Teacher Evaluation Model, the Focused School Leader Evaluation Model, and the District Leader Evaluation Model — were built as professional development models first. That origin matters.
Dr. Marzano’s work started with the question: what does effective instruction actually look like? The evaluation framework grew out of that answer, not the other way around. Around 2010, with Race to the Top requiring research-based evaluation, that professional development foundation became the structure for an evaluation system that could do both jobs: develop teachers and assess their effectiveness.
The model is also grounded in what the research tells us about how people learn. As the National Academies of Sciences, Engineering, and Medicine note in How People Learn II, learners need to be engaged in metacognition, have clear targets, and know what they’re expected to be able to do — all of which are built into the Marzano model (National Academies of Sciences, Engineering, and Medicine, 2018).
Here is what sets the Marzano model apart from most evaluation systems: simply doing what you’re supposed to do gets you to the developing level. To reach applying or effective, you have to demonstrate impact on a majority of students.
This shifts the entire conversation. Instead of “I got a 2” or “I got a 3,” the post-observation conversation becomes: Why did you make that instructional choice? How did it help? Did it go the way you expected? Which students got it, and which didn’t? That is a conversation about instruction. That is a conversation that leads somewhere.
The Research We’re Seeing Now
We’ll be releasing new research soon, and I want to share a preview of what it shows. We are seeing a 0.55 correlation between teachers evaluated using the Focused Teacher Evaluation Model’s 23 elements and their students’ ELA achievement — and a 0.52 correlation with math achievement.
That shouldn’t be surprising. When every observation is tied to student evidence, and every rating reflects how students responded to instruction, of course performance on district and state assessments follows. The lesson-by-lesson data and the annual data are measuring the same thing.
I don’t know of another evaluation system that can demonstrate that connection. And that connection is the point: evaluation should not be a separate process from the work of improving instruction. It should be the same process.
Where to Start
At the end of the webinar, I asked leaders to reflect on one question: where does your current system fall short? Not all five findings at once. Just one. The highest-leverage place to start.
For many, it was Finding 3 — the lack of meaningful support for struggling teachers. Support precedes accountability. If we don’t have a clear process for helping a teacher improve, we don’t have a basis for saying they can’t.
For others, it was Finding 1 — the disconnect between evaluation ratings and student outcomes. If 95% of your teachers are rated effective or highly effective, but only 60% of your students are proficient, that gap is data. It is telling you something about your evaluation system’s design, not about your teachers.
Wherever you start, the goal is the same: an evaluation system that can tell you something true and useful about what is happening in classrooms — and that connects what evaluators observe to what teachers actually need to grow.
We can’t change it if we can’t see it. The first step is building a system that lets us see clearly.
Wondering where your system stands?
The District Evaluation Diagnostic walks you through 18 questions across six areas of your evaluation system. It takes about 10 minutes and gives you a clear picture of where you are — and where the gaps are costing you.
Take the District Evaluation Diagnostic →
Or reach Lorah directly at [email protected] or 480-535-2868 to schedule a no-obligation conversation.
References
Muhammad, A. (2009). Transforming school culture: How to overcome staff division. Solution Tree Press.
National Academies of Sciences, Engineering, and Medicine. (2018). How people learn II: Learners, contexts, and cultures. National Academies Press. https://doi.org/10.17226/24783
Schwarz, R. (2013). Smart leaders, smarter teams: How you and your team get unstuck. Jossey-Bass.
TNTP. (2009). The widget effect: Our national failure to acknowledge and act on differences in teacher effectiveness. https://tntp.org/publications/view/the-widget-effect-failure-to-act-on-differences-in-teacher-effectiveness
About the Author
Lorah Neville is the Executive Director of the Marzano Evaluation Center. She brings more than two decades of K–12 experience as a teacher, principal, director of curriculum, and superintendent. She works directly with district and school leaders nationwide to implement the Marzano evaluation models and build evaluation systems that connect to meaningful instructional improvement. She holds a degree in Elementary Education and a Master’s in Educational Leadership, and has served as adjunct faculty at Northern Arizona University and Arizona State University.


