Time to Rethink the Test Pile
State and district assessment systems have grown larger and more complex over the last decade, yet districts that test the most do not see better results. So concludes a report published by my organization in 2026, in which my colleagues reviewed assessment policies across 42 states, audited local assessment systems in nearly 70 districts, and interviewed state and district leaders who designed and managed those systems.[1]
In several of the districts, students take as many as 88 standardized academic assessments before entering high school.[2] In the highest testing districts, students averaged seven such assessments per grade, with most administered one to three times per year. Roughly 60 percent of these assessments were selected and administered locally; 40 percent were required and administered by the state. So neither state overreach nor local excess is solely responsible for the accumulation of tests across the education system.
Neither state overreach nor local excess is solely responsible for the accumulation of tests across the education system.
English learners face an even heavier burden. Across kindergarten through eighth grade, English learners experienced an average of 47 additional hours of testing compared with their peers. This increase was driven largely by English language proficiency identification and annual testing requirements, along with additional screeners, diagnostics, and interim assessments. District leaders told researchers these measures were often required for compliance, even when the instructional value of the results was unclear or duplicative of data already collected.
Across kindergarten through eighth grade, English learners experienced an average of 47 additional hours of testing compared with their peers.
We examined the relationship between assessment volume and student outcomes and found no statistically significant relationship between testing volume and proficiency on state English language arts or math assessments. Nor was there a relationship between testing volume and growth in English or math proficiency.
Each additional assessment requires staff time for administration, data review, professional development and reporting, which often comes at the expense of instruction. The Achievement Network found that teachers are spending more than 40 hours per year reconciling and triangulating assessment data they often cannot use to inform instruction.[3]
Each additional assessment requires staff time for administration, data review, professional development and reporting, which often comes at the expense of instruction.
With so many tests yoked to so many purposes, state and district leaders also often lack a clear understanding of what each assessment is actually meant to do.
How and Why It Happens
Assessment systems rarely become bloated because of a single bad decision. They grow through layering. A new screener is added to comply with state literacy policy. A diagnostic is adopted to support Multi-Tiered Systems of Support. An interim assessment is introduced to monitor school progress or predict summative performance. Each tool serves a defensible purpose on its own. Together, they create crowded calendars and overlapping data streams that are difficult for educators to interpret or act on. “Every assessment has a rationale,” one district leader noted. “The problem is, no one ever asks which ones we should stop doing.”
State policy plays a role in this accumulation, often in ways that are easy to overlook. The study found that policies tied to screening requirements, intervention entry and exit rules, progress monitoring expectations, retention, and acceleration policies collectively shape large portions of local assessment systems, even when those policies are not explicit mandates.
State policy plays a role in this accumulation, often in ways that are easy to overlook.
Forty-two states now require universal reading screening, frequently paired with dyslexia-related mandates. Eighteen states have introduced math screening requirements, many within the last few years. Yet states vary widely in how these policies are designed. They differ in what gets measured, how often screening occurs, whether tools are consolidated or layered onto existing systems, and how prescriptive states are about tool selection.
Statewide interim assessments further complicate the picture. At least 20 states now offer statewide interim options, including through-year models that signal new expectations for monitoring progress. In some states, such as Indiana and Montana, these models are relevant to curriculum and have helped some districts replace lower quality commercial tools, reducing overall testing.[4] Elsewhere, statewide interims lack alignment with quality materials or offer poor reporting. Consequently, districts often retain existing assessments alongside state options, increasing burden instead of coherence.
At least 20 states now offer statewide interim options, including through-year models that signal new expectations for monitoring progress.
As assessment systems expand, so does the promise that each new tool will provide instructionally useful data. But the study found that many assessments are being used in ways that are inconsistent with their original design. Screeners intended for identification are treated as lesson planning tools. Interim benchmarks are interpreted as high-stakes indicators of instructional effectiveness. Progress monitoring tools are expected to diagnose the full range of student learning needs.
Vendors frequently market these tools as instructional supports, yet district leaders reported limited transparency about whether they help improve instruction. Technical documentation may demonstrate reliability or predictive validity but often offers little guidance on how results connect to curriculum, pacing, or specific instructional decisions. States, in turn, often approve or mandate tools without clearly defining evidence expectations for instructional use, leaving districts to interpret quality on their own.
States … often approve or mandate tools without clearly defining evidence expectations for instructional use, leaving districts to interpret quality on their own.
Teachers feel the consequences most directly, receiving data from multiple assessments that do not align with one another nor their curriculum. Results arrive on different timelines, using different benchmarks, and sending mixed signals about what mastery looks like. As one educator put it, “We’re drowning in data but still guessing about what to do next.”
What Can Be Done
First, assessment systems need explicit decision rules. State policymakers and state agency staff must clarify the purposes of each assessment to prevent tools from being misused as all-purpose instruments. Consequently, when they introduce a new assessment, districts should be required to identify which existing ones it replaces rather than piling on.
State policymakers and state agency staff must clarify the purposes of each assessment to prevent tools from being misused as all-purpose instruments.
Second, policymakers should strengthen evidence of expectations for instructional use. Assessments that claim to support instruction must meet higher standards, with clear evidence that results align to curricula, support grade-level learning, and inform specific decisions. Procurement and approval processes should reflect these standards so districts can accurately assess quality.
Finally, assessment policy must be aligned with instructional strategy. As investments in high-quality instructional materials grow, states must prioritize curriculum-aligned assessments and discourage tools that disrupt instructional coherence. While audits can identify excess, lasting change requires ongoing system design that integrates policy and evidence to ensure that assessment systems are intentionally built and properly timed.
Khaled Ismail is a principal at Education First, a mission-driven education consultancy that partners with funders and system leaders in all 50 states.
Notes
[1] Education First, “Rethinking the Test Pile,” report (February 2026).
[2] This number includes screeners, diagnostics, interims, and summative tests. It does not include end-of-unit, classroom, or teacher-created assessments, which would drive this number much higher.
[3] Achievement Network, “Missing Link,” white paper (2025).
[4] Tara Czupryk, Chawanna Chambers, and Senna Lamba, “Assessment to Action: Redesigning Indiana’s State Assessment for Instructional Impact,” report (Education First, N.d.); Aneesha Badrinarayan, Cedar Rose and Emma Fortier, “Moving toward Instructional Relevance: Recommendations for Through-Year Assessments Systems That Advance Teaching and Learning,” report (Education First, June 2026).
Also In this Issue
How State Board Leaders Can Build Next-Generation Accountability
By Morgan Scott Polikoff and Rachel AndersonCoherence in state policies plus aligned supports to districts pave the way forward.
Data for Improvement and Data for Accountability: Considerations for State Boards
By Elaine AllensworthCarefully designed accountability systems can spur schools to improve; others can undermine improvement and equity.
Making Room for What Matters: How State Boards Can Support Innovation in Accountability
By David Nitkin, Chelsea Waite, and Janette AvelarThe voices of school leaders who are most actively pushing the school improvement envelope can help shape the accountability conversation.
Measuring Academic Growth for School Accountability
By Christy HovanetzMore states should adopt a growth metric that compares students’ progress toward proficiency, not against other students.
How State Leaders Can Support Local Innovation and Improvement
By Jennifer Lin Russell, Donald J. Peurach, and Jennifer Zoltners ShererBy backing educational improvement networks, they can foster collaborative inquiry and experimentation.
i
i