Making Room for What Matters: How State Boards Can Support Innovation in Accountability
The voices of school leaders who are most actively pushing the school improvement envelope can help shape the accountability conversation.
Evidence abounds that schools are not meeting students’ needs. Nearly one in four is chronically absent, a rate still well above pre-pandemic levels.[1] Most describe themselves as disengaged from school—coasting, complying, or checked out—even when parents assume otherwise.[2] How students experience school predicts their grades, test scores, attendance, and behavior.[3] State boards of education can help ensure that the fixes to school systems are more than incremental.
Through a research collaboration called the Canopy Project, our team has been studying schools seeking to make learning purposeful and relevant, build durable skills alongside academic mastery, connect students to postsecondary pathways that match their aspirations, and design for students who are not well-served by traditional models of schooling.[4] Canopy schools are explicitly reimagining the purpose, structure, and experiences of school and are nominated to the project by experts familiar with the schools’ work. The sample is not representative of all public schools and is too small to disaggregate by state. However, the benefit of our deliberately focused sample is that it includes the educators most actively testing the limits of what current systems allow.
Canopy schools are explicitly reimagining the purpose, structure, and experiences of school and are nominated to the project by experts familiar with the schools’ work.
Across our annual surveys of hundreds of innovative school leaders, accountability has surfaced again and again as a key policy factor that affects their schools (figure 1). That finding led us to investigate in greater detail how innovative school leaders experience state accountability policy. Our 2025 analysis draws upon survey responses from 186 leaders, including 147 public school leaders of charters and traditional district schools, across 41 states, plus interviews with eight of those leaders.[5]

What Innovative School Leaders Say
Public district and charter school leaders are not simply “for or against” accountability. They neither want to eliminate it nor maintain the status quo. They describe accountability systems as doing some jobs well while also undermining other of their aims. That paradox points to exactly where state boards can act—protecting what accountability gets right while changing what holds innovation back.
Accountability policies help reveal performance gaps and focus attention on core academics. Some schools find this helpful. Sixty percent of leaders agreed that state accountability systems help identify which student groups in their state need more support. Leaders described using accountability data to run annual needs assessments in reading and math, reallocate funds toward math coaching, and look for patterns in student outcomes in the disaggregated data. One leader noted that the oversight function matters precisely because it catches failure a school might otherwise rationalize away. If the school disproportionately suspended students with disabilities or students of color, the leader said, “we would deserve to be flagged.” High-performing schools valued the comparability for a different reason: They can use an “apples-to-apples” comparison against every public school in the state to grow enrollment and build trust with existing families.
Sixty percent of leaders agreed that state accountability systems help identify which student groups in their state need more support.
Comparable, public, disaggregated measures of academic performance are a hard-won equity guardrail. Most school leaders recognize and value their importance. State accountability systems make disparities visible and make it harder for any school, district, or state to look away when students are not getting the support they need to succeed.
Accountability systems do not capture school quality, highlight which schools to learn from, or help them improve student outcomes. Some of the same leaders were skeptical that current accountability systems capture what makes a school good or surfaces schools worth learning from. About a third agreed that state ratings show which schools are serving students well, a third disagreed, and another third neither agreed nor disagreed. When asked what information they most value in judging their own schools, leaders ranked standardized test scores below performance assessments and authentic demonstrations of learning, family feedback, attendance, and school climate (figure 2). They do not dismiss test scores; they simply see them as part of a bigger picture.

Skepticism is particularly apparent among leaders of schools designed to serve specific student populations, such as transfer students, students with severe learning differences, or students new to the US. They describe feeling caught between their mission and the structure of the accountability system: “To align with existing school grading metrics, we would have to shift our focus away from the very population we were created to support,” one said. Our findings align with a separate national survey from 2024 in which two-thirds of school leaders doubted that tests accurately measure ability for students with disabilities and English learners.[6]
Skepticism is particularly apparent among leaders of schools designed to serve specific student populations.
Nor did many leaders find accountability data useful for improving outcomes. Fewer than a third said accountability helps them improve student results, and even fewer said it helps them improve outcomes for students with disabilities and English learners.
Timing compounds the problem. Results arrive too late to inform instruction, and score reports are often inscrutable or disconnected from instruction. The assessments that leaders do care about and rely on tend to be implemented locally and more closely connected to their curricula and instructional materials.
The assessments that leaders do care about and rely on tend to be implemented locally and more closely connected to their curricula and instructional materials.
Some leaders said that accountability measures capture student composition rather than growth: One leader of a school serving newly arrived immigrants was placed in corrective action after the state compared its newcomers’ scores with the statewide average for all Latine students. As the school’s newcomer population declined, its rating improved even though instruction had not changed, rendering the data useless for real improvement.
Accountability systems make it harder to pilot new approaches and personalize learning, especially in high schools. Canopy leaders described three ways that current accountability systems made it harder for them to pilot and scale new learning models: one-size-fits-all pacing, a narrow focus on foundational academics, and administrative burden.
First, the system assumes all students move at a single pace. Federal rules require students to be tested on the content of their enrolled grade, regardless of the grade level at which they are actually performing. A student enrolled in fifth grade but working on third-grade skills still sits for the fifth-grade test. The tension is not the expectation that all students should ultimately reach grade-level standards. But for competency- or mastery-based schools serving students with widely varying academic histories and needs, one leader said that the requirement for grade-level testing incentivizes schools to “jump around between [students’] skill level and their enrolled grade level” rather than tailoring a personalized acceleration plan for the student. The school’s goal is the same as the motivation for the federal rules: prepare every student to succeed on grade-level content or beyond. But accountability systems that assume students will reach grade-level expectations through the same pace and sequence can become a straitjacket.
Accountability systems that assume students will reach grade-level expectations through the same pace and sequence can become a straitjacket.
One high school in our study, built for rolling, self-paced course completion, ran into the same wall with fixed state testing windows. The grade-based “grammar of schooling” that many Canopy schools are trying to disrupt, like standardized pacing and age-based benchmarks, can prevent leaders from implementing personalized models that they believe will accelerate student learning toward and beyond grade-level standards.
Second, the narrow focus on reading and math crowds out everything else. Leaders described a constant pull toward covering tested standards at the expense of depth, application across disciplines, and the durable skills they believe matter most. Nontested subjects get squeezed. One language immersion school that already performs well on state tests decided to cut instruction time in the immersion language rather than risk its math and reading scores.
Leaders described a constant pull toward covering tested standards at the expense of depth, application across disciplines, and the durable skills they believe matter most.
The mismatch is sharpest in high schools, where school leaders report that postsecondary success, career and technical credentials, and long-term results are largely invisible in state ratings, even though many of the schools in our study are explicitly aiming for these outcomes. One leader said the school could move college matriculation to 100 percent and barely budge its rating “because it all comes back to state testing.” Another said the school earned “no credit” for students who graduate with an associate degree already in hand.
High school leaders in our sample were four times more likely than their elementary and middle school peers to say accountability hinders their work and to report that they rely far more on long-term outcomes data to judge their success (figure 3). Even though federal law gives high schools more flexibility than elementary or middle schools, leaders suggested that the gap between what accountability systems measure and what high schools are trying to accomplish dwarfs that flexibility.

Third, the administrative burden is heavy and underappreciated. Leaders estimated that state testing consumes roughly two weeks of learning time and far more adult time. The burden often lands hardest on administrators and special educators. During testing windows, their time spent on logistics, data entry, and proctoring comes at the expense of working with students. As one school leader said,
For the two-month window of testing, every single special educator, 75 percent of their job is just test administration, test logistics, [and] data entry. All the [testing] accommodations have to be listed in the state platform, so now I’m pulling my valuable special education teachers out of working with kids to input [data] for a minimum of thirty minutes per kid…. A lot of our [students with disabilities] are doing one-on-one testing, so when teachers are doing a four-hour test with them, that means the other 12 kids on their caseload are getting absolutely no support during the testing window.
How State Boards Can Respond
Only 10 percent of leaders supported eliminating accountability, and only 5 percent wanted to keep it exactly as is (figure 4). On the whole, school leaders prefer a clear middle ground: They do not want to eliminate accountability, but they would like it to be right-sized and revised. State boards thus have support and room for flexibility and change without abandoning the system altogether.

We see four particularly promising avenues for state boards:
- defend a statewide role for academic accountability;
- broaden and tailor measures of success;
- prioritize a lighter footprint over more testing, even if additional testing is instructionally useful; and
- prioritize research and development on new forms of assessment and accountability.
Defend a statewide role for academic accountability. Boards must hold the line on comparable, statewide academic performance data and publicly released, disaggregated results. Equity guardrails are worth protecting, and boards should say so plainly even as they push for change. They can also use their authority over performance standards to raise the bar where it has slipped: Many states have set proficiency cut scores well below what national benchmarks like the National Assessment of Educational Progress (NAEP) treat as meaningful readiness. Recent work in Virginia demonstrates how states can transparently recalibrate expectations rather than quietly lowering them.[7] Where data arrives too late to be useful, boards can press their chief and agency to release it faster.[8]
Boards must hold the line on comparable, statewide academic performance data and publicly released, disaggregated results.
Broaden and tailor measures of success. Leaders want systems that balance test data with other information about learning opportunities, experiences, and outcomes. Strong majorities favor including access to learning opportunities, school climate data, and local or alternative assessments in accountability frameworks, as well as reducing the weight of standardized tests (figure 5). Crucially, they call for reducing the weight of standardized tests, not eliminating those tests entirely. Preserving standardized exams, but with a lighter footprint, can fulfill the need for comparable data across districts, while also incorporating local assessment data that Canopy school leaders find more valuable.

Boards can encourage their agencies to incorporate well-chosen measures of student experience and engagement, which offer near-term, actionable signals. In one analysis, students who reported stronger school experiences had markedly higher GPAs and test scores and far lower chronic absence.[9]
Boards can encourage their agencies to incorporate well-chosen measures of student experience and engagement, which offer near-term, actionable signals.
Boards can also create room for specialized schools to be held to mission-aligned standards. Massachusetts did this for Map Academy, an alternative charter school that reengages students after long absences from school. By conventional metrics, including a daily attendance rate near 45 percent, Map Academy might appear like a school perpetually in need of intervention. The school worked with the state to build a framework that treats engagement as a prerequisite for academic progress, setting differentiated goals for students at different engagement levels rather than holding all students to the same year-end bar. Massachusetts later adopted elements of the model as its standard for alternative charter schools. Map Academy’s co-founder believes any school could negotiate mission-aligned metrics with the state but notes it took philanthropic funding, consultant support, and years to do so. State boards can create policies that allow specialized schools to develop differentiated, mission-aligned assessment frameworks without years of negotiation and philanthropic subsidy.
Boards can also modernize graduation requirements and content standards, balancing rigor with flexibility. For example, Indiana’s class of 2029 and beyond will earn credits under the state’s new diploma, which offers three “readiness seals” customizable to students’ interests in pursuing higher education, career pathways, and military service. New York recently updated its graduation standards to require that students demonstrate mastery of the newly adopted NYS Portrait of a Graduate and is developing a statewide competency-based diploma that aligns with that portrait. As states move toward more flexible systems, boards must ensure that clear guidance and equity guardrails prevent rigor from slipping during the transition. Standards should stay consistently high—especially when states, like New York, are fundamentally restructuring how that rigor is measured.
Boards can also modernize graduation requirements and content standards, balancing rigor with flexibility.
Where federal rules are the binding constraint, boards can press their agencies to pursue waivers or explore applications for the Innovative Assessment Demonstration Authority (IADA).[10] Based on the authors’ current knowledge, few states are currently applying for IADA and related waivers, which is a missed opportunity for assessment and accountability innovation.
Prioritize a lighter footprint over more testing, even if additional testing is instructionally useful. It is commonly theorized that the fix for low-value testing is more frequent, through-year testing that provides teachers with timely data. But leaders—including those who strongly support state testing—told us they would rather test less often, not more, and were skeptical that additional testing would produce data they would actually use (figure 6).

Before endorsing mandatory state-led, through-year, or interim assessments, boards can ask their agencies pointed questions: What is the real administrative burden? Will educators use the resulting data if they already have formative data from their own local assessments? Does more frequent testing align with our learning goals? Evidence suggests that less-frequent testing can meet state oversight needs while freeing up schools to pursue richer, performance-based work[11]—and federal waiver and IADA pathways exist to test such approaches.
Less-frequent testing can meet state oversight needs while freeing up schools to pursue richer, performance-based work.
State-provided assessments are also likely to be duplicative with assessments that districts and schools are already using: A recent national study found that districts administer as many as 88 distinct assessments across grades K–8, with little evidence that the sheer volume connects to better outcomes.[12] For the many districts that already have strong formative assessment systems in place, , there is little reason to think that additional assessments from the state will provide instructional value that outweighs their logistical burden. In contrast, districts without coherent assessment systems may be likely to layer state-created interims on top of an already-overwhelming and disjointed assessment mandate.
Prioritize research and development on new forms of assessment and accountability. Today’s tests were built to do an essential job: provide system-level visibility into student academic performance. What they cannot do is simultaneously provide low-burden statewide oversight and capture the durable skills, complex reasoning, and real-world application that schools increasingly aim to develop. That is not a political problem, it is a technical one, and it has a technical solution: new assessment tools built through serious, sustained R&D that can do both.
New AI tools offer promise here. They may eventually be able to assess complex competencies through portfolios, performance tasks, and even passive observation of authentic learning at costs and scales that were not previously feasible. But that future will not arrive on its own. It will require serious, sustained investment to develop, test, validate, and govern new assessment tools before they are ready to carry public stakes.
It will require serious, sustained investment to develop, test, validate, and govern new assessment tools before they are ready to carry public stakes.
Boards cannot run that R&D on their own, but they can demand it and resource it. The key is treating assessment R&D as public infrastructure rather than a grant-funded add-on. It will require dedicated staffing; clear public research questions around comparability, bias, validity, and human-AI balance; and a commitment that results from pilots actually shape future statewide policy.
Some states are already moving toward this future. North Carolina established a standing R&D function within its education agency and is partnering with ETS to pilot AI-enabled assessments of durable skills like critical thinking and collaboration using federal money from the Competitive Grants for State Assessment. Kentucky’s local laboratories of learning are generating evidence designed to inform future statewide approaches rather than simply creating local innovations that do not scale. Organizations like ETS, the Carnegie Foundation for the Advancement of Teaching, LearnerStudio, the Study Group, AERDF, Learning Data Insights, and Transcend are running pilots and building partnerships to advance the future of assessment. Boards can push their agencies to follow that posture and can use their authority to ask the pointed questions that keep that work honest. The recently published State R&D Playbook from the Alliance for Learning Innovation, Transcend, and Education Reimagined provides additional examples of what this could look like.[13]
A Moment of Opportunity
State boards have real influence on whether and how schools succeed in developing new approaches to learning. Through the standards they set, the assessment and accountability systems they oversee, and their public communication, state boards create conditions that either let these new models thrive or quietly undermine them. As the federal role in education recedes, the leverage state boards can wield is only growing.
State boards have real influence on whether and how schools succeed in developing new approaches to learning.
The school leaders in the Canopy Project are not asking to be let off the hook. They want accountability to work better, not disappear. They are building schools that take what students need seriously, but they are operating within systems that were designed for a different vision of what school is supposed to be. That is a solvable problem, but it requires political will. It requires holding equity guardrails firm, taking the feedback of practitioners closest to the work seriously, and using the authority boards actually have over standards, systems, and the questions they ask in public to build something better. The appetite for that leadership exists. The question is whether boards will satisfy it.
David Nitkin is a managing partner at Transcend Education, Chelsea Waite is research principal at the Center on Reinventing Public Education, and Janette Avelar is a research fellow for the Canopy Project and a researcher at the Multilingual Learning Research Center.
Notes
[1] Nat Malkus, “Lingering Absence in Public Schools: Tracking Post-Pandemic Chronic Absenteeism into 2024,” report (Washington, DC: American Enterprise Institute, June 2025).
[2] Rebecca Winthrop, Youssef Shoukry, and David Nitkin, “The Disengagement Gap: Why Student Engagement Isn’t What Parents Expect,” report (Washington, DC: Center for Universal Education, Brookings Institution, January 2025).
[3] Transcend Education, “The Relationship between Student Experiences and Outcomes,” research brief (Transcend, September 2025).
[4] The Canopy Project is a national effort co-led by the Center on Reinventing Public Education and Transcend that identifies schools that are building more empowering, engaging, effective learning environments. Organizations with expertise in school transformation nominate schools to participate in the project, and those schools’ leaders complete an annual survey.
[5] Chelsea Waite, David Nitkin, and Janette Avelar, “Making Room for What Matters: Innovative School Leaders Want Accountability but with a Lighter Footprint,” report (Center on Reinventing Public Education and Transcend, October 2025).
[6] Naaz Modan, “92% of School Leaders Concerned about Academic Recovery, NCES Survey Says,” K-12 Dive, April 16, 2024.
[7] On states setting proficiency cut scores below national benchmarks, see “Honesty Gap,” web page (AssessmentHQ, 2024). On Virginia’s recent work, see Graham Moomaw, “Youngkin’s Education Board Moves to Toughen Proficiency Standards,” Virginia Mercury, September 26, 2025.
[8] Chad Aldeman, “It’s Officially Fall. Have You Seen Your Child’s Spring Test Results?” EduProgress blog, September 2024.
[9] Transcend Education, “Better Experience, Better Outcomes,” Beyond Fine website (2025).
[10] The US Department of Education issued guidance in July 2025 on how states may request waivers, including through IADA. KnowledgeWorks offered suggestions for how states and districts can use waivers to pilot innovative assessments. Lillian Pace, “Six Targeted Federal Waiver Ideas for Advancing Student-Centered Learning,” article (KnowledgeWorks, 2025).
[11] Education First, “Rethinking the Test Pile: A National Study of K-8 Academic Assessments,” national study (2026); Transcend Education, “The Future of Assessment,” paper (2023).
[12] Education First, “Rethinking the Test Pile.”
[13] Alliance for Learning Innovation, Education Reimagined, and Transcend, “State Education R&D Playbook,” website.
Also In this Issue
How State Board Leaders Can Build Next-Generation Accountability
By Morgan Scott Polikoff and Rachel AndersonCoherence in state policies plus aligned supports to districts pave the way forward.
Data for Improvement and Data for Accountability: Considerations for State Boards
By Elaine AllensworthCarefully designed accountability systems can spur schools to improve; others can undermine improvement and equity.
Making Room for What Matters: How State Boards Can Support Innovation in Accountability
By David Nitkin, Chelsea Waite, and Janette AvelarThe voices of school leaders who are most actively pushing the school improvement envelope can help shape the accountability conversation.
Measuring Academic Growth for School Accountability
By Christy HovanetzMore states should adopt a growth metric that compares students’ progress toward proficiency, not against other students.
How State Leaders Can Support Local Innovation and Improvement
By Jennifer Lin Russell, Donald J. Peurach, and Jennifer Zoltners ShererBy backing educational improvement networks, they can foster collaborative inquiry and experimentation.
i
i