Form ED‑1thesolutionsite.com

One test is designed to be compared and one is not

NAEP samples nationally on a stable framework, which is why it can be tracked over decades where state tests cannot.

A grid of blank answer bubbles, macro
NAEP samples nationally on a stable framework, which is why it can be tracked over decades where state tests cannot.

American public education produces two kinds of test results that look similar on a spreadsheet but are built for entirely different purposes. One is designed to measure the same thing, the same way, year after year, so that a reading score in 2024 can be set honestly beside a reading score from 1992. The other is designed to certify that a particular state's students have met that state's particular standards — a job that makes cross-state and cross-decade comparison structurally difficult, sometimes impossible. Understanding why the two types exist, and why they behave so differently, requires following each back to the office that commissioned it.

How the two systems differFrom the page
NAEPfederal sample survey; no stakes for individual students; stable cross-decade framework; administered by NCES under an independent governing board
State assessmentsannual census tests; results used for school ratings and sometimes student promotion; keyed to each state's own standards; proficiency thresholds set by state legislatures
Matrix samplingNAEP technique in which each student answers only a subset of items; population picture is reconstructed statistically, allowing more total items than any one student could complete
ESSA (2015)the current federal authorisation requiring annual state testing; sets the mandate but not the content or cut scores
Proficiency gapthe documented divergence between what each state labels "proficient" and where that performance falls on the NAEP scale

The measure that was never meant to judge

The National Assessment of Educational Progress — NAEP — began in 1969, administered then under contract for the federal government. The Education Commission of the States ran the early programme; from 1988 onward, NAEP has operated under statutory mandate with the National Center for Education Statistics (NCES) holding responsibility and an independent governing board — the National Assessment Governing Board, created by Congress — setting framework and policy. That independence matters more than it might appear: the Governing Board is deliberately insulated from the US Department of Education's political leadership so that no administration can quietly adjust what the test rewards.

NAEP does not test every student. It is a sample survey, drawing from public and private schools across all fifty states and the District of Columbia, at grades four, eight and — less regularly — twelve. Because it samples rather than censuses, NAEP can use a wider matrix of questions than any single student ever sees. Psychometricians call this matrix sampling or booklet-spiralling: different students get different clusters of items, and statistical methods reconstruct a population-level picture from the combined responses. This approach makes the test administratively lean and — crucially — allows the framework to stay stable even as individual items are retired and replaced. No stakes attach to any individual student's performance because no individual student completes a scoreable set.

ChronologyFrom the page
  1. 1969NAEP first administered
  2. 1988Congress gives NAEP a statutory mandate and creates the National Assessment Governing Board
  3. 1992start of the NAEP trend line most commonly cited for reading and mathematics comparisons
  4. 2001No Child Left Behind makes annual state testing a federal condition of funding
  5. 2007Fordham Institute analysis documents wide variation in state proficiency cut scores relative to NAEP
  6. 2015Every Student Succeeds Act replaces NCLB; testing mandate retained, federal control of cut scores relinquished
  7. 2022post-pandemic NAEP results document historically significant drop in fourth-grade reading scores

The consequence of no stakes is freedom. Because NAEP results carry no consequence for a school's accreditation, a district's funding, or a student's graduation, there is no structural pressure on states to teach to it, on administrators to exempt low-scoring students from the sample, or on anyone to manipulate the score. That insulation is the source of NAEP's authority as a long-run yardstick. Reading and mathematics frameworks have been sufficiently stable that researchers can track national and state-level trends from the early 1990s to the present. The pandemic's documented drop in fourth-grade reading scores, widely reported after the 2022 results, was legible precisely because there was an unbroken baseline against which to measure it.

The measures that were meant to certify

State assessments work from the opposite logic. A state legislature mandates that students in grades three through eight — and once in high school — be tested against the state's own academic standards. Those standards differ across states. California's mathematics framework, Mississippi's English language arts standards, and Massachusetts's science standards are each the product of a distinct policy process, adopted in different years and revised on different schedules. A test keyed to California's standards cannot meaningfully report a Mississippi student's performance, because the two tests are not measuring the same thing by design.

A wall map of the world, faded, in a plain room

The federal requirement that states test annually in those grade bands comes from the Elementary and Secondary Education Act, reauthorised most recently as the Every Student Succeeds Act in 2015. ESSA requires annual testing and the public reporting of results disaggregated by subgroup — race and ethnicity, income, disability status, English-learner status — but it leaves the content of the standards and the cut scores that define proficiency entirely to the states. The result is that "proficiency" means something different in each of the fifty states. A 2007 analysis by the Thomas B. Fordham Institute, comparing state proficiency cut scores against the NAEP scale, found enormous variation: a student who would be labelled proficient on one state's assessment might fall well below proficiency on another's when the same performance was evaluated against NAEP's framework. That gap between state-labelled proficiency and NAEP-measured proficiency became one of the recurring diagnostics used in federal policy debates during the No Child Left Behind era.

The high stakes attached to state assessments — school ratings, corrective-action requirements, in some states individual student promotion and graduation consequences — also create systematic pressure on what is taught. This is not a moral failing of administrators; it is a rational response to incentives. When a particular test determines whether a school is placed on an improvement list, schools optimise for that test. The content and format of state assessments therefore shapes instruction in ways that NAEP, with no stakes attached, does not.

Why the gap between them is informative, not embarrassing

The divergence between NAEP results and state-test results is sometimes presented as evidence of fraud or manipulation. That framing misunderstands the design. The two instruments are answering different questions. A state assessment asks: has this student met the threshold our legislature set for this grade level, in this subject, under this framework? NAEP asks: how does this population's performance compare with other populations, and with the same population at earlier points in time?

A row of empty folding chairs facing a stage

The comparison becomes analytically useful when the two move in opposite directions. If a state's own assessment shows rising proficiency rates while the state's NAEP scores remain flat — or decline — that divergence is informative. It suggests either that the state's proficiency cut score is set at a level below NAEP's, that the state's test is becoming easier over time, or some combination. Researchers such as Daniel Koretz have used exactly this kind of cross-instrument comparison to interrogate whether accountability reforms produced real learning gains or statistical artefacts.

This is also why NAEP is sometimes called "the Nation's Report Card" — a label codified in federal statute. It functions as an external check on the proliferation of state-level metrics, not because it is more rigorous in some abstract sense, but because it operates outside the incentive structures that state assessments inevitably inhabit. James Coleman's 1966 report found, to the surprise of its sponsors, that school inputs mattered less to outcomes than family background; NAEP's longitudinal data has since become one of the primary databases through which researchers continue to test, extend and challenge that finding.

Neither instrument is redundant. State assessments do work NAEP cannot: they report results for every school, for every student subgroup, every year, against the specific standards a state has decided its students should meet. NAEP does work state assessments cannot: it holds a mirror steady enough that the reflection thirty years ago and the reflection today can be set side by side. The two tests are not rivals in the same competition. They are instruments calibrated for different distances — one for the room, one for the horizon.

Related pagesMeasure and elsewhere