The number of dual language immersion (DLI) programs in U.S. public schools more than tripled between 2010 and 2021, from roughly 1,000 to over 3,600, according to the American Councils Research Center’s national canvass.
That growth is good news. But a program’s existence tells you nothing about its quality, and the accountability systems built between 2022 and 2026 have mostly inherited a single blunt instrument: English standardized test scores. These scores are one indicator, and they are structurally biased against capturing what dual language education produces.
This article gives district leaders, board members, and funders a practical measurement toolkit: a clear definition of program quality, a four-indicator framework (we call it FASE), the leading research sources behind each indicator, and an honest account of what philanthropy can and can’t do to build evaluation capacity.
Quick Answer
How do you measure the quality of a dual language program?
Look beyond English test scores at four evidence-based indicator categories:
(1) Fidelity — whether the promised language allocation (e.g., 90/10 or 50/50) actually happens in classrooms, verified by observation rather than the master schedule;
(2) Access — whether English learners and low-income families are enrolled at rates matching or exceeding their share of the community;
(3) Staffing — the share of positions filled by teachers certified and proficient in the partner language, and their retention rate; and
(4) Evidence of biliteracy — assessed growth in both languages, cohort retention through upper grades, and Seal of Biliteracy attainment. Frameworks from the Children’s Equity Project, the U.S. Department of Education’s DLI Playbooks, and CAL’s Guiding Principles operationalize each category.
What “Quality” Means in a Dual Language Program
A dual language program is an instructional model in which students learn academic content in two languages, English and a partner language, with the explicit goals of bilingualism and biliteracy, high academic achievement and sociocultural competence.
Programs may be two-way (integrating English-dominant students and partner-language-dominant students) or one-way (serving primarily one language group), and typically allocate instruction on a 90/10 or 50/50 model.
Program quality is the degree to which a program delivers on these goals for all of its students, not just on how well the English learners aquire English. The field is not short on quality frameworks. The Center for Applied Linguistics’ Guiding Principles for Dual Language Education has anchored the field for two decades. Many states have their own dual language frameworks, and the now-defunct Office of English Language Acquisition in the U.S. Department of Education’s Office released Dual Language Immersion Playbooks to guide implementation.
In 2024, the Children’s Equity Project at Arizona State University published a research-informed framework built from a synthesis of 170 studies and case studies of 11 programs, organized into seven dimensions, including programmatic structures, language allocation, curriculum and pedagogy, assessment, and family engagement. And WestEd’s 2024 research brief synthesizes what current DLI research does and doesn’t establish.
What’s missing is not frameworks.
It’s measurement in practice: districts routinely lack the data systems, staff time, and evaluation expertise to apply any framework. That’s the gap addressed in this article and, we’d argue, where philanthropy can help.
Why This Matters Now (2022–2026)
Three forces have converged in the current accountability cycle:
Explosive growth without quality guardrails
More than 3,600 DLI programs now operate across at least 44 states, with California, Texas, New York, Utah, and North Carolina accounting for roughly 60% and Spanish representing about 80% of programs (American Councils, 2021).
Many launched quickly, often in response to parent demand, before the staffing and evaluation infrastructure they needed was in place. The bilingual teacher shortage guarantees that some of these programs are dual language in name and mostly English in practice.
An equity problem hiding inside a good-news story
The Century Foundation’s 2023 analysis estimates that fewer than 8% of the nation’s English learners are enrolled in dual language immersion, the model the research consensus identifies as most effective for them, while over 83% sit in English-only instruction.
Meanwhile, in gentrifying districts, DLI seats increasingly go to English-dominant families who arrive first in enrollment lotteries or buy a home in the school’s boundaries. A program can post excellent test scores while quietly failing the students it was built to serve
Post-ESSER budget scrutiny
As pandemic-era federal funds have expired, school boards and district offices are asking every program to justify its cost. DLI programs judged solely on English test scores are dangerously exposed because their strongest, best-documented effects appear late, and bilingually.
In the RAND lottery study of Portland’s programs, students randomly assigned to DLI outperformed peers in English reading by 13% of a standard deviation in grade 5 and 22% in grade 8, meaningful effects that compound over time and that a grade-3 snapshot will miss entirely.
Put simply, districts that can’t demonstrate quality will struggle to defend their budgets, and those that measure the wrong things may cut their best programs.
The Mistake: Confusing Existence With Quality and Test Scores With Evidence
Two failure modes dominate.
The existence fallacy. Strategic plans celebrate that DLI programs exist: three schools, five strands, a ribbon-cutting. Nobody checks whether the fourth-grade Spanish block quietly became English test prep in March. A dual language program can exist on paper and die in the master schedule.
The monolingual measurement fallacy. When districts do measure, they reach for the data they already have: English standardized tests. But judging a bilingual program only in English is like judging a decathlete only in the sprint.
It ignores half the program’s promised outcomes (partner-language literacy), penalizes the early grades (where partner-language instruction dominates and English effects haven’t yet compounded), and creates perverse incentives to dilute the model exactly when fidelity matters most.
Our view: in dual language education, implementation data is accountability data. Whether the promised minutes of partner-language instruction occurred, were taught by qualified instructors, and reached an equitably enrolled class is not a soft “process measure” — it is the leading indicator of every outcome the program exists to produce. Test scores are the lagging indicator.
Districts that only monitor the lagging indicator will discover problems years too late to fix.
The FASE Framework: Four Indicators Beyond Test Scores
We organize the field’s quality frameworks into four measurable categories that a district of any size can track and a funder can underwrite. We call it FASE: Fidelity, Access, Staffing, Evidence of biliteracy.
Fun fact: Fase is Spanish for “phase,” which is fitting, because quality measurement is a stage every maturing program must pass through.
F — Fidelity: Is the language allocation real?
Language allocation fidelity is the degree to which the instructional minutes actually delivered in each language match the program’s stated model (e.g., 90/10 or 50/50). It is the single most commonly broken promise in dual language education, usually invisibly, one substitute teacher or test-prep season at a time.
Measure it with:
• Scheduled vs. delivered minutes. Twice-yearly classroom walkthroughs (10–15 minutes, simple language-of-instruction tally) sampled across grades. The CEP framework treats language allocation as its own quality dimension for a reason.
• Materials parity. Are grade-level texts, science kits, and assessments available in the partner language, or are teachers translating worksheets? Count titles, not intentions.
• Erosion tracking. Chart the partner-language share by grade. Most programs erode toward English in grades 3–5 under testing pressure; many 90/10 programs shift towards a 50/50 model by third or fourth grade. Make that shift intentional to avoid the erosion of partner language in favor of high stakes testing.
A — Access: Who actually gets the seats?
A DLI program that mostly serves English-dominant, higher-income families is a lovely enrichment program, but it’s just that. Given that the model’s strongest documented benefits accrue to English learners, anything short of that is an equity failure.
GIt’s —and an equity failure. Measure it with:
• Enrollment parity ratios. ELs’ share of DLI seats ÷ ELs’ share of district (or attendance-zone) enrollment. Same for free/reduced-price lunch eligibility. A ratio near or above 1.0 is the target; TCF’s national analysis suggests most communities are far below it.
• Pipeline equity. Where do families learn about the program, in what languages, and how early? Lottery design details (sibling preferences, neighborhood set-asides, EL set-asides) determine access more than any brochure.
• Attrition equity. Who leaves the program between kindergarten and grade 5, and is exit concentrated among ELs, students with disabilities, or mobile families? A program that retains only its most advantaged students will show rising scores for the wrong reason.
S — Staffing: Can the program be taught as designed?
No indicator predicts program survival better than staffing. Every other quality dimension collapses without teachers who are qualified, proficient in the academic partner language, and staying.
Measure it with:
• Qualified-fill rate. Share of partner-language positions filled by teachers with both the bilingual credential or equivalent qualifications, and demonstrated academic-language proficiency (not just conversational fluency).
• Retention and pipeline depth. Three-year retention rate for DLI teachers, plus the number of candidates in local grow-your-own or residency pipelines. This is where our work on Latino educator recruitment and retention and the bilingual teacher shortage intersects directly with program quality.
• DLI-specific professional learning. Hours of PD tailored to biliteracy instruction, not generic PD delivered in English about teaching in Spanish. Even monolingual colleagues who share students need strategies for supporting multilingual learners.
E — Evidence of biliteracy: Are students becoming bilingual, biliterate learners?
Outcomes still matter, measured in both languages, over the right time horizon.
Measure it with:
• Dual-language assessment. Annual reading (and where feasible, writing) measures in the partner language alongside English. If the district assesses only in English, it has decided in advance that half the program’s mission doesn’t count.
• Longitudinal cohort tracking. Follow entering kindergarten cohorts through grade 5 and beyond. The RAND findings —English reading advantages growing from grade 5 to grade 8— argue for patience and for measuring long enough to see it.
• Biliteracy milestones. Middle- and high-school markers: continued enrollment in advanced partner-language coursework and, ultimately, Seal of Biliteracy attainment rates, disaggregated by student group.
• English learner progress. EL reclassification trajectories for DLI participants versus similar non-participants: the research suggests DLI students may reclassify somewhat later but at higher ultimate rates, so report trajectories rather than single-year snapshots.
What This Looks Like: A One-Page Scorecard
A midsize district we’d consider well-instrumented reports the following annually, per program, on one page:
| FASE Indicator | Example metric | Green looks like |
|---|---|---|
| Fidelity | Delivered partner-language minutes vs. model | Within 10% of the model, all grades |
| Fidelity | Grade-level materials in partner language | Full core coverage, K–5 |
| Access | EL enrollment parity ratio | ≥ 1.0 |
| Access | K–5 attrition gap (EL vs. non-EL) | No significant gap |
| Staffing | Qualified-fill rate | ≥ 90% |
| Staffing | 3-year DLI teacher retention | ≥ 80% |
| Evidence | Partner-language reading growth | On growth targets in both languages |
| Evidence | Seal of Biliteracy attainment (former DLI students) | Rising, disaggregated |
Nothing on that page requires a research university. It requires a decision that these numbers count and roughly 0.2 FTE of evaluation capacity, which most districts don’t have. Hold that thought.
A 12-Month Measurement Starter Plan
Months 1–2 — Baseline the promise. Write down (or excavate) the program model: allocation percentages by grade, staffing plan, enrollment goals. You cannot measure fidelity to a promise nobody wrote down.
Months 3–4 — Instrument access. Pull enrollment and attrition data by EL status and FRPL. Compute parity ratios. This is one afternoon with existing data and often the most eye-opening step.
Months 5–8 — Observe fidelity. Train two or three staff on a simple language-allocation walkthrough tool (the ED/OELA playbooks and CAL’s Guiding Principles self-assessment offer starting templates). Sample every DLI classroom once.
Months 9–12 — Add the second language to the assessment. Pilot a partner-language reading measure at two grade levels. Publish the first one-page scorecard, internally at minimum, and set one improvement target per FASE letter for year two.
What Most People Get Wrong
• “Our scores are up, so the program works.” Rising English scores can reflect enrollment skew (Access failure) rather than program effect. Check who’s enrolled and who’s leaving before celebrating.
• “We’ll evaluate it once it’s established.” Fidelity erodes fastest in years 2–4, exactly when nobody is looking. Measurement delayed is usually measurement of a different, weaker program.
• “Partner-language assessment is too expensive.” Spanish reading measures are widely available and cheap relative to program cost. For less-common partner languages, the challenge is real. Use curriculum-embedded assessments and writing samples scored with common rubrics rather than skipping the outcome entirely.
• “Quality measurement is a compliance exercise.” Done wrong, yes. Done right, it’s protective: the programs with credible multi-measure evidence are the ones that survive budget season. Measurement is armor, not paperwork.
Honest Caveats
• Small programs, small numbers. A single strand of 25 students per grade will produce noisy outcome data. Lean harder on FASE’s implementation indicators (F, A, S) and multi-year rolling cohorts; resist year-to-year score reactions.
• One-way vs. two-way programs need different access lenses. In a one-way program serving mostly ELs, the parity question inverts: are English-dominant families’ interests being used to justify converting the model? Context governs.
• Research humility. The strongest causal evidence (lottery studies like Portland’s) covers English reading effects; evidence on partner-language outcomes and long-run effects is thinner — WestEd’s brief is candid about these limits. Measure locally, precisely because the national literature can’t answer every question about your program.
• This toolkit evaluates programs, not teachers. Using fidelity walkthroughs punitively will destroy the candor that makes the data accurate. Keep quality measurement separate from individual evaluation.
How to Know Your Measurement System Is Working
1. Decisions cite the scorecard. Budget, hiring, and lottery-design conversations reference FASE data within a year.
2. Problems surface early. You learn about allocation erosion from a walkthrough in November, not a test score two years later.
3. Families see themselves in the data. Access metrics are shared publicly, in families’ languages, and enrollment parity moves toward 1.0.
4. Teachers trust it. DLI teachers describe walkthroughs as support, and materials-parity data turns into actual materials budgets.
5. The program’s story changes. Leaders stop saying “we have three dual language programs” and start saying “here’s what our three programs delivered this year, in both languages.”
Where Philanthropy Fits: Underwriting Evaluation Capacity
Here is the uncomfortable pattern we see as a funder: philanthropy loves to fund the launch, the ribbon-cutting, the classroom library, the first cohort, and almost never funds the measurement.
Yet evaluation capacity is among the highest-leverage grants in this field: a part-time evaluator, a partner-language assessment license, walk-through-tool training, or a university evaluation partnership can protect a multimillion-dollar program. If you fund dual language education without funding the ability to know whether it’s working, you are funding hope.
We think communities deserve evidence.
For how classroom-level practice connects to this system-level picture, see our piece on bridging theory and practice for multilingual learners.
The Bottom Line
The dual language field won the existence battle: 3,600+ programs and counting. The 2022–2026 accountability window is the deciding factor in whether it wins the quality battle.
English test scores alone will lose that fight. They measure half the mission, on the wrong timeline, for a skewed sample. Measure FASE instead: Fidelity to the language model, Access for the students the model serves best, Staffing that makes the model teachable, and Evidence of biliteracy in both languages.
We believe the next decade of dual language education will be defined less by how many programs open, and more defined by how they deliver. The districts and funders who build measurement muscle now will be the ones still standing when the counting stops.
FAQ
How do you measure the quality of a dual language program? Track four indicator categories beyond English test scores: language allocation fidelity (delivered vs. promised instructional minutes in each language), equitable access (EL and low-income enrollment parity and attrition), staffing (qualified-fill and retention rates for bilingual-certified teachers), and evidence of biliteracy (assessed growth in both languages, cohort retention, and Seal of Biliteracy attainment).
What is language allocation fidelity? Language allocation fidelity is the degree to which the instructional time actually delivered in each language matches the program’s stated model, such as 90/10 or 50/50. It is verified through classroom observation and materials audits, not the master schedule.
Why aren’t test scores enough to evaluate dual language programs? English-only tests measure half the program’s goals, miss partner-language literacy entirely, and understate effects that compound over time — the RAND Portland lottery study found English reading advantages of 13% of a standard deviation in grade 5 growing to 22% by grade 8.
How many dual language programs are there in the United States? The American Councils Research Center’s national canvass identified more than 3,600 dual language immersion programs as of the 2021–22 school year, up from roughly 1,000 in 2010, with about 80% taught in Spanish.
What percentage of English learners attend dual language programs? The Century Foundation’s 2023 analysis estimates fewer than 8% of U.S. English learners are enrolled in dual language immersion, while more than 83% receive English-only instruction. That said, Texas has one of the highest EL enrollment numbers in DL, serving about 20% of this population. In California, it’s only about 8%.
What are the main dual language program models? Two-way programs integrate English-dominant and partner-language-dominant students; one-way programs serve primarily one language group. Both typically follow a 90/10 allocation (starting with 90% partner-language instruction and shifting toward 50/50) or a 50/50 allocation throughout.
What quality frameworks exist for dual language education? The most widely used are CAL’s Guiding Principles for Dual Language Education (3rd edition), the U.S. Department of Education/OELA Dual Language Immersion Playbooks, and the Children’s Equity Project’s 2024 research-informed framework built on a synthesis of 170 studies.


