What Did the Children Actually Learn?

What Did the Children Actually Learn?

By: Matthew Jukes

New research from the Center for Global Development shows that standardised effect sizes promise comparability but often don’t deliver it. It proposes that researchers report raw effects and reference distributions alongside standardised estimates as a best practice. 

In 2021, an evaluation in East Africa reported reading and writing effect sizes that the authors described as “comparable to some of the largest measured in the literature”: 0.64 and 0.45 standard deviations (SD). In education research, this is considered a “large effect size,” and by this measure, the programme appears to have achieved landmark success improving children’s literacy skills.

The 0.64 SD effect on reading corresponded to a raw gain of only 1.8 additional correct words per minute. This modest improvement in reading translated into a large effect size only because the group they were being compared to made so little progress: 96 percent of children in control schools could not read a single word at endline.

Matthew Jukes, Senior Director of Impact, The Luminos Fund

Reporting standardised effect sizes as the headline figure for a study’s impact is standard practice. It provides a common metric to compare programs using different assessment tools implemented in different contexts. But, as described in a new paper by Jack Rossiter, David Evans, Susannah Hares, and Catherine Henny—which cites the East African example above—standardised effect sizes, reported in isolation, often tell us very little about what children actually learned.

Standardised effect sizes, reported in isolation, often tell us very little about what children actually learned.

The Exchange-Rate Problem

Effect sizes measure the magnitude of an intervention’s impact in standardised units. They are calculated by dividing the raw gain in scores from baseline to endline by the standard deviation, a measure of how much individual scores vary.

Because the standard deviation depends on the spread of scores in a comparison group, the same real-world learning gain (such as one additional word read per minute) can produce very different effect sizes. This means effect sizes are not the neutral cross-study currency researchers treat them as.

The Rossiter et al. paper demonstrates this with oral reading fluency data. Across 66 study-language-grade combinations, an improvement of just one correct word per minute is worth anywhere from 0.03 to 0.55 SD. The same learning gain can look up to seventeen times bigger in one study than another.

In low-literacy settings, where many children score zero, the distribution of scores is so compressed that even very small raw gains produce large effect sizes.

Effect sizes are a useful tool. The problem is reporting them without enough context to understand what children can actually do differently as a result of the programme.

Can the Children Read?

As the paper proposes, the fix is to report interpretable outcomes alongside standardised effects—to say whether children can now do the things that matter for learning to progress.

In some areas this is genuinely hard. Oral reading fluency benchmarks vary across languages and scripts, and calibrating them sensibly takes real work.  But that complexity shouldn’t obscure a simpler point: many early reading skills are perfectly interpretable without benchmarks at all.

Many early reading skills are perfectly interpretable without benchmarks at all.

For instance, children need to know the sounds represented by the letters of the alphabet reliably before they can move on to more advanced skills. A programme can report exactly what share of children have reached that threshold, rather than leaving readers to decode a standard deviation.

Reporting interpretable outcomes also tells you what children are ready to learn next. A child who has mastered foundational skills such as letter-sound knowledge is ready to begin decoding words. One who can decode accurately but haltingly needs additional fluency practice. One who can read connected text with some ease is ready for a stronger focus on comprehension.

Knowing where children fall along this progression tells us clearly what they have learned, what support they may need next, and whether a programme is effectively doing its job.

In a Grade 2 literacy programme, the key question is whether children have reached the level needed to benefit from Grade 3 instruction. In an accelerated learning programme, it is whether students are ready to re-enter mainstream schooling. A standard deviation cannot answer either question.

This kind of reporting is absent from a large share of published quantitative work on literacy.

What We See in Our Own Data

The Luminos Fund runs accelerated learning programmes that enable vulnerable and out-of-school children to cover three years of learning in one year. An independent randomised controlled trial (RCT) of the Luminos programme in Liberia found large effects of 1.2–1.4 SD on reading. In Ethiopia, a government-run accelerated learning programme implementing the Luminos model improved reading by 1.4 SD.

But it is in the raw data where the meaning sits.

Across our programmes in Liberia, Ghana, and Ethiopia in 2024–25, students entering with minimal prior schooling left scoring around 90 percent on a letter sounds task. Average oral reading fluency gains ranged from 22 to 47 correct words per minute across programmes, moving students from non-readers to readers of connected text in a single school year. In Liberia, students progressed from reading just 4 correct words per minute at baseline to 51 by endline—4.6 times the 2016 average for primary school-age children in the country.

Explore our Results

Independent evaluations of the Luminos programme in Ethiopia, Ghana, Liberia, and The Gambia during the 2024–25 school year show children who entered Luminos classrooms unable to read even a single word are now reading full sentences and solving math problems with confidence.

In reading comprehension, scores on five-item comprehension tasks ranged from roughly 25 to 45 percent, up from near zero at baseline in every case. 

We are pleased with these gains, though getting all children to understand what they read in a single year remains the harder challenge—particularly for children building comprehension in a language they don’t speak at home.

Effect sizes offer a measure of comparability; raw scores give us interpretability. That is the standard the field should hold itself to.

In each case, the gains in students’ literacy skills are simple to describe and easy to interpret. Effect sizes offer a measure of comparability; raw scores give us interpretability. That is the standard the field should hold itself to.

When a programme claims to have improved learning, it should be able to answer a basic question: what can children do now that they could not do before?

Luminos students in Liberia work on a literacy assignment in their notebooks writing vowel letters. (Photo: John Healey for the Luminos Fund) 

Matthew Jukes is Senior Director of Impact at the Luminos Fund where he oversees data systems and strategy to support program impact measurement, implementation, and maintenance. He plays a key role in ensuring data quality and in sharing data across the organization to inform Luminos’ programming and decision-making.

Luminos Fund supported the independent research paper discussed in this blog.

 

71 Commercial Street, #232 | Boston, MA 02109 |  USA
+1 781 333 8317   info@luminosfund.org

The Luminos Fund is a 501(c)(3), tax-exempt charitable organization registered in the United States (EIN 36-4817073).

Copyright © 2026 The Luminos Fund. Luminos Method® and Luminos Fund® are registered trademarks of The Luminos Fund in the United States. All rights reserved.

Privacy Policy

What does it take to really get children learning?
Read the new SSIR article by Kirsty Newman and George Werner.

X
We use cookies in order to give you the best possible experience on our website. By continuing to use this site, you agree to our use of cookies.
Accept
Reject