Child in silhouette reading

Education researchers should move beyond small, tightly controlled studies and focus much more heavily on how effective programmes can be implemented by governments at scale, according to a new analysis in Science by What Works Hub for Global Education researchers. 

The world has made huge progress in getting more children into school, but in many countries this has not actually improved literacy, numeracy and other learning outcomes. 

There is solid evidence now about certain interventions that can improve outcomes - but much less is known about how to put them into practice in the real world at scale. "The result is a persistent gap between evidence and practice", says lead author Noam Angrist. 

Approaches that produce strong results in small pilot studies can have much smaller effects when they are expanded and delivered through government systems. Yet studies often do not measure whether programmes were implemented as intended, how widely they were adopted, or what they actually cost. 

“An intervention that works in a controlled study may face very different challenges when implemented through government systems, across thousands of schools and teachers”, says Noam, who is Academic Director of the What Works Hub for Global Education, based at the Blavatnik School. “Implementation is not simply a final step after identifying an effective intervention. It is something that needs to be measured, understood, and improved. It needs its own evidence base.”  

The authors propose a framework for making what they call 'implementation science' a central part of global education research. This framework underpins the work of the What Works Hub for Global Education. It draws on experience from health research, where implementation science has developed partly in response to the long delays between scientific discoveries and their adoption in clinical practice. 

Their framework has four interconnected stages: efficacy trials, effectiveness trials, policy plans and practice at scale. Most education research currently sits at the first stage: testing whether an intervention can work under controlled conditions. The authors argue that researchers need to devote much more attention to the next stage: determining whether an intervention remains effective when delivered at large scale, by governments and across different populations and settings. 

Such trials should examine not just whether an intervention works, but whether its effects are widespread, efficient and equitable. That means testing delivery at scale, comparing the costs and effectiveness of different delivery models, and examining whether benefits vary between groups and contexts. 

The paper points to several examples of programmes that have already been taken to national or large-scale implementation, including Kenya’s Tusome structured pedagogy programme and Teaching at the Right Level in India. 

It also highlights the potential of more iterative approaches to research, including repeated A/B tests that allow programmes to be adapted as they are implemented. One example comes from Botswana, where 12 A/B tests were used to optimise a targeted instruction programme. Seven produced efficiency gains. The tests included changes designed to reduce costs as well as changes intended to increase learning impacts; involving caregivers, for example, more than doubled learning impacts at low marginal cost. 

The other side of the framework concerns the route from research evidence to government policy and then to what happens in classrooms. “Policymakers do not make decisions on evidence alone”, says Angrist. “Political priorities, institutional constraints and the visibility of different outcomes can all influence which policies are adopted.” Learning outcomes may, for example, receive less political attention than highly visible inputs such as building schools or hiring teachers. 

Measuring learning more quickly and transparently could make outcomes more politically salient, while education evidence units working with government officials could help strengthen the use of research in policymaking. 

Even when governments adopt evidence-based policies, implementation can falter. The paper cites data showing large gaps between education policy and practice across 50 countries. It argues that research should therefore examine the bureaucratic systems responsible for delivery, as well as the regulations and incentives shaping behaviour.

As one example, evidence from Indian states suggests that bureaucratic cultures which encourage deliberation and flexible problem-solving can support better education service delivery than systems focused primarily on legalistic compliance. 

A central message of the paper is that implementation itself must be measured. "Quantitative implementation metrics were formally accounted for in only 12.1% of education randomised controlled trials published between 2019 and 2022", says Noam. "Without such measurements, we don't know whether a disappointing result reflects a flawed intervention or simply a programme that was not properly implemented."

'An agenda for implementation science in education' is published in Science