Discussion about this post

User's avatar
Stephen Brien's avatar

Lant, interesting piece.

There is a version of this argument that goes deeper than external validity. RCTs or A/B testing in business contexts (and having run a fair number of these outside development settings, the pattern is consistent) are genuinely useful for calibration questions in relatively stable, mature environments where the goal is optimisation. Getting the dosage right. Tuning a known intervention within a functioning system. They perform much less well as systems design tools. The extrapolation problem isn't just geographic; it's categorical. A calibration finding doesn't tell you much about systems architecture.

The questions that actually determine whether countries build functioning education systems, such as how to form a professional teaching corps, how to create a self-sustaining curriculum architecture, and what relationship between central and local pedagogy produces coherent outcomes, are systems design questions. No accumulation of calibration findings can answer them. That's asking the tool to do something it wasn't built for.

Your point about Vietnam illustrates this. Vietnam's PISA performance reflects a systems-design achievement (teaching-profession formation, curriculum coherence, and community expectations), built over decades through processes that RCT methodology couldn't have designed or validated.

Ken Opalo has been making a related point that academic research and policy research are two different enterprises. But the calibration/systems design distinction suggests the problem runs even further into the tool than the institutional structure around it.

Much of the development challenge is systems design rather than calibration. The science foundation proposal you mentioned was implicitly treating a calibration tool as a substitute for systems design capacity. That may be the deepest credulity of all, because it's unseen.

Lev Heller's avatar

I came across this recently, but this it's fascinating to see more of an exploration of the mindset underlying the aggressive push for adoption of RCTs. One aspect I've found particularly interesting is the way that RCTs were adopted from other sectors in isolation of other evaluation components that are supposed to supplement them. To my understanding, businesses don't run RCTs and then trust that successfully applying the same practices in other contexts will get results. There's an absence of the kind of BI and adaptive management practices that are supposed to accompany implementation.

Ironically, those are probably more useful both for applying ideas at scale, and for iterative development of solutions like you propose in PDIA. There might be more to learn in thinking about how we can build the tech infrastructure that offers BI for implementation fidelity (not AI, quality data architecture for programs), or ways to produce leading indicators on site by site policy effectiveness.

2 more comments...

No posts

Ready for more?