During the last weeks, we had some heated conversations that are probably changing how we do measurement (and evidence generation) at Kabakoo.
Here is the situation: Over the past few months at Kabakoo, we have been exploring how to simplify our baseline survey. Used across our cohorts, the baseline establishes learners’ (and controls’) starting conditions against which we can later measure change and impact following participation in our programs.
A 35-40 minute questionnaire comes with real costs such as learner time, survey fatigue, and subsequently high drop-off. Obviously, the more we scale (think of the gov-backed national program above), the more this cost compounds.
So we went back to our existing data. We reviewed six years of baseline and follow-up data; we then tested how much information shorter versions could preserve.
One of the first things we learned was that reducing questions does not scale the way we expected. A 32-question questionnaire retained only about 54% of the information in the original full questionnaire (60-70 questions). At the same time, a 52-question version preserved more information, but exceeded our 15-minute target.
This pushed us to think (and debate) harder. What if we approached baseline surveying as a journey, not as a single-point observation? Access, platform fluency, completing a meaningful action, receiving feedback, revising work, collaborating with others, could all teach us something for evaluation.
Now, imagine a room of (product & learning) designers and (quantitative) researchers and engineers. Guess who was saying “we absolutely NEED these 70 questions!!!”
Our internal paradigm now is moving toward a system that asks little, very well, and shifts the rest to observation. This is consistent with the lesson our gender data taught us a couple of weeks ago. As we reported in our 07/2026 update, in the same sample of learners, barriers tied to social norms surfaced far more in video reflections than in written portfolio entries. Increasingly aligning measurement with observed behavior basically turns that finding into our evaluation design.
The discussion is still open on what this means for the richness of the counterfactual data the current system produces. Though our 05/2026 finding tempers the worry. Due to remarkable spillovers, only 96 of our 369 original controls were still “clean” at 12 months. The pristine comparison group our quant researchers are protecting is eroding anyway. Probably because we are running an intervention with a strong social layer, with some degree of success; strong spillovers are actually to be expected. Happy to have your thoughts on this!
Starting with the next cohort, instead of trying to capture dozens of signals in one large questionnaire, we will spread measurement across the learner’s journey. The baseline is now down to roughly 20 simple questions and takes less than 10 minutes, followed by short check-ins and follow-ups over time. Wherever behavior can be observed directly through learner activity, we aim to measure it.