Artificial Intelligence · Published 2025-12-05
When Machine Learning Fills Your Missing Data Better Than You
Let us let you in on one of the statistics little secret in official statistics: we have been faking it for decades. Not occasionally. Not “here and there.” Constantly. Relentlessly. With the desperate enthusiasm of a teenage boy…
Let us let you in on one of the statistics little secret in official statistics: we have been faking it for decades.
Not occasionally. Not “here and there.” Constantly. Relentlessly. With the desperate enthusiasm of a teenage boy discovering Photoshop. Every census, every survey, every labor-force report you have ever quoted in a parliamentary speech is, at its core, an elaborate fiction written on top of a mountain of missing values.
We called it “imputation,”! We dressed it up in respectable academic language (hot-deck, regression-based, nearest-neighbor) while secretly praying nobody would notice that 23% of the fertility module was invented by a econometricians because that is the only way to do that.
And then artificial intelligence walked in wearing nothing but a neural network and ruined traditional imputation forever.
Act I: The Old Orgy of Lies
Picture the scene: Kenya 2019 census. The tablets crash for three straight days in Turkana. When connectivity finally returns, enumerators discover that “submit later” never actually submitted anything. Result? X households with beautiful GPS coordinates and absolutely zero answers. The solution? A team of heroic (read: traumatized) data editors spent eight months copying age structures from neighboring villages that looked vaguely similar.
They called it “cold-deck imputation.” I call it statistical necromancy performed on the corpses of real data. The final report was published with a straight face. The World Bank applauded the “robust methodology.” Somewhere, a fertility rate quietly committed seppuku out of shame.
Act II: Enter the Generative Gods
Fast-forward to 2024. Nigeria is piloting a continuous population register using mobile-money metadata, DHS clusters, and whatever tax data the FIRS feels like sharing that week. Naturally, 38% of the education variables are missing because half the country never bothered to register for a BVN.
The old guard reaches for their cold-deck spreadsheets and… stop. A young woman from Lagos with purple braids and a dangerous glint in her eye uploads the mess to a conditional tabular GAN (generative adversarial network, for the boomers still reading this on Internet Explorer).
Six hours later the model has hallucinated 11.4 million plausible education histories that respect marginal distributions, geographic correlations, and even the weird spike in secondary-school completion every time a state governor builds a campaign school before elections. Validation against gold-standard surveys? Error reduced from 19% to 2.7%. Time saved: approximately the remaining lifespan of everyone who would have done it manually.
The model didn’t just fill holes. It seduced the data into coherence.
Act III: Diffusion Models Do It Slowly (And That’s the Point)
For the true connoisseurs of statistical thriller, nothing beats diffusion models (the same technology that turns “cat wearing a tuxedo” into photorealistic art). In plain English: you take your mangled dataset, add noise until it looks like TV static, then teach the model to denoise it step-by-step while respecting every known constraint.
One example of imputation in a national census context is the 2021 Canadian Census, which used administrative data and household imputation to address non-response in low-response areas, but not specifically diffusion-based imputation. South Africa’s 2001 Census used multiple imputation methods for missing income data, but again no diffusion model is mentioned. In sum, diffusion-based imputers are an emerging, promising approach documented in recent academic research but have not yet become a standard or documented method for income data imputation in national censuses outside experimental or pilot studies. The imputed income distribution can perfectly predicted last-mile mobile-money flows, satellite-detected night lights, and even the price of tomatoes in local markets. The model basically reverse-engineered poverty from first principles and then whispered the answers in Amharic.
Act IV: The Multiple Imputation That Actually Multiples
Old-school multiple imputation gave us five plausible datasets and told us to average them like cowards. Modern neural MI spits out five thousand plausible realities, runs every tabulation through all of them, and hands you confidence intervals so tight they could be used as tourniquets.
Several studies, including a 2024 World Bank publication and related research, document the use of survey-to-survey imputation methods for filling gaps in consumption data in Tanzania's LSMS-like household surveys. These methods improved poverty estimates and showed shifts in headcount rates, often within confidence intervals (95% intervals), consistent with the claim that observed shifts may be statistically insignificant yet politically impactful.
Documents also recount instances where statistical results clashed with political expectations, prompting tension between statisticians providing technical evidence that movements remained within uncertainty bounds and political leaders reacting to perceived changes in poverty statistics. This narrative matches the claim's scenario of a calm statistician general demonstrating the stability of results to the minister, who then accepted the findings.
Act V: The Dark Art of Outlier Seduction
And then there is outlier detection, the dominatrix of data cleaning. Traditional methods (three standard deviations, bless their cotton socks) flagged every Somali pastoralist as an anomaly because they refused to stay in one enumeration area like good little farmers.
A graph neural network trained on five years of mobile positioning data learned that “anomalous” movement is actually just transhumance, laughed gently, and left the pastoralists alone. It then turned its loving attention to the urban elite who claim to earn 40,000 shillings a month while topping up mobile money with millions. Those outliers were corrected with extreme prejudice.
Epilogue: The Afterglow
We used to treat missing data like a shameful family secret to be hidden behind weighted adjustments and tortured footnotes. Artificial intelligence treats it like sipping coffee.
The holes in your data are not flaws, darling. They are invitations. And the new generation of models accepts those invitations with an enthusiasm that borders on the obscene.
So next time someone asks how Africa can produce world-class statistics with third-world infrastructure, smile knowingly and say:
“We let the machines finish what humans were too tired, too broke, or too corrupt to complete. And oh, they finished beautifully.”
Next article: Predictive Demographics (or how AI finally fixed the fertility forecast that white demographers have been getting wrong since 1984).
Stay curious, am I talking to myself... please leave your comments or I will stop these hallucination's!
THE DAILY PULSE provides analytical commentary on health sector insights, development policy, and African research ecosystems.
About the Author
Dr. Julius Kirimi Sindi is a global expert in research funding, policy impact, and donor relations. With extensive experience in analyzing philanthropy, business, and science funding, Dr. Sindi fosters sustainable and inclusive research ecosystems. He has facilitated international business relationships across Africa, Europe, and Asia. His upcoming book, "The Blueprint of Life Well Lived," explores successful strategies for navigating complex business environments while achieving sustainable growth. He is the author of an upcoming book How Societies Change and Why Most Reforms Fail, which introduces an African Theory of Scaling rooted in emotional truth, political safety, and system coherence. He is also the creator of The Daily Pulse, a widely read LinkedIn newsletter offering sharp, human-centered analysis of policy, politics, and development.
Join the conversation
What did this article make you think about?
Thoughtful questions, reflections and respectful disagreement are welcome. First-time contributions are reviewed before publication.