Data Operations: International Institutes as Data Vendors

Rwanda's National Institute of Statistics has spent two decades running household, health, nutrition and agriculture surveys, much of it alongside the World Food Program and UNICEF. I worked on one of them, as an ethnographer on the WFP Comprehensive Food Security and Vulnerability Analysis. None of that data was collected with commercial model training in mind, and that is where the question starts.

The money that paid for this data is gone

Since 1984 the Demographic and Health Surveys program ran more than 450 surveys across 90 low and middle income countries and released the results free. USAID terminated it in early 2025. Eighteen countries lost a survey round, Rwanda among them, and Rwanda's malaria indicator survey was cancelled. The Gates Foundation now carries part of the platform. Who pays to collect this data has stopped being a theoretical question, and the countries that produced it are the ones being asked to answer it.

Why labs are buying

Epoch AI puts the stock of quality-adjusted public human text at roughly 300 trillion tokens, fully consumed between 2026 and 2032. That is why labs have started paying for data instead of only scraping it. Stack Overflow licensed its corpus through OverflowAPI and signed Google, then OpenAI, and I was a senior technical product manager there. What made that corpus sellable had little to do with how large it was. The provenance was documented, the structure held up over fifteen years, and one organization had the standing to grant permission.

What is already open, and what is not

NISR releases its data and analysis under CC BY 4.0, which permits commercial reuse with attribution, with anonymized microdata in a public catalog. There is nothing to sell there that has not already been given away. What has never been released is identifiable and linked microdata, geospatial enumeration files, the ability to follow the same households across rounds, and the Kinyarwanda field notes and recordings that never reached a published table. That material is harder to release and far more useful to a model than any published indicator, and it is also the kind of thing African language datasets are currently short of. One terminology fix while I am here: these are repeated cross-sectional surveys rather than longitudinal panels.

Does open publication give up the claim?

Genetic sequence data from the Global South sat freely downloadable for years, and countries were told they had waived any claim to it. At COP16 in 2024, parties to the Convention on Biological Diversity created the Cali Fund, which draws contributions from companies commercially benefiting from that data and returns benefit sharing, monetary and otherwise, to the countries and communities it came from. What changed was that the negotiation stopped being about access and started being about commercial use.

Wealthy countries already charge for governed access to national data. Finland's Findata charges permit and processing fees, UK Biobank charges tiered fees on a cost recovery basis, and the EU Data Governance Act allows cost-based fees for reuse of protected data. That capability is exactly what the countries sitting on the most survey data have the least of.

The objection that matters

Consent. People answered these surveys understanding they were for statistics and program evaluation, and they did not consent with this data to be used for commercial model training. Ownership comes second, since donors funded and co-designed many of these surveys and the rights are usually shared.

What would come back to the people who answered

The framework for this already exists and came out of Africa. The CARE Principles for Indigenous Data Governance, drafted in Gaborone in 2018, put collective benefit first and say plainly that data systems have to be built so the communities the data came from can use it, apply it to policy, and evaluate the services they receive.

Rwanda has already shown what that looks like. Digital Umuganda and Mozilla built the Kinyarwanda voice corpus on Common Voice past two thousand hours, which produced the Mbaza chatbot and Kinyarwanda voice services now running in national deployments. People contributed recordings and got back something they could talk to in their own language. Someone who cannot read can ask it about the weather or the harvest.

The people who answered these questionnaires gave up hours of their day, often to questions about their children's health that nobody enjoys being asked, and they answered because they were told it would improve the services they receive. It would be dishonest to suggest a licensing deal puts something in their hands.

What can come back is collective, and it only counts if it is written into the agreement rather than promised around it. That means a Kinyarwanda model handed back to the institute or country, compute and infrastructure, derived models deposited where local researchers can reach them, a hard prohibition on re-identification, and permissions that expire.

The data is obviously valuable. What nobody has built yet is the consent, provenance and access infrastructure that would make using it defensible, and that is the work I would encourage data operations teams to think about.

Subscribe to Everyone is searching

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe