Evaluating UNFPA’s support during the 2020 census round required managing a large multilingual dataset of more than 1,300 documents. By pairing AI tools with human oversight, the evaluation team scaled up data processing without compromising accuracy and depth.
The independent evaluation of UNFPA support to the 2020 round of population and housing censuses represents a key milestone in UNFPA’s responsible adoption of artificial intelligence (AI) in evaluation. It demonstrates how AI tools continue to streamline data extraction and analysis while upholding ethical standards. Covering UNFPA's census support from 2015 to 2024, the evaluation drew on nearly 1,300 documents, key informant interviews, 16 country case studies, a country office survey and a regional study of Latin America and the Caribbean.
A dedicated governance framework guided the use of AI throughout the census evaluation. Aligned with UNFPA's strategy for use of generative AI in evaluation, the United Nations Principles for the Ethical Use of Artificial Intelligence and UNFPA's information security policy, the approach prioritized data integrity from the outset. Enterprise-grade tools were deployed under strict zero-data-retention protocols, ensuring that no UNFPA data was stored or used to train external models.
Three specialized tools were deployed to support each stage of the document review, ensuring a seamless workflow from translation to synthesis. First, DeepL Pro translated Spanish and French documents from three large institutional datasets (annual narrative reports, country programme documents and country programme evaluations) into English. That meant the analysis could draw on the full multilingual evidence base rather than English-only sources. Another tool, InsightWise, then filtered these documents for census-relevant content. Documents from these three datasets were selected for their scale and relatively standardized formats, which suited AI-augmented extraction. The process produced three structured databases of extracted text, with each excerpt tagged by key analytical variables such as country and region.
Claude AI then ran two rounds of analysis on the structured databases. The first analysis, before fieldwork, generated preliminary syntheses linked to the evaluation questions, which informed data collection for country case studies. The second, during the reporting phase, ran targeted frequency analyses, such as counting references to Chief Technical Advisers, training topics and census technologies, to address specific aspects of the evaluation questions. The team deliberately kept that round narrow, prompting Claude AI for counts and lists rather than interpretive analysis. This way, AI handled the heavy lifting of data retrieval while the evaluators drove all data interpretation. Human judgement was applied throughout the process, with AI outputs validated against the original source documents before integration into the final analysis.
Using AI to expand the document review allowed the evaluation to triangulate findings from the country case studies and the regional analysis against UNFPA’s census support worldwide. This broader scope also addressed critical evidence gaps, enabling the evaluation to generalize findings across all 160 countries and territories supported by UNFPA during the 2020 census round rather than being limited to the 16 case study contexts.
Beyond scaling analytical reach, integrating AI also yielded notable gains in cost efficiency, time savings, and overall quality.
Furthermore, the analytical process underscored the critical role of precise prompt engineering. While AI handled translation, retrieval and structuring at scale with high accuracy, reliable outputs depended entirely on framing prompts around verifiable evidence, lists, and frequency counts rather than broad narrative synthesis or interpretation. This core principle now guides the Independent Evaluation Office (IEO) across all AI-supported evaluations.
The census evaluation is one of several exercises through which the IEO is carefully scaling AI use across the evaluation function. UNFPA's 2024 strategy on generative AI in evaluation sets out six principles and a phased plan for responsible AI use, backed by safeguards such as: ethical AI clauses in consultant contracts, AI disclaimers in reports, clear protocol for human verifications, and internal tracking for evaluation quality and efficiency gains. This strategy has also shaped the Evaluation Assistant, UNFPA's internal AI platform for evaluation evidence on demand. It accelerates evidence synthesis making evaluative insights directly accessible for timely decision-making, particularly facilitating the use of evaluations.
Beyond advancing AI integration, the IEO is continuously driving new approaches to maximize the reach and influence of its evaluations. One key way we are pushing forward with evaluation use and follow-up is through Evidence Utilization Labs. Following the presentation of the census evaluation and its management response to the Executive Board, the IEO partnered with the Programme Division's Data and Analytics Branch to convene a Lab to internalize the management response and promote implementation of its action points at all levels. The Lab also provided regional and country offices with a collaborative platform to discuss opportunities, challenges, and good practices as preparations begin for the 2030 census round.
Ultimately, whether through innovative evidence labs or AI-augmented evaluation workflows, the goal remains the same: putting timely, high-quality evidence into the hands of decision-makers.
This article was written with AI support with human authors in the lead.