Where health-economic evidence exists — and where it is missing

The evidence base at a glance

Studies in view
—
Countries in view
—
Model-based
—
Uses a QALY
—
Studies over time
Map — intensity by country
Composition of the filtered corpus
Countries — any two variables
Top countries in this view

Country landscape

This tab is unfiltered: it always shows the whole corpus and does not read or change the Explorer filters. Click a bubble to select a country and fill the two mix panels and the disease-group × evaluation-type grid below. Click a heatmap cell to read the studies behind it. You can also search for a country with the picker below.

Disease burden × health spending × income group
Evaluation-type mix —
Disease-group mix —
Disease group × evaluation type —

Each cell is a share of its disease-group row, so a row with any studies sums to 100% across the evaluation types. A study can sit in several disease groups, so rows can overlap and their totals exceed the country’s study count. Click a cell to open the studies behind it.

Eligible studies
—
Per 100k DALYs
—
Per US$1bn spending
—
Per million people
—
Output over time — country vs income-group mean
Disease mix — country vs income group
Evaluation-method mix — country vs income group
Approach & reporting — country vs income group
Where the country sits among its income peers
What counts as an eligible study

The map covers full economic evaluations of health-care delivery published from 1 January 2010 onward: studies that compare two or more ways of delivering care on both costs and consequences. The unit of analysis is the study; the analysis is descriptive and, where it compares research against disease burden, ecological.

Title/abstract screening was machine-learning based, with a measured sensitivity of 0.959 against the current corrected human reference vintage (93/97; 95% CI 0.899–0.984). Partial evaluations (cost-minimisation, budget-impact, cost-of-illness) and records unclear at extraction are excluded, so the analysis population is full evaluations by construction. Studies whose consequence is a non-clinical metric (e.g. diagnostic accuracy, utilisation) are kept and flagged. Reviews synthesising published economic evaluations are retained by extraction design (framework v6.8 codes them as literature/registry synthesis); 1,602 of the 38,199 (4.2%) are title-flagged reviews, 65% of which name no single country, so country-level panels draw mostly on articles.

Selection funnelRecords
Data sources
InputSource
Study metadata & bibliometrics CrossRef + OpenAlex + PubMed + NLM MeSH harvest (~93% coverage of the eligible population)
Income groupsWorld Bank API classification
Population World Bank SP.POP.TOTL (latest observation, mostly 2025)
Disease burden IHME GBD 2023 DALYs (country totals + 22 Level-2 causes)
Health spending IHME Health Spending 1995–2022 (total health expenditure per capita, PPP 2022; national total = per-capita × latest population)
Disease coding MeSH C-tree any-topic headings (2026 vocabulary, incl. the C12.050 pregnancy-complications branch and F03) mapped to 12 groups, crosswalked to GBD Level-2 causes
Research topics OpenAlex topic taxonomy plus an emergent theme model (abstract embeddings → UMAP → HDBSCAN)

Intervention type and care setting were extracted for every eligible study by a single-pass language model (validated against a keyword lexicon, the emergent-topic clusters and a second model lineage; no human validation available).

Reference data are fetched from dated sources, never bundled snapshots.

Rules that keep the numbers honest

Extracted geography only. The author-affiliation proxy is accurate where it applies, but its availability is income-biased — merging it would inflate the headline burden gap from 17.8× to 25.8×, so it is excluded from every income figure.

Multi-country studies count once per country in the country panels, and every figure declares whether its denominator is studies or study–country pairs.

Countries without a World Bank income group are visibly absent, never zero-filled as “0 studies”.

Burden proportionality is a reference point, not a normative target: research allocation also reflects tractability, intervention cost and the stock of existing evidence.

Abstract decision-signals measure what abstracts report, not what studies did — “not reported” is not “not done”.

MeSH uses any-topic headings: major-topic marking jumps at the 2020 harvest boundary and would manufacture a trend.

Limitations

The corpus is effectively English-only (99.8%). Geography is missing for ~33% of studies; country-level figures cover the placeable subset. Harvest coverage is partial (47–93% depending on field) and declared per figure. Comparisons of research to burden are ecological. Extraction fields disagree for 8.3% of studies, handled by conservative disqualification. The 50,661 records unsure at title/abstract are held out of the population pending a handling policy.