Benchmark Datasets#
CausalML ships loaders for three standard causal-inference benchmarks. Each is downloaded from its original source the first time it is used and cached on disk; none is redistributed with the package.
from causalml.dataset import fetch_lalonde, fetch_ihdp, fetch_twins
lalonde = fetch_lalonde() # Bunch
X, y, treatment = fetch_ihdp(replication=0, return_X_y_t=True)
Files are cached in ~/causalml-data, overridable with the CAUSALML_DATA
environment variable or the data_home argument. Every download is verified
against a SHA256 digest, so a file that has been replaced or truncated raises
instead of being parsed as data. download_if_missing=False refuses to reach
the network, and causalml.dataset.clear_data_dir() empties the cache.
Which metric applies to which dataset#
What can be measured is a property of how the data was made, not a choice:
Dataset |
Ground truth |
Metrics that apply |
|---|---|---|
|
simulated per-unit |
|
IHDP |
simulated |
|
Twins |
both twins observed |
|
LaLonde |
experimental ATE only |
|
Criteo, Hillstrom |
none |
AUUC, Qini, RATE |
PEHE is only computable where the counterfactual was constructed. A ranking of estimators on simulated outcomes is a statement about that simulation.
Datasets with a loader#
LaLonde (National Supported Work)#
A randomized job-training experiment: 445 rows, 185 treated, outcome re78
(1978 earnings in dollars). Its experimental difference in means, about $1,794,
is the yardstick observational estimators are judged against
[20]. The sample is the Dehejia-Wahba one
[11].
Source: the causaldata package (MIT), read at a pinned revision.
IHDP#
Covariates from the Infant Health and Development Program randomized trial with
outcomes simulated on response surface B [14]. The file holds
100 replications of the same 747 units. Each replication simulates its own
outcomes and draws its own 672 / 75 train-test split, so row i is a
different unit in each one and the treated count varies with it.
Results on IHDP are reported as a mean and standard error across replications:
scores = [pehe(fetch_ihdp(replication=r).tau, predict(r)) for r in range(100)]
A single replication is not comparable to a published IHDP number. Neither, exactly, is a mean over this file: the same release also exists in a 1,000-replication version, and that is what the published tables of the CEVAE [22] and DragonNet [29] papers aggregate over (with the 672 further split 63/27 into train and validation). The data-generating process and split geometry match; the replication count does not.
Source: the clinicalml/cfrnet lineage
(MIT), files hosted at fredjo.com.
Twins#
Same-sex twin births from the NBER linked birth / infant death records: 11,400 pairs, 30 covariates. Treatment is being the heavier twin and the outcome is one-year mortality. Because both twins are observed, both potential outcomes are measured rather than simulated — the only dataset here whose ground truth is not a modelling assumption [22].
Two things to know before using it. The raw outcome columns hold days survived
with 9999 standing for “survived the year”, so mortality is
outcome < 9999; read as a number instead, the column averages about 8,000 and
means nothing. And revealing one twin per pair is what makes this observational —
fetch_twins assigns at random, giving a trial with known counterfactuals,
while the confounded variants in the literature assign from a covariate and
differ between papers. Both potential outcomes are returned as y0 and y1
so a caller can construct their own and say which.
Source: NBER linked birth / infant death data (US federal, public domain), via the van der Schaar lab mirror at a pinned revision.
Datasets without a loader#
The remaining benchmarks named in the v1.0 roadmap are not shipped. The rule the loaders follow is to ship one where the source carries an explicit permissive license or is public-domain government data, fetch it at runtime, and never vendor or mirror it. These do not meet that bar today:
Dataset |
Source |
Status |
|---|---|---|
ACIC (2016-2019 competitions) |
No license file. Available through the R packages in that repository. |
|
Jobs |
LaLonde lineage, distributed with the Shalit et al. code |
No license file. The experimental subset overlaps |
Criteo-Uplift |
Non-commercial license. |
|
Hillstrom |
No license file. |
No license is not the same as no permission, and these files have been redistributed in the literature for years. It does mean CausalML does not redistribute them. For Criteo and Hillstrom, scikit-uplift provides loaders.
Loading one yourself takes the same shape as the built-in loaders:
import pandas as pd
df = pd.read_csv("your_download.csv")
X = df[feature_columns].to_numpy()
y = df["outcome"].to_numpy()
treatment = df["treatment"].to_numpy()
learner.fit(X=X, treatment=treatment, y=y)