statspai.datasets¶
datasets ¶
Canonical econometrics datasets with documented expected estimates.
This subpackage provides deterministic, reproducible datasets used
throughout the causal-inference literature, consolidated under a
single import path sp.datasets:
import statspai as sp df = sp.datasets.nsw_lalonde() # real extract, no network df.shape (614, 11) sp.datasets.nsw_lalonde(simulated=True).attrs['expected_experimental_att'] 1794
One rule: if StatsPAI ships the real published data for a dataset,
a bare call returns it. simulated=True asks for the calibrated
replica instead. list_datasets()['source'] says which you get.
Nothing here touches the network — the whole catalogue works offline.
Each function returns a pd.DataFrame with:
- A fully-deterministic DGP (fixed seed, BSD/MIT redistributable).
df.attrscontaining:'paper'— original paper citation'expected_*'— theoretically-anchored estimates (from the published paper on the ORIGINAL data, not necessarily the simulated replica)'notes'— what to be careful about when using this replica
The simulated replicas are designed to match the structure and
summary statistics of the original datasets. For numerical R /
Stata parity against the original data, see
tests/external_parity/PUBLISHED_REFERENCE_VALUES.md.
Available datasets
DID / panel
mpdta() — Callaway-Sant'Anna teen employment
teen_employment() — alias of mpdta()
RD
lee_2008_senate() — US Senate RD (Lee 2008)
IV
card_1995() — IV returns-to-schooling
angrist_krueger_1991() — quarter-of-birth IV
Matching / SOO
nsw_lalonde() — LaLonde NSW job training (real MatchIt
extract n=614)
nsw_dw() — Dehejia-Wahba NSW + PSID comparison
Synthetic control
california_prop99() — ADH tobacco (re-exported from synth)
basque_terrorism() — Abadie-Gardeazabal (re-exported)
german_reunification() — ADH 2015 (re-exported)
Public health / epidemiology (REAL data)
nhefs() — Hernán-Robins What If NHEFS (g-methods
canon: quit-smoking → weight / mortality)
load_nhefs() — alias of nhefs()
basque_terrorism ¶
Basque Country terrorism dataset (simulated).
Returns a balanced panel of GDP per capita (thousands of 1986 USD) for 17 Spanish regions, 1955--1997. The Basque Country is the treated unit; treatment begins in 1970 (onset of ETA terrorism).
The simulated data reproduce the gradual widening of an approximately 10 % GDP gap between the Basque Country and its synthetic counterfactual after 1970.
References
Abadie, A. & Gardeazabal, J. (2003). "The Economic Costs of Conflict: A Case Study of the Basque Country." American Economic Review, 93(1), 113--132. [@abadie2003economic]
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns: |
Examples:
german_reunification ¶
German reunification dataset (simulated).
Returns a balanced panel of GDP per capita for 17 OECD countries, 1960--2003. West Germany is the treated unit; treatment begins in 1990 (reunification).
The simulated trajectories reproduce the key stylised facts: Luxembourg has the highest GDP per capita (~40 000), Portugal the lowest (~10 000), and all countries share a common upward growth trend. Post-1990, West Germany exhibits an approximately 1 500 GDP-per-capita decline relative to its synthetic counterfactual.
References
Abadie, A., Diamond, A. & Hainmueller, J. (2015). "Comparative Politics and the Synthetic Control Method." American Journal of Political Science, 59(2), 495--510. [@abadie2015comparative]
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns: |
Examples:
angrist_krueger_1991 ¶
Simulated replica of Angrist-Krueger (1991) quarter-of-birth IV.
Classical weak-instrument case. Quarter of birth predicts years of schooling because compulsory-schooling laws tie entry age to calendar date (Q1 borns are slightly older at entry so can drop out with fewer years of school). First-stage F is a few dozen on several million observations; point estimates are unstable on subsets.
Our replica uses n=5,000 (the original is ~329k). Published IV returns-to-schooling on the original: 0.08-0.11 depending on controls and birth cohort.
Returns:
| Type | Description |
|---|---|
pd.DataFrame with columns: lwage, educ, q1, q2, q3, q4, year_of_birth.
|
|
References
Angrist, J. & Krueger, A. (1991). Does Compulsory School Attendance Affect Schooling and Earnings? QJE 106(4), 979-1014. [@angrist1991does]
card_1995 ¶
Card (1995) NLS Young Men data — simulated replica or real extract.
Card uses proximity to a 4-year college (nearc4) as an
instrument for years of education in a wage equation. Published
OLS and IV point estimates (Card 1995 Table 2):
- OLS: β_educ ≈ 0.075 (col 2)
- IV (nearc4): β_educ ≈ 0.132 (col 5)
IV exceeds OLS — the "Card puzzle". The LATE interpretation is for compliers on the margin of attending college because of proximity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int
|
RNG seed for the simulated DGP (ignored when |
42
|
simulated
|
bool
|
If True, return a deterministic simulated replica calibrated so
StatsPAI estimators recover OLS ≈ 0.11 and IV ≈ 0.142.
If False, load the real NLSYM extract bundled in
|
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Simulated columns: |
References
Card, D. (1995). Using Geographic Variation in College Proximity to Estimate the Return to Schooling. In Christofides et al. (eds.), Aspects of Labour Market Behaviour. [@card1995using]
castle_doctrine ¶
castle_doctrine(region_year_fe: bool = False, state_trends: bool = False, event_time: bool = False) -> DataFrame
Cheng & Hoekstra (2013) castle-doctrine panel — real data.
The canonical staggered difference-in-differences teaching dataset: 50 US states x 11 years (2000-2010). Between 2005 and 2009, 21 states expanded "castle doctrine" self-defence law (no duty to retreat); 29 states never did, giving a clean never-treated control group. Cheng & Hoekstra (2013) find these expansions raised homicide by roughly 8 log points rather than deterring crime.
This is the dataset behind Chapter 9 of Cunningham's Causal Inference: The Mixtape [@cunningham2021causal], where it motivates the Goodman-Bacon decomposition and the modern staggered-adoption estimators.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
region_year_fe
|
bool
|
Append the 44 |
False
|
state_trends
|
bool
|
Append the 51 |
False
|
event_time
|
bool
|
Append |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
550 rows x 29 base columns. Key modelling variables:
|
Notes
post is not 1{year >= effyear}. Cheng & Hoekstra code
post = 1{year > effyear}: the adoption year itself is coded
untreated because the law was in force for only part of it. The
fractional exposure in that year lives in cdl (e.g. Alabama 2006
= 0.5808). Reconstructing the treatment as year >= effyear
silently changes 21 observations and moves the headline estimate.
This also makes the adoption cohort ambiguous for group-time
estimators: gvar = effyear keeps a clean pre-treatment base
period but counts the partially-exposed adoption year as treated,
while gvar = effyear + 1 (consistent with post) pushes that
partial year into the base period instead. On this panel the choice
moves the Callaway-Sant'Anna simple ATT from 0.1104 to 0.0194 — see
sp.replicate('castle_2013') for the full discussion.
Reference values verified against Stata 18 MP (see
tests/reference_parity/test_castle_stata_parity.py):
==================================== ========== ========== Specification beta SE ==================================== ========== ========== TWFE, unweighted 0.0693984 0.0558596 TWFE, aweight=popwt 0.0755332 0.0331936 TWFE, weighted + controls 0.0796349 0.0308756 TWFE, full (region x year + trends) 0.0769490 0.0339377 Callaway-Sant'Anna, gvar=effyear 0.1103830 0.0387242 ==================================== ========== ==========
Examples:
>>> import statspai as sp
>>> df = sp.datasets.castle_doctrine()
>>> r = sp.feols('l_homicide ~ post | sid + year', data=df,
... weights='popwt', vcov={'CRV1': 'sid'})
References
Cheng, C. & Hoekstra, M. (2013). Does Strengthening Self-Defense Law Deter Crime or Escalate Violence?: Evidence from Expansions to Castle Doctrine. Journal of Human Resources, 48(3), 821-853. [@cheng2013does]
Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press. [@cunningham2021causal]
lee_2008_senate ¶
Lee (2008) US Senate RD — simulated replica or real extract.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int
|
RNG seed for the simulated DGP (ignored when |
42
|
simulated
|
bool
|
Default False: this dataset ships the real extract and hands back
the real extract, so |
False
|
Notes
The real-data branch lets you reproduce Lee (2008) Table 1 /
CCT (2014) Table 4 numbers exactly. StatsPAI's
sp.rdrobust(df, y='y', x='x', c=0, kernel='triangular',
bwselect='cct') recovers Conventional ≈ 7.41 and Robust ≈ 7.51
on this dataset (paper headline ≈ 7.99).
Returns:
| Type | Description |
|---|---|
DataFrame
|
Simulated columns: |
References
Lee, D. (2008). Randomized experiments from non-random selection in U.S. House elections. Journal of Econometrics 142, 675-697. [@lee2008randomized] Calonico, S., Cattaneo, M.D. & Titiunik, R. (2014). Robust nonparametric confidence intervals for regression-discontinuity designs. Econometrica 82(6), 2295-2326. [@calonico2014robust]
mpdta ¶
Simulated replica of the mpdta dataset from R's did package.
The original mpdta is a county-year panel of log teen-employment
(2003-2007) where some counties raise their minimum wage in 2004,
2006, or 2007 (staggered adoption).
Our replica preserves:
- 500 counties × 5 years = 2500 rows
- Three treatment cohorts: 2004, 2006, 2007 + never-treated
- Negative homogeneous ATT ≈ -0.04 log points (matches the published
R did::att_gt aggregated ATT of roughly -0.045 on the original)
- County-level clustering in residuals
Returns:
| Type | Description |
|---|---|
pd.DataFrame with columns: countyreal, year, lemp, first_treat, treat
|
|
Notes
df.attrs['expected_simple_att'] = -0.040 (R did::att_gt +
aggte(type='simple') on the original data returns -0.03995
with est_method='reg', never-treated controls, and analytic
SEs -- see tests/orig_parity/02_mpdta_original; our replica's
target is -0.04).
Because this is a simulated DGP, numerical values will not match
R did::att_gt to high precision on the original data; but they
match the sign, order of magnitude, and aggregation pattern.
References
Callaway, B. & Sant'Anna, P.H.C. (2021). Difference-in-Differences with Multiple Time Periods. Journal of Econometrics 225(2), 200-230. [@callaway2021difference]
nhefs ¶
NHEFS — the canonical dataset of Hernán & Robins, Causal Inference: What If (2020), bundled as real, public-domain data for exact replication of the book's g-methods examples.
The National Health and Nutrition Examination Survey I (NHANES I)
Epidemiologic Followup Study (NHEFS) follows US adults from a
1971-1975 baseline to a 1982 re-examination. The book uses it
throughout Part II to estimate the average causal effect of
quitting smoking (qsmk) on 10-year weight change
(wt82_71, kg) and on 10-year mortality (death).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
complete_case
|
bool
|
If False, return the full NHEFS extract (n=1629, 67 columns).
If True, restrict to subjects with a non-missing 1982 weight
( |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
67 columns. Key modelling variables used in the book:
|
Notes
This is real data (df.attrs['data_source'] == 'real'), unlike
the simulated econometrics replicas in this module. Because the data
are the genuine book extract, StatsPAI reproduces the book's published
numbers — not merely their neighbourhood:
- Crude (unadjusted) mean weight-change difference, quitters vs non-quitters: 2.54 kg (book §12.2; StatsPAI 2.5406).
- IP-weighted average treatment effect (stabilized weights,
Chapter 12 MSM): 3.4 kg, 95% CI (2.4, 4.5) (book Program 12.4;
StatsPAI
sp.ipw3.48, gold statsmodels MSM 3.44). - Parametric g-formula / standardization (Chapter 13): 3.5 kg.
- G-estimation of a structural nested mean model (Chapter 14):
psi≈ 3.4.
Strict numerical reproductions of the full chapter programs live in
tests/external_parity/test_whatif_nhefs.py and the public-health
validation notebooks under examples/public_health/.
Provenance & licence
NHEFS is a US Federal public-use survey (NCHS / NIH) and is therefore
in the public domain as a US Government work. The specific analytic
extract bundled here (n=1629 × 67) is the one distributed by Hernán &
Robins with the book and re-packaged in the MIT-licensed causaldata
package (Huntington-Klein); it is byte-reproducible from
causaldata.nhefs. Redistribution here is consistent with both the
public-domain status of the underlying survey and StatsPAI's policy of
only bundling freely redistributable datasets.
References
Hernán, M.A. & Robins, J.M. (2020). Causal Inference: What If. Boca Raton: Chapman & Hall/CRC. [@hernan2020causal]
nsw_dw ¶
Dehejia-Wahba NSW + PSID-1 non-experimental comparison.
Combines the 185 NSW treated (from the experiment) with 2,490 non-experimental PSID males as the comparison group — the classic observational-vs-experimental benchmark.
A naive OLS on re78 ~ treat (no covariates) yields strongly negative estimates (~-$8,500) because the PSID controls are much better-off on average. With PSM on rich covariates, the estimate should return to the experimental benchmark of ≈ $1,794.
Returns:
| Type | Description |
|---|---|
pd.DataFrame with columns: treat, age, education, black, hispanic,
|
married, nodegree, re74, re75, re78. Treated units (185) are the NSW experimental cohort; controls (2,490) are PSID. |
References
Dehejia, R. & Wahba, S. (1999). Causal Effects in Nonexperimental Studies. JASA 94(448), 1053-1062. [@dehejia1999causal]
nsw_lalonde ¶
LaLonde NSW data — real MatchIt extract (default) or simulated replica.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int
|
RNG seed for the simulated replica (ignored when |
42
|
simulated
|
bool
|
If False (the default since 1.21.0), load the real
The default flipped from True to False in 1.21.0 so the honest path is the default one — see MIGRATION.md. |
False
|
Notes
The bundled real data is MatchIt::lalonde (n=614), NOT the
larger DW (1999) NSW + PSID-1 sample (n=2,675). On this smaller
subset, naive OLS gives ATT roughly -\(635 (less negative than DW
Table 3's headline -\)8,498, which uses the full PSID-1). For
the headline naive-bias demonstration, use the simulated
nsw_dw() panel instead.
Simulated replica calibration
sasp_panel ¶
SASP sex-worker session panel — real data.
The Survey of Adult Service Providers, collected by Cunningham &
Kendall [@cunningham2011prostitution]: sex workers reporting on
individual client sessions, with the log hourly wage (lnw) and
whether the session involved unprotected sex (unsafe). Chapter 8
of Cunningham's Causal Inference: The Mixtape
[@cunningham2021causal] uses it to teach the within (fixed-effects)
transformation: pooled OLS conflates who a provider is with what
she does, and only within-provider variation identifies the
compensating differential for unsafe sex.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
analytic_sample
|
bool
|
If False, return the full survey (1787 rows). If True, apply the book's recipe: drop rows missing any modelling variable, then keep only providers with exactly four sessions — 1028 rows, 257 providers, balanced. Use this to reproduce the published numbers. |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
31 columns. Key modelling variables:
|
Notes
Twelve of the 25 controls are provider-level and drop out of the
within estimator — see :data:SASP_TIME_INVARIANT. Stata's
xtreg, fe omits them silently; StatsPAI raises
NumericalInstability if you hand demeaned constants to
sp.regress, which is the intended behaviour: a regressor with no
within variation is not identified, and saying so is better than
quietly dropping it.
On the analytic sample, all three routes agree with Stata 18 MP to
~5e-10 (see tests/reference_parity/test_sasp_within_parity.py):
============================== =========== =========== specification b(unsafe) SE ============================== =========== =========== pooled OLS, HC1 0.0134074 0.0283001 within (FE), cluster by id 0.0510339 0.0282831 manual demeaning, cluster id 0.0510339 0.0282831 ============================== =========== ===========
Pooled OLS puts the unsafe-sex premium at 1.3 log points; controlling
for the provider raises it to 5.1 — the selection runs the other way
from the naive reading, which is the chapter's point. The manual
demeaning row reproduces xtreg, fe exactly, standard error
included, which is the algebraic identity the chapter is teaching.
Examples:
>>> import statspai as sp
>>> df = sp.datasets.sasp_panel(analytic_sample=True)
>>> fe = sp.feols('lnw ~ unsafe + llength + reg | id', data=df,
... vcov={'CRV1': 'id'})
References
Cunningham, S. & Kendall, T. D. (2011). Prostitution 2.0: The changing face of sex work. Journal of Urban Economics, 69(3), 273-287. [@cunningham2011prostitution]
Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press. [@cunningham2021causal]
texas_prison ¶
Texas 1993 prison-capacity expansion panel — real data.
51 US states (50 + DC) x 16 years (1985-2000). Starting in 1993 Texas
expanded operational prison capacity by roughly 35% per year for three
years, approximately doubling it — a natural experiment used throughout
Chapter 10 of Cunningham's Causal Inference: The Mixtape
[@cunningham2021causal] to teach the synthetic control method. The
outcome is the count of Black male prisoners (bmprison); Texas is
statefip == 48.
Returns:
| Type | Description |
|---|---|
DataFrame
|
816 rows x 24 columns. Key modelling variables:
|
Notes
Classic SCM does not reproduce across implementations here, and that
is the point of shipping it. The book's recipe uses four lagged
outcomes among its predictors, which leaves the predictor-weight matrix
V weakly identified (Kaul et al. 2022). The resulting nested V-W
problem is non-convex, and Stata's synth and StatsPAI settle on
different local optima:
===================== ================================== =============
implementation donor weights mean gap 1994-2000
===================== ================================== =============
Stata synth CA .408 IL .360 LA .122 FL .109 23073.70
StatsPAI sp.synth FL .436 NY .311 IL .253 23779.41
===================== ================================== =============
StatsPAI attains the lower pre-treatment RMSE (865.3 vs 1227.0), so
neither is "wrong" on its own objective. The estimated effect agrees
to about 3% despite entirely different donor sets — the effect is
identified far more robustly than the weights are. Do not read the
donor weights as a finding. See sp.replicate('texas_1993').
Examples:
>>> import statspai as sp
>>> df = sp.datasets.texas_prison()
>>> sc = sp.synth(data=df, outcome='bmprison', unit='state', time='year',
... treated_unit='Texas', treatment_time=1993)
References
Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press. [@cunningham2021causal]
Cunningham's data readme attributes the Texas natural experiment to a "Cornwell and Cunningham (2016)" manuscript and to Perkinson (2010), Texas Tough: The Rise of America's Prison Empire. The Perkinson book is verifiable; no "Cornwell and Cunningham (2016)" record could be located in Crossref, so it is reported here as the upstream author's own attribution rather than asserted as a publication.
thornton_hiv ¶
Thornton (2008) HIV-incentive experiment — real data.
A randomized experiment in rural Malawi [@thornton2008demand]: respondents who had been tested for HIV were randomly offered a cash incentive to collect their results at a nearby centre. The randomization makes this the Mixtape's worked example for randomization inference — the Fisherian sharp-null test that needs no asymptotics and no distributional assumption, only the randomization that actually happened.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
complete_case
|
bool
|
If False, return all 4820 survey rows. If True, restrict to rows
with both |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
17 columns, a modelling subset of the 133 in the published file:
|
Notes
Reference values verified against Stata 18 MP on the complete-case
sample (see tests/reference_parity/test_thornton_ri_parity.py):
========================== ============ ======
quantity value n
========================== ============ ======
mean got | any = 1 0.789235640 2211
mean got | any = 0 0.338683788 623
simple difference (SDO) 0.450551852 2834
OLS got ~ any, HC1 0.450551852 SE 0.020857971
========================== ============ ======
The effect is enormous — a 45-point increase on a 34-point base — so
randomization inference returns p = 0 at any practical
permutation count: no reassignment of the treatment label comes close
to the observed difference. That is the honest answer, not a
rounding artefact, and it makes the dataset a poor place to study
borderline p-values but an excellent one for seeing what RI is.
Examples:
>>> import statspai as sp
>>> df = sp.datasets.thornton_hiv(complete_case=True)
>>> ri = sp.ri_test(df, y='got', treat='any', n_perms=1000, seed=42)
>>> round(ri['observed'], 6)
0.450552
References
Thornton, R. L. (2008). The Demand for, and Impact of, Learning HIV Status. American Economic Review, 98(5), 1829-1863. [@thornton2008demand]
Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press. [@cunningham2021causal]
from_fred ¶
from_fred(payload: Union[JSONLike, Mapping[str, JSONLike]], *, series_id: Optional[str] = None) -> DataFrame
Normalise FRED series observations to a tidy time series.
Accepts:
- a single series'
{"observations": [{"date", "value"}, ...]}dict, - just the observations list,
- a mapping
{series_id: observations}for several series, merged ondateinto a wide frame (one column per series).
FRED's missing-value sentinel "." becomes NaN; dates are parsed to
datetime64.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict or list or mapping of series_id -> observations
|
|
required |
series_id
|
str
|
Column name for the value series when a single series is passed
(default |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Examples:
from_sdmx ¶
Normalise an SDMX-JSON payload (OECD / Eurostat / IMF) to long form.
SDMX-JSON encodes each observation by integer indices into per-dimension
code lists. This expands those indices back into human-readable dimension
columns plus a value column — the shape sp.detect_design can read.
Accepts the SDMX-JSON 1.0 structure ({"dataSets": [...],
"structure": {"dimensions": {...}}}) or a pre-tidied list of record
dicts (returned as a DataFrame unchanged).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict or list or DataFrame
|
|
required |
value_name
|
str
|
|
"value"
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One column per SDMX dimension (e.g. |
Examples:
>>> payload = {
... "dataSets": [{"series": {"0:0": {"observations": {"0": [3.2]}}}}],
... "structure": {"dimensions": {
... "series": [
... {"id": "LOCATION", "values": [{"id": "USA", "name": "USA"}]},
... {"id": "SUBJECT", "values": [{"id": "UNR", "name": "Unemp"}]}],
... "observation": [
... {"id": "TIME_PERIOD", "values": [{"id": "2020", "name": "2020"}]}]
... }}}
>>> import statspai as sp
>>> df = sp.from_sdmx(payload)
>>> df.loc[0, "LOCATION"], df.loc[0, "TIME_PERIOD"], df.loc[0, "value"]
('USA', '2020', 3.2)
from_worldbank ¶
Normalise a World Bank Indicators API payload to a tidy panel.
Accepts any of the shapes a World Bank MCP / the v2 REST API returns:
- the raw
[metadata, rows]two-element list, - just the
rowslist of observation dicts, - a
DataFramealready close to tidy.
Each observation dict looks like::
{"indicator": {"id": "NY.GDP.PCAP.KD", "value": "GDP per capita"},
"country": {"id": "US", "value": "United States"},
"countryiso3code": "USA", "date": "2020", "value": 63027.7}
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
list or dict or DataFrame
|
|
required |
wide
|
bool
|
If True, pivot indicators to columns indexed by (iso3, year) — handy when several indicators were fetched and you want one regression frame. |
False
|
value_name
|
str
|
Name of the value column in long form. |
"value"
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Long form columns: |
Examples:
>>> rows = [{"indicator": {"id": "NY.GDP", "value": "GDP"},
... "country": {"id": "US", "value": "United States"},
... "countryiso3code": "USA", "date": "2020", "value": 1.0}]
>>> import statspai as sp
>>> df = sp.from_worldbank(rows)
>>> list(df.columns)
['country', 'iso3', 'indicator', 'indicator_id', 'year', 'value']
>>> int(df.loc[0, 'year'])
2020
california_prop99 ¶
California Proposition 99 panel (Abadie-Diamond-Hainmueller 2010).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
simulated
|
bool
|
If True, return the simulated covariate-rich replica from
|
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns (both branches): |
References
Abadie, A., Diamond, A. & Hainmueller, J. (2010). Synthetic Control Methods for Comparative Case Studies. Journal of the American Statistical Association 105(490), 493-505. [@abadie2010synthetic]
list_datasets ¶
Return a DataFrame describing all available datasets.
Columns: name, design, n_obs, paper, paper_original, expected_main, source.
Every dataset here loads from disk — none of them touches the
network, so the whole catalogue works offline. source says which
variant a bare name() call returns: "bundled CSV" is the real
extract shipped in datasets/data/, "simulated" a
deterministic replica calibrated near the published estimate.
paper_originalis the headline number from the published paper on the ORIGINAL data (what readers expect to see).expected_mainis what the canonical estimator recovers from the variant a barename()call returns (what users will actually observe), and it is prefixedREAL data:wherever that default is the bundled extract rather than a replica. For a"simulated"row the two columns differ because the replica is a deterministic DGP calibrated to the neighbourhood of the published value, not the original data. For a"bundled CSV"row any gap is an estimator or extract difference and is named as such.
For the strict numerical neighbourhood proofs see
tests/external_parity/test_published_replications.py and
tests/external_parity/PUBLISHED_REFERENCE_VALUES.md.