Skip to content

Cross-language parity matrix

Auto-generated — do not hand-edit. Regenerate with python scripts/build_parity_index.py. Every row traces to a committed test artifact; nothing here is asserted from memory.

StatsPAI's promise is that every number it reports is either aligned with an external reference implementation, recovered against a known truth, or honestly marked as neither. This page makes that promise auditable function-by-function, and keeps the three cases apart — a method with no Stata or R sibling can reach the second and never the first, and saying so is the point. Query any function programmatically:

import statspai as sp
sp.parity_status("feols")     # one function
sp.parity_matrix()            # the whole matrix
sp.parity_summary()           # honest coverage counts

Taxonomy

grade meaning
bit-exact matches a named R/Stata reference to machine tolerance (headline relative error ≤ 1e-6)
aligned matches a named reference within a documented, pre-registered looser tolerance (cross-fit / convention disagreement)
analytical-only recovers a known population parameter on a deterministic DGP, or a closed-form identity (no cross-package reference)
external-replication reproduces published-paper numbers on a calibrated replica
unverified registered public API, no qualifying numerical-parity evidence attached yet — the honest gap

Coverage at a glance

Read the two evidence kinds separately. Only the first answers "does StatsPAI agree with Stata/R"; the second answers "does StatsPAI recover the right answer", which is a different — and for methods with no Stata/R sibling, the only available — question. Summing them into one "verified" figure would let the smaller claim borrow the authority of the larger one, so this page does not print that total.

evidence kind grade functions
Compared against R/Stata (T2) bit-exact 362
aligned 53
subtotal 415
No external software reference analytical-only (T1) 130
external-replication (published numbers) 2
subtotal 132
No numerical evidence yet unverified 651

Honest denominators

The all-registered denominator understates coverage: it counts result and exception classes, which can never carry a parity grade, and infrastructure functions that render tables, draw plots, build agent schemas or load data. The estimator denominator is the number to drive release over release.

denominator cross-language any evidence total cross-lang share
estimator callables 415 546 786 52.8%
infrastructure (parity N/A) 0 0 125 0.0%
result / exception classes 0 1 287 0.0%
all registered 415 547 1198 34.6%

Coverage by estimator family

Families with zero cross-language rows are the highest-leverage targets when a reference implementation exists, and the honest ceiling when one does not — a method with no Stata/R sibling can reach analytical-only and no further. This table is generated from the same records as the rest of the page, so it cannot drift from them.

family cross-language any evidence estimator callables
causal 147 213 344
regression 32 36 37
spatial 28 29 34
panel 27 28 30
decomposition 20 21 29
network 23 24 25
inference 18 21 23
mendelian 18 19 23
diagnostics 17 18 22
epi 16 17 17
dag 0 0 15
bayes 0 7 14
postestimation 6 6 12
timeseries 11 12 12
neural_causal 0 0 11
power 6 7 11
conformal_causal 0 3 10
structural 5 7 10
frontier 5 7 9
survival 8 8 8
robustness 3 4 7
interference 0 1 7
survey 6 6 6
target_trial 0 6 6
transport 3 5 6
fairness 0 6 6
other 2 3 6
longitudinal 0 5 5
experimental 3 3 5
causal_llm 0 0 4
bartik 4 4 4
nonparametric 2 4 4
causal_discovery 0 0 3
surrogate 0 3 3
causal_rl 0 0 3
assimilation 0 3 3
gformula 1 2 2
ope 0 2 2
causal_text 0 0 2
missing 1 2 2
mediation 2 2 2
censoring 1 1 1
synth 0 1 1

bit-exact — 362 functions

Machine-tolerance agreement with a named R/Stata reference.

function reference versions tolerance rel err (R / Stata) test
absorb_ols fixest::feols 0.14.0 and Stata reghdfe (Track A 03_hdfe / 15_hdfe_cluster goldens); Stata reghdfe with aweights, singleton dropping, two-way clustering R 4.5.2; fixest 0.14.0; Stata 18 MP coefficients rtol 1e-12 (observed 2e-15), iid SEs 1e-12 (observed 8e-15), clustered SEs 1e-10 (observed 5.6e-11) — / — test_panel_absorb_ols_parity.py
adjust_pvalues base R stats::p.adjust (bonferroni/holm/BH) R 4.5.2 exact (atol 1e-15; observed 0) — / — test_mht_parity.py (+1)
aipw Stata teffects aipw (ATE, POmeans); R AIPW::AIPW 0.6.9.3 stratified_fit(k_split = 1) R 4.5.2; AIPW 0.6.9.3; SuperLearner 2.0.40; Stata 18 1e-10 rel on ATE, potential-outcome means and SEs (observed <= 1.4e-14) — / — test_teffects_R_parity.py (+2)
ancova Stata regress, vce(robust) / vce(cluster) Stata 18 1e-9 rel — / — test_misc_sens_stata_parity.py
anderson_rubin_ci R ivmodel::AR.test confidence set ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 Both endpoints at 5e-15 after the boundary bisection replaced the grid-point endpoints (previously 8.1e-3 / 4.9e-3). A reference-free test also asserts the two AR entry points agree with each other and that neither endpoint lands exactly on a grid node. — / — test_weakiv_meta_parity.py
anderson_rubin_test R ivmodel::AR.test ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 Statistic 1.1e-15, p-value 1.2e-13, degrees of freedom exact, and the analytic AR confidence set 1.2e-14. — / — test_weakiv_meta_parity.py
arima stats::arima R 4.5.2; stats 4.5.2 rel_est<=1e-06, rel_se<=1e-06 7.4e-07 / 9.3e-09 39_arima.py (+2)
assortativity R igraph::assortativity_degree igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 1.2e-16 on karate. — / — test_network_parity.py
attributable_risk base-R closed form (attributable fraction exposed + PAF) R 4.5.2 AFE + PAF point estimates 1e-12 abs (observed 0); CI not pinned — / — test_epi_extra_parity.py (+1)
attrition_bounds Stata leebounds (thresholds held exactly) leebounds 1.5; Stata 18 1e-12 rel — / — test_misc_sens_stata_parity.py
attrition_test Stata tabulate, chi2; regress Stata 18 1e-10 rel — / — test_misc_sens_stata_parity.py
auc R pROC::auc; Stata roctab R 4.5.2; pROC 1.19.0.1; Stata 18 MP 1e-10 rel (observed 2e-16) — / — test_survival_epi_R_parity.py (+2)
average_treatment_effect grf::average_treatment_effect 2.6.1 (target.sample all, treated, control, overlap), forest outputs held fixed R 4.5.2; grf 2.6.1 estimate and std.err 1e-10 rel (observed 2.0e-15) — / — test_ml_causal_R_parity.py (+1)
bacon_decomposition bacondecomp::bacon R 4.5.2; bacondecomp 0.1.1 rel_est<=1e-06, rel_se<=1e-06 5.6e-16 / 9.6e-09 20_bacon.py (+2)
balance_check Stata iebaltab (ietoolkit) ietoolkit 7.5; Stata 18 1e-9 rel — / — test_misc_sens_stata_parity.py
balance_panel base R counts == n_periods R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 0 / 0 69_balance_panel.py (+2)
bartik 2SLS: R AER::ivreg + sandwich (HC1 / classical), Stata ivregress 2sls, vce(robust) small / small; Rotemberg weights: R bartik.weight::bw and Stata bartik_weight (Goldsmith-Pinkham, Sorkin & Swift) R 4.5.2; AER 1.2.16; sandwich 3.1.1; bartik.weight 0.1.0 (GitHub paulgp/bartik-weight@722ceb85484d6a2bf77985edf2403515eacd1770, R-code/pkg); Stata 18 MP; bartik_weight GitHub paulgp/bartik-weight@722ceb85484d6a2bf77985edf2403515eacd1770 code/bartik_weight.ado 1e-9 rel on all coefficients and SEs (observed <= 5.4e-15) and on Rotemberg alpha_k / beta_k (observed 2.5e-12) — / — test_did_synth_shiftshare_parity.py (+2)
bauer_sinning Stata mvdcmp (Powers, Yoshioka & Yun), logit Stata 18 MP; mvdcmp SSC Logit: explained, unexplained, gap, Yun-weighted detailed explained and unexplained terms, their 6x6 delta-method covariance and the aggregate SEs at 1e-9. Probit aligned at 2e-6 (reference probit not fully converged; see note). — / — test_decomp_R_parity.py
benjamini_hochberg base R stats::p.adjust(method='BH') R 4.5.2 exact (atol 1e-15; observed 0) — / — test_mht_parity.py (+1)
betareg betareg::betareg(link.phi="log") R 4.5.2; betareg 3.2.4 rel_est<=1e-06, rel_se<=0.01 2.2e-12 / 2.1e-08 61_betareg.py (+2)
betweenness_centrality R igraph::betweenness igraph 2.3.3 Normalised and raw on karate, raw on the directed graph, all at 1e-10. — / — test_network_parity.py (+1)
bias_factor R EValue::multi_bound(confounding(), RRAUc, RRUcY) R 4.5.2; EValue 4.1.4 1e-14 rel (observed 0) — / — test_inference_sens_R_parity.py (+1)
biprobit R VGAM::vglm(binom2.rho) bivariate probit R 4.5.2; VGAM 1.1.14 coef / rho 1e-6 abs (observed <= 2e-7); logLik 1e-6 rel — / — test_biprobit_parity.py (+1)
bjs didimputation::did_imputation R 4.5.2; didimputation 0.5.1 rel_est<=1e-06, rel_se<=1e-06 4.8e-08 / 3.5e-07 16_bjs.py (+3)
block_weights spdep::nb2blocknb(NULL, ID) + nb2listw(style = 'W') R 4.5.2; spdep 1.4.2 neighbour sets exact; weights 1e-12 rel (observed 0) — / — test_spatial_survey_R_parity.py (+1)
blp_test GenericML::BLP 0.2.3 (vcovHC const and HC1; with and without the baseline proxy) R 4.5.2; GenericML 0.2.3; sandwich 3.1.1 beta and SE 1e-10 rel (observed 2.3e-15); one-sided p 1e-8 rel — / — test_ml_causal_R_parity.py (+1)
bonacich_power R igraph::power_centrality and sna::bonpow igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 3.9e-16 against igraph and 5.8e-16 against sna at beta = 0.1. — / — test_network_parity.py
bonferroni base R stats::p.adjust(method='bonferroni') R 4.5.2 exact (atol 1e-15; observed 0) — / — test_mht_parity.py (+1)
borusyak_jaravel_spiess didimputation::did_imputation R 4.5.2; didimputation 0.5.1 rel_est<=1e-06, rel_se<=1e-06 4.8e-08 / 3.5e-07 16_bjs.py (+3)
breslow_day_test R DescTools::BreslowDayTest (correct=FALSE/TRUE); Stata cc, by() bd tarone R 4.5.2; DescTools 0.99.60; Stata 18 MP 1e-10 rel (observed 1.3e-14) — / — test_survival_epi_R_parity.py (+2)
calibrate_confounding_strength R sensemakr::ovb_bounds (kd = ky = k) sensemakr 0.1.6 1e-12 rel (CI 1e-10) — / — test_misc_sens_R_parity.py
calibration_test grf::test_calibration 2.6.1 (vcov.type HC3 default and HC1), forest outputs held fixed R 4.5.2; grf 2.6.1; sandwich 3.1.1 coef, se, t 1e-10 rel (observed 1.1e-15); one-sided p-value 1e-8 rel (observed 1.6e-13) — / — test_ml_causal_R_parity.py (+1)
callaway_santanna did::att_gt + aggte R 4.5.2; did 2.3.0 rel_est<=1e-06, rel_se<=1e-09 1.3e-15 / 1.3e-15 04_csdid.py (+2)
cgs_continuous_did contdid::cont_did R 4.5.2 rel_est<=1e-06 2.4e-14 / — 80_contdid.py (+1)
clogit survival::clogit R 4.5.2; survival 3.8.3 rel_est<=1e-06, rel_se<=1e-06 1.3e-08 / 1.3e-08 46_clogit.py (+2)
closeness_centrality Wasserman-Faust closeness from R igraph::distances igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Exact (0.0) on a disconnected graph with three blocks and three isolates -- the case the correction exists for; the connected-graph values also match igraph::closeness(normalized = TRUE) to 1e-10. — / — test_network_parity.py
cluster_robust_se R sandwich::vcovCL (HC1, cadjust; HC0; two-way multi0=FALSE); Stata regress, vce(cluster) R 4.5.2; sandwich 3.1.1; Stata 18 SE 1e-10 rel (observed 2.1e-15 R, 9.8e-16 Stata) — / — test_inference_sens_R_parity.py (+3)
clustering R igraph::transitivity(type = 'local', isolates = 'zero') igraph 2.3.3 Exact on karate and on a disconnected graph whose isolates and degree-1 nodes score 0 on both sides. — / — test_network_parity.py (+1)
cohen_kappa base-R closed form (Cohen's kappa point estimate) R 4.5.2 kappa + agreements 1e-12 abs (observed ~1e-16); SE not pinned — / — test_epi_extra_parity.py (+1)
conformal_synth scinference (Chernozhukov-Wuthrich-Zhu authors' package, GitHub kwuthrich/scinference 567c688): estimation_method='sc', permutation_method='mb' R 4.5.2; scinference 0.0.0.9000 567c6889ce0a1d269a62d415b88aa6baf723a3fe; limSolve 2.0.3; quadprog 1.5.8 p-values exact (rank statistics); ATT 1e-9 rel (observed 9e-16); CI end points exact on the grid — / — test_synth_rest_R_parity.py (+1)
conley Stata acreg (Colella, Lalive, Sakalli & Thoenig) Stata 18 MP; acreg 1.1.0 SE 1e-9 rel (observed ~5e-15); absorbed-IV spatial 1e-10 — / — test_conley_acreg_spacetime_parity.py (+1)
continuous_did fixest::feols(y ~ dose:post id + time) (method='twfe') R 4.5.2; fixest 0.14.0 slope & SE 1e-9 rel (observed 2.2e-15), iid / ~id / ~region, balanced + unbalanced — / —
contrast Stata 18 margins r.g / rb3.g / ar.g / gw.g Stata 18 1e-10 rel (observed <= 7.9e-15 est / 4.9e-15 SE) — / — test_r2_postest_parity.py (+1)
cox survival::coxph R 4.5.2; survival 3.8.3 rel_est<=1e-06, rel_se<=1e-06 8.4e-16 / 2.1e-10 24_coxph.py (+2)
cr2_se clubSandwich::vcovCR(type="CR2"/"CR3") R 4.5.2; clubSandwich 0.6.2 rel_est<=1e-06, rel_se<=1e-06 1.8e-08 / 2.2e-08 53_cr2.py (+2)
cr3_jackknife_vcov R sandwich::vcovJK(center='estimate'), summclust (CV3); Stata regress, vce(jackknife, cluster() mse double) R 4.5.2; sandwich 3.1.1; summclust 0.7.0; Stata 18 vcov 1e-10 rel (observed 6.5e-14 R, SE 4.3e-15 Stata) — / — test_inference_sens_R_parity.py (+3)
cuminc R cmprsk::cuminc (estimate, var, Tests); Stata stcompet (ci, se, hi, lo) R 4.5.2; cmprsk 2.2.12; Stata 18 MP; stcompet 1.0.7 (06nov2012) 1e-10 rel (observed: CIF 6e-16, Gray variance 3e-13, delta SE 2e-16, Gray test 9e-15) — / — test_survival_epi_R_parity.py (+2)
cusum_test strucchange::efp(type = 'Rec-CUSUM') + sctest 1.5.4; Stata 18 estat sbcusum R 4.5.2; strucchange 1.5.4; Stata 18 process, statistic and p-value 1e-10 rel (observed 7.1e-14); Stata boundary constants 3e-6 — / — test_timeseries_R_parity.py (+2)
das_gupta R DasGuptR::dgnpop (product rate function, summed over strata) DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 Das Gupta's Table 2.1 (two factors) and Table 6.5 (four factors x six age groups), every factor effect and both crude rates at 1e-10. — / — test_decomp_R_parity.py
ddd Stata 18 MP regress [aw=w], robust (aweight HC1) Stata 18 MP b / se 1e-12 abs (observed <= 3e-15) — / — test_did2x2_ddd_weighted_robust_parity.py
ddd_heterogeneous triplediff::ddd + agg_ddd R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 5.7e-15 / — 77_ddd.py (+1)
decompose oaxaca::oaxaca R 4.5.2; oaxaca 0.1.5 rel_est<=1e-06, rel_se<=0.05 6.3e-16 / 1.3e-16 30_oaxaca.py (+2)
degree_centrality R igraph::degree igraph 2.3.3 Normalised on karate and raw in / out / all modes on a 40-node directed graph, exact. — / — test_network_parity.py (+1)
demean textbook mean-within (algorithmic) R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 3.5e-15 / 8.8e-15 68_demean_within.py (+2)
dfl_decompose ddecompose::dfl_decompose R 4.5.2; ddecompose 1.0.0 rel_est<=1e-06, rel_se<=1e-06 1.2e-09 / 1.8e-13 31_dfl.py (+3)
diagnostic_test R epiR::epi.tests; Stata diagti R 4.5.2; epiR 2.0.94; Stata 18 MP; diagt 2.032 (diagti 2.053) 1e-12 rel (observed 2e-15) — / — test_survival_epi_R_parity.py (+2)
did_2stage did2s::did2s R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 4.8e-08 / 2.4e-12 73_did2s.py (+3)
did_2x2 Stata 18 MP regress [aw=w], robust (aweight HC1) Stata 18 MP b / se 1e-12 abs (observed <= 3e-16) — / — test_did2x2_ddd_weighted_robust_parity.py
did_estimate synthdid::did_estimate 0.0.9 R 4.5.2; synthdid 0.0.9 estimate and jackknife SE at 1e-9 on the Prop. 99 replica and a five-treated panel (observed est 3.1e-15, jackknife SE 5.7e-16); R's 40 placebo and 40 bootstrap replications replayed at 1e-9 (observed 1.9e-13) — / — test_did_synth_R_parity.py (+1)
did_imputation didimputation::did_imputation R 4.5.2; didimputation 0.5.1 rel_est<=1e-06, rel_se<=1e-06 4.8e-08 / 3.5e-07 16_bjs.py (+2)
did_multiplegt DIDmultiplegt::did_multiplegt (archived 0.1.4) R 4.5.2 rel_est<=1e-06 3.9e-15 / 3.2e-15 81_didm.py (+2)
did_multiplegt_dyn DIDmultiplegtDYN::did_multiplegt_dyn R 4.5.2 rel_est<=1e-06 3.3e-15 / 2.1e-15 78_multiplegt_dyn.py (+2)
did_timevarying_covariates ptetools::pte_default(d_outcome=TRUE, est_method='reg') and did::att_gt(est_method='reg') R 4.5.2; ptetools 1.0.1; did 2.3.0; DRDID 1.2.3 ATT(g,t) and overall ATT 1e-9 rel (observed 3.8e-15) — / — test_did_synth_didvar_parity.py (+1)
direct_method Open Bandit Pipeline (obp) 0.5.7 DirectMethod (same reward-model matrix via q_hat=) obp 0.5.7 value 1e-12 rel (observed 0.0); SE identity 1e-10 — / — test_ml_causal_obp_parity.py (+1)
direct_standardize R epitools::ageadjust.direct; Stata dstdize R 4.5.2; epitools 0.5.10.1; Stata 18 MP 1e-10 rel (observed 3e-13) — / — test_survival_epi_R_parity.py (+2)
discos DiSCos::DiSCo 0.1.4 (Gunsilius distributional synthetic controls), mixture = FALSE; mixture = TRUE vs GLPK on DiSCo's LP R 4.5.2; DiSCos 0.1.4; pracma 2.4.6; quadprog 1.5.8; CVXR 1.8.2; Rglpk 0.6.5.1 quantile weights 1e-11 abs (observed 4.9e-13), counterfactual quantile functions 1e-10 rel, quantile effects 1e-10 abs (observed 1.9e-12); mixture LP weights vs GLPK 1e-12 abs, vs DiSCo's SCS solution 5e-6 abs (observed 7.1e-7) — / — test_did_synth_synthvar_parity.py (+1)
distance_band R spdep::dnearneigh spdep 1.4.2; spatialreg 1.4.3 Neighbour sets identical for all 120 points at a 0.25 radius. — / — test_spdep_parity.py
distributional_did didFF::distDD 0.1.0 (Roth & Sant'Anna) R 4.5.2; didFF 0.1.0; did 2.3.0 per-bin effect & SE 1e-9 (observed est 2.6e-12 abs, SE 2.5e-10 rel) — / — test_did_synth_didvar_parity.py (+3)
dml DoubleML::DoubleMLPLR R 4.5.2; DoubleML 1.0.2 rel_est<=1e-10, rel_se<=1e-10 0 / 3.7e-15 08_dml.py (+2)
dml_panel fixest::demean 0.14.0 (unit and unit+time absorption, balanced and unbalanced) and ddml::ddml_plm 0.3.1 (OLS learner, unit clusters, shared folds) R 4.5.2; fixest 0.14.0; ddml 0.3.1; sandwich 3.1.1; doubleml 0.11.3; scikit-learn 1.6.1 within transform 1e-10 abs (observed 2.6e-14); estimate 1e-10 rel vs ddml and DoubleML (observed 4.1e-16); SE 1e-10 rel vs DoubleML (observed 1.4e-15) and vs ddml after the CR1 factor — / — test_ml_causal_dml_parity.py (+2)
dml_sensitivity doubleml (Python) DoubleML.sensitivity_analysis — bias_bound and adjusted theta bounds 1e-12 (observed 2.5e-15); RV 1e-6 (observed 9.2e-8); RVa is a documented convention gap (<5e-3, observed 1.4e-3) because StatsPAI exhausts theta -z*se with the unadjusted SE while doubleml lets the SE move with the confounding scenario
dose_response Stata doseresponse / gpscore (Hirano-Imbens normal GPS, quadratic T and GPS with interaction) Stata 18; doseresponse SSC 1e-9 rel on the dose-response function at 5 doses (observed 5.1e-10) — / — test_teffects_R_parity.py (+2)
doubly_robust Open Bandit Pipeline (obp) 0.5.7 DoublyRobust (lambda_ inf and 2; same reward model via q_hat=) obp 0.5.7 value 1e-12 rel (observed 1.8e-16); SE identity 1e-10 — / — test_ml_causal_obp_parity.py (+1)
drdid DRDID::drdid_imp_panel R 4.5.2; DRDID 1.2.3 rel_est<=1e-06, rel_se<=1e-06 2.6e-15 / 2.2e-16 38_drdid.py (+2)
dyadic_regression R dyadRobust (Aronow-Samii-Assenova dyadic-robust variance) igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Coefficients 7e-16 and standard errors 2e-15 on undirected and directed dyads, after the 1.28.0 fix to the shared-member weighting; also asserted against a brute-force construction of the definition. — / — test_network_parity.py
ebalance ebal::ebalance 0.2.1 (Hainmueller 2012) — ATT rel <= 1e-5 (observed 3.2e-7); moment gap <= 1e-10 — / — test_matching_r_parity.py (+1)
effective_f_test Stata weakivtest (Montiel Olea & Pflueger) after ivreg2; R ivDiag::eff_F 1.0.6 Stata 18 MP; weakivtest 10/28/2020; ivreg2 4.1.12; R 4.5.2; ivDiag 1.0.6 F_eff rel 1e-9 on 6 designs (observed 3.7e-14 vs Stata, 1.3e-12 vs ivDiag for k = 1) — / — test_rd_iv_R_parity.py (+2)
eigenvector_centrality R sna::evcent (unit L2 norm, as here) sna 2.8; igraph 2.3.3 5e-11 undirected, 8e-11 directed -- power-iteration tolerance on both sides. igraph::eigen_centrality max-scales instead (and 2.x ignores scale = FALSE), so it agrees only up to one scalar; the earlier note claiming the igraph convention was wrong. — / — test_network_parity.py (+1)
engle_granger egranger 1.0.6 (Stata SSC); urca::ur.df on lm residuals; aTSA::coint.test R 4.5.2; urca 1.3.4; aTSA 3.1.2.1; Stata 18; egranger 1.0.6 Z(t) and step-1 coefficients 1e-10 rel (observed 2.1e-13); MacKinnon (2010) critical values 1e-12 vs egranger — / — test_timeseries_R_parity.py (+2)
ergm R ergm::ergm(estimate = 'MPLE') igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Coefficients 2e-16 to 7e-13 for edges + triangle + nodematch + nodecov + absdiff (undirected) and edges + mutual (directed); standard errors 2e-8 (directed) and <= 3.2e-7 (undirected), inside the 1e-6 budget. — / — test_network_parity.py
etregress Stata 18 MP official etregress (Maddala 1983 model) Stata 18 MP Two-step: 5e-9 on every coefficient and every standard error, including the Heckman correction for the estimated first stage. ML: the likelihood, score and observed information are pinned at 9e-11 -- our Hessian reproduces Stata's reported standard errors when evaluated at Stata's own parameter vector, which is independent of either optimiser. At our own optimum the parameters sit within 2e-5 of Stata's; that gap is the two optimisers' stopping points, not a formula difference, and StatsPAI's stops at the HIGHER log-likelihood with a gradient ~300x smaller (asserted, so a regression that makes our optimum worse fails even though the 1e-4 parity assertions would still pass). vce(robust) carries Stata's N/(N-1) meat factor and vce(cluster) its g/(g-1). — / — test_etregress_stata_parity.py
etwfe etwfe::etwfe + emfx R 4.5.2; etwfe 0.6.2 rel_est<=1e-06, rel_se<=0.001 1.8e-13 / 3.9e-14 17_etwfe.py (+2)
etwfe_emfx etwfe::etwfe + emfx R 4.5.2; etwfe 0.6.2 rel_est<=1e-06, rel_se<=0.001 1.8e-13 / 3.9e-14 17_etwfe.py (+2)
evalue EValue::evalues.RR R 4.2.3; EValue 4.1.4 rel_est<=1e-06, rel_se<=1e-06 5.8e-14 / 1.2e-16 23_evalue.py (+2)
evalue_from_result R EValue::twoXtwoRR -> evalues.RR R 4.5.2; EValue 4.1.4 1e-13 rel (observed 3.3e-16) — / — test_inference_sens_R_parity.py (+1)
evalue_rd R EValue::evalues.RD R 4.5.2; EValue 4.1.4 1e-12 rel (observed 5.8e-14: grid built by seq vs np.arange) — / — test_inference_sens_R_parity.py (+1)
evalue_rr R EValue::evalues.RR EValue 4.1.4 Point and CI E-values at 1e-12 across ten cases, including RR < 1 and CIs crossing the null. — / — test_evalue_rr_parity.py
event_study fixest::feols(y ~ i(rel, treat, ref=-1) R 4.5.2; fixest 0.14.0 rel_est<=1e-09, rel_se<=1e-09 3.2e-13 / 1.5e-14 85_twfe_event_study.py (+2)
fairlie Stata fairlie 1.0.7 (Jann, SSC) Stata 18.0 MP; fairlie 1.0.7 16jun2008 (SSC) tightly converged logit/probit: contributions and SEs 1e-12 rel (observed 4.5e-15 / 5.4e-14); at Stata's default logit tolerance contributions 1e-8 and SEs 1e-5 (observed 1.6e-10 / 1.6e-6) — / — test_decomp_qte_parity.py (+1)
fect fect::fect(Y ~ D + X1 + X2, method=, force="two-way", se=FALSE, CV=FALSE, tol=1e-12, max.iteration=20000); Stata side uses the authors' fect_stata (GitHub, installed into a local ado path) R 4.5.2; fect 2.4.1 rel_est<=1e-06, rel_se<=1e-06 1.8e-13 / 9.8e-10 86_fect.py (+2)
feglm fixest::feglm (family="logit") / fixest::fepois R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=5e-05 9.7e-09 / 1.8e-09 67_panel_glm.py (+2)
feols fixest::feols R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=1e-06 5.2e-15 / 2.9e-15 03_hdfe.py (+2)
fepois fixest::feglm (family="logit") / fixest::fepois R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=5e-05 9.7e-09 / 1.8e-09 67_panel_glm.py (+2)
ffl_decompose R ddecompose::ob_decompose(reweighting = TRUE) rifreg 1.1.0; dineq 0.1.0; ddecompose 1.0.0 Observed difference, composition, structure, specification and reweighting errors at 1e-9 (1e-12 absolute floor; logit MLE in the path), both reference directions, for the variance, Gini (exact RIF supplied as custom_rif_function) and 10th/50th/90th percentiles. — / — test_decomp_R_parity.py
finegray R cmprsk::crr (coef, var, invinf, loglik); Stata stcrreg R 4.5.2; cmprsk 2.2.12; Stata 18 MP R 1e-10 rel (observed 6e-15); Stata coefficients 1e-7 and SEs 1e-8 (Stata's ml stops 5e-9 away), log likelihood 1e-10 — / — test_survival_epi_R_parity.py (+2)
fisher_exact R ri2::conduct_ri (randomizr full enumeration); Stata ritest over the full assignment set R 4.5.2; ri2 0.5.0; Stata 18; ritest 1.1.7 p exact (k / N_assignments); statistic 1e-12 rel vs R, 2^-23 vs Stata — / — test_inference_sens_R_parity.py (+3)
florentine_families R ergm flomarriage igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 The 20 marriage ties are identical edge for edge; the Pucci isolate is omitted (15 nodes against 16), which is documented. — / — test_network_parity.py
four_way_decomposition Stata med4way (yreg/mreg linear); CMAverse::cmest 0.1.0 (rb, paramfunc, delta) R 4.5.2; CMAverse 0.1.0; Stata 18 components 1e-12 rel; delta SEs 1e-12 rel vs CMAverse (vcov='ols'); variances 5e-9 rel vs med4way (vcov='ml') — / — test_teffects_R_parity.py (+2)
fracreg stats::glm(quasibinomial('logit')) [fractional response] R 4.5.2 coefficients 1e-10 abs (observed ~8e-15) — / — test_glm_ext_parity.py (+1)
frontier sfaR::sfacross R 4.5.2; sfaR 1.0.1 rel_est<=1e-06, rel_se<=5e-05 4.1e-08 / 4.0e-08 28_frontier.py (+2)
g_computation base R stats::lm g-formula standardization (Robins 1986) — psi 1e-8 (observed <= 7e-16; bootstrap SE pinned loosely +/-25%) — / — test_gformula_parity.py (+1)
g_estimation DTRreg::DTRreg 2.4 (method = 'gest', treat.type = 'bin', weight = 'none') R 4.5.2; DTRreg 2.4 1e-10 rel on each stage psi (observed 4.6e-15) — / — test_teffects_R_parity.py (+2)
gap_closing R ddecompose::dfl_decompose (method='ipw') and ob_decompose (method='regression') DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 Observed, counterfactual and closed gaps at 1e-9 for IPW in both directions (logit MLE in the path) and 1e-10 for regression. method='aipw' has no reference and is checked for double robustness on a known-truth DGP (T1). — / — test_decomp_R_parity.py
gardner_did did2s::did2s R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 4.8e-08 / 2.4e-12 73_did2s.py (+2)
gate_test GenericML::GATES 0.2.3 (monotonize = FALSE) with GenericML::quantile_group membership R 4.5.2; GenericML 0.2.3; sandwich 3.1.1 gamma_k, SE and gamma_K - gamma_1 1e-10 rel (observed 1.9e-15); group membership identical — / — test_ml_causal_R_parity.py (+1)
geary R spdep::geary.test spdep 1.4.2; spatialreg 1.4.3 C 2.0e-15. The closed-form variance and z are new in 1.27.0 (they were NaN whenever permutations=0) and match both spdep nulls at 1.5e-14: randomisation (with the m4/m2^2 term) and normality. — / — test_spdep_parity.py
gelbach Stata b1x2 (Gelbach's own command), robust and homoskedastic Stata 18 MP; b1x2 SSC Per-variable contributions at 1e-10; their full covariance, SEs and the total-change variance at 1e-9, robust and homoskedastic. — / — test_decomp_R_parity.py
geographic_rd rdmulti::rdms; Stata side uses the rdms ado from the rdpackages GitHub mirror (rdmulti is not on SSC: ssc describe rdmulti returns r(601)), installed into a local gitignored ado path R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 1.2e-12 / 5.4e-10 89_rdms.py (+3)
getis_ord_g R spdep::globalG.test (binary weights) spdep 1.4.2; spatialreg 1.4.3 G 5.4e-16. Binary weights, which is what spdep recommends for this statistic. — / — test_spdep_parity.py
getis_ord_local R spdep::localG spdep 1.4.2; spatialreg 1.4.3 Gi 1.7e-14; Gi 4.3e-13 after the star=False branch stopped borrowing Gi's standardisation. — / — test_spdep_parity.py
gformula_ice_fn ltmle::ltmle 1.3.0 (gcomp = TRUE, SL.library = list(Q = 'SL.lm')) point estimate; base-R lm() ICE with geex::m_estimate 1.1.1 sandwich SE R 4.5.2; ltmle 1.3.0; SuperLearner 2.0.40; geex 1.1.1 point 1e-10 rel vs ltmle (observed 2.1e-11), 1e-12 vs hand lm (observed 2.4e-15); sandwich SE 1e-9 rel (observed 7.1e-12) — / — test_r2_teffects_parity.py (+3)
glm base R stats::glm (binomial logit + Poisson log) R 4.5.2 coef / logLik / AIC 1e-8 abs (observed <= 5e-13); SE ~1e-3 rel — / — test_glm_parity.py (+1)
gmm Stata 18 gmm (linear and exponential-mean IV; twostep, igmm, onestep); R gmm::gmm Stata 18 MP; R 4.5.2; gmm 1.9.1 vs Stata estimates rtol 1e-10 (observed 1.2e-12), SEs 1e-9 linear / 1e-6 nonlinear (Stata's numerical Jacobian; observed 2.5e-7), J rtol 1e-10; vs R rtol 1e-6 (observed 1.8e-7) — / — test_panel_gmm_stata_parity.py (+1)
granger_causality Stata 18 vargranger; vars::causality 1.6.1 R 4.5.2; vars 1.6.1; Stata 18 chi2 and F 1e-10 rel (observed 4.1e-15), p-values 1e-8 — / — test_timeseries_R_parity.py (+2)
gsynth gsynth::gsynth R 4.5.2; gsynth 1.4.0 rel_est<=1e-06, rel_se<=1e-06 7.7e-14 / 2.2e-15 19_gsynth.py (+2)
gwr GWmodel::gwr.basic 2.4.1 R 4.5.2; GWmodel 2.4.1 local betas and SEs 1e-9 rel (observed 1.3e-11 / 3.5e-14); RSS, AIC, AICc, BIC, enp, edf 1e-10 rel — / — test_spatial_survey_R_parity.py (+1)
gwr_bandwidth GWmodel::bw.gwr 2.4.1 (golden section gold(); gwr.aic / gwr.cv) R 4.5.2; GWmodel 2.4.1 selected bandwidth 1e-10 rel (observed 0 on 8 configurations); criterion values 1e-11 rel (observed 7.0e-15) — / — test_spatial_survey_R_parity.py (+1)
harvest_did R did::att_gt(control_group='notyettreated', base_period='universal') for every (cohort, horizon) cell; did::aggte(type='dynamic') for the event study under weighting='n_treated' did 2.3.0; DRDID 1.2.3 Every 2x2 cell (ATT and SE) and every event-study horizon (ATT and SE) at 1e-9 relative; observed agreement 1.7e-14. The inverse-variance aggregate over horizons (the default headline) has no reference and is checked by identity only. — / — test_did_synth_misc_parity.py
hdfe_ols fixest::feols R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=1e-06 5.2e-15 / 2.9e-15 03_hdfe.py (+3)
heckman sampleSelection::heckit R 4.5.2; sampleSelection 1.2.14 rel_est<=1e-06, rel_se<=0.0005 1.0e-11 / 1.0e-11 43_heckman.py (+2)
het_test lmtest::bptest (studentized Breusch-Pagan) R 4.5.2; lmtest 0.9.40 statistic & p-value 1e-10 rel (observed ~1e-13) — / — test_diagnostics_parity.py (+1)
heterogeneity_of_effect R metafor::rma(method = 'DL') metafor 5.0.1 1e-12 rel — / — test_misc_sens_R_parity.py
holm base R stats::p.adjust(method='holm') R 4.5.2 exact (atol 1e-15; observed 0) — / — test_mht_parity.py (+1)
honest_did HonestDiD::createSensitivityResults_relativeMagnitudes R 4.5.2; HonestDiD 0.2.8 abs_est<=1e-06, abs_se<=1e-06 4.4e-16 / 5.6e-17 21_honest_relmags.py (+2)
horowitz_manski Stata tebounds 1.8 worst-case bounds (identical estimand for ATE worst-case bounds) Stata 18; tebounds 1.8 (SJ15-2 st0386) 1e-12 rel / 1e-15 abs on both bounds — / — test_teffects_R_parity.py (+2)
hurdle pscl::hurdle(dist='poisson', zero.dist='binomial') R 4.5.2; pscl 1.5.9 count + zero coefficients 1e-6 abs (observed ~2e-8) — / — test_glm_ext_parity.py (+1)
icc Stata 18 estat icc after mixed (ML, REML) and melogit; performance::icc; psych::ICC (balanced ANOVA identity) Stata 18 MP; R 4.5.2; performance 0.16.0; psych 2.6.5 estimate / SE rtol 1e-6, logit-scale CI rtol 2e-6 (observed <= 2.4e-7 / 7e-7 / 1.2e-6); balanced REML ICC = ANOVA ICC(1) rtol 1e-9 — / — test_panel_icc_lrtest_parity.py
impacts R spatialreg::impacts on a lagsarlm fit spdep 1.4.2; spatialreg 1.4.3 Direct, indirect and total at 1e-6, inheriting the SAR rho's own agreement. — / — test_spdep_parity.py
incidence_rate_ratio base-R closed form (rate ratio + conditional-binomial exact CI) R 4.5.2 estimate 1e-12; exact CI 1e-10 abs (observed ~3e-15) — / — test_epi_parity.py (+1)
indirect_standardize Stata istdize (exact CI); R epitools::ageadjust.indirect (log-normal CI) R 4.5.2; epitools 0.5.10.1; Stata 18 MP 1e-10 rel (observed 6e-16) — / — test_survival_epi_R_parity.py (+2)
inequality_index base-R closed form (Gini/Theil-T/Theil-L/Atkinson; = ineq) R 4.5.2 all indices 1e-12 abs (observed ~2e-16) — / — test_inequality_parity.py (+1)
interactive_fe Stata regife (SSC, Gomez) ..., noconstant; R phtt::Eup(additive.effects = 'none') Stata 18 MP; regife 2026-03-30 SSC; R 4.5.2; phtt 3.1.2 slopes rtol 1e-9 vs both (observed <= 2e-11); SEs rtol 1e-9 vs regife with dof='regife' (homoskedastic and cluster), phtt SE reconstructed rtol 1e-9 — / — test_panel_ife_parity.py
interflex interflex::interflex(vartype="delta", vcov.type="robust", neval=5, nbins=3, bw=1); Stata side uses the SSC interflex command R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 4.0e-15 / 1.5e-14 87_interflex.py (+2)
ipcw survival::coxph(ties = 'breslow') + basehaz(centered = FALSE) R 4.5.2; survival 3.8.3 1e-9 rel on every weight (observed 2.0e-11) — / — test_teffects_R_parity.py (+2)
ips Open Bandit Pipeline (obp) 0.5.7 InverseProbabilityWeighting (lambda_ inf and 2) obp 0.5.7 value 1e-12 rel (observed 0.0); SE identity 1e-10 — / — test_ml_causal_obp_parity.py (+1)
ipw base R stats::glm(binomial) + hand-rolled Hajek weighted means — Hajek ATE/ATT estimate 1e-9 (observed <= 2e-15; SE not pinned) — / — test_ipw_parity.py (+1)
irf vars::irf 1.6.1; Stata 18 irf create R 4.5.2; vars 1.6.1; Stata 18 1e-10 rel vs vars, 1e-9 vs Stata irf file (observed 2.6e-14) — / — test_timeseries_R_parity.py (+2)
its lm + sandwich::NeweyWest 3.1.1; Stata 18 newey; itsa 1.0.0 (SSC) R 4.5.2; sandwich 3.1.1; Stata 18; itsa 1.0.0 coefficients and Newey-West SE 1e-10 rel (observed 7.9e-14); itsa 1e-6 (glm2 IRLS) — / — test_timeseries_R_parity.py (+2)
iv AER::ivreg R 4.5.2; AER 1.2.16 rel_est<=1e-06, rel_se<=1e-06 1.1e-11 / 1.1e-11 02_iv.py (+3)
iv_diag R ivDiag::ivDiag 1.0.6 (analytic block) R 4.5.2; ivDiag 1.0.6; lfe 3.1.1 2SLS / OLS coefficients and SEs, classical first-stage F, effective F, tF critical value and interval: rel 1e-9 on 6 designs (observed <= 6e-12) — / — test_rd_iv_R_parity.py (+1)
ivreg AER::ivreg R 4.5.2; AER 1.2.16 rel_est<=1e-06, rel_se<=1e-06 1.1e-11 / 1.1e-11 02_iv.py (+2)
jackknife_se R sandwich::vcovJK(center='mean'); Stata regress, vce(jackknife, cluster() double) R 4.5.2; sandwich 3.1.1; Stata 18 SE / CI 1e-10 rel, p 1e-9 (observed SE 2.0e-15, p 1.4e-14, CI 9.2e-12) — / — test_inference_sens_R_parity.py (+3)
jive Stata jive 1.0.2 (Stata Journal st0108) ujive1 / ujive2 Stata 18 MP; jive 1.0.2 coefficients and SEs (default and robust) rel 1e-9 (observed 3.0e-13 / 4.4e-13) — / — test_rd_iv_R_parity.py (+1)
johansen urca::ca.jo 1.3.4; Stata 18 vecrank R 4.5.2; urca 1.3.4; Stata 18 eigenvalues, trace & max-eigenvalue statistics 1e-11 rel vs ca.jo, 1e-10 vs vecrank (observed 8.2e-14); Osterwald-Lenum table equal to Stata _vecgetcv cell by cell — / — test_timeseries_R_parity.py (+2)
join_counts R spdep::joincount.multi (binary weights) spdep 1.4.2; spatialreg 1.4.3 BB, WW and BW all exact. A reference-free guard also asserts BB + WW + BW = S0/2, the identity the BW defect violated (70.75 against 50). — / — test_spdep_parity.py
kaplan_meier survival::survfit R 4.5.2; survival 3.8.3 S(t) at every event time 1e-12 (observed ~3e-17); median exact — / — test_survival_km_parity.py (+1)
karate_club R igraph::make_graph('Zachary') igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Adjacency matrix identical. — / — test_network_parity.py
katz_centrality R igraph::alpha_centrality igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 7.1e-16 with normalized = False (normalized = True L2-scales the same vector). — / — test_network_parity.py
kdensity Stata kdensity (at(), bwidth(), kernel()); R bw.nrd0 / bw.SJ R 4.5.2; Stata 18 MP density and default widths 1e-10 rel (observed 9e-16); Sheather-Jones 1e-6 vs bw.SJ(nb=1e7, tol=1e-14) (observed 2e-7, R's pair-count binning) — / — test_survival_epi_R_parity.py (+2)
kernel_weights spdep::nb2listwdist(type = 'dpd', alpha = 2) on dnearneigh(0, h); GWmodel::gw.weight R 4.5.2; spdep 1.4.2; GWmodel 2.4.1 weights 1e-12 rel / 1e-15 abs (observed 6.9e-16 rel, 3.3e-16 abs) — / — test_spatial_survey_R_parity.py (+1)
kitagawa_decompose R DasGuptR::dgnpop with ratefunction sum(size*rate)/sum(size) DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 Rate and composition effects on Das Gupta's Table 5.1 at 1e-10; interaction exactly 0. — / — test_decomp_R_parity.py
knn_weights R spdep::knearneigh + knn2nb spdep 1.4.2; spatialreg 1.4.3 Neighbour sets identical for all 120 points, k=4, on a random point set chosen so no distance ties make the answer non-unique. — / — test_spdep_parity.py
lee_bounds Stata leebounds 1.5 (Tauchmann), vce(analytic) Stata 18; leebounds 1.5 (2013-07-17, Tauchmann) 1e-10 rel on bounds and analytic variances — / — test_teffects_R_parity.py (+2)
liml ivmodel::LIML R 4.5.2; ivmodel 1.9.1 rel_est<=1e-06, rel_se<=1e-06 1.7e-15 / 9.7e-16 59_liml.py (+2)
linear_calibration survey::calibrate(calfun='linear', unbounded); Stata svycal regress R 4.5.2; survey 4.5; Stata 18 MP calibrated weights 1e-12 rel (observed 5.5e-15) — / — test_survey_calib_R_parity.py (+2)
lm_tests R spdep::lm.RStests spdep 1.4.2; spatialreg 1.4.3 All five statistics and their p-values at 1e-9. Before the fix: LM_err 39.47 against 19.58, and Robust_LM_err 20.49 (p=6e-6) against 0.0397 (p=0.84) -- the Anselin lag-vs-error decision rule, reversed. — / — test_spdep_parity.py
local_projections lpirfs::lp_lin R 4.5.2; lpirfs 0.2.5 rel_est<=1e-06, rel_se<=1e-06 5.0e-15 / 4.4e-15 34_lp.py (+2)
logit stats::glm(family=binomial("logit")) R 4.5.2; stats 4.5.2 rel_est<=1e-06, rel_se<=1e-06 2.7e-11 / 2.7e-11 57_logit.py (+2)
logrank_test survival::survdiff R 4.5.2; survival 3.8.3 chi-square 1e-10 rel (observed ~8e-16); p-value 1e-10 abs — / — test_survival_km_parity.py (+1)
lp_did direct transcription (no LP-DiD R package installed); Stata side uses the authors' lpdid R 4.5.2 rel_est<=1e-10, rel_se<=1e-10 5.0e-15 / 2.5e-15 83_lpdid.py (+2)
lpoly Stata lpoly (at(), bwidth(), degree(), kernel(), se(), pwidth()) Stata 18 MP 1e-10 rel (observed 1.6e-12) — / — test_survival_epi_R_parity.py (+2)
lrtest Stata 18 lrtest; R anova() on lme4 ML fits Stata 18 MP; R 4.5.2; lme4 2.0.1 chi2 rtol 1e-6, df exact, p rtol 1e-5 (observed chi2 <= 1e-9) — / — test_panel_icc_lrtest_parity.py
ltmle ltmle::ltmle 1.3-0 (glm, gbounds c(0.01, 1), variance.method = 'ic') R 4.5.2; ltmle 1.3.0 psi1 / psi0 / ATE / SE 1e-9 rel without censoring (observed 5e-13); 1e-8 with censoring nodes (observed 5e-10) — / — test_teffects_R_parity.py (+2)
manski_bounds Stata tebounds 1.8 (SJ15-2 st0386), erates(0) Stata 18; tebounds 1.8 (SJ15-2 st0386) 1e-12 rel / 1e-15 abs on both bounds — / — test_teffects_R_parity.py (+2)
mantel_haenszel base-R closed form (Robins-Breslow-Greenland MH; = epiR) R 4.5.2 estimate, se_log, CI 1e-12 abs (observed 0) — / — test_epi_parity.py (+1)
margins_at Stata 18 margins, at(...); R marginaleffects::avg_predictions Stata 18; R 4.5.2; marginaleffects 0.32.0 1e-10 rel (observed <= 3.5e-15 est / SE; marginaleffects SE 1.8e-10, numerical Jacobian) — / — test_r2_postest_parity.py (+2)
markup Stata markupest 1.0.1 (Rovigatti), method(dlw) pmethod(lp) Stata 18; markupest 1.0.1 10May2020 1e-10 rel (observed 8.1e-13 corrected, 3.4e-15 uncorrected) — / — test_frontier_struct_R_parity.py (+1)
match MatchIt::matchit 4.7.2 (nearest, glm/logit distance) — 1:1 and 2:1 PS matching without replacement: rel <= 1e-9. Mahalanobis metric pinned against MatchIt:::mahalanobis_dist (rel <= 1e-10); greedy m_order='data'/'closest' rel <= 1e-9. — / — test_matching_r_parity.py (+1)
mc_panel MCPanel::mcnnm_fit (Athey, Bayati, Doudchenko, Imbens & Khosravi; github.com/susanathey/MCPanel) and fect::fect(method = "mc") R 4.5.2; MCPanel 0.0 @ 6b2706fd7c35f3266048ceb22a7e9a61ae1774da; fect 2.4.1; gsynth 1.4.0 ATT and fitted untreated matrix 1e-9 rel at fixed lambda (observed <= 2.6e-13), four fixed-effect modes x two lambdas — / — test_did_synth_mc_parity.py (+1)
mc_synth MCPanel::mcnnm_fit (Athey, Bayati, Doudchenko, Imbens & Khosravi; github.com/susanathey/MCPanel) and fect::fect(method = "mc") R 4.5.2; MCPanel 0.0 @ 6b2706fd7c35f3266048ceb22a7e9a61ae1774da; fect 2.4.1 ATT and fitted untreated matrix 1e-9 rel at fixed lambda (observed <= 6.0e-13), two-way and no-FE x two lambdas — / — test_did_synth_mc_parity.py (+1)
mccrary_test R rdd::DCdensity 0.57 (CRAN archive) R 4.5.2; rdd 0.57 theta, se, z rel 1e-9; p rel 1e-8; bin width 1e-12 (observed 6.5e-12) — / — test_rd_iv_rd_R_parity.py (+1)
mde base-R closed form (RCT minimum detectable effect) R 4.5.2 effect size 1e-6 abs (output rounded to 6 dp; observed ~2e-8) — / — test_power_extra_parity.py (+1)
mediate mediation::mediate R 4.5.2; mediation 4.5.1 rel_est<=1e-06, rel_se<=0.1 6.7e-15 / 3.6e-15 36_mediation.py (+3)
mediate_interventional CMAverse::cmest 0.1.0 (gformula with postc: rpnde / rpnie / te; rb paramfunc without) R 4.5.2; CMAverse 0.1.0 IIE / IDE / total 1e-9 rel (observed 2.3e-15) — / — test_teffects_R_parity.py (+2)
mediate_sensitivity R mediation::medsens (lm/lm, rho.by = 0.1) mediation 4.5.1 1e-10 rel — / — test_misc_sens_R_parity.py
mediation mediation::mediate R 4.5.2; mediation 4.5.1 rel_est<=1e-06, rel_se<=0.1 6.7e-15 / 3.6e-15 36_mediation.py (+2)
mediation_decompose Stata paramed (Liu & Emsley, SSC), yreg(linear) mreg(linear) with interaction Stata 18.0 MP; paramed SSC CDE/NDE/NIE/total and delta-method SEs 1e-12 rel (observed 9.5e-15) — / — test_decomp_qte_parity.py (+1)
megamma glmmTMB 1.1.14 Gamma(link = 'log') (Laplace, AD Hessian); Stata 18 meglm, family(gamma) link(log) intmethod(laplace) R 4.5.2; glmmTMB 1.1.14; TMB 1.9.25; Stata 18 MP vs glmmTMB estimates / SEs rtol 1e-6 (observed 6.9e-11 / 3.8e-8); vs Stata Laplace estimates rtol 1e-6 (observed 2.2e-7); objective identity rtol 1e-11 — / — test_panel_glmm_parity.py
meglm Stata 18 meglm (gaussian; binomial with binomial()); lme4::lmer(REML = FALSE); lme4::glmer Stata 18 MP; R 4.5.2; lme4 2.0.1 gaussian vs Stata estimates / SEs rtol 1e-6 (observed 4.2e-9 / 2.8e-7), vs lmer ML estimates 6.1e-11; binomial-trials vs Stata 2.4e-9 / 6.6e-8 — / — test_panel_glmm_parity.py
melly_decompose Stata cdeco, method(qr) 1.0.2 (Chernozhukov, Fernandez-Val & Melly; bmelly/Stata counterfactual) Stata 18.0 MP; cdeco 1.0.2 01mar2023 (bmelly/Stata counterfactual, Distribution-Date 20220803) fitted and counterfactual quantiles 1e-10 rel (observed 6.2e-16); QR process coefficients 1e-12 abs (observed 7.5e-15) — / — test_decomp_qte_parity.py (+1)
menbreg glmmTMB 1.1.14 nbinom2 (Laplace, AD Hessian); Stata 18 menbreg intmethod(laplace) / mcaghermite(7) R 4.5.2; glmmTMB 1.1.14; TMB 1.9.25; lme4 2.0.1; Stata 18 MP vs glmmTMB estimates / SEs rtol 1e-6 (observed 1.7e-11 / 2.3e-7); vs Stata Laplace estimates rtol 1e-6 (observed 7.5e-8); objective identity at Stata's and glmmTMB's estimates rtol 1e-11 — / — test_panel_glmm_parity.py
mendelian_randomization R MendelianRandomization::mr_allmethods(method = 'main') MendelianRandomization 0.10.0 1e-12 rel — / — test_misc_sens_R_parity.py
meologit Stata 18 meologit intmethod(mcaghermite) intpoints(7) / intmethod(laplace); ordinal::clmm Stata 18 MP; R 4.5.2; ordinal 2025.12.29 AGHQ-7 vs Stata estimates rtol 1e-6 (observed 3.8e-7), SEs incl. cutpoints rtol 2e-6 (observed 1.3e-6; 5.7e-7 at Stata's estimates); objective identity vs Stata and clmm rtol 1e-11 (observed <= 1.4e-11) — / — test_panel_glmm_parity.py
mepoisson Stata 18 mepoisson, intmethod(laplace) / intmethod(mcaghermite) intpoints(7); lme4::glmer(nAGQ = 1 / 7) Stata 18 MP; R 4.5.2; lme4 2.0.1 objective identity at Stata's estimates rtol 1e-11 (observed 6.6e-12); Laplace estimates / SEs vs Stata rtol 1e-6 (observed 9.0e-11 / 1.2e-7); AGHQ-7 vs lme4 rtol 1e-6 (observed 4.4e-8 / 9.3e-8) — / — test_panel_glmm_parity.py
meta_analysis R metafor::rma (method='FE' and 'DL') ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 6.2e-16 across all nine reported quantities: fixed and random pooled effect and standard error, tau^2, Cochran Q and its p-value, I^2 and H^2. REML is not implemented here, which is a capability gap rather than a disagreement. — / — test_weakiv_meta_parity.py
metalearner econml.metalearners SLearner / TLearner / XLearner — S / T / X conditional-average-treatment-effect vectors match econml elementwise to 1e-12 absolute (observed <= 1.1e-15) when both fitting stages use the same base learner — / — test_metalearner_econml_parity.py
mgwr PySAL mgwr 2.2.1 Sel_BW(multi=True).search() + MGWR.fit() (authors' implementation); fixed-bandwidth back-fitting also GWmodel::gwr.multiscale 2.4.1 python 3.13.9; mgwr 2.2.1; numpy 2.2.6; R 4.5.2; GWmodel 2.4.1 bandwidths, initial bandwidth, bandwidth history and iteration count exact; betas 1e-9 rel (observed 3.0e-12); SEs, ENP_j, tr(S), sigma2, AICc/AIC/BIC 1e-10 rel (observed <= 1e-15); SOC path 1e-9 (observed 8e-13) — / — test_r2_spatial_parity.py (+3)
mi_estimate R mice::pool; Stata mi estimate: regress mice 3.19.0; Stata 18 1e-9 rel (p, CI via t quantiles at fractional df); 1e-12 est/SE/df — / — test_misc_sens_R_parity.py (+1)
mixed lme4::lmer R 4.5.2; lme4 2.0.1 rel_est<=1e-06, rel_se<=1e-06 2.6e-12 / 8.3e-11 25_lmm.py (+2)
mixlogit Stata mixlogit 1.4.0 (SSC, Hole), nrep(50) burn(15), on identical Halton draws Stata 18 MP; mixlogit 1.4.0 means / SDs / Sigma / SEs rtol 1e-6 (observed <= 2.2e-7), log-likelihood rtol 1e-10 (observed 5e-13) — / — test_panel_mixlogit_parity.py
mlogit nnet::multinom R 4.5.2; nnet 7.3.20 rel_est<=1e-06, rel_se<=5e-05 2.6e-07 / 1.5e-11 44_mlogit.py (+2)
moran R spdep::moran.test (randomisation null) spdep 1.4.2; spatialreg 1.4.3 I 1.9e-15, expectation, variance and z all at 1e-15 on the row-standardised lattice. — / — test_spdep_parity.py
moran_local R spdep::localmoran spdep 1.4.2; spatialreg 1.4.3 Every Ii at 8.1e-15. — / — test_spdep_parity.py
moran_residuals R spdep::lm.morantest spdep 1.4.2; spatialreg 1.4.3 Statistic 5e-16; the p-value at 1e-7 once X is supplied so the Cliff-Ord regression-residual null can be formed. Both spdep alternatives are recorded because lm.morantest defaults to one-sided. — / — test_spdep_parity.py
mr R MendelianRandomization::mr_ivw through the sp.mr dispatcher MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 IVW estimate and default random-effects SE, 1e-10. — / — test_mr_R_parity.py
mr_bma Zuber et al. summary_mvMR_BF (GitHub verena-zuber/demo_AMD 4981b5a) summary_mvMR_BF.R demo_AMD@4981b5a 1e-12 rel — / — test_misc_sens_R_parity.py
mr_cml R MendelianRandomization::mr_cML (DP = FALSE, n = 17723) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Estimate and SE for every K = 0..6, the BIC-selected fit with its invalid set {12, 14}, and the MA-BIC average, all at 1e-9 (both sides iterate to d theta <= 1e-7).
mr_egger R MendelianRandomization::mr_egger, TwoSampleMR::mr_egger_regression MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Slope, intercept and both SEs at 1e-10. — / — test_mr_R_parity.py
mr_f_statistic R MendelianRandomization::mr_ivw @Fstat MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Mean F statistic at 1e-10. — / — test_mr_R_parity.py
mr_heterogeneity R TwoSampleMR::mr_ivw / mr_egger_regression (Q, Q_df, Q_pval) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 IVW and Egger (Ruecker) Q, degrees of freedom and p-values at 1e-10. — / — test_mr_R_parity.py
mr_ivw R MendelianRandomization::mr_ivw (default / fixed / random), TwoSampleMR::mr_ivw MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Estimate, SE under all three models, RSE and Cochran's Q at 1e-10 (observed <= 1e-15). — / — test_mr_R_parity.py
mr_leave_one_out R MendelianRandomization::mr_ivw on each leave-one-out subset MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 All 28 leave-one-out estimates and default-model SEs at 1e-10. — / — test_mr_R_parity.py
mr_median R MendelianRandomization::mr_median (weighted / simple / penalized) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Point estimates for all three weightings at 1e-10. The bootstrap SE is Monte Carlo on both sides and is not compared (T3). — / — test_mr_R_parity.py
mr_mediation R MendelianRandomization::mr_ivw + mr_mvivw MendelianRandomization 0.10.0 1e-12 rel — / — test_misc_sens_R_parity.py
mr_mode R MendelianRandomization::mr_mbe (weighted / unweighted, stderror = simple) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Point estimates for both weightings at 1e-10 -- the same point of the same 512-point density grid. The bootstrap SE is Monte Carlo on both sides and is not compared (T3). — / — test_mr_R_parity.py
mr_multivariable R MendelianRandomization::mr_mvivw (default random effects) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Three direct effects and SEs at 1e-10. — / — test_mr_R_parity.py
mr_pleiotropy_egger R TwoSampleMR::mr_egger_regression (intercept test) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Intercept, SE and t(n - 2) p-value at 1e-10. — / — test_mr_R_parity.py
mr_presso R MRPRESSO::mr_presso (NbDistribution = 2000) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Raw estimate and SE, observed RSS, and the outlier-corrected estimate and SE at 1e-10; outlier set {12, 14} identical. The simulated p-values are Monte Carlo on both sides (T3) and follow the reference's k / B convention. — / — test_mr_R_parity.py
mr_radial R RadialMR::ivw_radial (alpha = 0.05, no Bonferroni) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Square-root weights, per-variant Q contributions and total Q at 1e-10; the outlier set is identical with bonferroni=False (StatsPAI's default applies Bonferroni). — / — test_mr_R_parity.py
mr_steiger R TwoSampleMR::mr_steiger with r from get_r_from_bsen MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 R^2 on both traits and the direction at 1e-10; the p-value (1.8e-73) at 1e-12. — / — test_mr_R_parity.py
msm ipw::ipwtm + lm / glm(quasibinomial) + sandwich::vcovCL(type = 'HC1') R 4.5.2; ipw 1.3.0; sandwich 3.1.1 coefficients and cluster SEs 1e-9 rel (observed 1.5e-13); Stata regress / logit [pw] 1e-7 — / — test_teffects_R_parity.py (+2)
multi_cutoff_rd rdmulti::rdmc 2.0.0 (Cattaneo, Titiunik, Vazquez-Bare & Keele) R 4.5.2; rdmulti 2.0.0 identical to sp.rdmc on the fixture (exact); per-cutoff and pooled estimates vs R 1e-9 rel — / — test_rdmulti_parity.py (+1)
multi_outcome_synth augsynth::augsynth_multiout 0.2.0 (progfunc='None', scm=TRUE, combine_method 'concat' / 'avg'); synth_qp re-run at OSQP eps 1e-12 R 4.5.2; augsynth 0.2.0; osqp 1.0.0 weights 1e-10 abs (observed <= 9e-15); per-outcome ATT 1e-9 rel — / — test_synth_rest_R_parity.py (+1)
multi_treatment Stata teffects aipw with a multivalued treatment (mlogit propensity) Stata 18 1e-10 rel on both contrasts, potential-outcome means and sandwich SEs — / — test_teffects_R_parity.py (+2)
multiway_cluster_vcov sandwich::vcovCL(cluster=~g1+g2+g3) R 4.5.2; sandwich 3.1.1 rel_est<=1e-06, rel_se<=1e-06 2.1e-15 / 2.1e-15 56_multiway_cluster.py (+2)
nbreg MASS::glm.nb R 4.5.2; MASS 7.3.65 rel_est<=1e-06, rel_se<=0.005 5.1e-10 / 1.0e-11 42_nbreg.py (+2)
negd Stata regress, vce(robust) Stata 18 1e-9 rel — / — test_misc_sens_stata_parity.py
netlm R sna::netlm igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Coefficients 2e-15 directed and undirected. QAP p-values are permutation draws and are not compared. — / — test_network_parity.py
netlogit R sna::netlogit igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Coefficients 1.4e-9 (IRLS on both sides). QAP p-values are permutation draws and are not compared. — / — test_network_parity.py
network_components R igraph::components igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Counts and sizes exact on a disconnected graph; weak and strong counts on the directed graph. — / — test_network_parity.py
network_modularity R igraph::modularity igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Exact for a fixed split and for igraph's own fast-greedy partition. — / — test_network_parity.py
network_summary R igraph edge_density / diameter / mean_distance / transitivity igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Density, diameter, mean path length, transitivity and assortativity exact; average clustering matches igraph::transitivity(type = 'average', isolates = 'zero'), the convention used here. — / — test_network_parity.py
number_needed_to_treat base-R closed form (NNT = 1/risk difference) R 4.5.2 estimate 1e-12 abs (observed 0); CI not pinned — / — test_epi_parity.py (+1)
oaxaca oaxaca::oaxaca R 4.5.2; oaxaca 0.1.5 rel_est<=1e-06, rel_se<=0.05 6.3e-16 / 1.3e-16 30_oaxaca.py (+3)
odds_ratio base-R closed form (Woolf logit; = epiR::epi.2by2) R 4.5.2 estimate, se_log, CI 1e-12 abs (observed 0) — / — test_epi_parity.py (+1)
ologit MASS::polr(method="logistic") R 4.5.2; MASS 7.3.65 rel_est<=1e-06, rel_se<=1e-05 1.5e-07 / 3.4e-07 45_ologit.py (+2)
oprobit MASS::polr(method="probit") R 4.5.2; MASS 7.3.65 rel_est<=1e-06, rel_se<=1e-06 3.7e-07 / 2.1e-10 49_oprobit.py (+2)
oster_bounds Stata psacalc (Oster); R robomit::o_delta / o_beta Stata 18; psacalc 2.1; R 4.5.2; robomit 1.0.7 1e-12 rel vs psacalc (observed 4.4e-14); 5e-7 abs vs robomit (it rounds to 6 dp) — / — test_inference_sens_R_parity.py (+3)
oster_delta Stata psacalc (Oster); R robomit::o_delta / o_beta Stata 18; psacalc 2.1; R 4.5.2; robomit 1.0.7 1e-12 rel vs psacalc (observed 4.4e-14); 5e-7 abs vs robomit — / — test_inference_sens_R_parity.py (+3)
overlap_weights WeightIt::weightit 1.7.0 (method='glm'), R 4.5.2 R 4.5.2; WeightIt 1.7.0 All four estimands of the shared-propensity family (Li, Li & Li 2019 Table 1) relative to WeightIt: ATO 2.5e-14, ATE 4.1e-14, ATT 2.3e-14, ATC 4.1e-14. The propensity score itself matches R glm(family=binomial) to 2.6e-14 absolute. — / — test_overlap_weights_r_parity.py (+1)
pagerank R igraph::page_rank (damping 0.85) igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 5e-12 undirected, 1e-12 directed. — / — test_network_parity.py
panel plm::plm + plm::phtest R 4.5.2; plm 2.6.7 rel_est<=1e-06, rel_se<=0.001 4.7e-14 / 1.5e-15 35_panel.py (+2)
panel_fgls Stata 18 xtgls, panels(hetero) Stata 18 MP 6.3e-16 on every coefficient and 7.0e-16 on every standard error against the two-step default, and 4.7e-08 against xtgls, igls for the iterated variant. A reference-free test also asserts the two are distinct estimators, so a silent return to iterating fails. — / — test_panel_stata_parity.py
panel_logit Stata 18 xtlogit, re Stata 18 MP Graded by CONVERGENCE rather than a fixed tolerance: Stata integrates adaptively and StatsPAI does not, so the honest claim is that agreement improves as the Gauss-Hermite rule is refined. Observed 3.8e-04 at 12 points and 2.4e-07 at 60, with the log-likelihood at 2.1e-08 -- the sharpest single check, since it is the same objective evaluated at the same optimum. sigma_u and rho at 1e-4. — / — test_panel_stata_parity.py
panel_probit Stata 18 xtprobit, re Stata 18 MP Same convergence grading: 8.0e-05 at 12 quadrature points and 4.0e-08 at 60, log-likelihood at 1.7e-09, sigma_u and rho at 1e-4. — / — test_panel_stata_parity.py
panel_qtet qte::panel.qtet 1.3.1 (Callaway & Li 2019) — all 19 quantiles: abs < 1e-8 (observed 6.8e-12); ATT abs < 1e-6. panel.qtet composes ordinary ecdf evaluations and type-7 quantiles, both of which have exact numpy equivalents, so this is machine-precision agreement rather than a tolerance band. — / — test_panel_qtet_parity.py (+1)
panel_unitroot plm::purtest 2.6.7; Stata 18 xtunitroot R 4.5.2; plm 2.6.7; Stata 18 all statistics 1e-10 rel (observed 9.1e-15) — / — test_timeseries_R_parity.py (+2)
poisson stats::glm(family=poisson()) R 4.5.2; stats 4.5.2 rel_est<=1e-06, rel_se<=1e-06 9.2e-15 / 8.7e-12 58_poisson.py (+2)
policy_tree policytree::policy_tree R 4.5.2; policytree 1.2.4 rel_est<=1e-06, rel_se<=1e-06 9.6e-16 / 1.4e-16 70_policy_tree.py (+2)
policy_value grf::average_treatment_effect(subset = policy == 1) x mean(policy) on grf get_scores; policytree::double_robust_scores reward contrast R 4.5.2; grf 2.6.1; policytree 1.2.4 1e-12 rel (observed 6.7e-16) — / — test_r2_teffects_parity.py (+1)
power_case_control Stata power twoproportions; R stats::power.prop.test(strict=TRUE) R 4.5.2; Stata 18 MP 1e-11 rel vs Stata (observed 4e-13, Stata's normal CDF); 1e-10 vs R (observed 6e-16) — / — test_survival_epi_R_parity.py (+2)
power_cluster_rct base-R closed form (design-effect-inflated z-approx power) R 4.5.2 power 1e-12 abs (observed ~2e-16) — / — test_power_extra_parity.py (+1)
power_logrank base-R closed form (Schoenfeld log-rank power) R 4.5.2 power 1e-12 abs (observed ~2e-16) — / — test_power_parity.py (+1)
power_rct base-R closed form (two-sample pooled-sigma z-approx power) R 4.5.2 power 1e-12 abs (observed ~2e-16) — / — test_power_parity.py (+1)
power_two_proportions base-R closed form (unpooled Wald two-proportion z-approx) R 4.5.2 power 1e-12 abs (observed ~2e-16) — / — test_power_parity.py (+1)
ppmlhdfe fixest::fepois R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=2e-06 4.9e-13 / 2.2e-15 37_ppmlhdfe.py (+2)
prevalence_ratio base-R closed form (Katz-log; = epiR::epi.2by2) R 4.5.2 estimate, se_log, CI 1e-12 abs (observed ~2e-16) — / — test_epi_parity.py (+1)
principal_strat AER::ivreg 1.2.16 (Wald LATE); SACE bounds via Stata leebounds R 4.5.2; AER 1.2.16 1e-10 rel on the monotonicity complier LATE — / — test_teffects_R_parity.py (+2)
probit stats::glm(family=binomial("probit")) R 4.5.2; stats 4.5.2 rel_est<=1e-06, rel_se<=0.01 3.1e-07 / 1.6e-08 48_probit.py (+2)
psm MatchIt::matchit R 4.5.2; MatchIt 4.7.2 rel_est<=1e-06, rel_se<=1e-06 1.2e-15 / 2.0e-16 11_psm.py (+2)
psmatch2 Stata 18 MP + psmatch2 4.0.12 / pstest 4.2.2 (Leuven & Sianesi 2003) Stata 18 MP; psmatch2 4.0.12 Observed py<->Stata relative gaps on the committed fixtures: nearest-neighbour ATT 1.2e-16 and its analytic SE exactly 0 (as is _weight per row); radius ATT 2.8e-16; Abadie-Imbens ai(1)/ai(2) SE 3.0e-15/1.7e-15; PSM-DID across all five weight regimes 2.0e-14; pstest per-covariate rows 1.3e-14; Mahalanobis ATT 1.2e-13; llr ATT 2.1e-10 (worst of four kernels); kernel ATT 9.1e-10; pstest summary block 1.3e-9. The two loosest rows are bounded by fixture precision, not by the estimator: pstest accumulates MeanBias/MedBias in a Stata float (2.6e-9) and r(seatt) for the llr reroute was captured at 8 significant digits (9.9e-9). — / — test_psmatch2_parity.py (+3)
pwcompare Stata 18 margins g, pwcompare(effects) mcompare(noadjust bonferroni sidak); R emmeans pairs(adjust=) Stata 18; R 4.5.2; emmeans 2.0.3 1e-10 rel (observed <= 7.9e-15 diff / 4.9e-15 SE)
qreg quantreg::rq R 4.5.2; quantreg 6.1 rel_est<=1e-06, rel_se<=0.1 3.3e-15 / 4.4e-15 40_qreg.py (+2)
queen_weights spdep::poly2nb(queen = TRUE) + nb2listw(style = 'W' / 'S' / 'U') R 4.5.2; spdep 1.4.2; sf 1.1.1; spData 2.3.5 neighbour sets exact; weights 1e-12 rel (observed 0) — / — test_spatial_survey_R_parity.py (+1)
rake survey::rake (to its fixed point) and survey::calibrate(calfun='raking'); Stata svycal rake R 4.5.2; survey 4.5; Stata 18 MP calibrated weight shares 1e-12 rel at tol=1e-14 (observed 1.5e-15); 1e-8 at the default tol=1e-10 (observed 1.0e-10) — / — test_survey_calib_R_parity.py (+2)
rd2d R rd2d::rd2d / rd2d.distance 1.0.0 (Cattaneo, Titiunik & Yu) R; rd2d 1.0.0; sandwich 3.1.1 estimate.p/q, std.err.p/q, t, CI, cross-point covariance rel 1e-9 (observed 1.6e-11); bandwidths rel 1e-8 (observed 1.2e-12); N exact — / — test_rd_open_R_parity.py (+1)
rd2d_bw R rd2d::rdbw2d / rdbw2d.distance 1.0.0 R; rd2d 1.0.0 per-point bandwidths rel 1e-8 (observed 5.0e-13) — / — test_rd_open_R_parity.py (+1)
rd_honest RDHonest::RDHonest 1.0.1.9000 (Armstrong & Kolesar) R 4.5.2; RDHonest 1.0.1.9000 estimate / std.error / maximum.bias / conf.low / conf.high 1e-9 rel at fixed bandwidth; 1e-6 rel when the bandwidth and M are selected — / — test_rdhonest_parity.py (+1)
rdbwhte R rdhte::rdbwhte 0.2.0 R 4.5.2; rdhte 0.2.0; rdrobust 4.0.0 bandwidths rel 1e-8 (continuous and per-subgroup) — / — test_rd_iv_rd_R_parity.py (+1)
rdbwselect rdrobust::rdbwselect; Stata side uses the authors' rdbwselect ado. certwo is R-only: Stata rdbwselect 10.0.0 exits r(3200) on it, including on the package's own rdrobust_senate.dta R 4.5.2; rdrobust 3.0.0 rel_est<=1e-06 9.4e-13 / 3.5e-09 88_rdbwselect.py (+2)
rdd rdrobust::rdrobust R 4.5.2; rdrobust 3.0.0 rel_est<=1e-06, rel_se<=0.1 2.5e-14 / 9.4e-11 06_rd.py (+3)
rddensity rddensity::rddensity R 4.5.2; rddensity 2.6 rel_est<=1e-06, rel_se<=1e-06 9.3e-12 / 1.8e-11 09_rddensity.py (+2)
rdhte R rdhte::rdhte 0.2.0 (sandwich 3.1.1) R 4.5.2; rdhte 0.2.0; sandwich 3.1.1; rdrobust 4.0.0 coef, coef.bc, se.rb, vcov rel 1e-9 (observed 2.2e-12); bandwidths 1e-8 — / — test_rd_iv_rd_R_parity.py (+1)
rdhte_lincom R rdhte::rdhte_lincom 0.2.0 R 4.5.2; rdhte 0.2.0 estimate, z, CI, joint chi-square rel 1e-9; p 1e-8 — / — test_rd_iv_rd_R_parity.py (+1)
rdmc rdmulti::rdmc 2.0.0 (Cattaneo, Titiunik, Vazquez-Bare & Keele) R 4.5.2; rdmulti 2.0.0 per-cutoff coefficients, robust coefficients, robust SEs and the pooled weighted estimate 1e-9 rel; selected bandwidths 1e-5 rel — / — test_rdmulti_parity.py (+1)
rdms rdmulti::rdms; Stata side uses the rdms ado from the rdpackages GitHub mirror (rdmulti is not on SSC: ssc describe rdmulti returns r(601)), installed into a local gitignored ado path R 4.5.2 rel_est<=1e-06, rel_se<=1e-06 1.2e-12 / 5.4e-10 89_rdms.py (+2)
rdplot R rdrobust::rdplot 4.0.0 R 4.5.2; rdrobust 4.0.0 J / J_IMSE / J_MV and bin counts exact; bin means, SEs, t intervals and polynomial values rel 1e-9 (observed 4.0e-12); p = 4 global polynomial coefficients rel 1e-8 (observed 1.6e-10, raw-scale conditioning) — / — test_rd_iv_rd_R_parity.py (+1)
rdplotdensity R rddensity::rdplotdensity 2.6 (lpdensity 2.5) R 4.5.2; rddensity 2.6; lpdensity 2.5 f_p, f_q, se_p, se_q at every grid point rel 1e-8 (observed 8.0e-12); nh exact — / — test_rd_iv_rd_R_parity.py (+1)
rdpower rdpower::rdpower 3.0 (Cattaneo, Titiunik & Vazquez-Bare) R 4.5.2; rdpower 3.0 robust bias-corrected SE & power 1e-8 rel (observed 4.2e-14) — / — test_rdlocrand_parity.py (+1)
rdrandinf rdlocrand::rdrandinf 2.0 (Cattaneo, Titiunik & Vazquez-Bare) R 4.5.2; rdlocrand 2.0 observed statistic & asymptotic p-value 1e-8 rel (observed 2.3e-15) — / — test_rdlocrand_parity.py (+1)
rdrobust rdrobust::rdrobust R 4.5.2; rdrobust 3.0.0 rel_est<=1e-06, rel_se<=0.1 2.5e-14 / 9.4e-11 06_rd.py (+2)
rdsampsi rdpower::rdsampsi 3.0 (Cattaneo, Titiunik & Vazquez-Bare) R 4.5.2; rdpower 3.0 required sample sizes n_left / n_right / n_total asserted as exact integer equality (no tolerance) — / — test_rdlocrand_parity.py (+1)
rdwinselect rdlocrand::rdwinselect 2.0 (Cattaneo, Titiunik & Vazquez-Bare) R 4.5.2; rdlocrand 2.0 window grid 1e-12 rel (observed 0); per-window counts Nl / Nr asserted as exact integer equality — / — test_rdlocrand_parity.py (+1)
reciprocity R igraph::reciprocity igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Exact on a 40-node directed graph. — / — test_network_parity.py
regress lm + sandwich::vcovHC R 4.5.2; sandwich 3.1.1 rel_est<=1e-06, rel_se<=1e-06 1.1e-12 / 1.3e-12 01_ols.py (+2)
relative_risk base-R closed form (Katz-log; = epiR::epi.2by2 / Stata epitab) R 4.5.2 estimate, se_log, CI 1e-12 abs (observed 0) — / — test_epi_parity.py (+1)
reset_test lmtest::resettest(power=2:3, type='fitted') R 4.5.2; lmtest 0.9.40 F-statistic & p-value 1e-10 rel (observed ~1e-13) — / — test_diagnostics_parity.py (+1)
ri_test R ri2::conduct_ri (randomizr full enumeration); Stata ritest over the full assignment set R 4.5.2; ri2 0.5.0; Stata 18; ritest 1.1.7 p exact (k / N_assignments); statistic 1e-12 rel vs R, 2^-23 vs Stata — / — test_inference_sens_R_parity.py (+3)
rif_decomposition dineq::rif + manual OLS R 4.5.2; dineq 0.1.0 rel_est<=1e-06, rel_se<=1e-06 2.2e-15 / 1.4e-16 32_rif.py (+2)
rifreg R rifreg::rifreg (variance, quantiles) and dineq::rif + lm (Gini) rifreg 1.1.0; dineq 0.1.0; ddecompose 1.0.0 Coefficients at 1e-10 for the variance and the 10th/50th/90th percentiles (quantile_convention='rifreg') and for the Gini against dineq's exact RIF; the stock rifreg Gini, which integrates the Lorenz curve numerically, agrees to 1e-4. — / — test_decomp_R_parity.py
risk_difference base-R closed form (Wald; = epiR::epi.2by2 / Stata epitab) R 4.5.2 estimate, se, CI 1e-12 abs (observed 0) — / — test_epi_parity.py (+1)
rkd R rdrobust::rdrobust(deriv = 1, vce = 'hc1') 4.0.0; Stata rdrobust, deriv(1) 11.1.0 R 4.5.2; rdrobust 4.0.0; Stata 18 MP; rdrobust (Stata) 11.1.0 conventional and robust estimate / SE, bandwidth: rel 1e-9 (observed 5.2e-12 vs R, 6.0e-11 vs Stata) — / — test_rd_iv_rd_R_parity.py (+2)
rlasso R hdm::rlasso 0.3.2 R 4.5.2; hdm 0.3.2 support exact; beta / sigma / loadings / residuals atol 1e-6, lambda0 rtol 1e-8 (observed rel <= 1.8e-12) — / — test_rlasso_parity.py (+1)
rlasso_effect R hdm::rlassoEffect 0.3.2 R 4.5.2; hdm 0.3.2 alpha / se atol 1e-6 (observed rel <= 3.3e-15), incl. hdm's GrowthData vignette — / — test_rlasso_parity.py (+2)
rlasso_effects R hdm::rlassoEffects 0.3.2 R 4.5.2; hdm 0.3.2 alpha / se rtol 1e-9 on cps2012 (observed 2.8e-12); atol 1e-6 on the synthetic fixture — / — test_rlasso_parity.py (+2)
rlasso_iv R hdm::rlassoIV 0.3.2 R 4.5.2; hdm 0.3.2 coef / se atol 1e-6 (observed <= 3e-15); EminentDomain atol 1e-4 (observed 6.1e-9: pseudo-inverse of a rank-deficient control block) — / — test_rlasso_parity.py (+2)
rlassologit_effect R hdm::rlassologitEffect 0.3.2 R 4.5.2; hdm 0.3.2 alpha / se atol 1e-6 (observed rel 1.1e-15 / 5.3e-14 post; se 3.2e-7 with post=False) — / — test_rlassologit_effect_parity.py (+1)
rlassologit_effects R hdm::rlassologitEffects 0.3.2 R 4.5.2; hdm 0.3.2 coef / se atol 1e-6 (observed rel 8.6e-16 / 2.5e-14) — / — test_rlassologit_effect_parity.py (+1)
robust_synth scpi::scest(w.constr = list(name = 'ols')) with scdata(constant = TRUE) and stats::lm (unconstrained SC with intercept); glmnet (ridge / lasso / elastic net, unpenalised intercept) R 4.5.2; scpi 4.0.1; glmnet 4.1.10 OLS weights / intercept / fitted path 1e-10 rel (observed 8.4e-13); penalised paths 1e-8 rel, weights atol 1e-10 (observed 3.4e-11 abs, glmnet's coordinate-descent stop) — / — test_did_synth_synthvar_parity.py (+1)
roc_curve R pROC::roc/auc/var/ci.auc (DeLong); Stata roctab (default and hanley) R 4.5.2; pROC 1.19.0.1; Stata 18 MP 1e-10 rel (observed 6e-15) — / — test_survival_epi_R_parity.py (+2)
rook_weights spdep::poly2nb(queen = FALSE) + nb2listw(style = 'W') R 4.5.2; spdep 1.4.2; sf 1.1.1; spData 2.3.5 neighbour sets exact; weights 1e-12 rel (observed 0) — / — test_spatial_survey_R_parity.py (+1)
rosenbaum_bounds R DOS2::senWilcox (Rosenbaum); Stata rbounds; R stats::binom.test (sign test); R rbounds::psens (zero_method='wilcox', 4-dp) R 4.5.2; DOS2 0.5.2; rbounds 2.2; Stata 18; rbounds (Stata) 1.1.6 bounding p-values 1e-12 rel (observed 4.6e-15); Stata sig- 1e-15 abs; psens at its own 4-dp rounding — / — test_inference_sens_R_parity.py (+3)
rosenbaum_gamma R DOS2::senWilcox (Rosenbaum); Stata rbounds R 4.5.2; DOS2 0.5.2; Stata 18; rbounds (Stata) 1.1.6 bounding p-values 1e-12 rel (observed 4.6e-15) — / — test_inference_sens_R_parity.py (+1)
sac R spatialreg::sacsarlm spdep 1.4.2; spatialreg 1.4.3 rho and lambda at 1e-5, slope coefficients at 1e-6 -- a bounded two-parameter ML line search on both sides. — / — test_spdep_parity.py
sar spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) R 4.5.2; spatialreg 1.4.3 rel_est<=1e-06, rel_se<=1e-06 8.3e-08 / 5.1e-08 65_spatial.py (+2)
sar_gmm spatialreg::stsls(W2X=FALSE) / spatialreg::GMerrorsar R 4.5.2; spatialreg 1.4.3 rel_est<=1e-06, rel_se<=1e-06 4.6e-08 / 7.3e-16 66_spatial_gmm.py (+2)
sbw sbw::sbw 1.2 (Zubizarreta 2015), quadprog solver — ATT rel <= 1e-8 (observed <= 4e-10) under both standardisation conventions: tolerance_scale='target' == bal_std='target' and 'group' == bal_std='group', at bal_tol 0.05 and 0.02. — / — test_matching_r_parity.py (+1)
sc_estimate synthdid::sc_estimate 0.0.9 (Arkhangelsky, Athey, Hirshberg, Imbens & Wager) R 4.5.2; synthdid 0.0.9 estimate, unit/time weights and jackknife SE at 1e-9 on the Prop. 99 replica and a five-treated panel (observed est 8.1e-15, jackknife SE 4.8e-15); each of R's 40 placebo and 40 bootstrap replications replayed at 1e-9 (observed 2.1e-11) — / — test_did_synth_R_parity.py (+1)
scdata R scpi::scdata (features = outcome, no cov.adj, constant = FALSE) scpi 4.0.1; CVXR 1.9.2; ECOSolveR 0.6.1; Qtools 1.6.0; quantreg 6.1 A, B, P matrices identical (0 difference) on scpi_germany; donor order = R sort(B.names). — / — test_did_synth_scpi_parity.py
sdid synthdid::synthdid_estimate R 4.5.2; synthdid 0.0.9 rel_est<=1e-06, rel_se<=1e-06 7.8e-16 / 7.2e-08 12_sdid.py (+2)
sdm spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) R 4.5.2; spatialreg 1.4.3 rel_est<=1e-06, rel_se<=1e-06 8.3e-08 / 5.1e-08 65_spatial.py (+2)
sem spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) R 4.5.2; spatialreg 1.4.3 rel_est<=1e-06, rel_se<=1e-06 8.3e-08 / 5.1e-08 65_spatial.py (+2)
sem_gmm spatialreg::stsls(W2X=FALSE) / spatialreg::GMerrorsar R 4.5.2; spatialreg 1.4.3 rel_est<=1e-06, rel_se<=1e-06 4.6e-08 / 7.3e-16 66_spatial_gmm.py (+2)
sensemakr sensemakr::sensemakr R 4.5.2; sensemakr 0.1.6 rel_est<=1e-06, rel_se<=1e-06 5.0e-08 / 5.0e-08 22_sensemakr.py (+2)
sensitivity_specificity R epiR::epi.tests (method wilson / exact); Stata diagti R 4.5.2; epiR 2.0.94; Stata 18 MP; diagt 2.032 (diagti 2.053) 1e-12 rel on intervals, 1e-10 on ratios (observed 2e-15) — / — test_survival_epi_R_parity.py (+2)
shift_share_political AER::ivreg + sandwich HC1; ShiftShareSE::ivreg_ss (EHW / AKM / AKM0); bartik.weight::bw; anova(lm) share balance R 4.5.2; AER 1.2.16; sandwich 3.1.1; ShiftShareSE 1.1.0; bartik.weight 0.1.0 estimate / SEs / Rotemberg / F 1e-9 rel (observed <= 5e-15); AKM p-value 1e-7 rel (ShiftShareSE uses 2*(1-pnorm), cancellation at p ~ 1e-9) — / — test_synth_rest_R_parity.py (+1)
shift_share_political_panel fixest::feols 0.14.0 (ssc adj=FALSE, cluster.adj=FALSE; unit / time / two-way clusters; unit, time, two-way FE; unbalanced panel); ShiftShareSE::ivreg_ss with FE dummies (AKM); bartik.weight::bw with FE dummies; AER + HC0 per period R 4.5.2; fixest 0.14.0; ShiftShareSE 1.1.0; bartik.weight 0.1.0; AER 1.2.16; sandwich 3.1.1 estimate / SEs / Rotemberg / first-stage F 1e-9 rel (observed <= 1.1e-14) — / — test_synth_rest_R_parity.py (+1)
shift_share_se R ShiftShareSE::ivreg_ss (Adao, Kolesar & Morales), AKM row; Stata SSC ivreg_ss R 4.5.2; ShiftShareSE 1.1.0; Stata 18 MP; ivreg_ss SSC 20241116 1e-9 rel on beta and the AKM / AKM0 / EHW / Homoscedastic SEs; observed <= 6e-14 — / — test_did_synth_shiftshare_parity.py (+2)
slx R spatialreg::lmSLX spdep 1.4.2; spatialreg 1.4.3 Every coefficient at 1e-10. — / — test_spdep_parity.py
snips Open Bandit Pipeline (obp) 0.5.7 SelfNormalizedInverseProbabilityWeighting obp 0.5.7 value 1e-12 rel (observed 0.0); delta-method SE identity 1e-10 and equal to sp.ope.snips — / — test_ml_causal_obp_parity.py (+1)
source_decompose Stata descogini (Lerman-Yitzhaki) Stata 18 MP; descogini SSC With gini='population': total Gini, and each source's S_k, G_k, R_k and share of the total at 1e-12. The default n/(n-1)-corrected Gini differs by exactly that factor; S_k, R_k and shares are identical. — / — test_decomp_R_parity.py
spatial_iv sphet::spreg(model = 'lag', het = TRUE) 2.1.1 R 4.5.2; sphet 2.1.1; spdep 1.4.2 coefficients and HC0 SEs 1e-10 rel (observed 9.5e-14 / 6.8e-14) — / — test_spatial_survey_R_parity.py (+1)
spillover_did R did::att_gt(control_group='nevertreated') + did::aggte(type='simple') per group (direct / ring r, ring cohort = exposure onset); single cohort also fixest::feols(dbar ~ treat + ring1 + ring2, vcov='hetero', ssc(adj=FALSE)) did 2.3.0; DRDID 1.2.3; fixest 0.14.0 Direct and ring effects, their SEs and every (group, onset cohort, period) cell at 1e-9 relative on a single-cohort and a staggered spatial panel; observed agreement 1e-14. The ring construction is recomputed independently in R from the coordinates. — / — test_did_synth_misc_parity.py
sqreg R quantreg::rq (Barrodale-Roberts), Koenker 2005 quantreg see sqreg_R.json provenance Coefficients 3.5e-14 against quantreg::rq at tau = 0.25 / 0.50 / 0.75 -- both sides minimise the same pinball loss with the same simplex. Standard errors differ from R's se='iid' by ONE SCALAR PER QUANTILE, constant across coefficients to 6e-16: the sandwich is identical and only the sparsity estimate 1/f(0) differs (Powell kernel here, Koenker-Bassett with a Siddiqui/Hall-Sheather bandwidth there). The test asserts the ratio's constancy rather than a numerical band, which a structural difference could not satisfy. R's default se='nid' (Hendricks-Koenker, also Stata qreg's) is a third convention and is recorded as one. — / — test_sqreg_parity.py
ssaggregate R ShiftShareSE::ivreg_ss / reg_ss (Adao, Kolesar & Morales) and ssaggregate (Borusyak, Hull & Jaravel; R kylebutts/ssaggregate + AER::ivreg/sandwich HC0); Stata SSC ivreg_ss / reg_ss and ssaggregate + ivreg2, robust R 4.5.2; ShiftShareSE 1.1.0; ssaggregate (R) 0.0.0.9000 (GitHub kylebutts/ssaggregate@22df93980250891a0cc247f6020136cd33c65ba2); AER 1.2.16; sandwich 3.1.1; Stata 18 MP; reg_ss / ivreg_ss SSC 20241116; ssaggregate (Stata) SSC 1.2.2 (20200826) 1e-9 rel on beta, every SE row (Homoscedastic, EHW, Reg. cluster, AKM, AKM0), AKM/AKM0 CIs and the shock-level frame; p-values also atol 1e-15 (references use 2*(1-Phi)); observed <= 6e-14 (frame 2.9e-13) — / — test_did_synth_shiftshare_parity.py (+2)
stabilized_weights ipw::ipwtm 1.3.0 (type = 'all'; binomial logit; gaussian via geepack::geeglm) R 4.5.2; ipw 1.3.0; geepack 1.3.13 1e-10 rel on every weight (observed 1.5e-12) — / — test_teffects_R_parity.py (+2)
stacked_did hand-written stack + fixest::feols R 4.5.2; fixest 0.14.0 rel_est<=1e-06 3.9e-13 / 7.1e-13 75_stacked.py (+2)
staggered_cs staggered::staggered_cs 1.2.2 (Roth & Sant'Anna) R 4.5.2; staggered 1.2.2 estimate, Neyman SE and adjusted SE at abs 1e-9 on mpdta, a randomised rollout and a null panel x simple/cohort/calendar (observed rel 6.6e-14) — / — test_staggered_extended_parity.py (+1)
staggered_rollout staggered::staggered / staggered_cs / staggered_sa (1.2.2) R 4.5.2 rel_est<=1e-10 9.1e-16 / 5.7e-16 82_staggered.py (+2)
staggered_sa staggered::staggered_sa 1.2.2 (Roth & Sant'Anna) R 4.5.2; staggered 1.2.2 estimate, Neyman SE and adjusted SE at abs 1e-9 on mpdta, a randomised rollout and a null panel x simple/cohort/calendar (observed rel 6.6e-14) — / — test_staggered_extended_parity.py (+1)
staggered_synth augsynth::multisynth 0.2.0 (Ben-Michael, Feller & Rothstein; partially pooled SCM for staggered adoption) R 4.5.2; augsynth 0.2.0; osqp 1.0.0 ATT, per-unit / per-cohort ATT, event-time ATT and jackknife SE 1e-8 rel (observed <= 2.6e-11); weights 1e-8 abs (observed 2.1e-10); nu 1e-8, imbalance norms 1e-6 rel — / — test_did_synth_synthvar_parity.py (+1)
structural_break strucchange::Fstats + sctest(type = 'supF') / breakpoints; mbreaks::dosequa (Bai-Perron sequential); Stata estat sbsingle R 4.5.2; strucchange 1.5.4; mbreaks 1.0.1; Stata 18 statistics 1e-10 rel (observed <= 1e-14); sup-F p-values 1e-10 rel for p > 1e-6, 1e-13 abs below (R uses 1 - pchisq); break dates exact — / — test_r2_ts_parity.py (+3)
subcluster_wild_bootstrap Stata boottest, bootcluster() (CRVE clustered at g6, signs flipped at s12) Stata 18; boottest 4.5.3 p exact (multiple of 1/4096); t 1e-10 rel (observed 1.8e-14) — / — test_inference_sens_stata_parity.py (+1)
subgroup_analysis Stata regress + testparm Stata 18 1e-9 rel — / — test_misc_sens_stata_parity.py
subgroup_decompose Stata ineqdeco, bygroup() Stata 18 MP; ineqdeco SSC GE(0), GE(1), GE(2) totals and within components at 1e-12, between components at 1e-11. The Gini path (Dagum) has no ineqdeco counterpart and keeps its analytical checks. — / — test_decomp_R_parity.py
sun_abraham fixest::sunab R 4.5.2; fixest 0.14.0 rel_est<=1e-06, rel_se<=0.03 2.8e-11 / 2.7e-11 05_sunab.py (+2)
sureg systemfit::systemfit(method="SUR", noDfCor) R 4.5.2; systemfit 1.1.30 rel_est<=1e-06, rel_se<=1e-06 1.5e-14 / 1.5e-15 60_sureg.py (+2)
survivor_average_causal_effect Stata leebounds 1.5 (Zhang-Rubin SACE bounds = Lee bounds under monotonicity) Stata 18; leebounds 1.5 (2013-07-17, Tauchmann) 1e-10 rel (upper bound; both bounds equal sp.lee_bounds) — / — test_teffects_R_parity.py (+2)
svydesign survey::svydesign + svymean/svytotal/svyglm/degf (strata, nested PSUs, fpc, survey.lonely.psu); Stata svyset + svy: mean/total/regress/logit/poisson + estat effects R 4.5.2; survey 4.5; Stata 18 MP estimates, SEs, DEFF, CI bounds 1e-10 rel, p-values 1e-9 rel (observed <= 4e-14; p 6e-13) — / — test_survey_design_R_parity.py (+2)
svyglm survey::svyglm (design-based GLM + linearization SE) R 4.5.2 coefficients + SE 1e-10 abs (observed ~2e-15 / 6e-15) — / — test_survey_parity.py (+1)
svymean survey::svymean (Horvitz-Thompson/Hajek + Taylor SE) R 4.5.2 estimate + SE 1e-10 abs (observed ~5e-15 / 8e-17) — / — test_survey_parity.py (+1)
svytotal survey::svytotal (Horvitz-Thompson + Taylor SE) R 4.5.2 estimate 1e-12 rel; SE 1e-10 rel (observed ~2e-12 / 1e-14) — / — test_survey_parity.py (+1)
synth Synth::synth R 4.5.2; Synth 1.1.10 rel_est<=1e-06, rel_se<=1e-06 7.8e-08 / 7.7e-08 52_scm_unique.py (+2)
synth_donor_sensitivity Synth::synth 1.1.10 on the replayed numpy donor subsets (custom.v matched; quadprog exact) R 4.5.2; Synth 1.1.10; quadprog 1.5.8 ATT and pre-RMSPE 1e-9 rel (observed <= 1e-13) — / — test_synth_rest_R_parity.py (+1)
synth_loo Synth::synth 1.1.10 (per fit, custom.v matched; QP solved exactly by quadprog::solve.QP) R 4.5.2; Synth 1.1.10; quadprog 1.5.8; kernlab 0.9.33 ATT and pre-RMSPE 1e-9 rel (observed <= 3e-14) — / — test_synth_rest_R_parity.py (+1)
synth_rmspe_filter SCtools::mspe.test 0.3.3.1 (discard.extreme, mspe.limit 2/5/20) on Synth fits R 4.5.2; SCtools 0.3.3.1; Synth 1.1.10; quadprog 1.5.8 placebo pre-RMSPE and ratios 1e-9 rel (observed <= 1e-13); p-values and kept counts exact — / — test_synth_rest_R_parity.py (+1)
synth_time_placebo Synth::synth 1.1.10 (per placebo time, custom.v matched; quadprog exact) R 4.5.2; Synth 1.1.10; quadprog 1.5.8 placebo ATT 1e-9 rel (observed <= 1e-13) on identified placebo times — / — test_synth_rest_R_parity.py (+1)
synthdid_estimate synthdid::synthdid_estimate R 4.5.2; synthdid 0.0.9 rel_est<=1e-06, rel_se<=1e-06 7.8e-16 / 7.2e-08 12_sdid.py (+3)
synthesise_evidence R metafor::rma(method = 'FE') metafor 5.0.1 1e-12 rel — / — test_misc_sens_R_parity.py
tF_adjustment R ivDiag::tF 1.0.6 (LMMP 2022 tF table) R 4.5.2; ivDiag 1.0.6 rel 1e-12 on the same 26 F values (observed 0) — / — test_rd_iv_R_parity.py (+1)
tF_critical_value R ivDiag::tF 1.0.6 (LMMP 2022 tF table) R 4.5.2; ivDiag 1.0.6 critical value at 26 F values from 4 to 1e4: rel 1e-12 (observed 0) — / — test_rd_iv_R_parity.py (+1)
test_calibration grf::test_calibration 2.6.1 (vcov.type HC3 default and HC1), forest outputs held fixed R 4.5.2; grf 2.6.1; sandwich 3.1.1 coef, se, t 1e-10 rel (observed 1.1e-15); one-sided p-value 1e-8 rel (observed 1.6e-13) — / — test_ml_causal_R_parity.py (+1)
three_sls R systemfit::systemfit(method='3SLS') R 4.5.2; systemfit 1.1.30 coef 1e-9 abs (observed <= 1e-15); SE ~5e-3 rel — / — test_threesls_parity.py (+1)
tmle tmle::tmle R 4.5.2; tmle 2.1.1 rel_est<=1e-06, rel_se<=1e-06 1.9e-09 / 1.9e-09 72_tmle.py (+2)
tobit censReg::censReg R 4.5.2; censReg 0.5.38 rel_est<=1e-06, rel_se<=1e-05 2.8e-08 / 2.8e-08 41_tobit.py (+2)
transitivity R igraph::transitivity(type = 'global') igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Exact on karate. — / — test_network_parity.py
transport_weights_fn R glm + quantile(type = 7) + sandwich::vcovHC(HC0) sandwich 3.1.1 1e-12 rel — / — test_misc_sens_R_parity.py
truncreg truncreg::truncreg(method="NR") R 4.5.2; truncreg 0.2.5 rel_est<=1e-06, rel_se<=1e-06 3.5e-10 / 7.9e-08 62_truncreg.py (+2)
twoway_cluster sandwich::vcovCL(cluster=~g1+g2) R 4.5.2; sandwich 3.1.1 rel_est<=1e-06, rel_se<=1e-06 7.8e-16 / 7.8e-16 54_twoway_cluster.py (+2)
unified_sensitivity R EValue::evalues.OLS; sensemakr::sensemakr (rv_q, rv_qa) EValue 4.1.4; sensemakr 0.1.6 1e-10 rel — / — test_misc_sens_R_parity.py
var vars::VAR R 4.5.2; vars 1.6.1 rel_est<=1e-06, rel_se<=1e-06 3.1e-15 / 6.6e-15 33_var.py (+2)
vif R car::vif ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 2.0e-16 on every variance inflation factor, once the returned values stopped being rounded to two decimals. — / — test_weakiv_meta_parity.py
wild_cluster_boot R fwildclusterboot::boottest; Stata boottest (WCR, Rademacher, full enumeration) R 4.5.2; fwildclusterboot 0.14.3; Stata 18; boottest 4.5.3 p exact (multiple of 1/4096); t 1e-10 rel (observed 1.9e-14) — / — test_inference_sens_R_parity.py (+3)
wild_cluster_bootstrap R fwildclusterboot::boottest; Stata boottest (WCR, Rademacher, full enumeration) R 4.5.2; fwildclusterboot 0.14.3; Stata 18; boottest 4.5.3 p exact (multiple of 1/4096); t 1e-10 rel (observed 1.9e-14) — / — test_inference_sens_R_parity.py (+3)
wild_cluster_ci_inv R fwildclusterboot::boottest confidence interval (uniroot tol 1e-13) R 4.5.2; fwildclusterboot 0.14.3 CI endpoints 1e-9 rel (observed 4.8e-12) — / — test_inference_sens_R_parity.py (+2)
wooldridge_did etwfe::etwfe + emfx R 4.5.2; etwfe 0.6.2 rel_est<=1e-06, rel_se<=0.001 1.8e-13 / 3.9e-14 17_etwfe.py (+2)
xtabond plm::pgmm R 4.5.2; plm 2.6.7 rel_est<=1e-06, rel_se<=1e-06 9.0e-16 / 1.4e-15 50_xtabond.py (+2)
xtdpdsys Stata 18 xtdpdsys (built-in); xtabond2 iv(x, eq(diff)) h(2) Stata 18 MP; xtabond2 SSC 03.07.00 abdata, n on L.n (and w k): coefficients and SEs at 1e-9 (observed <= 6e-12) for one-step robust, two-step Windmeijer and classical one-step, instrument count equal; xtabond2 with iv(w k, eq(diff)) h(2) reproduces the same numbers. — / — test_dynpanel_abdata_parity.py
xtlsdvc Stata xtlsdvc V1.0.4 (Bruno 2005), SSC Stata 18 MP; xtlsdvc 1.0.4 abdata: bias-corrected coefficients for initial(ab/ah/bb) x bias(1/2/3) and the AR(1)-only model at rtol 1e-7 (observed 4e-14 to 1.6e-9, the looser end through the Anderson-Hsiao initialiser). Standard errors are excluded by design: xtlsdvc reports the uncorrected LSDV ones, which sp.xtlsdvc reproduces and warns about. — / — test_lsdvc_parity.py
xtnbreg Stata 18 xtnbreg, fe / re (Hausman-Hall-Griliches); pglm::pglm(family = negbin, model = 'within' / 'random') Stata 18 MP; R 4.5.2; pglm 0.2.4 coefficients / SEs rtol 1e-7 (observed <= 2.7e-8: Stata fe's score at its own estimate is 3e-7, pglm's gradtol), log-likelihood rtol 1e-12 — / — test_panel_xtnbreg_parity.py
yu_elwert_decompose R cdgd::cdgd0_manual on independently fitted within-cell lm / within-group glm nuisances DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 method='efficient': disparity, baseline, prevalence, effect, selection and their EIF standard errors at 1e-9. method='plugin' has no reference implementation and is covered by its exact additivity identity. — / — test_decomp_R_parity.py
zip_model pscl::zeroinfl(dist="poisson") R 4.5.2; pscl 1.5.9 rel_est<=1e-06, rel_se<=0.0001 1.2e-07 / 2.8e-08 63_zip.py (+2)

aligned — 53 functions

Agreement within a documented, pre-registered looser tolerance.

function reference versions tolerance rel err (R / Stata) test
ackerberg_caves_frazer Stata prodest, method(lp) acf valueadded; R prodest::prodestACF Stata 18 MP; prodest SSC; R prodest 1.0.2 StatsPAI returns an exact root of the just-identified moment conditions (criterion < 1e-25), the one nearest the stage-1 coefficients; R stops 2.4e-6 (relative) from that root with criterion 9e-16. Stata's Nelder-Mead stops at non-roots (criterion 4e-6 / 7e-6), which the test records. The parity panel has three roots, all reported in diagnostics['acf_roots']. — / — test_prodest_parity.py
aft survival::survreg (Weibull AFT) R 4.5.2; survival 3.8.3 coefficients & log-scale 5e-5 abs (observed ~1e-5) — / — test_aft_parity.py (+1)
augsynth augsynth::augsynth R 4.5.2; augsynth 0.2.0 rel_est<=2e-05, rel_se<=1e-06 7.9e-06 / 1.7e-08 18_augsynth.py (+2)
blp pyblp 1.2.0 (Conlon & Gortmaker), identical Halton nodes via agent_data pyblp 1.2.0; numpy 2.2.6; python 3.10.20 beta, sigma, SEs, objective/N, own elasticities 1e-6 rel (observed <= 1.6e-8) — / — test_blp_pyblp_parity.py (+1)
breakdown_m HonestDiD::findOptimalFLCI 0.2.8 (Rambachan & Roth), breakdown by uniroot on the bound facing zero R 4.5.2; HonestDiD 0.2.8; CVXR 1.8.2 vs HonestDiD with its Monte-Carlo folded-normal quantile replaced by the exact one: 1e-9 (observed 6.1e-11); vs HonestDiD as shipped: 1e-3 (observed 4.3e-4), the simulation error of .qfoldednormal (1e6 draws, seed 0; 1.96224 vs exact 1.95996 at mu = 0) — / — test_did_synth_R_parity.py (+1)
cate_eval grf::rank_average_treatment_effect 2.6.1 (AUTOC, QINI, TOC), fed grf's nuisances; shares sp.rate's operator R 4.5.2; grf 2.6.1 Point estimate (AUTOC, QINI) and TOC curve on q = 0.1..1 exact at 1e-10 rel (observed 4.6e-15), including a 30-group tied-priority case against rank_average_treatment_effect.fit. SE is T3: the analytic rank-corrected influence-function SE is compared with grf's half-sample bootstrap (R = 2000) at 6% rel (observed 2.6% AUTOC, 1.7% QINI). — / — test_ml_causal_R_parity.py (+1)
causal_forest grf::causal_forest R 4.5.2; grf 2.6.1 rel_est<=0.01, rel_se<=0.05 2.4e-03 / — 13_causal_forest.py (+1)
cbps CBPS::CBPS 0.24 (Imai & Ratkovic 2014) — ATE over/exact and ATT exact: rel <= 5e-3 (R's optimiser slack). ATT over is NOT pinned to R -- CBPS's ATT gradient mis-scales the balance block by n/n_1 and stops off-stationarity; StatsPAI is asserted to attain strictly better covariate balance instead. — / — test_matching_r_parity.py (+1)
centrality R igraph degree / betweenness / closeness / page_rank / eigen_centrality igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 degree, betweenness (normalised and raw), closeness and PageRank at 1e-10 on Zachary's karate club. The eigenvector column is L2-normalised (networkx) where igraph max-scales it: the ratio is constant across nodes to 1e-14, so it is the same vector under a documented normalisation. — / — test_network_parity.py
cfm_decompose Stata cdeco, method(logit) 1.0.2 with drprocess (Chernozhukov, Fernandez-Val & Melly) Stata 18.0 MP; cdeco 1.0.2 01mar2023 (bmelly/Stata counterfactual, Distribution-Date 20220803) CDFs at the thresholds 1e-8 abs (observed 2.4e-10: Stata's logit stops at its default tolerance); quantiles 1e-12 abs (observed 0) — / — test_decomp_qte_parity.py (+1)
cic qte::CiC R 4.5.2 rel_est<=1e-06 2.4e-15 / 6.0e-03 74_cic.py (+2)
cloglog stats::glm(binomial('cloglog')) R 4.5.2 coefficients 5e-5 abs (observed ~1e-5; IRLS convergence) — / — test_glm_ext_parity.py (+1)
community_detection R igraph::cluster_louvain (T3: randomised on both sides) igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 T3, not T2: 200 seeded runs on karate have mean modularity within four combined standard errors of igraph's 200 runs (0.4157 vs 0.4145), and both reach the same maximum, 0.41979, the known optimum for this graph. — / — test_network_parity.py
conditional_lr_ci R ivmodel::CLR ivmodel 1.9.1 Monte Carlo by construction and graded as such: ivmodel integrates Moreira's conditional distribution while StatsPAI simulates it, so the two cannot agree deterministically. The test asserts the error SHRINKS with n_sim rather than pinning a number -- 3.8e-3 at n_sim=5,000 and 1.7e-4 at 200,000 -- which is the only honest statement about a simulated critical value. The endpoints themselves are bisected off-grid, so the residual is the critical value and not the grid. — / — test_weakiv_meta_parity.py
cox_frailty R survival::coxph(... + frailty(id, theta=, sparse=FALSE)); Stata stcox, shared() R 4.5.2; survival 3.8.3; Stata 18 MP fixed theta: beta and SE 1e-9, integrated log likelihood 1e-10 (observed 4e-14); theta maximiser 1e-6 vs R optimize (observed 5e-8) and 5e-5 vs Stata e(theta) (observed 7.5e-6) — / — test_survival_epi_R_parity.py (+2)
demeaned_synth augsynth::augsynth(progfunc = 'None', fixedeff = TRUE) 0.2.0 (de-meaned SCM) R 4.5.2; augsynth 0.2.0; osqp 1.0.0 gap path and ATT 1e-7 rel (observed 1.2e-9); weights 1e-8 abs — / — test_did_synth_synthvar_parity.py (+1)
dml_model_averaging ddml::ddml_plm 0.3.1 (shortstack = TRUE, ensemble_type = 'nnls1'), OLS candidates, shared folds R 4.5.2; ddml 0.3.1; sandwich 3.1.1 short-stacking weights 1e-10 (observed 4.4e-14); ddml's final lm(y_r ~ d_r) with intercept and HC1 SE rebuilt from StatsPAI's stacked residuals at 1e-10 (observed 4.8e-16) — / — test_ml_causal_dml_parity.py (+2)
functional_form_test didFF::didFF R 4.5.2 rel_est<=0.001 1.3e-14 / — 79_didff.py (+1)
garch Stata 18 arch; rugarch::ugarchfit 1.5.6 R 4.5.2; rugarch 1.5.6; Stata 18 log-likelihood at the reference optimum 1e-12 (observed 8.5e-16); vs Stata: b 1e-5, SE 2e-5 (observed 8.2e-6 / 1e-5); vs rugarch: params 5e-4, SE at their parameters 1e-2 — / — test_timeseries_R_parity.py (+2)
genmatch Matching::Match 4.10-15 (Weight = 3, Weight.matrix) — Deterministic kernel only: given the same diagonal W, the 1-NN assignment agrees with Matching::Match on all 163 uniquely matched treated units on MatchIt::lalonde. — / — test_matching_r_parity.py (+1)
grapple GRAPPLE::grappleRobustEst (GitHub jingshuw/GRAPPLE 317e837) GRAPPLE 0.2.2 sandwich at GRAPPLE's estimates 1e-6 (l2 1e-12); fitted beta 1e-3, tau2 / SE 1e-4 — / — test_misc_sens_R_parity.py
hits R igraph::hits_scores igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 Hub and authority vectors are igraph's up to normalisation: L1 here (documented), max = 1 in igraph; the ratio is constant across nodes to 1e-11. — / — test_network_parity.py
lcsf sfaR::sfalcmcross 1.0.1 (2 classes, half-normal) R 4.5.2; sfaR 1.0.1 estimates 1e-6 rel (observed 1.1e-8); OIM SEs 1e-6 rel (observed 8.6e-8) — / — test_frontier_struct_R_parity.py (+1)
levinsohn_petrin Stata prodest, method(lp) valueadded (Rovigatti & Mollisi); R prodest::prodestLP Stata 18 MP; prodest SSC; R prodest 1.0.2 As olley_pakes with materials as the proxy: free-input coefficient at 1e-12 against both; state coefficient at the minimiser, R 1.3e-6 and Stata up to 1.8e-3 away where their optimisers stop. — / — test_prodest_parity.py
lincom Stata 18 lincom (after regress / ivregress / logit / poisson) Stata 18 1e-6 rel (observed <= 2.3e-15 on linear fits, <= 2.3e-8 overall) — / — test_postestimation_stata_parity.py (+1)
malmquist sfaR::sfacross 1.0.1 per period + sfaR::efficiencies (teBC / teJLMS); TC from sfaR betas R 4.5.2; sfaR 1.0.1; metafrontier 0.3.1 EC / TC / M 1e-6 rel (observed 2.7e-7); metafrontier::malmquist_meta(method='sfa') EC_group and 1/TC_group 1e-4 (observed 1.8e-5, reference optim BFGS with finite-difference gradients) — / — test_r2_frontier_parity.py (+1)
margins Stata 18 margins, dydx(*) Stata 18 1e-6 rel (observed 2.7e-7 AME / 7.5e-8 SE on probit and logit-with-interaction; <= 1e-10 otherwise) — / — test_postestimation_stata_parity.py (+1)
melogit lme4::glmer R 4.5.2; lme4 2.0.1 rel_est<=0.0002, rel_se<=0.01 1.7e-04 / 1.4e-08 26_glmm_logit.py (+2)
model_averaging_dml ddml::ddml_plm 0.3.1 (shortstack = TRUE, ensemble_type = 'nnls1'), OLS candidates, shared folds R 4.5.2; ddml 0.3.1; sandwich 3.1.1 short-stacking weights 1e-10 (observed 4.4e-14); ddml's final lm(y_r ~ d_r) with intercept and HC1 SE rebuilt from StatsPAI's stacked residuals at 1e-10 (observed 4.8e-16) — / — test_ml_causal_dml_parity.py (+2)
mr_raps R mr.raps 0.4.3 (simple / overdispersed / overdispersed.robust) MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 Simple and L2-overdispersed fits (beta, SE, tau2) at 1e-8. Robust Huber / Tukey: the sandwich reproduces R's SEs at 1e-10 when evaluated at R's own (beta, tau2) and integrate() moments; the fitted values agree to 5e-5 (beta), 5e-4 (SE), 1.5e-3 (tau2) because R stops uniroot at its default tolerance and integrate() at 1.2e-4 -- a test shows StatsPAI's root satisfies the estimating equation more tightly than R's. — / — test_mr_R_parity.py
olley_pakes Stata prodest, method(op) valueadded (Rovigatti & Mollisi); R prodest::prodestOP Stata 18 MP; prodest SSC; R prodest 1.0.2 Simulated unbalanced panel with 31 calendar gaps. Free-input coefficient (stage-1 OLS) at 1e-12 against both, polynomial degree 2 and 3. State coefficient: StatsPAI solves the first-order condition and has a sum of squares no larger than at either reference estimate; the references stop their optimisers early (R BFGS 5.9e-6 away, Stata Nelder-Mead at tolerance 1e-5 up to 4.3e-3 away). — / — test_prodest_parity.py
optimal_match optmatch::pairmatch 0.10.8 on a logit propensity score — Total matched distance <= optmatch's (1 + 1e-6). The matched pairs are not pinned: the assignment problem is degenerate on this data, so equally optimal solutions report different ATTs. — / — test_matching_r_parity.py (+1)
pretrends_power pretrends::pretrends / pretrends::slope_for_power (GitHub, not CRAN) R 4.5.2 rel_est<=0.001 4.0e-05 / 1.4e-04 76_pretrends.py (+2)
pretrends_slope_for_power pretrends::pretrends / pretrends::slope_for_power (GitHub, not CRAN) R 4.5.2 rel_est<=0.001 4.0e-05 / 1.4e-04 76_pretrends.py (+2)
qdid qte::QDiD 1.3.1 — max deviation / scale < 0.08, sign agreement on the large effects and correlation > 0.999. R's quantiles come from BMisc::weighted_quantile (stats::optimize on a piecewise-linear check function, which has plateaus where every point is a minimiser) while sp.qdid interpolates the empirical inverse CDF; on this fixture the gap is at most ~152 currency units against effects running to ~8900. — / — test_qdid_parity.py (+1)
qte qte::ci.qte / qte::ci.qtet 1.3.1 (Firpo 2007) — max relative deviation < 0.01 on lalonde.exp / lalonde.psid. Both sides minimise the same weighted check function; R's BMisc::weighted_quantile uses a golden-section search whose answer on a plateau is an optimiser artifact, so point-value equality is not asserted. — / — test_firpo_qte_parity.py (+1)
rate grf::rank_average_treatment_effect and rank_average_treatment_effect.fit 2.6.1 (AUTOC, QINI, TOC), forest outputs held fixed R 4.5.2; grf 2.6.1 Point estimate (AUTOC, QINI) and TOC curve on q = 0.1..1 exact at 1e-10 rel (observed 4.6e-15), including a 30-group tied-priority case against rank_average_treatment_effect.fit. SE is T3: the analytic rank-corrected influence-function SE is compared with grf's half-sample bootstrap (R = 2000) at 6% rel (observed 2.6% AUTOC, 1.7% QINI). — / — test_ml_causal_R_parity.py (+1)
rd_bias_aware_fuzzy R RDHonest 1.0.1.9000, RDHonest(y d ~ x) R 4.5.2; RDHonest 1.0.1.9000 estimate, std.error, maximum.bias, conf.low/high, M, first stage: rel 1e-9 at fixed h and M (observed 1.1e-14), 1e-6 with selected h (observed 3.5e-8, optimiser) — / —
rd_discrete R RDHonest::RDHonest / RDHonestBME 1.0.1.9000 (Kolesar) R; RDHonest 1.0.1.9000 estimate, std.error, maximum.bias, conf.low/high rel 1e-9 at fixed h and for BME; 1e-6 when RDHonest selects h and M (optimiser) — / — test_rd_open_R_parity.py (+1)
rdrbounds R rdlocrand::rdrbounds 2.0 R 4.5.2; rdlocrand 2.0 upper and lower bounds within 4 pooled binomial SEs at 4000 draws (Monte Carlo on both sides, not parity) — / — test_rd_iv_rd_R_parity.py (+1)
rdsensitivity R rdlocrand::rdsensitivity / rdrandinf 2.0 R 4.5.2; rdlocrand 2.0 per-window estimates rel 1e-9 (rdrandinf observed statistic); randomization p-values within 4 pooled binomial SEs at 4000 draws (Monte Carlo on both sides, not parity) — / — test_rd_iv_rd_R_parity.py (+1)
rlassologit R hdm::rlassologit 0.3.2 (glmnet 4.1.10 engine) R 4.5.2; hdm 0.3.2; glmnet 4.1.10 support exact; post-Lasso atol 1e-5 (observed rel 1.7e-10); post=False atol 1e-4 (observed 1.3e-6: glmnet coordinate-descent convergence) — / — test_rlassologit_parity.py (+1)
sarar_gmm spatialreg::gstsls 1.4.3 (Kelejian-Prucha GS2SLS) R 4.5.2; spatialreg 1.4.3; spdep 1.4.2 beta, rho and SEs 1e-9 rel (observed 3.8e-11 / 3.0e-11); lambda 1e-7 rel (observed 6.5e-9) — / — test_spatial_survey_R_parity.py (+1)
scest R scpi::scest (w.constr simplex / lasso / ridge / ols / L1-L2, V = 'separate') scpi 4.0.1; CVXR 1.9.2; clarabel 0.11.2; osqp 1.0.0 ols and lasso weights at 1e-9 abs; ridge Q / lambda and L1-L2 Q2 at 1e-10. simplex / ridge / L1-L2 weights are bounded by the conic solver: R's objective exceeds StatsPAI's exact optimum by <= 1e-7 relative (CLARABEL gap 1e-8), R's point is feasible, and w_R - w_py
scpi R scpi::scpi (effect = 'unit-time', u.missp, u.sigma = HC1, u.order = e.order = 1, rho = type-2, e.method = all) scpi 4.0.1; CVXR 1.9.2; ECOSolveR 0.6.1; Qtools 1.6.0; quantreg 6.1 On R's weights: rho, Q.star, u.mean, Omega, Sigma, e.mean at 1e-9; out-of-sample e.var and gaussian / ls / qreg bounds at 1e-9 against R with rrq(method = 'br') (exact LP) and 5e-4 abs against the default Frisch-Newton rrq. In-sample simulation fed R's draws: per-draw median <= 1e-6 and max <= 2e-4 vs ECOS at 1e-12, quantile bounds at 1e-5; vs default ECOS (1e-8) bounds within 2e-3 abs. Average-effect CI = scdataMulti(effect = 'unit') within 5e-4. J > T0 (California, 38 donors) covered. — / — test_did_synth_scpi_parity.py
shapley_inequality Stata shapley2 1.5 (Chavez Juarez) over regress + ineqdeco (Jenkins) Stata 18.0 MP; shapley2 1.5 10jun15 (SSC) value function v(S) 1e-11 rel (observed 1.4e-13); Shapley values 1e-6 rel (observed 2.7e-7, shapley2 routes v(S) through float variables); totals 1e-12 — / — test_decomp_qte_parity.py (+1)
spatial_panel splm::spml(model = 'within') 1.6.5; Stata xsmle 1.4.5 fe type(ind) vce(oim) R 4.5.2; splm 1.6.5; plm 2.6.7; Stata 18; xsmle version 1.4.5 5jun2017 vs splm: estimates and SEs 1e-7 rel (observed 7.2e-8 / 1.3e-8, the splm optimize floor); vs xsmle (tightened ml tolerances): estimates and beta SEs 1e-9 (observed 2.5e-13 / 8.5e-11), spatial-parameter SE 1e-8 (observed 3.0e-9) — / — test_spatial_survey_R_parity.py (+2)
survreg survival::survreg (Weibull AFT) R 4.5.2; survival 3.8.3 coefficients & log-scale 5e-5 abs (observed ~1e-5) — / — test_aft_parity.py (+1)
test Stata 18 test (after regress / ivregress / logit / poisson) Stata 18 1e-6 rel (observed <= 2.3e-15 on linear fits, <= 8e-10 on ML fits, far-tail p <= 4.7e-8) — / — test_postestimation_stata_parity.py (+1)
weakrobust R ivmodel::CLR 1.9.1; Stata weakiv 2.4.07 (md small) R 4.5.2; ivmodel 1.9.1; Stata 18 MP; weakiv 2.4.07 CLR statistic and p-value vs ivmodel rel 1e-9 (observed 1.2e-11); CLR set vs ivmodel 5e-5 (observed 1.2e-5, ivmodel's uniroot default tolerance); CLR / K / AR vs Stata weakiv 1e-6 (observed 1.1e-7 / 1.1e-7 / 2.1e-8, not bisected further) — / — test_rd_iv_R_parity.py (+2)
xtfrontier frontier::sfa R 4.5.2; frontier 1.1.8 rel_est<=0.001, rel_se<=0.05 2.8e-06 / 1.8e-06 29_panel_sfa.py (+2)
zinb pscl::zeroinfl(dist="negbin") R 4.5.2; pscl 1.5.9 rel_est<=1e-05, rel_se<=0.001 9.5e-07 / 4.5e-11 64_zinb.py (+2)
zisf Stata chks 1.1 (estimation(zsf) eoption(ml)); R sfa::zsfm 1.2.0 (ZISF / ZISF_Z, likelihood at its optimum) R 4.5.2; sfa 1.2.0; numDeriv 2016.8.1.1; stata 18; chks 1.1 (chks.pkg dated 20190320) estimates and OIM SEs 1e-6 rel (observed chks 8.6e-8 / 9.4e-8; sfa likelihood at its optimum 1.3e-8 / 5.4e-8); sfa's reported L-BFGS-B point 5e-5 / 5e-4 (observed 1.6e-5 / 1.8e-4) — / — test_r2_frontier_parity.py (+2)

external-replication — 2 functions

Reproduces published-paper numbers; sources in tests/external_parity/PUBLISHED_REFERENCE_VALUES.md.

function test
aggte test_honest_did_paper_parity.py (+1)
parallel_trends_robustness test_rebel_canal_published.py

analytical-only — 130 functions

Recovers a known DGP truth / closed-form identity within tolerance; no cross-package reference. See tests/reference_parity/REFERENCES.md.

function test
W test_spdep_parity.py
aggte_from_influence test_aggte_r_did_parity.py
always_treat test_longitudinal_parity.py
assimilative_causal test_assimilation_parity.py
auto_cate test_ml_causal_recovery_parity.py
auto_cate_tuned test_ml_causal_recovery_parity_round2.py
bayes_iv test_bayes_diagnostics_parity.py
bayes_rd test_bayes_diagnostics_parity.py
bayes_synth test_bayes_synth_parity.py
bcf test_bcf_parity.py
bcf_factor_exposure test_bcf_factor_exposure_parity.py
beyond_average_late test_beyond_average_late_parity.py (+2)
bidirectional_pci test_proximal_parity.py
bootstrap test_bootstrap_parity.py
boundary_rd test_rd_open_R_parity.py
bradford_hill test_bradford_hill_parity.py
breakdown_frontier test_breakdown_frontier_parity.py
bunching test_bunching_parity.py
bvar test_timeseries_R_parity.py (+1)
calibrate_cate test_panel_forest_recovery.py
cardinality_match test_cardinality_match_parity.py
causal_impact test_did_synth_misc_parity.py
causal_kalman test_assimilation_parity.py
causal_policy_forest test_ope_parity.py
check_absorbing test_absorbing_reference.py
clone_censor_weight test_target_trial_parity.py
cluster_cate test_ml_causal_recovery_parity_round2.py
cluster_cross_interference test_cluster_cross_interference_parity.py
conformal_cate test_conformal_causal_parity.py
conformal_fair_ite test_conformal_fair_ite_parity.py
conformal_ite test_conformal_ite_parity.py
conformal_ite_interval test_conformal_causal_parity.py
continuous_iv_late test_continuous_iv_late_parity.py
counterfactual_fairness test_fairness_parity.py
dag test_misc_sens_R_parity.py
demographic_parity test_fairness_parity.py
did test_cs_weighted_parity.py (+2)
did_balance test_did_balance_parity.py
did_forest test_panel_forest_recovery.py
did_had test_did_had_parity.py
dist_iv test_decomp_qte_parity.py (+1)
distributional_te test_distributional_te_inference.py (+1)
dml_diagnostics test_ml_causal_dml_parity.py
dynamic_dml test_dynamic_dml_econml_parity.py
equalized_odds test_fairness_parity.py
evidence_without_injustice test_fairness_parity.py
fairness_audit test_fairness_parity.py
fci test_fci_parity.py
focal_cate test_ml_causal_recovery_parity_round2.py
forest_diagnostics test_ml_causal_R_parity.py
forest_group_effects test_fe_forest_imputation_recovery.py
forest_policy_tree test_fe_forest_policy_recovery.py
fortified_pci test_proximal_parity.py
front_door test_front_door_parity.py
frontdoor test_frontdoor_parity.py
general_bunching test_bunching_parity.py
geolift test_geolift_parity.py
ges test_ges_parity.py
gformula_mc test_gformula_family_parity.py (+1)
hal_tmle test_ml_causal_recovery_parity_round2.py
hausman_test test_diag_recovery_parity.py
honest_variance test_forest_rate_honest_parity.py (+1)
identify_transport test_misc_sens_R_parity.py
immortal_time_check test_target_trial_parity.py
influence_functions test_aggte_r_did_parity.py
interference test_interference_parity.py
ivqreg test_ivqreg_parity.py (+1)
kan_dlate test_dist_iv_parity.py
lasso_iv test_lasso_iv_parity.py
lasso_select test_lasso_select_parity.py
lingam test_causal_discovery_parity.py
long_term_from_short test_surrogate_parity.py
longitudinal_analyze test_longitudinal_parity.py
longitudinal_contrast test_longitudinal_parity.py
lpbwselect_mse_dpi test_did_had_parity.py (+1)
lprobust_at_point test_lprobust_parity.py
ltmle_survival test_ml_causal_recovery_parity_round2.py
machado_mata test_decomp_qte_parity.py
matrix_completion test_matrix_completion_parity.py
metafrontier test_frontier_efficiency_parity.py (+1)
mice test_imputation_parity.py
mr_lap test_mr_lap_parity.py
network_exposure test_interference_parity.py
network_graph test_network_parity.py
never_treat test_longitudinal_parity.py
notch test_notch_parity.py
notears test_causal_discovery_parity.py
orthogonal_to_bias test_fairness_parity.py
particle_filter test_assimilation_parity.py
pate test_pate_parity.py
pc_algorithm test_causal_discovery_parity.py
peer_effects test_peer_effects_parity.py
policy_weight_ate test_policy_weight_parity.py
policy_weight_marginal test_policy_weight_parity.py
policy_weight_observed_prte test_policy_weight_parity.py
policy_weight_subsidy test_policy_weight_parity.py
power_ols test_recovery_batch_parity.py
pretrends_equivalence test_fect_equivalence_parity.py
prod_fn test_structural_parity.py
proximal test_proximal_parity.py
proximal_surrogate_index test_surrogate_parity.py
qte_hd_panel test_hd_panel_qte.py
quasi_untreated_test test_did_had_parity.py
rate_split test_fe_forest_rate_recovery.py
regime test_longitudinal_parity.py
romano_wolf test_romano_wolf_parity.py
selection_bounds test_selection_bounds_parity.py
sequential_sdid test_synth_rest_R_parity.py
sharp_ope_unobserved test_ope_parity.py
spatial_did test_spatial_models_parity.py (+1)
spillover test_interference_parity.py
stepwise test_stepwise_parity.py
stochastic_dominance test_distributional_te_parity.py
super_learner test_ml_causal_recovery_parity.py
surrogate_index test_surrogate_parity.py
survival_sensitivity test_misc_sens_R_parity.py
synth_experimental_design test_synth_rest_R_parity.py
synth_power test_synth_rest_R_parity.py
synth_survival test_synth_rest_R_parity.py
target_trial_checklist test_target_trial_parity.py
target_trial_emulate test_target_trial_parity.py
target_trial_protocol test_target_trial_parity.py
target_trial_report test_target_trial_parity.py
translog_design test_translog_design_parity.py
transport_generalize test_transport_parity.py
twfe_decomposition test_did_synth_R_parity.py
weighted_conformal_prediction test_conformal_causal_parity.py
wooldridge_prod test_prodest_parity.py
xlearner test_ml_causal_recovery_parity.py
yatchew_linearity_test test_did_had_parity.py

unverified — 651 functions

These are registered public functions with no cross-language or published-reference parity evidence attached yet. This is the honest coverage gap, not a claim of incorrectness — many are frontier methods with no Stata/R sibling to align against. Query any of them with sp.parity_status(name); the closing roadmap lives in docs/dev/parity_status_roadmap.md.