Cross-language parity matrix¶
Auto-generated — do not hand-edit. Regenerate with
python scripts/build_parity_index.py. Every row traces to a committed test artifact; nothing here is asserted from memory.
StatsPAI's promise is that every number it reports is either aligned with an external reference implementation, recovered against a known truth, or honestly marked as neither. This page makes that promise auditable function-by-function, and keeps the three cases apart — a method with no Stata or R sibling can reach the second and never the first, and saying so is the point. Query any function programmatically:
import statspai as sp
sp.parity_status("feols") # one function
sp.parity_matrix() # the whole matrix
sp.parity_summary() # honest coverage counts
Taxonomy¶
| grade | meaning |
|---|---|
bit-exact |
matches a named R/Stata reference to machine tolerance (headline relative error ≤ 1e-6) |
aligned |
matches a named reference within a documented, pre-registered looser tolerance (cross-fit / convention disagreement) |
analytical-only |
recovers a known population parameter on a deterministic DGP, or a closed-form identity (no cross-package reference) |
external-replication |
reproduces published-paper numbers on a calibrated replica |
unverified |
registered public API, no qualifying numerical-parity evidence attached yet — the honest gap |
Coverage at a glance¶
Read the two evidence kinds separately. Only the first answers "does StatsPAI agree with Stata/R"; the second answers "does StatsPAI recover the right answer", which is a different — and for methods with no Stata/R sibling, the only available — question. Summing them into one "verified" figure would let the smaller claim borrow the authority of the larger one, so this page does not print that total.
| evidence kind | grade | functions |
|---|---|---|
| Compared against R/Stata (T2) | bit-exact | 362 |
| aligned | 53 | |
| subtotal | 415 | |
| No external software reference | analytical-only (T1) | 130 |
| external-replication (published numbers) | 2 | |
| subtotal | 132 | |
| No numerical evidence yet | unverified | 651 |
Honest denominators¶
The all-registered denominator understates coverage: it counts result and exception classes, which can never carry a parity grade, and infrastructure functions that render tables, draw plots, build agent schemas or load data. The estimator denominator is the number to drive release over release.
| denominator | cross-language | any evidence | total | cross-lang share |
|---|---|---|---|---|
| estimator callables | 415 | 546 | 786 | 52.8% |
| infrastructure (parity N/A) | 0 | 0 | 125 | 0.0% |
| result / exception classes | 0 | 1 | 287 | 0.0% |
| all registered | 415 | 547 | 1198 | 34.6% |
Coverage by estimator family¶
Families with zero cross-language rows are the highest-leverage targets when a reference implementation exists, and the honest ceiling when one does not — a method with no Stata/R sibling can reach analytical-only and no further. This table is generated from the same records as the rest of the page, so it cannot drift from them.
| family | cross-language | any evidence | estimator callables |
|---|---|---|---|
| causal | 147 | 213 | 344 |
| regression | 32 | 36 | 37 |
| spatial | 28 | 29 | 34 |
| panel | 27 | 28 | 30 |
| decomposition | 20 | 21 | 29 |
| network | 23 | 24 | 25 |
| inference | 18 | 21 | 23 |
| mendelian | 18 | 19 | 23 |
| diagnostics | 17 | 18 | 22 |
| epi | 16 | 17 | 17 |
| dag | 0 | 0 | 15 |
| bayes | 0 | 7 | 14 |
| postestimation | 6 | 6 | 12 |
| timeseries | 11 | 12 | 12 |
| neural_causal | 0 | 0 | 11 |
| power | 6 | 7 | 11 |
| conformal_causal | 0 | 3 | 10 |
| structural | 5 | 7 | 10 |
| frontier | 5 | 7 | 9 |
| survival | 8 | 8 | 8 |
| robustness | 3 | 4 | 7 |
| interference | 0 | 1 | 7 |
| survey | 6 | 6 | 6 |
| target_trial | 0 | 6 | 6 |
| transport | 3 | 5 | 6 |
| fairness | 0 | 6 | 6 |
| other | 2 | 3 | 6 |
| longitudinal | 0 | 5 | 5 |
| experimental | 3 | 3 | 5 |
| causal_llm | 0 | 0 | 4 |
| bartik | 4 | 4 | 4 |
| nonparametric | 2 | 4 | 4 |
| causal_discovery | 0 | 0 | 3 |
| surrogate | 0 | 3 | 3 |
| causal_rl | 0 | 0 | 3 |
| assimilation | 0 | 3 | 3 |
| gformula | 1 | 2 | 2 |
| ope | 0 | 2 | 2 |
| causal_text | 0 | 0 | 2 |
| missing | 1 | 2 | 2 |
| mediation | 2 | 2 | 2 |
| censoring | 1 | 1 | 1 |
| synth | 0 | 1 | 1 |
bit-exact — 362 functions¶
Machine-tolerance agreement with a named R/Stata reference.
| function | reference | versions | tolerance | rel err (R / Stata) | test |
|---|---|---|---|---|---|
absorb_ols |
fixest::feols 0.14.0 and Stata reghdfe (Track A 03_hdfe / 15_hdfe_cluster goldens); Stata reghdfe with aweights, singleton dropping, two-way clustering | R 4.5.2; fixest 0.14.0; Stata 18 MP | coefficients rtol 1e-12 (observed 2e-15), iid SEs 1e-12 (observed 8e-15), clustered SEs 1e-10 (observed 5.6e-11) | — / — | test_panel_absorb_ols_parity.py |
adjust_pvalues |
base R stats::p.adjust (bonferroni/holm/BH) | R 4.5.2 | exact (atol 1e-15; observed 0) | — / — | test_mht_parity.py (+1) |
aipw |
Stata teffects aipw (ATE, POmeans); R AIPW::AIPW 0.6.9.3 stratified_fit(k_split = 1) | R 4.5.2; AIPW 0.6.9.3; SuperLearner 2.0.40; Stata 18 | 1e-10 rel on ATE, potential-outcome means and SEs (observed <= 1.4e-14) | — / — | test_teffects_R_parity.py (+2) |
ancova |
Stata regress, vce(robust) / vce(cluster) | Stata 18 | 1e-9 rel | — / — | test_misc_sens_stata_parity.py |
anderson_rubin_ci |
R ivmodel::AR.test confidence set | ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 | Both endpoints at 5e-15 after the boundary bisection replaced the grid-point endpoints (previously 8.1e-3 / 4.9e-3). A reference-free test also asserts the two AR entry points agree with each other and that neither endpoint lands exactly on a grid node. | — / — | test_weakiv_meta_parity.py |
anderson_rubin_test |
R ivmodel::AR.test | ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 | Statistic 1.1e-15, p-value 1.2e-13, degrees of freedom exact, and the analytic AR confidence set 1.2e-14. | — / — | test_weakiv_meta_parity.py |
arima |
stats::arima | R 4.5.2; stats 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 7.4e-07 / 9.3e-09 | 39_arima.py (+2) |
assortativity |
R igraph::assortativity_degree | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | 1.2e-16 on karate. | — / — | test_network_parity.py |
attributable_risk |
base-R closed form (attributable fraction exposed + PAF) | R 4.5.2 | AFE + PAF point estimates 1e-12 abs (observed 0); CI not pinned | — / — | test_epi_extra_parity.py (+1) |
attrition_bounds |
Stata leebounds (thresholds held exactly) | leebounds 1.5; Stata 18 | 1e-12 rel | — / — | test_misc_sens_stata_parity.py |
attrition_test |
Stata tabulate, chi2; regress | Stata 18 | 1e-10 rel | — / — | test_misc_sens_stata_parity.py |
auc |
R pROC::auc; Stata roctab | R 4.5.2; pROC 1.19.0.1; Stata 18 MP | 1e-10 rel (observed 2e-16) | — / — | test_survival_epi_R_parity.py (+2) |
average_treatment_effect |
grf::average_treatment_effect 2.6.1 (target.sample all, treated, control, overlap), forest outputs held fixed | R 4.5.2; grf 2.6.1 | estimate and std.err 1e-10 rel (observed 2.0e-15) | — / — | test_ml_causal_R_parity.py (+1) |
bacon_decomposition |
bacondecomp::bacon | R 4.5.2; bacondecomp 0.1.1 | rel_est<=1e-06, rel_se<=1e-06 | 5.6e-16 / 9.6e-09 | 20_bacon.py (+2) |
balance_check |
Stata iebaltab (ietoolkit) | ietoolkit 7.5; Stata 18 | 1e-9 rel | — / — | test_misc_sens_stata_parity.py |
balance_panel |
base R counts == n_periods | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 0 / 0 | 69_balance_panel.py (+2) |
bartik |
2SLS: R AER::ivreg + sandwich (HC1 / classical), Stata ivregress 2sls, vce(robust) small / small; Rotemberg weights: R bartik.weight::bw and Stata bartik_weight (Goldsmith-Pinkham, Sorkin & Swift) | R 4.5.2; AER 1.2.16; sandwich 3.1.1; bartik.weight 0.1.0 (GitHub paulgp/bartik-weight@722ceb85484d6a2bf77985edf2403515eacd1770, R-code/pkg); Stata 18 MP; bartik_weight GitHub paulgp/bartik-weight@722ceb85484d6a2bf77985edf2403515eacd1770 code/bartik_weight.ado | 1e-9 rel on all coefficients and SEs (observed <= 5.4e-15) and on Rotemberg alpha_k / beta_k (observed 2.5e-12) | — / — | test_did_synth_shiftshare_parity.py (+2) |
bauer_sinning |
Stata mvdcmp (Powers, Yoshioka & Yun), logit | Stata 18 MP; mvdcmp SSC | Logit: explained, unexplained, gap, Yun-weighted detailed explained and unexplained terms, their 6x6 delta-method covariance and the aggregate SEs at 1e-9. Probit aligned at 2e-6 (reference probit not fully converged; see note). | — / — | test_decomp_R_parity.py |
benjamini_hochberg |
base R stats::p.adjust(method='BH') | R 4.5.2 | exact (atol 1e-15; observed 0) | — / — | test_mht_parity.py (+1) |
betareg |
betareg::betareg(link.phi="log") | R 4.5.2; betareg 3.2.4 | rel_est<=1e-06, rel_se<=0.01 | 2.2e-12 / 2.1e-08 | 61_betareg.py (+2) |
betweenness_centrality |
R igraph::betweenness | igraph 2.3.3 | Normalised and raw on karate, raw on the directed graph, all at 1e-10. | — / — | test_network_parity.py (+1) |
bias_factor |
R EValue::multi_bound(confounding(), RRAUc, RRUcY) | R 4.5.2; EValue 4.1.4 | 1e-14 rel (observed 0) | — / — | test_inference_sens_R_parity.py (+1) |
biprobit |
R VGAM::vglm(binom2.rho) bivariate probit | R 4.5.2; VGAM 1.1.14 | coef / rho 1e-6 abs (observed <= 2e-7); logLik 1e-6 rel | — / — | test_biprobit_parity.py (+1) |
bjs |
didimputation::did_imputation | R 4.5.2; didimputation 0.5.1 | rel_est<=1e-06, rel_se<=1e-06 | 4.8e-08 / 3.5e-07 | 16_bjs.py (+3) |
block_weights |
spdep::nb2blocknb(NULL, ID) + nb2listw(style = 'W') | R 4.5.2; spdep 1.4.2 | neighbour sets exact; weights 1e-12 rel (observed 0) | — / — | test_spatial_survey_R_parity.py (+1) |
blp_test |
GenericML::BLP 0.2.3 (vcovHC const and HC1; with and without the baseline proxy) | R 4.5.2; GenericML 0.2.3; sandwich 3.1.1 | beta and SE 1e-10 rel (observed 2.3e-15); one-sided p 1e-8 rel | — / — | test_ml_causal_R_parity.py (+1) |
bonacich_power |
R igraph::power_centrality and sna::bonpow | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | 3.9e-16 against igraph and 5.8e-16 against sna at beta = 0.1. | — / — | test_network_parity.py |
bonferroni |
base R stats::p.adjust(method='bonferroni') | R 4.5.2 | exact (atol 1e-15; observed 0) | — / — | test_mht_parity.py (+1) |
borusyak_jaravel_spiess |
didimputation::did_imputation | R 4.5.2; didimputation 0.5.1 | rel_est<=1e-06, rel_se<=1e-06 | 4.8e-08 / 3.5e-07 | 16_bjs.py (+3) |
breslow_day_test |
R DescTools::BreslowDayTest (correct=FALSE/TRUE); Stata cc, by() bd tarone | R 4.5.2; DescTools 0.99.60; Stata 18 MP | 1e-10 rel (observed 1.3e-14) | — / — | test_survival_epi_R_parity.py (+2) |
calibrate_confounding_strength |
R sensemakr::ovb_bounds (kd = ky = k) | sensemakr 0.1.6 | 1e-12 rel (CI 1e-10) | — / — | test_misc_sens_R_parity.py |
calibration_test |
grf::test_calibration 2.6.1 (vcov.type HC3 default and HC1), forest outputs held fixed | R 4.5.2; grf 2.6.1; sandwich 3.1.1 | coef, se, t 1e-10 rel (observed 1.1e-15); one-sided p-value 1e-8 rel (observed 1.6e-13) | — / — | test_ml_causal_R_parity.py (+1) |
callaway_santanna |
did::att_gt + aggte | R 4.5.2; did 2.3.0 | rel_est<=1e-06, rel_se<=1e-09 | 1.3e-15 / 1.3e-15 | 04_csdid.py (+2) |
cgs_continuous_did |
contdid::cont_did | R 4.5.2 | rel_est<=1e-06 | 2.4e-14 / — | 80_contdid.py (+1) |
clogit |
survival::clogit | R 4.5.2; survival 3.8.3 | rel_est<=1e-06, rel_se<=1e-06 | 1.3e-08 / 1.3e-08 | 46_clogit.py (+2) |
closeness_centrality |
Wasserman-Faust closeness from R igraph::distances | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Exact (0.0) on a disconnected graph with three blocks and three isolates -- the case the correction exists for; the connected-graph values also match igraph::closeness(normalized = TRUE) to 1e-10. | — / — | test_network_parity.py |
cluster_robust_se |
R sandwich::vcovCL (HC1, cadjust; HC0; two-way multi0=FALSE); Stata regress, vce(cluster) | R 4.5.2; sandwich 3.1.1; Stata 18 | SE 1e-10 rel (observed 2.1e-15 R, 9.8e-16 Stata) | — / — | test_inference_sens_R_parity.py (+3) |
clustering |
R igraph::transitivity(type = 'local', isolates = 'zero') | igraph 2.3.3 | Exact on karate and on a disconnected graph whose isolates and degree-1 nodes score 0 on both sides. | — / — | test_network_parity.py (+1) |
cohen_kappa |
base-R closed form (Cohen's kappa point estimate) | R 4.5.2 | kappa + agreements 1e-12 abs (observed ~1e-16); SE not pinned | — / — | test_epi_extra_parity.py (+1) |
conformal_synth |
scinference (Chernozhukov-Wuthrich-Zhu authors' package, GitHub kwuthrich/scinference 567c688): estimation_method='sc', permutation_method='mb' | R 4.5.2; scinference 0.0.0.9000 567c6889ce0a1d269a62d415b88aa6baf723a3fe; limSolve 2.0.3; quadprog 1.5.8 | p-values exact (rank statistics); ATT 1e-9 rel (observed 9e-16); CI end points exact on the grid | — / — | test_synth_rest_R_parity.py (+1) |
conley |
Stata acreg (Colella, Lalive, Sakalli & Thoenig) | Stata 18 MP; acreg 1.1.0 | SE 1e-9 rel (observed ~5e-15); absorbed-IV spatial 1e-10 | — / — | test_conley_acreg_spacetime_parity.py (+1) |
continuous_did |
fixest::feols(y ~ dose:post | id + time) (method='twfe') | R 4.5.2; fixest 0.14.0 | slope & SE 1e-9 rel (observed 2.2e-15), iid / ~id / ~region, balanced + unbalanced | — / — |
contrast |
Stata 18 margins r.g / rb3.g / ar.g / gw.g | Stata 18 | 1e-10 rel (observed <= 7.9e-15 est / 4.9e-15 SE) | — / — | test_r2_postest_parity.py (+1) |
cox |
survival::coxph | R 4.5.2; survival 3.8.3 | rel_est<=1e-06, rel_se<=1e-06 | 8.4e-16 / 2.1e-10 | 24_coxph.py (+2) |
cr2_se |
clubSandwich::vcovCR(type="CR2"/"CR3") | R 4.5.2; clubSandwich 0.6.2 | rel_est<=1e-06, rel_se<=1e-06 | 1.8e-08 / 2.2e-08 | 53_cr2.py (+2) |
cr3_jackknife_vcov |
R sandwich::vcovJK(center='estimate'), summclust (CV3); Stata regress, vce(jackknife, cluster() mse double) | R 4.5.2; sandwich 3.1.1; summclust 0.7.0; Stata 18 | vcov 1e-10 rel (observed 6.5e-14 R, SE 4.3e-15 Stata) | — / — | test_inference_sens_R_parity.py (+3) |
cuminc |
R cmprsk::cuminc (estimate, var, Tests); Stata stcompet (ci, se, hi, lo) | R 4.5.2; cmprsk 2.2.12; Stata 18 MP; stcompet 1.0.7 (06nov2012) | 1e-10 rel (observed: CIF 6e-16, Gray variance 3e-13, delta SE 2e-16, Gray test 9e-15) | — / — | test_survival_epi_R_parity.py (+2) |
cusum_test |
strucchange::efp(type = 'Rec-CUSUM') + sctest 1.5.4; Stata 18 estat sbcusum | R 4.5.2; strucchange 1.5.4; Stata 18 | process, statistic and p-value 1e-10 rel (observed 7.1e-14); Stata boundary constants 3e-6 | — / — | test_timeseries_R_parity.py (+2) |
das_gupta |
R DasGuptR::dgnpop (product rate function, summed over strata) | DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 | Das Gupta's Table 2.1 (two factors) and Table 6.5 (four factors x six age groups), every factor effect and both crude rates at 1e-10. | — / — | test_decomp_R_parity.py |
ddd |
Stata 18 MP regress [aw=w], robust (aweight HC1) | Stata 18 MP | b / se 1e-12 abs (observed <= 3e-15) | — / — | test_did2x2_ddd_weighted_robust_parity.py |
ddd_heterogeneous |
triplediff::ddd + agg_ddd | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 5.7e-15 / — | 77_ddd.py (+1) |
decompose |
oaxaca::oaxaca | R 4.5.2; oaxaca 0.1.5 | rel_est<=1e-06, rel_se<=0.05 | 6.3e-16 / 1.3e-16 | 30_oaxaca.py (+2) |
degree_centrality |
R igraph::degree | igraph 2.3.3 | Normalised on karate and raw in / out / all modes on a 40-node directed graph, exact. | — / — | test_network_parity.py (+1) |
demean |
textbook mean-within (algorithmic) | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 3.5e-15 / 8.8e-15 | 68_demean_within.py (+2) |
dfl_decompose |
ddecompose::dfl_decompose | R 4.5.2; ddecompose 1.0.0 | rel_est<=1e-06, rel_se<=1e-06 | 1.2e-09 / 1.8e-13 | 31_dfl.py (+3) |
diagnostic_test |
R epiR::epi.tests; Stata diagti | R 4.5.2; epiR 2.0.94; Stata 18 MP; diagt 2.032 (diagti 2.053) | 1e-12 rel (observed 2e-15) | — / — | test_survival_epi_R_parity.py (+2) |
did_2stage |
did2s::did2s | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 4.8e-08 / 2.4e-12 | 73_did2s.py (+3) |
did_2x2 |
Stata 18 MP regress [aw=w], robust (aweight HC1) | Stata 18 MP | b / se 1e-12 abs (observed <= 3e-16) | — / — | test_did2x2_ddd_weighted_robust_parity.py |
did_estimate |
synthdid::did_estimate 0.0.9 | R 4.5.2; synthdid 0.0.9 | estimate and jackknife SE at 1e-9 on the Prop. 99 replica and a five-treated panel (observed est 3.1e-15, jackknife SE 5.7e-16); R's 40 placebo and 40 bootstrap replications replayed at 1e-9 (observed 1.9e-13) | — / — | test_did_synth_R_parity.py (+1) |
did_imputation |
didimputation::did_imputation | R 4.5.2; didimputation 0.5.1 | rel_est<=1e-06, rel_se<=1e-06 | 4.8e-08 / 3.5e-07 | 16_bjs.py (+2) |
did_multiplegt |
DIDmultiplegt::did_multiplegt (archived 0.1.4) | R 4.5.2 | rel_est<=1e-06 | 3.9e-15 / 3.2e-15 | 81_didm.py (+2) |
did_multiplegt_dyn |
DIDmultiplegtDYN::did_multiplegt_dyn | R 4.5.2 | rel_est<=1e-06 | 3.3e-15 / 2.1e-15 | 78_multiplegt_dyn.py (+2) |
did_timevarying_covariates |
ptetools::pte_default(d_outcome=TRUE, est_method='reg') and did::att_gt(est_method='reg') | R 4.5.2; ptetools 1.0.1; did 2.3.0; DRDID 1.2.3 | ATT(g,t) and overall ATT 1e-9 rel (observed 3.8e-15) | — / — | test_did_synth_didvar_parity.py (+1) |
direct_method |
Open Bandit Pipeline (obp) 0.5.7 DirectMethod (same reward-model matrix via q_hat=) | obp 0.5.7 | value 1e-12 rel (observed 0.0); SE identity 1e-10 | — / — | test_ml_causal_obp_parity.py (+1) |
direct_standardize |
R epitools::ageadjust.direct; Stata dstdize | R 4.5.2; epitools 0.5.10.1; Stata 18 MP | 1e-10 rel (observed 3e-13) | — / — | test_survival_epi_R_parity.py (+2) |
discos |
DiSCos::DiSCo 0.1.4 (Gunsilius distributional synthetic controls), mixture = FALSE; mixture = TRUE vs GLPK on DiSCo's LP | R 4.5.2; DiSCos 0.1.4; pracma 2.4.6; quadprog 1.5.8; CVXR 1.8.2; Rglpk 0.6.5.1 | quantile weights 1e-11 abs (observed 4.9e-13), counterfactual quantile functions 1e-10 rel, quantile effects 1e-10 abs (observed 1.9e-12); mixture LP weights vs GLPK 1e-12 abs, vs DiSCo's SCS solution 5e-6 abs (observed 7.1e-7) | — / — | test_did_synth_synthvar_parity.py (+1) |
distance_band |
R spdep::dnearneigh | spdep 1.4.2; spatialreg 1.4.3 | Neighbour sets identical for all 120 points at a 0.25 radius. | — / — | test_spdep_parity.py |
distributional_did |
didFF::distDD 0.1.0 (Roth & Sant'Anna) | R 4.5.2; didFF 0.1.0; did 2.3.0 | per-bin effect & SE 1e-9 (observed est 2.6e-12 abs, SE 2.5e-10 rel) | — / — | test_did_synth_didvar_parity.py (+3) |
dml |
DoubleML::DoubleMLPLR | R 4.5.2; DoubleML 1.0.2 | rel_est<=1e-10, rel_se<=1e-10 | 0 / 3.7e-15 | 08_dml.py (+2) |
dml_panel |
fixest::demean 0.14.0 (unit and unit+time absorption, balanced and unbalanced) and ddml::ddml_plm 0.3.1 (OLS learner, unit clusters, shared folds) | R 4.5.2; fixest 0.14.0; ddml 0.3.1; sandwich 3.1.1; doubleml 0.11.3; scikit-learn 1.6.1 | within transform 1e-10 abs (observed 2.6e-14); estimate 1e-10 rel vs ddml and DoubleML (observed 4.1e-16); SE 1e-10 rel vs DoubleML (observed 1.4e-15) and vs ddml after the CR1 factor | — / — | test_ml_causal_dml_parity.py (+2) |
dml_sensitivity |
doubleml (Python) DoubleML.sensitivity_analysis | — | bias_bound and adjusted theta bounds 1e-12 (observed 2.5e-15); RV 1e-6 (observed 9.2e-8); RVa is a documented convention gap (<5e-3, observed 1.4e-3) because StatsPAI exhausts | theta | -z*se with the unadjusted SE while doubleml lets the SE move with the confounding scenario |
dose_response |
Stata doseresponse / gpscore (Hirano-Imbens normal GPS, quadratic T and GPS with interaction) | Stata 18; doseresponse SSC | 1e-9 rel on the dose-response function at 5 doses (observed 5.1e-10) | — / — | test_teffects_R_parity.py (+2) |
doubly_robust |
Open Bandit Pipeline (obp) 0.5.7 DoublyRobust (lambda_ inf and 2; same reward model via q_hat=) | obp 0.5.7 | value 1e-12 rel (observed 1.8e-16); SE identity 1e-10 | — / — | test_ml_causal_obp_parity.py (+1) |
drdid |
DRDID::drdid_imp_panel | R 4.5.2; DRDID 1.2.3 | rel_est<=1e-06, rel_se<=1e-06 | 2.6e-15 / 2.2e-16 | 38_drdid.py (+2) |
dyadic_regression |
R dyadRobust (Aronow-Samii-Assenova dyadic-robust variance) | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Coefficients 7e-16 and standard errors 2e-15 on undirected and directed dyads, after the 1.28.0 fix to the shared-member weighting; also asserted against a brute-force construction of the definition. | — / — | test_network_parity.py |
ebalance |
ebal::ebalance 0.2.1 (Hainmueller 2012) | — | ATT rel <= 1e-5 (observed 3.2e-7); moment gap <= 1e-10 | — / — | test_matching_r_parity.py (+1) |
effective_f_test |
Stata weakivtest (Montiel Olea & Pflueger) after ivreg2; R ivDiag::eff_F 1.0.6 | Stata 18 MP; weakivtest 10/28/2020; ivreg2 4.1.12; R 4.5.2; ivDiag 1.0.6 | F_eff rel 1e-9 on 6 designs (observed 3.7e-14 vs Stata, 1.3e-12 vs ivDiag for k = 1) | — / — | test_rd_iv_R_parity.py (+2) |
eigenvector_centrality |
R sna::evcent (unit L2 norm, as here) | sna 2.8; igraph 2.3.3 | 5e-11 undirected, 8e-11 directed -- power-iteration tolerance on both sides. igraph::eigen_centrality max-scales instead (and 2.x ignores scale = FALSE), so it agrees only up to one scalar; the earlier note claiming the igraph convention was wrong. | — / — | test_network_parity.py (+1) |
engle_granger |
egranger 1.0.6 (Stata SSC); urca::ur.df on lm residuals; aTSA::coint.test | R 4.5.2; urca 1.3.4; aTSA 3.1.2.1; Stata 18; egranger 1.0.6 | Z(t) and step-1 coefficients 1e-10 rel (observed 2.1e-13); MacKinnon (2010) critical values 1e-12 vs egranger | — / — | test_timeseries_R_parity.py (+2) |
ergm |
R ergm::ergm(estimate = 'MPLE') | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Coefficients 2e-16 to 7e-13 for edges + triangle + nodematch + nodecov + absdiff (undirected) and edges + mutual (directed); standard errors 2e-8 (directed) and <= 3.2e-7 (undirected), inside the 1e-6 budget. | — / — | test_network_parity.py |
etregress |
Stata 18 MP official etregress (Maddala 1983 model) |
Stata 18 MP | Two-step: 5e-9 on every coefficient and every standard error, including the Heckman correction for the estimated first stage. ML: the likelihood, score and observed information are pinned at 9e-11 -- our Hessian reproduces Stata's reported standard errors when evaluated at Stata's own parameter vector, which is independent of either optimiser. At our own optimum the parameters sit within 2e-5 of Stata's; that gap is the two optimisers' stopping points, not a formula difference, and StatsPAI's stops at the HIGHER log-likelihood with a gradient ~300x smaller (asserted, so a regression that makes our optimum worse fails even though the 1e-4 parity assertions would still pass). vce(robust) carries Stata's N/(N-1) meat factor and vce(cluster) its g/(g-1). | — / — | test_etregress_stata_parity.py |
etwfe |
etwfe::etwfe + emfx | R 4.5.2; etwfe 0.6.2 | rel_est<=1e-06, rel_se<=0.001 | 1.8e-13 / 3.9e-14 | 17_etwfe.py (+2) |
etwfe_emfx |
etwfe::etwfe + emfx | R 4.5.2; etwfe 0.6.2 | rel_est<=1e-06, rel_se<=0.001 | 1.8e-13 / 3.9e-14 | 17_etwfe.py (+2) |
evalue |
EValue::evalues.RR | R 4.2.3; EValue 4.1.4 | rel_est<=1e-06, rel_se<=1e-06 | 5.8e-14 / 1.2e-16 | 23_evalue.py (+2) |
evalue_from_result |
R EValue::twoXtwoRR -> evalues.RR | R 4.5.2; EValue 4.1.4 | 1e-13 rel (observed 3.3e-16) | — / — | test_inference_sens_R_parity.py (+1) |
evalue_rd |
R EValue::evalues.RD | R 4.5.2; EValue 4.1.4 | 1e-12 rel (observed 5.8e-14: grid built by seq vs np.arange) | — / — | test_inference_sens_R_parity.py (+1) |
evalue_rr |
R EValue::evalues.RR | EValue 4.1.4 | Point and CI E-values at 1e-12 across ten cases, including RR < 1 and CIs crossing the null. | — / — | test_evalue_rr_parity.py |
event_study |
fixest::feols(y ~ i(rel, treat, ref=-1) | R 4.5.2; fixest 0.14.0 | rel_est<=1e-09, rel_se<=1e-09 | 3.2e-13 / 1.5e-14 | 85_twfe_event_study.py (+2) |
fairlie |
Stata fairlie 1.0.7 (Jann, SSC) | Stata 18.0 MP; fairlie 1.0.7 16jun2008 (SSC) | tightly converged logit/probit: contributions and SEs 1e-12 rel (observed 4.5e-15 / 5.4e-14); at Stata's default logit tolerance contributions 1e-8 and SEs 1e-5 (observed 1.6e-10 / 1.6e-6) | — / — | test_decomp_qte_parity.py (+1) |
fect |
fect::fect(Y ~ D + X1 + X2, method=, force="two-way", se=FALSE, CV=FALSE, tol=1e-12, max.iteration=20000); Stata side uses the authors' fect_stata (GitHub, installed into a local ado path) | R 4.5.2; fect 2.4.1 | rel_est<=1e-06, rel_se<=1e-06 | 1.8e-13 / 9.8e-10 | 86_fect.py (+2) |
feglm |
fixest::feglm (family="logit") / fixest::fepois | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=5e-05 | 9.7e-09 / 1.8e-09 | 67_panel_glm.py (+2) |
feols |
fixest::feols | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=1e-06 | 5.2e-15 / 2.9e-15 | 03_hdfe.py (+2) |
fepois |
fixest::feglm (family="logit") / fixest::fepois | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=5e-05 | 9.7e-09 / 1.8e-09 | 67_panel_glm.py (+2) |
ffl_decompose |
R ddecompose::ob_decompose(reweighting = TRUE) | rifreg 1.1.0; dineq 0.1.0; ddecompose 1.0.0 | Observed difference, composition, structure, specification and reweighting errors at 1e-9 (1e-12 absolute floor; logit MLE in the path), both reference directions, for the variance, Gini (exact RIF supplied as custom_rif_function) and 10th/50th/90th percentiles. | — / — | test_decomp_R_parity.py |
finegray |
R cmprsk::crr (coef, var, invinf, loglik); Stata stcrreg | R 4.5.2; cmprsk 2.2.12; Stata 18 MP | R 1e-10 rel (observed 6e-15); Stata coefficients 1e-7 and SEs 1e-8 (Stata's ml stops 5e-9 away), log likelihood 1e-10 | — / — | test_survival_epi_R_parity.py (+2) |
fisher_exact |
R ri2::conduct_ri (randomizr full enumeration); Stata ritest over the full assignment set | R 4.5.2; ri2 0.5.0; Stata 18; ritest 1.1.7 | p exact (k / N_assignments); statistic 1e-12 rel vs R, 2^-23 vs Stata | — / — | test_inference_sens_R_parity.py (+3) |
florentine_families |
R ergm flomarriage | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | The 20 marriage ties are identical edge for edge; the Pucci isolate is omitted (15 nodes against 16), which is documented. | — / — | test_network_parity.py |
four_way_decomposition |
Stata med4way (yreg/mreg linear); CMAverse::cmest 0.1.0 (rb, paramfunc, delta) | R 4.5.2; CMAverse 0.1.0; Stata 18 | components 1e-12 rel; delta SEs 1e-12 rel vs CMAverse (vcov='ols'); variances 5e-9 rel vs med4way (vcov='ml') | — / — | test_teffects_R_parity.py (+2) |
fracreg |
stats::glm(quasibinomial('logit')) [fractional response] | R 4.5.2 | coefficients 1e-10 abs (observed ~8e-15) | — / — | test_glm_ext_parity.py (+1) |
frontier |
sfaR::sfacross | R 4.5.2; sfaR 1.0.1 | rel_est<=1e-06, rel_se<=5e-05 | 4.1e-08 / 4.0e-08 | 28_frontier.py (+2) |
g_computation |
base R stats::lm g-formula standardization (Robins 1986) | — | psi 1e-8 (observed <= 7e-16; bootstrap SE pinned loosely +/-25%) | — / — | test_gformula_parity.py (+1) |
g_estimation |
DTRreg::DTRreg 2.4 (method = 'gest', treat.type = 'bin', weight = 'none') | R 4.5.2; DTRreg 2.4 | 1e-10 rel on each stage psi (observed 4.6e-15) | — / — | test_teffects_R_parity.py (+2) |
gap_closing |
R ddecompose::dfl_decompose (method='ipw') and ob_decompose (method='regression') | DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 | Observed, counterfactual and closed gaps at 1e-9 for IPW in both directions (logit MLE in the path) and 1e-10 for regression. method='aipw' has no reference and is checked for double robustness on a known-truth DGP (T1). | — / — | test_decomp_R_parity.py |
gardner_did |
did2s::did2s | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 4.8e-08 / 2.4e-12 | 73_did2s.py (+2) |
gate_test |
GenericML::GATES 0.2.3 (monotonize = FALSE) with GenericML::quantile_group membership | R 4.5.2; GenericML 0.2.3; sandwich 3.1.1 | gamma_k, SE and gamma_K - gamma_1 1e-10 rel (observed 1.9e-15); group membership identical | — / — | test_ml_causal_R_parity.py (+1) |
geary |
R spdep::geary.test | spdep 1.4.2; spatialreg 1.4.3 | C 2.0e-15. The closed-form variance and z are new in 1.27.0 (they were NaN whenever permutations=0) and match both spdep nulls at 1.5e-14: randomisation (with the m4/m2^2 term) and normality. | — / — | test_spdep_parity.py |
gelbach |
Stata b1x2 (Gelbach's own command), robust and homoskedastic | Stata 18 MP; b1x2 SSC | Per-variable contributions at 1e-10; their full covariance, SEs and the total-change variance at 1e-9, robust and homoskedastic. | — / — | test_decomp_R_parity.py |
geographic_rd |
rdmulti::rdms; Stata side uses the rdms ado from the rdpackages GitHub mirror (rdmulti is not on SSC: ssc describe rdmulti returns r(601)), installed into a local gitignored ado path | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 1.2e-12 / 5.4e-10 | 89_rdms.py (+3) |
getis_ord_g |
R spdep::globalG.test (binary weights) | spdep 1.4.2; spatialreg 1.4.3 | G 5.4e-16. Binary weights, which is what spdep recommends for this statistic. | — / — | test_spdep_parity.py |
getis_ord_local |
R spdep::localG | spdep 1.4.2; spatialreg 1.4.3 | Gi 1.7e-14; Gi 4.3e-13 after the star=False branch stopped borrowing Gi's standardisation. | — / — | test_spdep_parity.py |
gformula_ice_fn |
ltmle::ltmle 1.3.0 (gcomp = TRUE, SL.library = list(Q = 'SL.lm')) point estimate; base-R lm() ICE with geex::m_estimate 1.1.1 sandwich SE | R 4.5.2; ltmle 1.3.0; SuperLearner 2.0.40; geex 1.1.1 | point 1e-10 rel vs ltmle (observed 2.1e-11), 1e-12 vs hand lm (observed 2.4e-15); sandwich SE 1e-9 rel (observed 7.1e-12) | — / — | test_r2_teffects_parity.py (+3) |
glm |
base R stats::glm (binomial logit + Poisson log) | R 4.5.2 | coef / logLik / AIC 1e-8 abs (observed <= 5e-13); SE ~1e-3 rel | — / — | test_glm_parity.py (+1) |
gmm |
Stata 18 gmm (linear and exponential-mean IV; twostep, igmm, onestep); R gmm::gmm | Stata 18 MP; R 4.5.2; gmm 1.9.1 | vs Stata estimates rtol 1e-10 (observed 1.2e-12), SEs 1e-9 linear / 1e-6 nonlinear (Stata's numerical Jacobian; observed 2.5e-7), J rtol 1e-10; vs R rtol 1e-6 (observed 1.8e-7) | — / — | test_panel_gmm_stata_parity.py (+1) |
granger_causality |
Stata 18 vargranger; vars::causality 1.6.1 | R 4.5.2; vars 1.6.1; Stata 18 | chi2 and F 1e-10 rel (observed 4.1e-15), p-values 1e-8 | — / — | test_timeseries_R_parity.py (+2) |
gsynth |
gsynth::gsynth | R 4.5.2; gsynth 1.4.0 | rel_est<=1e-06, rel_se<=1e-06 | 7.7e-14 / 2.2e-15 | 19_gsynth.py (+2) |
gwr |
GWmodel::gwr.basic 2.4.1 | R 4.5.2; GWmodel 2.4.1 | local betas and SEs 1e-9 rel (observed 1.3e-11 / 3.5e-14); RSS, AIC, AICc, BIC, enp, edf 1e-10 rel | — / — | test_spatial_survey_R_parity.py (+1) |
gwr_bandwidth |
GWmodel::bw.gwr 2.4.1 (golden section gold(); gwr.aic / gwr.cv) | R 4.5.2; GWmodel 2.4.1 | selected bandwidth 1e-10 rel (observed 0 on 8 configurations); criterion values 1e-11 rel (observed 7.0e-15) | — / — | test_spatial_survey_R_parity.py (+1) |
harvest_did |
R did::att_gt(control_group='notyettreated', base_period='universal') for every (cohort, horizon) cell; did::aggte(type='dynamic') for the event study under weighting='n_treated' | did 2.3.0; DRDID 1.2.3 | Every 2x2 cell (ATT and SE) and every event-study horizon (ATT and SE) at 1e-9 relative; observed agreement 1.7e-14. The inverse-variance aggregate over horizons (the default headline) has no reference and is checked by identity only. | — / — | test_did_synth_misc_parity.py |
hdfe_ols |
fixest::feols | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=1e-06 | 5.2e-15 / 2.9e-15 | 03_hdfe.py (+3) |
heckman |
sampleSelection::heckit | R 4.5.2; sampleSelection 1.2.14 | rel_est<=1e-06, rel_se<=0.0005 | 1.0e-11 / 1.0e-11 | 43_heckman.py (+2) |
het_test |
lmtest::bptest (studentized Breusch-Pagan) | R 4.5.2; lmtest 0.9.40 | statistic & p-value 1e-10 rel (observed ~1e-13) | — / — | test_diagnostics_parity.py (+1) |
heterogeneity_of_effect |
R metafor::rma(method = 'DL') | metafor 5.0.1 | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
holm |
base R stats::p.adjust(method='holm') | R 4.5.2 | exact (atol 1e-15; observed 0) | — / — | test_mht_parity.py (+1) |
honest_did |
HonestDiD::createSensitivityResults_relativeMagnitudes | R 4.5.2; HonestDiD 0.2.8 | abs_est<=1e-06, abs_se<=1e-06 | 4.4e-16 / 5.6e-17 | 21_honest_relmags.py (+2) |
horowitz_manski |
Stata tebounds 1.8 worst-case bounds (identical estimand for ATE worst-case bounds) | Stata 18; tebounds 1.8 (SJ15-2 st0386) | 1e-12 rel / 1e-15 abs on both bounds | — / — | test_teffects_R_parity.py (+2) |
hurdle |
pscl::hurdle(dist='poisson', zero.dist='binomial') | R 4.5.2; pscl 1.5.9 | count + zero coefficients 1e-6 abs (observed ~2e-8) | — / — | test_glm_ext_parity.py (+1) |
icc |
Stata 18 estat icc after mixed (ML, REML) and melogit; performance::icc; psych::ICC (balanced ANOVA identity) | Stata 18 MP; R 4.5.2; performance 0.16.0; psych 2.6.5 | estimate / SE rtol 1e-6, logit-scale CI rtol 2e-6 (observed <= 2.4e-7 / 7e-7 / 1.2e-6); balanced REML ICC = ANOVA ICC(1) rtol 1e-9 | — / — | test_panel_icc_lrtest_parity.py |
impacts |
R spatialreg::impacts on a lagsarlm fit | spdep 1.4.2; spatialreg 1.4.3 | Direct, indirect and total at 1e-6, inheriting the SAR rho's own agreement. | — / — | test_spdep_parity.py |
incidence_rate_ratio |
base-R closed form (rate ratio + conditional-binomial exact CI) | R 4.5.2 | estimate 1e-12; exact CI 1e-10 abs (observed ~3e-15) | — / — | test_epi_parity.py (+1) |
indirect_standardize |
Stata istdize (exact CI); R epitools::ageadjust.indirect (log-normal CI) | R 4.5.2; epitools 0.5.10.1; Stata 18 MP | 1e-10 rel (observed 6e-16) | — / — | test_survival_epi_R_parity.py (+2) |
inequality_index |
base-R closed form (Gini/Theil-T/Theil-L/Atkinson; = ineq) | R 4.5.2 | all indices 1e-12 abs (observed ~2e-16) | — / — | test_inequality_parity.py (+1) |
interactive_fe |
Stata regife (SSC, Gomez) ..., noconstant; R phtt::Eup(additive.effects = 'none') | Stata 18 MP; regife 2026-03-30 SSC; R 4.5.2; phtt 3.1.2 | slopes rtol 1e-9 vs both (observed <= 2e-11); SEs rtol 1e-9 vs regife with dof='regife' (homoskedastic and cluster), phtt SE reconstructed rtol 1e-9 | — / — | test_panel_ife_parity.py |
interflex |
interflex::interflex(vartype="delta", vcov.type="robust", neval=5, nbins=3, bw=1); Stata side uses the SSC interflex command | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 4.0e-15 / 1.5e-14 | 87_interflex.py (+2) |
ipcw |
survival::coxph(ties = 'breslow') + basehaz(centered = FALSE) | R 4.5.2; survival 3.8.3 | 1e-9 rel on every weight (observed 2.0e-11) | — / — | test_teffects_R_parity.py (+2) |
ips |
Open Bandit Pipeline (obp) 0.5.7 InverseProbabilityWeighting (lambda_ inf and 2) | obp 0.5.7 | value 1e-12 rel (observed 0.0); SE identity 1e-10 | — / — | test_ml_causal_obp_parity.py (+1) |
ipw |
base R stats::glm(binomial) + hand-rolled Hajek weighted means | — | Hajek ATE/ATT estimate 1e-9 (observed <= 2e-15; SE not pinned) | — / — | test_ipw_parity.py (+1) |
irf |
vars::irf 1.6.1; Stata 18 irf create | R 4.5.2; vars 1.6.1; Stata 18 | 1e-10 rel vs vars, 1e-9 vs Stata irf file (observed 2.6e-14) | — / — | test_timeseries_R_parity.py (+2) |
its |
lm + sandwich::NeweyWest 3.1.1; Stata 18 newey; itsa 1.0.0 (SSC) | R 4.5.2; sandwich 3.1.1; Stata 18; itsa 1.0.0 | coefficients and Newey-West SE 1e-10 rel (observed 7.9e-14); itsa 1e-6 (glm2 IRLS) | — / — | test_timeseries_R_parity.py (+2) |
iv |
AER::ivreg | R 4.5.2; AER 1.2.16 | rel_est<=1e-06, rel_se<=1e-06 | 1.1e-11 / 1.1e-11 | 02_iv.py (+3) |
iv_diag |
R ivDiag::ivDiag 1.0.6 (analytic block) | R 4.5.2; ivDiag 1.0.6; lfe 3.1.1 | 2SLS / OLS coefficients and SEs, classical first-stage F, effective F, tF critical value and interval: rel 1e-9 on 6 designs (observed <= 6e-12) | — / — | test_rd_iv_R_parity.py (+1) |
ivreg |
AER::ivreg | R 4.5.2; AER 1.2.16 | rel_est<=1e-06, rel_se<=1e-06 | 1.1e-11 / 1.1e-11 | 02_iv.py (+2) |
jackknife_se |
R sandwich::vcovJK(center='mean'); Stata regress, vce(jackknife, cluster() double) | R 4.5.2; sandwich 3.1.1; Stata 18 | SE / CI 1e-10 rel, p 1e-9 (observed SE 2.0e-15, p 1.4e-14, CI 9.2e-12) | — / — | test_inference_sens_R_parity.py (+3) |
jive |
Stata jive 1.0.2 (Stata Journal st0108) ujive1 / ujive2 | Stata 18 MP; jive 1.0.2 | coefficients and SEs (default and robust) rel 1e-9 (observed 3.0e-13 / 4.4e-13) | — / — | test_rd_iv_R_parity.py (+1) |
johansen |
urca::ca.jo 1.3.4; Stata 18 vecrank | R 4.5.2; urca 1.3.4; Stata 18 | eigenvalues, trace & max-eigenvalue statistics 1e-11 rel vs ca.jo, 1e-10 vs vecrank (observed 8.2e-14); Osterwald-Lenum table equal to Stata _vecgetcv cell by cell | — / — | test_timeseries_R_parity.py (+2) |
join_counts |
R spdep::joincount.multi (binary weights) | spdep 1.4.2; spatialreg 1.4.3 | BB, WW and BW all exact. A reference-free guard also asserts BB + WW + BW = S0/2, the identity the BW defect violated (70.75 against 50). | — / — | test_spdep_parity.py |
kaplan_meier |
survival::survfit | R 4.5.2; survival 3.8.3 | S(t) at every event time 1e-12 (observed ~3e-17); median exact | — / — | test_survival_km_parity.py (+1) |
karate_club |
R igraph::make_graph('Zachary') | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Adjacency matrix identical. | — / — | test_network_parity.py |
katz_centrality |
R igraph::alpha_centrality | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | 7.1e-16 with normalized = False (normalized = True L2-scales the same vector). | — / — | test_network_parity.py |
kdensity |
Stata kdensity (at(), bwidth(), kernel()); R bw.nrd0 / bw.SJ | R 4.5.2; Stata 18 MP | density and default widths 1e-10 rel (observed 9e-16); Sheather-Jones 1e-6 vs bw.SJ(nb=1e7, tol=1e-14) (observed 2e-7, R's pair-count binning) | — / — | test_survival_epi_R_parity.py (+2) |
kernel_weights |
spdep::nb2listwdist(type = 'dpd', alpha = 2) on dnearneigh(0, h); GWmodel::gw.weight | R 4.5.2; spdep 1.4.2; GWmodel 2.4.1 | weights 1e-12 rel / 1e-15 abs (observed 6.9e-16 rel, 3.3e-16 abs) | — / — | test_spatial_survey_R_parity.py (+1) |
kitagawa_decompose |
R DasGuptR::dgnpop with ratefunction sum(size*rate)/sum(size) | DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 | Rate and composition effects on Das Gupta's Table 5.1 at 1e-10; interaction exactly 0. | — / — | test_decomp_R_parity.py |
knn_weights |
R spdep::knearneigh + knn2nb | spdep 1.4.2; spatialreg 1.4.3 | Neighbour sets identical for all 120 points, k=4, on a random point set chosen so no distance ties make the answer non-unique. | — / — | test_spdep_parity.py |
lee_bounds |
Stata leebounds 1.5 (Tauchmann), vce(analytic) | Stata 18; leebounds 1.5 (2013-07-17, Tauchmann) | 1e-10 rel on bounds and analytic variances | — / — | test_teffects_R_parity.py (+2) |
liml |
ivmodel::LIML | R 4.5.2; ivmodel 1.9.1 | rel_est<=1e-06, rel_se<=1e-06 | 1.7e-15 / 9.7e-16 | 59_liml.py (+2) |
linear_calibration |
survey::calibrate(calfun='linear', unbounded); Stata svycal regress | R 4.5.2; survey 4.5; Stata 18 MP | calibrated weights 1e-12 rel (observed 5.5e-15) | — / — | test_survey_calib_R_parity.py (+2) |
lm_tests |
R spdep::lm.RStests | spdep 1.4.2; spatialreg 1.4.3 | All five statistics and their p-values at 1e-9. Before the fix: LM_err 39.47 against 19.58, and Robust_LM_err 20.49 (p=6e-6) against 0.0397 (p=0.84) -- the Anselin lag-vs-error decision rule, reversed. | — / — | test_spdep_parity.py |
local_projections |
lpirfs::lp_lin | R 4.5.2; lpirfs 0.2.5 | rel_est<=1e-06, rel_se<=1e-06 | 5.0e-15 / 4.4e-15 | 34_lp.py (+2) |
logit |
stats::glm(family=binomial("logit")) | R 4.5.2; stats 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 2.7e-11 / 2.7e-11 | 57_logit.py (+2) |
logrank_test |
survival::survdiff | R 4.5.2; survival 3.8.3 | chi-square 1e-10 rel (observed ~8e-16); p-value 1e-10 abs | — / — | test_survival_km_parity.py (+1) |
lp_did |
direct transcription (no LP-DiD R package installed); Stata side uses the authors' lpdid | R 4.5.2 | rel_est<=1e-10, rel_se<=1e-10 | 5.0e-15 / 2.5e-15 | 83_lpdid.py (+2) |
lpoly |
Stata lpoly (at(), bwidth(), degree(), kernel(), se(), pwidth()) | Stata 18 MP | 1e-10 rel (observed 1.6e-12) | — / — | test_survival_epi_R_parity.py (+2) |
lrtest |
Stata 18 lrtest; R anova() on lme4 ML fits | Stata 18 MP; R 4.5.2; lme4 2.0.1 | chi2 rtol 1e-6, df exact, p rtol 1e-5 (observed chi2 <= 1e-9) | — / — | test_panel_icc_lrtest_parity.py |
ltmle |
ltmle::ltmle 1.3-0 (glm, gbounds c(0.01, 1), variance.method = 'ic') | R 4.5.2; ltmle 1.3.0 | psi1 / psi0 / ATE / SE 1e-9 rel without censoring (observed 5e-13); 1e-8 with censoring nodes (observed 5e-10) | — / — | test_teffects_R_parity.py (+2) |
manski_bounds |
Stata tebounds 1.8 (SJ15-2 st0386), erates(0) | Stata 18; tebounds 1.8 (SJ15-2 st0386) | 1e-12 rel / 1e-15 abs on both bounds | — / — | test_teffects_R_parity.py (+2) |
mantel_haenszel |
base-R closed form (Robins-Breslow-Greenland MH; = epiR) | R 4.5.2 | estimate, se_log, CI 1e-12 abs (observed 0) | — / — | test_epi_parity.py (+1) |
margins_at |
Stata 18 margins, at(...); R marginaleffects::avg_predictions | Stata 18; R 4.5.2; marginaleffects 0.32.0 | 1e-10 rel (observed <= 3.5e-15 est / SE; marginaleffects SE 1.8e-10, numerical Jacobian) | — / — | test_r2_postest_parity.py (+2) |
markup |
Stata markupest 1.0.1 (Rovigatti), method(dlw) pmethod(lp) | Stata 18; markupest 1.0.1 10May2020 | 1e-10 rel (observed 8.1e-13 corrected, 3.4e-15 uncorrected) | — / — | test_frontier_struct_R_parity.py (+1) |
match |
MatchIt::matchit 4.7.2 (nearest, glm/logit distance) | — | 1:1 and 2:1 PS matching without replacement: rel <= 1e-9. Mahalanobis metric pinned against MatchIt:::mahalanobis_dist (rel <= 1e-10); greedy m_order='data'/'closest' rel <= 1e-9. | — / — | test_matching_r_parity.py (+1) |
mc_panel |
MCPanel::mcnnm_fit (Athey, Bayati, Doudchenko, Imbens & Khosravi; github.com/susanathey/MCPanel) and fect::fect(method = "mc") | R 4.5.2; MCPanel 0.0 @ 6b2706fd7c35f3266048ceb22a7e9a61ae1774da; fect 2.4.1; gsynth 1.4.0 | ATT and fitted untreated matrix 1e-9 rel at fixed lambda (observed <= 2.6e-13), four fixed-effect modes x two lambdas | — / — | test_did_synth_mc_parity.py (+1) |
mc_synth |
MCPanel::mcnnm_fit (Athey, Bayati, Doudchenko, Imbens & Khosravi; github.com/susanathey/MCPanel) and fect::fect(method = "mc") | R 4.5.2; MCPanel 0.0 @ 6b2706fd7c35f3266048ceb22a7e9a61ae1774da; fect 2.4.1 | ATT and fitted untreated matrix 1e-9 rel at fixed lambda (observed <= 6.0e-13), two-way and no-FE x two lambdas | — / — | test_did_synth_mc_parity.py (+1) |
mccrary_test |
R rdd::DCdensity 0.57 (CRAN archive) | R 4.5.2; rdd 0.57 | theta, se, z rel 1e-9; p rel 1e-8; bin width 1e-12 (observed 6.5e-12) | — / — | test_rd_iv_rd_R_parity.py (+1) |
mde |
base-R closed form (RCT minimum detectable effect) | R 4.5.2 | effect size 1e-6 abs (output rounded to 6 dp; observed ~2e-8) | — / — | test_power_extra_parity.py (+1) |
mediate |
mediation::mediate | R 4.5.2; mediation 4.5.1 | rel_est<=1e-06, rel_se<=0.1 | 6.7e-15 / 3.6e-15 | 36_mediation.py (+3) |
mediate_interventional |
CMAverse::cmest 0.1.0 (gformula with postc: rpnde / rpnie / te; rb paramfunc without) | R 4.5.2; CMAverse 0.1.0 | IIE / IDE / total 1e-9 rel (observed 2.3e-15) | — / — | test_teffects_R_parity.py (+2) |
mediate_sensitivity |
R mediation::medsens (lm/lm, rho.by = 0.1) | mediation 4.5.1 | 1e-10 rel | — / — | test_misc_sens_R_parity.py |
mediation |
mediation::mediate | R 4.5.2; mediation 4.5.1 | rel_est<=1e-06, rel_se<=0.1 | 6.7e-15 / 3.6e-15 | 36_mediation.py (+2) |
mediation_decompose |
Stata paramed (Liu & Emsley, SSC), yreg(linear) mreg(linear) with interaction | Stata 18.0 MP; paramed SSC | CDE/NDE/NIE/total and delta-method SEs 1e-12 rel (observed 9.5e-15) | — / — | test_decomp_qte_parity.py (+1) |
megamma |
glmmTMB 1.1.14 Gamma(link = 'log') (Laplace, AD Hessian); Stata 18 meglm, family(gamma) link(log) intmethod(laplace) | R 4.5.2; glmmTMB 1.1.14; TMB 1.9.25; Stata 18 MP | vs glmmTMB estimates / SEs rtol 1e-6 (observed 6.9e-11 / 3.8e-8); vs Stata Laplace estimates rtol 1e-6 (observed 2.2e-7); objective identity rtol 1e-11 | — / — | test_panel_glmm_parity.py |
meglm |
Stata 18 meglm (gaussian; binomial with binomial()); lme4::lmer(REML = FALSE); lme4::glmer | Stata 18 MP; R 4.5.2; lme4 2.0.1 | gaussian vs Stata estimates / SEs rtol 1e-6 (observed 4.2e-9 / 2.8e-7), vs lmer ML estimates 6.1e-11; binomial-trials vs Stata 2.4e-9 / 6.6e-8 | — / — | test_panel_glmm_parity.py |
melly_decompose |
Stata cdeco, method(qr) 1.0.2 (Chernozhukov, Fernandez-Val & Melly; bmelly/Stata counterfactual) | Stata 18.0 MP; cdeco 1.0.2 01mar2023 (bmelly/Stata counterfactual, Distribution-Date 20220803) | fitted and counterfactual quantiles 1e-10 rel (observed 6.2e-16); QR process coefficients 1e-12 abs (observed 7.5e-15) | — / — | test_decomp_qte_parity.py (+1) |
menbreg |
glmmTMB 1.1.14 nbinom2 (Laplace, AD Hessian); Stata 18 menbreg intmethod(laplace) / mcaghermite(7) | R 4.5.2; glmmTMB 1.1.14; TMB 1.9.25; lme4 2.0.1; Stata 18 MP | vs glmmTMB estimates / SEs rtol 1e-6 (observed 1.7e-11 / 2.3e-7); vs Stata Laplace estimates rtol 1e-6 (observed 7.5e-8); objective identity at Stata's and glmmTMB's estimates rtol 1e-11 | — / — | test_panel_glmm_parity.py |
mendelian_randomization |
R MendelianRandomization::mr_allmethods(method = 'main') | MendelianRandomization 0.10.0 | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
meologit |
Stata 18 meologit intmethod(mcaghermite) intpoints(7) / intmethod(laplace); ordinal::clmm | Stata 18 MP; R 4.5.2; ordinal 2025.12.29 | AGHQ-7 vs Stata estimates rtol 1e-6 (observed 3.8e-7), SEs incl. cutpoints rtol 2e-6 (observed 1.3e-6; 5.7e-7 at Stata's estimates); objective identity vs Stata and clmm rtol 1e-11 (observed <= 1.4e-11) | — / — | test_panel_glmm_parity.py |
mepoisson |
Stata 18 mepoisson, intmethod(laplace) / intmethod(mcaghermite) intpoints(7); lme4::glmer(nAGQ = 1 / 7) | Stata 18 MP; R 4.5.2; lme4 2.0.1 | objective identity at Stata's estimates rtol 1e-11 (observed 6.6e-12); Laplace estimates / SEs vs Stata rtol 1e-6 (observed 9.0e-11 / 1.2e-7); AGHQ-7 vs lme4 rtol 1e-6 (observed 4.4e-8 / 9.3e-8) | — / — | test_panel_glmm_parity.py |
meta_analysis |
R metafor::rma (method='FE' and 'DL') | ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 | 6.2e-16 across all nine reported quantities: fixed and random pooled effect and standard error, tau^2, Cochran Q and its p-value, I^2 and H^2. REML is not implemented here, which is a capability gap rather than a disagreement. | — / — | test_weakiv_meta_parity.py |
metalearner |
econml.metalearners SLearner / TLearner / XLearner | — | S / T / X conditional-average-treatment-effect vectors match econml elementwise to 1e-12 absolute (observed <= 1.1e-15) when both fitting stages use the same base learner | — / — | test_metalearner_econml_parity.py |
mgwr |
PySAL mgwr 2.2.1 Sel_BW(multi=True).search() + MGWR.fit() (authors' implementation); fixed-bandwidth back-fitting also GWmodel::gwr.multiscale 2.4.1 | python 3.13.9; mgwr 2.2.1; numpy 2.2.6; R 4.5.2; GWmodel 2.4.1 | bandwidths, initial bandwidth, bandwidth history and iteration count exact; betas 1e-9 rel (observed 3.0e-12); SEs, ENP_j, tr(S), sigma2, AICc/AIC/BIC 1e-10 rel (observed <= 1e-15); SOC path 1e-9 (observed 8e-13) | — / — | test_r2_spatial_parity.py (+3) |
mi_estimate |
R mice::pool; Stata mi estimate: regress | mice 3.19.0; Stata 18 | 1e-9 rel (p, CI via t quantiles at fractional df); 1e-12 est/SE/df | — / — | test_misc_sens_R_parity.py (+1) |
mixed |
lme4::lmer | R 4.5.2; lme4 2.0.1 | rel_est<=1e-06, rel_se<=1e-06 | 2.6e-12 / 8.3e-11 | 25_lmm.py (+2) |
mixlogit |
Stata mixlogit 1.4.0 (SSC, Hole), nrep(50) burn(15), on identical Halton draws | Stata 18 MP; mixlogit 1.4.0 | means / SDs / Sigma / SEs rtol 1e-6 (observed <= 2.2e-7), log-likelihood rtol 1e-10 (observed 5e-13) | — / — | test_panel_mixlogit_parity.py |
mlogit |
nnet::multinom | R 4.5.2; nnet 7.3.20 | rel_est<=1e-06, rel_se<=5e-05 | 2.6e-07 / 1.5e-11 | 44_mlogit.py (+2) |
moran |
R spdep::moran.test (randomisation null) | spdep 1.4.2; spatialreg 1.4.3 | I 1.9e-15, expectation, variance and z all at 1e-15 on the row-standardised lattice. | — / — | test_spdep_parity.py |
moran_local |
R spdep::localmoran | spdep 1.4.2; spatialreg 1.4.3 | Every Ii at 8.1e-15. | — / — | test_spdep_parity.py |
moran_residuals |
R spdep::lm.morantest | spdep 1.4.2; spatialreg 1.4.3 | Statistic 5e-16; the p-value at 1e-7 once X is supplied so the Cliff-Ord regression-residual null can be formed. Both spdep alternatives are recorded because lm.morantest defaults to one-sided. | — / — | test_spdep_parity.py |
mr |
R MendelianRandomization::mr_ivw through the sp.mr dispatcher | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | IVW estimate and default random-effects SE, 1e-10. | — / — | test_mr_R_parity.py |
mr_bma |
Zuber et al. summary_mvMR_BF (GitHub verena-zuber/demo_AMD 4981b5a) | summary_mvMR_BF.R demo_AMD@4981b5a | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
mr_cml |
R MendelianRandomization::mr_cML (DP = FALSE, n = 17723) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Estimate and SE for every K = 0..6, the BIC-selected fit with its invalid set {12, 14}, and the MA-BIC average, all at 1e-9 (both sides iterate to | d theta | <= 1e-7). |
mr_egger |
R MendelianRandomization::mr_egger, TwoSampleMR::mr_egger_regression | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Slope, intercept and both SEs at 1e-10. | — / — | test_mr_R_parity.py |
mr_f_statistic |
R MendelianRandomization::mr_ivw @Fstat | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Mean F statistic at 1e-10. | — / — | test_mr_R_parity.py |
mr_heterogeneity |
R TwoSampleMR::mr_ivw / mr_egger_regression (Q, Q_df, Q_pval) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | IVW and Egger (Ruecker) Q, degrees of freedom and p-values at 1e-10. | — / — | test_mr_R_parity.py |
mr_ivw |
R MendelianRandomization::mr_ivw (default / fixed / random), TwoSampleMR::mr_ivw | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Estimate, SE under all three models, RSE and Cochran's Q at 1e-10 (observed <= 1e-15). | — / — | test_mr_R_parity.py |
mr_leave_one_out |
R MendelianRandomization::mr_ivw on each leave-one-out subset | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | All 28 leave-one-out estimates and default-model SEs at 1e-10. | — / — | test_mr_R_parity.py |
mr_median |
R MendelianRandomization::mr_median (weighted / simple / penalized) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Point estimates for all three weightings at 1e-10. The bootstrap SE is Monte Carlo on both sides and is not compared (T3). | — / — | test_mr_R_parity.py |
mr_mediation |
R MendelianRandomization::mr_ivw + mr_mvivw | MendelianRandomization 0.10.0 | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
mr_mode |
R MendelianRandomization::mr_mbe (weighted / unweighted, stderror = simple) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Point estimates for both weightings at 1e-10 -- the same point of the same 512-point density grid. The bootstrap SE is Monte Carlo on both sides and is not compared (T3). | — / — | test_mr_R_parity.py |
mr_multivariable |
R MendelianRandomization::mr_mvivw (default random effects) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Three direct effects and SEs at 1e-10. | — / — | test_mr_R_parity.py |
mr_pleiotropy_egger |
R TwoSampleMR::mr_egger_regression (intercept test) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Intercept, SE and t(n - 2) p-value at 1e-10. | — / — | test_mr_R_parity.py |
mr_presso |
R MRPRESSO::mr_presso (NbDistribution = 2000) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Raw estimate and SE, observed RSS, and the outlier-corrected estimate and SE at 1e-10; outlier set {12, 14} identical. The simulated p-values are Monte Carlo on both sides (T3) and follow the reference's k / B convention. | — / — | test_mr_R_parity.py |
mr_radial |
R RadialMR::ivw_radial (alpha = 0.05, no Bonferroni) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Square-root weights, per-variant Q contributions and total Q at 1e-10; the outlier set is identical with bonferroni=False (StatsPAI's default applies Bonferroni). | — / — | test_mr_R_parity.py |
mr_steiger |
R TwoSampleMR::mr_steiger with r from get_r_from_bsen | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | R^2 on both traits and the direction at 1e-10; the p-value (1.8e-73) at 1e-12. | — / — | test_mr_R_parity.py |
msm |
ipw::ipwtm + lm / glm(quasibinomial) + sandwich::vcovCL(type = 'HC1') | R 4.5.2; ipw 1.3.0; sandwich 3.1.1 | coefficients and cluster SEs 1e-9 rel (observed 1.5e-13); Stata regress / logit [pw] 1e-7 | — / — | test_teffects_R_parity.py (+2) |
multi_cutoff_rd |
rdmulti::rdmc 2.0.0 (Cattaneo, Titiunik, Vazquez-Bare & Keele) | R 4.5.2; rdmulti 2.0.0 | identical to sp.rdmc on the fixture (exact); per-cutoff and pooled estimates vs R 1e-9 rel | — / — | test_rdmulti_parity.py (+1) |
multi_outcome_synth |
augsynth::augsynth_multiout 0.2.0 (progfunc='None', scm=TRUE, combine_method 'concat' / 'avg'); synth_qp re-run at OSQP eps 1e-12 | R 4.5.2; augsynth 0.2.0; osqp 1.0.0 | weights 1e-10 abs (observed <= 9e-15); per-outcome ATT 1e-9 rel | — / — | test_synth_rest_R_parity.py (+1) |
multi_treatment |
Stata teffects aipw with a multivalued treatment (mlogit propensity) | Stata 18 | 1e-10 rel on both contrasts, potential-outcome means and sandwich SEs | — / — | test_teffects_R_parity.py (+2) |
multiway_cluster_vcov |
sandwich::vcovCL(cluster=~g1+g2+g3) | R 4.5.2; sandwich 3.1.1 | rel_est<=1e-06, rel_se<=1e-06 | 2.1e-15 / 2.1e-15 | 56_multiway_cluster.py (+2) |
nbreg |
MASS::glm.nb | R 4.5.2; MASS 7.3.65 | rel_est<=1e-06, rel_se<=0.005 | 5.1e-10 / 1.0e-11 | 42_nbreg.py (+2) |
negd |
Stata regress, vce(robust) | Stata 18 | 1e-9 rel | — / — | test_misc_sens_stata_parity.py |
netlm |
R sna::netlm | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Coefficients 2e-15 directed and undirected. QAP p-values are permutation draws and are not compared. | — / — | test_network_parity.py |
netlogit |
R sna::netlogit | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Coefficients 1.4e-9 (IRLS on both sides). QAP p-values are permutation draws and are not compared. | — / — | test_network_parity.py |
network_components |
R igraph::components | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Counts and sizes exact on a disconnected graph; weak and strong counts on the directed graph. | — / — | test_network_parity.py |
network_modularity |
R igraph::modularity | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Exact for a fixed split and for igraph's own fast-greedy partition. | — / — | test_network_parity.py |
network_summary |
R igraph edge_density / diameter / mean_distance / transitivity | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Density, diameter, mean path length, transitivity and assortativity exact; average clustering matches igraph::transitivity(type = 'average', isolates = 'zero'), the convention used here. | — / — | test_network_parity.py |
number_needed_to_treat |
base-R closed form (NNT = 1/risk difference) | R 4.5.2 | estimate 1e-12 abs (observed 0); CI not pinned | — / — | test_epi_parity.py (+1) |
oaxaca |
oaxaca::oaxaca | R 4.5.2; oaxaca 0.1.5 | rel_est<=1e-06, rel_se<=0.05 | 6.3e-16 / 1.3e-16 | 30_oaxaca.py (+3) |
odds_ratio |
base-R closed form (Woolf logit; = epiR::epi.2by2) | R 4.5.2 | estimate, se_log, CI 1e-12 abs (observed 0) | — / — | test_epi_parity.py (+1) |
ologit |
MASS::polr(method="logistic") | R 4.5.2; MASS 7.3.65 | rel_est<=1e-06, rel_se<=1e-05 | 1.5e-07 / 3.4e-07 | 45_ologit.py (+2) |
oprobit |
MASS::polr(method="probit") | R 4.5.2; MASS 7.3.65 | rel_est<=1e-06, rel_se<=1e-06 | 3.7e-07 / 2.1e-10 | 49_oprobit.py (+2) |
oster_bounds |
Stata psacalc (Oster); R robomit::o_delta / o_beta | Stata 18; psacalc 2.1; R 4.5.2; robomit 1.0.7 | 1e-12 rel vs psacalc (observed 4.4e-14); 5e-7 abs vs robomit (it rounds to 6 dp) | — / — | test_inference_sens_R_parity.py (+3) |
oster_delta |
Stata psacalc (Oster); R robomit::o_delta / o_beta | Stata 18; psacalc 2.1; R 4.5.2; robomit 1.0.7 | 1e-12 rel vs psacalc (observed 4.4e-14); 5e-7 abs vs robomit | — / — | test_inference_sens_R_parity.py (+3) |
overlap_weights |
WeightIt::weightit 1.7.0 (method='glm'), R 4.5.2 | R 4.5.2; WeightIt 1.7.0 | All four estimands of the shared-propensity family (Li, Li & Li 2019 Table 1) relative to WeightIt: ATO 2.5e-14, ATE 4.1e-14, ATT 2.3e-14, ATC 4.1e-14. The propensity score itself matches R glm(family=binomial) to 2.6e-14 absolute. | — / — | test_overlap_weights_r_parity.py (+1) |
pagerank |
R igraph::page_rank (damping 0.85) | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | 5e-12 undirected, 1e-12 directed. | — / — | test_network_parity.py |
panel |
plm::plm + plm::phtest | R 4.5.2; plm 2.6.7 | rel_est<=1e-06, rel_se<=0.001 | 4.7e-14 / 1.5e-15 | 35_panel.py (+2) |
panel_fgls |
Stata 18 xtgls, panels(hetero) | Stata 18 MP | 6.3e-16 on every coefficient and 7.0e-16 on every standard error against the two-step default, and 4.7e-08 against xtgls, igls for the iterated variant. A reference-free test also asserts the two are distinct estimators, so a silent return to iterating fails. |
— / — | test_panel_stata_parity.py |
panel_logit |
Stata 18 xtlogit, re | Stata 18 MP | Graded by CONVERGENCE rather than a fixed tolerance: Stata integrates adaptively and StatsPAI does not, so the honest claim is that agreement improves as the Gauss-Hermite rule is refined. Observed 3.8e-04 at 12 points and 2.4e-07 at 60, with the log-likelihood at 2.1e-08 -- the sharpest single check, since it is the same objective evaluated at the same optimum. sigma_u and rho at 1e-4. | — / — | test_panel_stata_parity.py |
panel_probit |
Stata 18 xtprobit, re | Stata 18 MP | Same convergence grading: 8.0e-05 at 12 quadrature points and 4.0e-08 at 60, log-likelihood at 1.7e-09, sigma_u and rho at 1e-4. | — / — | test_panel_stata_parity.py |
panel_qtet |
qte::panel.qtet 1.3.1 (Callaway & Li 2019) | — | all 19 quantiles: abs < 1e-8 (observed 6.8e-12); ATT abs < 1e-6. panel.qtet composes ordinary ecdf evaluations and type-7 quantiles, both of which have exact numpy equivalents, so this is machine-precision agreement rather than a tolerance band. | — / — | test_panel_qtet_parity.py (+1) |
panel_unitroot |
plm::purtest 2.6.7; Stata 18 xtunitroot | R 4.5.2; plm 2.6.7; Stata 18 | all statistics 1e-10 rel (observed 9.1e-15) | — / — | test_timeseries_R_parity.py (+2) |
poisson |
stats::glm(family=poisson()) | R 4.5.2; stats 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 9.2e-15 / 8.7e-12 | 58_poisson.py (+2) |
policy_tree |
policytree::policy_tree | R 4.5.2; policytree 1.2.4 | rel_est<=1e-06, rel_se<=1e-06 | 9.6e-16 / 1.4e-16 | 70_policy_tree.py (+2) |
policy_value |
grf::average_treatment_effect(subset = policy == 1) x mean(policy) on grf get_scores; policytree::double_robust_scores reward contrast | R 4.5.2; grf 2.6.1; policytree 1.2.4 | 1e-12 rel (observed 6.7e-16) | — / — | test_r2_teffects_parity.py (+1) |
power_case_control |
Stata power twoproportions; R stats::power.prop.test(strict=TRUE) | R 4.5.2; Stata 18 MP | 1e-11 rel vs Stata (observed 4e-13, Stata's normal CDF); 1e-10 vs R (observed 6e-16) | — / — | test_survival_epi_R_parity.py (+2) |
power_cluster_rct |
base-R closed form (design-effect-inflated z-approx power) | R 4.5.2 | power 1e-12 abs (observed ~2e-16) | — / — | test_power_extra_parity.py (+1) |
power_logrank |
base-R closed form (Schoenfeld log-rank power) | R 4.5.2 | power 1e-12 abs (observed ~2e-16) | — / — | test_power_parity.py (+1) |
power_rct |
base-R closed form (two-sample pooled-sigma z-approx power) | R 4.5.2 | power 1e-12 abs (observed ~2e-16) | — / — | test_power_parity.py (+1) |
power_two_proportions |
base-R closed form (unpooled Wald two-proportion z-approx) | R 4.5.2 | power 1e-12 abs (observed ~2e-16) | — / — | test_power_parity.py (+1) |
ppmlhdfe |
fixest::fepois | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=2e-06 | 4.9e-13 / 2.2e-15 | 37_ppmlhdfe.py (+2) |
prevalence_ratio |
base-R closed form (Katz-log; = epiR::epi.2by2) | R 4.5.2 | estimate, se_log, CI 1e-12 abs (observed ~2e-16) | — / — | test_epi_parity.py (+1) |
principal_strat |
AER::ivreg 1.2.16 (Wald LATE); SACE bounds via Stata leebounds | R 4.5.2; AER 1.2.16 | 1e-10 rel on the monotonicity complier LATE | — / — | test_teffects_R_parity.py (+2) |
probit |
stats::glm(family=binomial("probit")) | R 4.5.2; stats 4.5.2 | rel_est<=1e-06, rel_se<=0.01 | 3.1e-07 / 1.6e-08 | 48_probit.py (+2) |
psm |
MatchIt::matchit | R 4.5.2; MatchIt 4.7.2 | rel_est<=1e-06, rel_se<=1e-06 | 1.2e-15 / 2.0e-16 | 11_psm.py (+2) |
psmatch2 |
Stata 18 MP + psmatch2 4.0.12 / pstest 4.2.2 (Leuven & Sianesi 2003) | Stata 18 MP; psmatch2 4.0.12 | Observed py<->Stata relative gaps on the committed fixtures: nearest-neighbour ATT 1.2e-16 and its analytic SE exactly 0 (as is _weight per row); radius ATT 2.8e-16; Abadie-Imbens ai(1)/ai(2) SE 3.0e-15/1.7e-15; PSM-DID across all five weight regimes 2.0e-14; pstest per-covariate rows 1.3e-14; Mahalanobis ATT 1.2e-13; llr ATT 2.1e-10 (worst of four kernels); kernel ATT 9.1e-10; pstest summary block 1.3e-9. The two loosest rows are bounded by fixture precision, not by the estimator: pstest accumulates MeanBias/MedBias in a Stata float (2.6e-9) and r(seatt) for the llr reroute was captured at 8 significant digits (9.9e-9). | — / — | test_psmatch2_parity.py (+3) |
pwcompare |
Stata 18 margins g, pwcompare(effects) mcompare(noadjust | bonferroni | sidak); R emmeans pairs(adjust=) | Stata 18; R 4.5.2; emmeans 2.0.3 | 1e-10 rel (observed <= 7.9e-15 diff / 4.9e-15 SE) |
qreg |
quantreg::rq | R 4.5.2; quantreg 6.1 | rel_est<=1e-06, rel_se<=0.1 | 3.3e-15 / 4.4e-15 | 40_qreg.py (+2) |
queen_weights |
spdep::poly2nb(queen = TRUE) + nb2listw(style = 'W' / 'S' / 'U') | R 4.5.2; spdep 1.4.2; sf 1.1.1; spData 2.3.5 | neighbour sets exact; weights 1e-12 rel (observed 0) | — / — | test_spatial_survey_R_parity.py (+1) |
rake |
survey::rake (to its fixed point) and survey::calibrate(calfun='raking'); Stata svycal rake | R 4.5.2; survey 4.5; Stata 18 MP | calibrated weight shares 1e-12 rel at tol=1e-14 (observed 1.5e-15); 1e-8 at the default tol=1e-10 (observed 1.0e-10) | — / — | test_survey_calib_R_parity.py (+2) |
rd2d |
R rd2d::rd2d / rd2d.distance 1.0.0 (Cattaneo, Titiunik & Yu) | R; rd2d 1.0.0; sandwich 3.1.1 | estimate.p/q, std.err.p/q, t, CI, cross-point covariance rel 1e-9 (observed 1.6e-11); bandwidths rel 1e-8 (observed 1.2e-12); N exact | — / — | test_rd_open_R_parity.py (+1) |
rd2d_bw |
R rd2d::rdbw2d / rdbw2d.distance 1.0.0 | R; rd2d 1.0.0 | per-point bandwidths rel 1e-8 (observed 5.0e-13) | — / — | test_rd_open_R_parity.py (+1) |
rd_honest |
RDHonest::RDHonest 1.0.1.9000 (Armstrong & Kolesar) | R 4.5.2; RDHonest 1.0.1.9000 | estimate / std.error / maximum.bias / conf.low / conf.high 1e-9 rel at fixed bandwidth; 1e-6 rel when the bandwidth and M are selected | — / — | test_rdhonest_parity.py (+1) |
rdbwhte |
R rdhte::rdbwhte 0.2.0 | R 4.5.2; rdhte 0.2.0; rdrobust 4.0.0 | bandwidths rel 1e-8 (continuous and per-subgroup) | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdbwselect |
rdrobust::rdbwselect; Stata side uses the authors' rdbwselect ado. certwo is R-only: Stata rdbwselect 10.0.0 exits r(3200) on it, including on the package's own rdrobust_senate.dta | R 4.5.2; rdrobust 3.0.0 | rel_est<=1e-06 | 9.4e-13 / 3.5e-09 | 88_rdbwselect.py (+2) |
rdd |
rdrobust::rdrobust | R 4.5.2; rdrobust 3.0.0 | rel_est<=1e-06, rel_se<=0.1 | 2.5e-14 / 9.4e-11 | 06_rd.py (+3) |
rddensity |
rddensity::rddensity | R 4.5.2; rddensity 2.6 | rel_est<=1e-06, rel_se<=1e-06 | 9.3e-12 / 1.8e-11 | 09_rddensity.py (+2) |
rdhte |
R rdhte::rdhte 0.2.0 (sandwich 3.1.1) | R 4.5.2; rdhte 0.2.0; sandwich 3.1.1; rdrobust 4.0.0 | coef, coef.bc, se.rb, vcov rel 1e-9 (observed 2.2e-12); bandwidths 1e-8 | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdhte_lincom |
R rdhte::rdhte_lincom 0.2.0 | R 4.5.2; rdhte 0.2.0 | estimate, z, CI, joint chi-square rel 1e-9; p 1e-8 | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdmc |
rdmulti::rdmc 2.0.0 (Cattaneo, Titiunik, Vazquez-Bare & Keele) | R 4.5.2; rdmulti 2.0.0 | per-cutoff coefficients, robust coefficients, robust SEs and the pooled weighted estimate 1e-9 rel; selected bandwidths 1e-5 rel | — / — | test_rdmulti_parity.py (+1) |
rdms |
rdmulti::rdms; Stata side uses the rdms ado from the rdpackages GitHub mirror (rdmulti is not on SSC: ssc describe rdmulti returns r(601)), installed into a local gitignored ado path | R 4.5.2 | rel_est<=1e-06, rel_se<=1e-06 | 1.2e-12 / 5.4e-10 | 89_rdms.py (+2) |
rdplot |
R rdrobust::rdplot 4.0.0 | R 4.5.2; rdrobust 4.0.0 | J / J_IMSE / J_MV and bin counts exact; bin means, SEs, t intervals and polynomial values rel 1e-9 (observed 4.0e-12); p = 4 global polynomial coefficients rel 1e-8 (observed 1.6e-10, raw-scale conditioning) | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdplotdensity |
R rddensity::rdplotdensity 2.6 (lpdensity 2.5) | R 4.5.2; rddensity 2.6; lpdensity 2.5 | f_p, f_q, se_p, se_q at every grid point rel 1e-8 (observed 8.0e-12); nh exact | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdpower |
rdpower::rdpower 3.0 (Cattaneo, Titiunik & Vazquez-Bare) | R 4.5.2; rdpower 3.0 | robust bias-corrected SE & power 1e-8 rel (observed 4.2e-14) | — / — | test_rdlocrand_parity.py (+1) |
rdrandinf |
rdlocrand::rdrandinf 2.0 (Cattaneo, Titiunik & Vazquez-Bare) | R 4.5.2; rdlocrand 2.0 | observed statistic & asymptotic p-value 1e-8 rel (observed 2.3e-15) | — / — | test_rdlocrand_parity.py (+1) |
rdrobust |
rdrobust::rdrobust | R 4.5.2; rdrobust 3.0.0 | rel_est<=1e-06, rel_se<=0.1 | 2.5e-14 / 9.4e-11 | 06_rd.py (+2) |
rdsampsi |
rdpower::rdsampsi 3.0 (Cattaneo, Titiunik & Vazquez-Bare) | R 4.5.2; rdpower 3.0 | required sample sizes n_left / n_right / n_total asserted as exact integer equality (no tolerance) | — / — | test_rdlocrand_parity.py (+1) |
rdwinselect |
rdlocrand::rdwinselect 2.0 (Cattaneo, Titiunik & Vazquez-Bare) | R 4.5.2; rdlocrand 2.0 | window grid 1e-12 rel (observed 0); per-window counts Nl / Nr asserted as exact integer equality | — / — | test_rdlocrand_parity.py (+1) |
reciprocity |
R igraph::reciprocity | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Exact on a 40-node directed graph. | — / — | test_network_parity.py |
regress |
lm + sandwich::vcovHC | R 4.5.2; sandwich 3.1.1 | rel_est<=1e-06, rel_se<=1e-06 | 1.1e-12 / 1.3e-12 | 01_ols.py (+2) |
relative_risk |
base-R closed form (Katz-log; = epiR::epi.2by2 / Stata epitab) | R 4.5.2 | estimate, se_log, CI 1e-12 abs (observed 0) | — / — | test_epi_parity.py (+1) |
reset_test |
lmtest::resettest(power=2:3, type='fitted') | R 4.5.2; lmtest 0.9.40 | F-statistic & p-value 1e-10 rel (observed ~1e-13) | — / — | test_diagnostics_parity.py (+1) |
ri_test |
R ri2::conduct_ri (randomizr full enumeration); Stata ritest over the full assignment set | R 4.5.2; ri2 0.5.0; Stata 18; ritest 1.1.7 | p exact (k / N_assignments); statistic 1e-12 rel vs R, 2^-23 vs Stata | — / — | test_inference_sens_R_parity.py (+3) |
rif_decomposition |
dineq::rif + manual OLS | R 4.5.2; dineq 0.1.0 | rel_est<=1e-06, rel_se<=1e-06 | 2.2e-15 / 1.4e-16 | 32_rif.py (+2) |
rifreg |
R rifreg::rifreg (variance, quantiles) and dineq::rif + lm (Gini) | rifreg 1.1.0; dineq 0.1.0; ddecompose 1.0.0 | Coefficients at 1e-10 for the variance and the 10th/50th/90th percentiles (quantile_convention='rifreg') and for the Gini against dineq's exact RIF; the stock rifreg Gini, which integrates the Lorenz curve numerically, agrees to 1e-4. | — / — | test_decomp_R_parity.py |
risk_difference |
base-R closed form (Wald; = epiR::epi.2by2 / Stata epitab) | R 4.5.2 | estimate, se, CI 1e-12 abs (observed 0) | — / — | test_epi_parity.py (+1) |
rkd |
R rdrobust::rdrobust(deriv = 1, vce = 'hc1') 4.0.0; Stata rdrobust, deriv(1) 11.1.0 | R 4.5.2; rdrobust 4.0.0; Stata 18 MP; rdrobust (Stata) 11.1.0 | conventional and robust estimate / SE, bandwidth: rel 1e-9 (observed 5.2e-12 vs R, 6.0e-11 vs Stata) | — / — | test_rd_iv_rd_R_parity.py (+2) |
rlasso |
R hdm::rlasso 0.3.2 | R 4.5.2; hdm 0.3.2 | support exact; beta / sigma / loadings / residuals atol 1e-6, lambda0 rtol 1e-8 (observed rel <= 1.8e-12) | — / — | test_rlasso_parity.py (+1) |
rlasso_effect |
R hdm::rlassoEffect 0.3.2 | R 4.5.2; hdm 0.3.2 | alpha / se atol 1e-6 (observed rel <= 3.3e-15), incl. hdm's GrowthData vignette | — / — | test_rlasso_parity.py (+2) |
rlasso_effects |
R hdm::rlassoEffects 0.3.2 | R 4.5.2; hdm 0.3.2 | alpha / se rtol 1e-9 on cps2012 (observed 2.8e-12); atol 1e-6 on the synthetic fixture | — / — | test_rlasso_parity.py (+2) |
rlasso_iv |
R hdm::rlassoIV 0.3.2 | R 4.5.2; hdm 0.3.2 | coef / se atol 1e-6 (observed <= 3e-15); EminentDomain atol 1e-4 (observed 6.1e-9: pseudo-inverse of a rank-deficient control block) | — / — | test_rlasso_parity.py (+2) |
rlassologit_effect |
R hdm::rlassologitEffect 0.3.2 | R 4.5.2; hdm 0.3.2 | alpha / se atol 1e-6 (observed rel 1.1e-15 / 5.3e-14 post; se 3.2e-7 with post=False) | — / — | test_rlassologit_effect_parity.py (+1) |
rlassologit_effects |
R hdm::rlassologitEffects 0.3.2 | R 4.5.2; hdm 0.3.2 | coef / se atol 1e-6 (observed rel 8.6e-16 / 2.5e-14) | — / — | test_rlassologit_effect_parity.py (+1) |
robust_synth |
scpi::scest(w.constr = list(name = 'ols')) with scdata(constant = TRUE) and stats::lm (unconstrained SC with intercept); glmnet (ridge / lasso / elastic net, unpenalised intercept) | R 4.5.2; scpi 4.0.1; glmnet 4.1.10 | OLS weights / intercept / fitted path 1e-10 rel (observed 8.4e-13); penalised paths 1e-8 rel, weights atol 1e-10 (observed 3.4e-11 abs, glmnet's coordinate-descent stop) | — / — | test_did_synth_synthvar_parity.py (+1) |
roc_curve |
R pROC::roc/auc/var/ci.auc (DeLong); Stata roctab (default and hanley) | R 4.5.2; pROC 1.19.0.1; Stata 18 MP | 1e-10 rel (observed 6e-15) | — / — | test_survival_epi_R_parity.py (+2) |
rook_weights |
spdep::poly2nb(queen = FALSE) + nb2listw(style = 'W') | R 4.5.2; spdep 1.4.2; sf 1.1.1; spData 2.3.5 | neighbour sets exact; weights 1e-12 rel (observed 0) | — / — | test_spatial_survey_R_parity.py (+1) |
rosenbaum_bounds |
R DOS2::senWilcox (Rosenbaum); Stata rbounds; R stats::binom.test (sign test); R rbounds::psens (zero_method='wilcox', 4-dp) | R 4.5.2; DOS2 0.5.2; rbounds 2.2; Stata 18; rbounds (Stata) 1.1.6 | bounding p-values 1e-12 rel (observed 4.6e-15); Stata sig- 1e-15 abs; psens at its own 4-dp rounding | — / — | test_inference_sens_R_parity.py (+3) |
rosenbaum_gamma |
R DOS2::senWilcox (Rosenbaum); Stata rbounds | R 4.5.2; DOS2 0.5.2; Stata 18; rbounds (Stata) 1.1.6 | bounding p-values 1e-12 rel (observed 4.6e-15) | — / — | test_inference_sens_R_parity.py (+1) |
sac |
R spatialreg::sacsarlm | spdep 1.4.2; spatialreg 1.4.3 | rho and lambda at 1e-5, slope coefficients at 1e-6 -- a bounded two-parameter ML line search on both sides. | — / — | test_spdep_parity.py |
sar |
spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) | R 4.5.2; spatialreg 1.4.3 | rel_est<=1e-06, rel_se<=1e-06 | 8.3e-08 / 5.1e-08 | 65_spatial.py (+2) |
sar_gmm |
spatialreg::stsls(W2X=FALSE) / spatialreg::GMerrorsar | R 4.5.2; spatialreg 1.4.3 | rel_est<=1e-06, rel_se<=1e-06 | 4.6e-08 / 7.3e-16 | 66_spatial_gmm.py (+2) |
sbw |
sbw::sbw 1.2 (Zubizarreta 2015), quadprog solver | — | ATT rel <= 1e-8 (observed <= 4e-10) under both standardisation conventions: tolerance_scale='target' == bal_std='target' and 'group' == bal_std='group', at bal_tol 0.05 and 0.02. | — / — | test_matching_r_parity.py (+1) |
sc_estimate |
synthdid::sc_estimate 0.0.9 (Arkhangelsky, Athey, Hirshberg, Imbens & Wager) | R 4.5.2; synthdid 0.0.9 | estimate, unit/time weights and jackknife SE at 1e-9 on the Prop. 99 replica and a five-treated panel (observed est 8.1e-15, jackknife SE 4.8e-15); each of R's 40 placebo and 40 bootstrap replications replayed at 1e-9 (observed 2.1e-11) | — / — | test_did_synth_R_parity.py (+1) |
scdata |
R scpi::scdata (features = outcome, no cov.adj, constant = FALSE) | scpi 4.0.1; CVXR 1.9.2; ECOSolveR 0.6.1; Qtools 1.6.0; quantreg 6.1 | A, B, P matrices identical (0 difference) on scpi_germany; donor order = R sort(B.names). | — / — | test_did_synth_scpi_parity.py |
sdid |
synthdid::synthdid_estimate | R 4.5.2; synthdid 0.0.9 | rel_est<=1e-06, rel_se<=1e-06 | 7.8e-16 / 7.2e-08 | 12_sdid.py (+2) |
sdm |
spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) | R 4.5.2; spatialreg 1.4.3 | rel_est<=1e-06, rel_se<=1e-06 | 8.3e-08 / 5.1e-08 | 65_spatial.py (+2) |
sem |
spatialreg::lagsarlm / spatialreg::errorsarlm / spatialreg::lagsarlm(Durbin=TRUE) | R 4.5.2; spatialreg 1.4.3 | rel_est<=1e-06, rel_se<=1e-06 | 8.3e-08 / 5.1e-08 | 65_spatial.py (+2) |
sem_gmm |
spatialreg::stsls(W2X=FALSE) / spatialreg::GMerrorsar | R 4.5.2; spatialreg 1.4.3 | rel_est<=1e-06, rel_se<=1e-06 | 4.6e-08 / 7.3e-16 | 66_spatial_gmm.py (+2) |
sensemakr |
sensemakr::sensemakr | R 4.5.2; sensemakr 0.1.6 | rel_est<=1e-06, rel_se<=1e-06 | 5.0e-08 / 5.0e-08 | 22_sensemakr.py (+2) |
sensitivity_specificity |
R epiR::epi.tests (method wilson / exact); Stata diagti | R 4.5.2; epiR 2.0.94; Stata 18 MP; diagt 2.032 (diagti 2.053) | 1e-12 rel on intervals, 1e-10 on ratios (observed 2e-15) | — / — | test_survival_epi_R_parity.py (+2) |
shift_share_political |
AER::ivreg + sandwich HC1; ShiftShareSE::ivreg_ss (EHW / AKM / AKM0); bartik.weight::bw; anova(lm) share balance | R 4.5.2; AER 1.2.16; sandwich 3.1.1; ShiftShareSE 1.1.0; bartik.weight 0.1.0 | estimate / SEs / Rotemberg / F 1e-9 rel (observed <= 5e-15); AKM p-value 1e-7 rel (ShiftShareSE uses 2*(1-pnorm), cancellation at p ~ 1e-9) | — / — | test_synth_rest_R_parity.py (+1) |
shift_share_political_panel |
fixest::feols 0.14.0 (ssc adj=FALSE, cluster.adj=FALSE; unit / time / two-way clusters; unit, time, two-way FE; unbalanced panel); ShiftShareSE::ivreg_ss with FE dummies (AKM); bartik.weight::bw with FE dummies; AER + HC0 per period | R 4.5.2; fixest 0.14.0; ShiftShareSE 1.1.0; bartik.weight 0.1.0; AER 1.2.16; sandwich 3.1.1 | estimate / SEs / Rotemberg / first-stage F 1e-9 rel (observed <= 1.1e-14) | — / — | test_synth_rest_R_parity.py (+1) |
shift_share_se |
R ShiftShareSE::ivreg_ss (Adao, Kolesar & Morales), AKM row; Stata SSC ivreg_ss | R 4.5.2; ShiftShareSE 1.1.0; Stata 18 MP; ivreg_ss SSC 20241116 | 1e-9 rel on beta and the AKM / AKM0 / EHW / Homoscedastic SEs; observed <= 6e-14 | — / — | test_did_synth_shiftshare_parity.py (+2) |
slx |
R spatialreg::lmSLX | spdep 1.4.2; spatialreg 1.4.3 | Every coefficient at 1e-10. | — / — | test_spdep_parity.py |
snips |
Open Bandit Pipeline (obp) 0.5.7 SelfNormalizedInverseProbabilityWeighting | obp 0.5.7 | value 1e-12 rel (observed 0.0); delta-method SE identity 1e-10 and equal to sp.ope.snips | — / — | test_ml_causal_obp_parity.py (+1) |
source_decompose |
Stata descogini (Lerman-Yitzhaki) | Stata 18 MP; descogini SSC | With gini='population': total Gini, and each source's S_k, G_k, R_k and share of the total at 1e-12. The default n/(n-1)-corrected Gini differs by exactly that factor; S_k, R_k and shares are identical. | — / — | test_decomp_R_parity.py |
spatial_iv |
sphet::spreg(model = 'lag', het = TRUE) 2.1.1 | R 4.5.2; sphet 2.1.1; spdep 1.4.2 | coefficients and HC0 SEs 1e-10 rel (observed 9.5e-14 / 6.8e-14) | — / — | test_spatial_survey_R_parity.py (+1) |
spillover_did |
R did::att_gt(control_group='nevertreated') + did::aggte(type='simple') per group (direct / ring r, ring cohort = exposure onset); single cohort also fixest::feols(dbar ~ treat + ring1 + ring2, vcov='hetero', ssc(adj=FALSE)) | did 2.3.0; DRDID 1.2.3; fixest 0.14.0 | Direct and ring effects, their SEs and every (group, onset cohort, period) cell at 1e-9 relative on a single-cohort and a staggered spatial panel; observed agreement 1e-14. The ring construction is recomputed independently in R from the coordinates. | — / — | test_did_synth_misc_parity.py |
sqreg |
R quantreg::rq (Barrodale-Roberts), Koenker 2005 | quantreg see sqreg_R.json provenance | Coefficients 3.5e-14 against quantreg::rq at tau = 0.25 / 0.50 / 0.75 -- both sides minimise the same pinball loss with the same simplex. Standard errors differ from R's se='iid' by ONE SCALAR PER QUANTILE, constant across coefficients to 6e-16: the sandwich is identical and only the sparsity estimate 1/f(0) differs (Powell kernel here, Koenker-Bassett with a Siddiqui/Hall-Sheather bandwidth there). The test asserts the ratio's constancy rather than a numerical band, which a structural difference could not satisfy. R's default se='nid' (Hendricks-Koenker, also Stata qreg's) is a third convention and is recorded as one. | — / — | test_sqreg_parity.py |
ssaggregate |
R ShiftShareSE::ivreg_ss / reg_ss (Adao, Kolesar & Morales) and ssaggregate (Borusyak, Hull & Jaravel; R kylebutts/ssaggregate + AER::ivreg/sandwich HC0); Stata SSC ivreg_ss / reg_ss and ssaggregate + ivreg2, robust | R 4.5.2; ShiftShareSE 1.1.0; ssaggregate (R) 0.0.0.9000 (GitHub kylebutts/ssaggregate@22df93980250891a0cc247f6020136cd33c65ba2); AER 1.2.16; sandwich 3.1.1; Stata 18 MP; reg_ss / ivreg_ss SSC 20241116; ssaggregate (Stata) SSC 1.2.2 (20200826) | 1e-9 rel on beta, every SE row (Homoscedastic, EHW, Reg. cluster, AKM, AKM0), AKM/AKM0 CIs and the shock-level frame; p-values also atol 1e-15 (references use 2*(1-Phi)); observed <= 6e-14 (frame 2.9e-13) | — / — | test_did_synth_shiftshare_parity.py (+2) |
stabilized_weights |
ipw::ipwtm 1.3.0 (type = 'all'; binomial logit; gaussian via geepack::geeglm) | R 4.5.2; ipw 1.3.0; geepack 1.3.13 | 1e-10 rel on every weight (observed 1.5e-12) | — / — | test_teffects_R_parity.py (+2) |
stacked_did |
hand-written stack + fixest::feols | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06 | 3.9e-13 / 7.1e-13 | 75_stacked.py (+2) |
staggered_cs |
staggered::staggered_cs 1.2.2 (Roth & Sant'Anna) | R 4.5.2; staggered 1.2.2 | estimate, Neyman SE and adjusted SE at abs 1e-9 on mpdta, a randomised rollout and a null panel x simple/cohort/calendar (observed rel 6.6e-14) | — / — | test_staggered_extended_parity.py (+1) |
staggered_rollout |
staggered::staggered / staggered_cs / staggered_sa (1.2.2) | R 4.5.2 | rel_est<=1e-10 | 9.1e-16 / 5.7e-16 | 82_staggered.py (+2) |
staggered_sa |
staggered::staggered_sa 1.2.2 (Roth & Sant'Anna) | R 4.5.2; staggered 1.2.2 | estimate, Neyman SE and adjusted SE at abs 1e-9 on mpdta, a randomised rollout and a null panel x simple/cohort/calendar (observed rel 6.6e-14) | — / — | test_staggered_extended_parity.py (+1) |
staggered_synth |
augsynth::multisynth 0.2.0 (Ben-Michael, Feller & Rothstein; partially pooled SCM for staggered adoption) | R 4.5.2; augsynth 0.2.0; osqp 1.0.0 | ATT, per-unit / per-cohort ATT, event-time ATT and jackknife SE 1e-8 rel (observed <= 2.6e-11); weights 1e-8 abs (observed 2.1e-10); nu 1e-8, imbalance norms 1e-6 rel | — / — | test_did_synth_synthvar_parity.py (+1) |
structural_break |
strucchange::Fstats + sctest(type = 'supF') / breakpoints; mbreaks::dosequa (Bai-Perron sequential); Stata estat sbsingle | R 4.5.2; strucchange 1.5.4; mbreaks 1.0.1; Stata 18 | statistics 1e-10 rel (observed <= 1e-14); sup-F p-values 1e-10 rel for p > 1e-6, 1e-13 abs below (R uses 1 - pchisq); break dates exact | — / — | test_r2_ts_parity.py (+3) |
subcluster_wild_bootstrap |
Stata boottest, bootcluster() (CRVE clustered at g6, signs flipped at s12) | Stata 18; boottest 4.5.3 | p exact (multiple of 1/4096); t 1e-10 rel (observed 1.8e-14) | — / — | test_inference_sens_stata_parity.py (+1) |
subgroup_analysis |
Stata regress + testparm | Stata 18 | 1e-9 rel | — / — | test_misc_sens_stata_parity.py |
subgroup_decompose |
Stata ineqdeco, bygroup() | Stata 18 MP; ineqdeco SSC | GE(0), GE(1), GE(2) totals and within components at 1e-12, between components at 1e-11. The Gini path (Dagum) has no ineqdeco counterpart and keeps its analytical checks. | — / — | test_decomp_R_parity.py |
sun_abraham |
fixest::sunab | R 4.5.2; fixest 0.14.0 | rel_est<=1e-06, rel_se<=0.03 | 2.8e-11 / 2.7e-11 | 05_sunab.py (+2) |
sureg |
systemfit::systemfit(method="SUR", noDfCor) | R 4.5.2; systemfit 1.1.30 | rel_est<=1e-06, rel_se<=1e-06 | 1.5e-14 / 1.5e-15 | 60_sureg.py (+2) |
survivor_average_causal_effect |
Stata leebounds 1.5 (Zhang-Rubin SACE bounds = Lee bounds under monotonicity) | Stata 18; leebounds 1.5 (2013-07-17, Tauchmann) | 1e-10 rel (upper bound; both bounds equal sp.lee_bounds) | — / — | test_teffects_R_parity.py (+2) |
svydesign |
survey::svydesign + svymean/svytotal/svyglm/degf (strata, nested PSUs, fpc, survey.lonely.psu); Stata svyset + svy: mean/total/regress/logit/poisson + estat effects | R 4.5.2; survey 4.5; Stata 18 MP | estimates, SEs, DEFF, CI bounds 1e-10 rel, p-values 1e-9 rel (observed <= 4e-14; p 6e-13) | — / — | test_survey_design_R_parity.py (+2) |
svyglm |
survey::svyglm (design-based GLM + linearization SE) | R 4.5.2 | coefficients + SE 1e-10 abs (observed ~2e-15 / 6e-15) | — / — | test_survey_parity.py (+1) |
svymean |
survey::svymean (Horvitz-Thompson/Hajek + Taylor SE) | R 4.5.2 | estimate + SE 1e-10 abs (observed ~5e-15 / 8e-17) | — / — | test_survey_parity.py (+1) |
svytotal |
survey::svytotal (Horvitz-Thompson + Taylor SE) | R 4.5.2 | estimate 1e-12 rel; SE 1e-10 rel (observed ~2e-12 / 1e-14) | — / — | test_survey_parity.py (+1) |
synth |
Synth::synth | R 4.5.2; Synth 1.1.10 | rel_est<=1e-06, rel_se<=1e-06 | 7.8e-08 / 7.7e-08 | 52_scm_unique.py (+2) |
synth_donor_sensitivity |
Synth::synth 1.1.10 on the replayed numpy donor subsets (custom.v matched; quadprog exact) | R 4.5.2; Synth 1.1.10; quadprog 1.5.8 | ATT and pre-RMSPE 1e-9 rel (observed <= 1e-13) | — / — | test_synth_rest_R_parity.py (+1) |
synth_loo |
Synth::synth 1.1.10 (per fit, custom.v matched; QP solved exactly by quadprog::solve.QP) | R 4.5.2; Synth 1.1.10; quadprog 1.5.8; kernlab 0.9.33 | ATT and pre-RMSPE 1e-9 rel (observed <= 3e-14) | — / — | test_synth_rest_R_parity.py (+1) |
synth_rmspe_filter |
SCtools::mspe.test 0.3.3.1 (discard.extreme, mspe.limit 2/5/20) on Synth fits | R 4.5.2; SCtools 0.3.3.1; Synth 1.1.10; quadprog 1.5.8 | placebo pre-RMSPE and ratios 1e-9 rel (observed <= 1e-13); p-values and kept counts exact | — / — | test_synth_rest_R_parity.py (+1) |
synth_time_placebo |
Synth::synth 1.1.10 (per placebo time, custom.v matched; quadprog exact) | R 4.5.2; Synth 1.1.10; quadprog 1.5.8 | placebo ATT 1e-9 rel (observed <= 1e-13) on identified placebo times | — / — | test_synth_rest_R_parity.py (+1) |
synthdid_estimate |
synthdid::synthdid_estimate | R 4.5.2; synthdid 0.0.9 | rel_est<=1e-06, rel_se<=1e-06 | 7.8e-16 / 7.2e-08 | 12_sdid.py (+3) |
synthesise_evidence |
R metafor::rma(method = 'FE') | metafor 5.0.1 | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
tF_adjustment |
R ivDiag::tF 1.0.6 (LMMP 2022 tF table) | R 4.5.2; ivDiag 1.0.6 | rel 1e-12 on the same 26 F values (observed 0) | — / — | test_rd_iv_R_parity.py (+1) |
tF_critical_value |
R ivDiag::tF 1.0.6 (LMMP 2022 tF table) | R 4.5.2; ivDiag 1.0.6 | critical value at 26 F values from 4 to 1e4: rel 1e-12 (observed 0) | — / — | test_rd_iv_R_parity.py (+1) |
test_calibration |
grf::test_calibration 2.6.1 (vcov.type HC3 default and HC1), forest outputs held fixed | R 4.5.2; grf 2.6.1; sandwich 3.1.1 | coef, se, t 1e-10 rel (observed 1.1e-15); one-sided p-value 1e-8 rel (observed 1.6e-13) | — / — | test_ml_causal_R_parity.py (+1) |
three_sls |
R systemfit::systemfit(method='3SLS') | R 4.5.2; systemfit 1.1.30 | coef 1e-9 abs (observed <= 1e-15); SE ~5e-3 rel | — / — | test_threesls_parity.py (+1) |
tmle |
tmle::tmle | R 4.5.2; tmle 2.1.1 | rel_est<=1e-06, rel_se<=1e-06 | 1.9e-09 / 1.9e-09 | 72_tmle.py (+2) |
tobit |
censReg::censReg | R 4.5.2; censReg 0.5.38 | rel_est<=1e-06, rel_se<=1e-05 | 2.8e-08 / 2.8e-08 | 41_tobit.py (+2) |
transitivity |
R igraph::transitivity(type = 'global') | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Exact on karate. | — / — | test_network_parity.py |
transport_weights_fn |
R glm + quantile(type = 7) + sandwich::vcovHC(HC0) | sandwich 3.1.1 | 1e-12 rel | — / — | test_misc_sens_R_parity.py |
truncreg |
truncreg::truncreg(method="NR") | R 4.5.2; truncreg 0.2.5 | rel_est<=1e-06, rel_se<=1e-06 | 3.5e-10 / 7.9e-08 | 62_truncreg.py (+2) |
twoway_cluster |
sandwich::vcovCL(cluster=~g1+g2) | R 4.5.2; sandwich 3.1.1 | rel_est<=1e-06, rel_se<=1e-06 | 7.8e-16 / 7.8e-16 | 54_twoway_cluster.py (+2) |
unified_sensitivity |
R EValue::evalues.OLS; sensemakr::sensemakr (rv_q, rv_qa) | EValue 4.1.4; sensemakr 0.1.6 | 1e-10 rel | — / — | test_misc_sens_R_parity.py |
var |
vars::VAR | R 4.5.2; vars 1.6.1 | rel_est<=1e-06, rel_se<=1e-06 | 3.1e-15 / 6.6e-15 | 33_var.py (+2) |
vif |
R car::vif | ivmodel 1.9.1; car 3.1.5; metafor 5.0.1 | 2.0e-16 on every variance inflation factor, once the returned values stopped being rounded to two decimals. | — / — | test_weakiv_meta_parity.py |
wild_cluster_boot |
R fwildclusterboot::boottest; Stata boottest (WCR, Rademacher, full enumeration) | R 4.5.2; fwildclusterboot 0.14.3; Stata 18; boottest 4.5.3 | p exact (multiple of 1/4096); t 1e-10 rel (observed 1.9e-14) | — / — | test_inference_sens_R_parity.py (+3) |
wild_cluster_bootstrap |
R fwildclusterboot::boottest; Stata boottest (WCR, Rademacher, full enumeration) | R 4.5.2; fwildclusterboot 0.14.3; Stata 18; boottest 4.5.3 | p exact (multiple of 1/4096); t 1e-10 rel (observed 1.9e-14) | — / — | test_inference_sens_R_parity.py (+3) |
wild_cluster_ci_inv |
R fwildclusterboot::boottest confidence interval (uniroot tol 1e-13) | R 4.5.2; fwildclusterboot 0.14.3 | CI endpoints 1e-9 rel (observed 4.8e-12) | — / — | test_inference_sens_R_parity.py (+2) |
wooldridge_did |
etwfe::etwfe + emfx | R 4.5.2; etwfe 0.6.2 | rel_est<=1e-06, rel_se<=0.001 | 1.8e-13 / 3.9e-14 | 17_etwfe.py (+2) |
xtabond |
plm::pgmm | R 4.5.2; plm 2.6.7 | rel_est<=1e-06, rel_se<=1e-06 | 9.0e-16 / 1.4e-15 | 50_xtabond.py (+2) |
xtdpdsys |
Stata 18 xtdpdsys (built-in); xtabond2 iv(x, eq(diff)) h(2) | Stata 18 MP; xtabond2 SSC 03.07.00 | abdata, n on L.n (and w k): coefficients and SEs at 1e-9 (observed <= 6e-12) for one-step robust, two-step Windmeijer and classical one-step, instrument count equal; xtabond2 with iv(w k, eq(diff)) h(2) reproduces the same numbers. | — / — | test_dynpanel_abdata_parity.py |
xtlsdvc |
Stata xtlsdvc V1.0.4 (Bruno 2005), SSC | Stata 18 MP; xtlsdvc 1.0.4 | abdata: bias-corrected coefficients for initial(ab/ah/bb) x bias(1/2/3) and the AR(1)-only model at rtol 1e-7 (observed 4e-14 to 1.6e-9, the looser end through the Anderson-Hsiao initialiser). Standard errors are excluded by design: xtlsdvc reports the uncorrected LSDV ones, which sp.xtlsdvc reproduces and warns about. | — / — | test_lsdvc_parity.py |
xtnbreg |
Stata 18 xtnbreg, fe / re (Hausman-Hall-Griliches); pglm::pglm(family = negbin, model = 'within' / 'random') | Stata 18 MP; R 4.5.2; pglm 0.2.4 | coefficients / SEs rtol 1e-7 (observed <= 2.7e-8: Stata fe's score at its own estimate is 3e-7, pglm's gradtol), log-likelihood rtol 1e-12 | — / — | test_panel_xtnbreg_parity.py |
yu_elwert_decompose |
R cdgd::cdgd0_manual on independently fitted within-cell lm / within-group glm nuisances | DasGuptR 2.2.0; ddecompose 1.0.0; cdgd 1.0.1 | method='efficient': disparity, baseline, prevalence, effect, selection and their EIF standard errors at 1e-9. method='plugin' has no reference implementation and is covered by its exact additivity identity. | — / — | test_decomp_R_parity.py |
zip_model |
pscl::zeroinfl(dist="poisson") | R 4.5.2; pscl 1.5.9 | rel_est<=1e-06, rel_se<=0.0001 | 1.2e-07 / 2.8e-08 | 63_zip.py (+2) |
aligned — 53 functions¶
Agreement within a documented, pre-registered looser tolerance.
| function | reference | versions | tolerance | rel err (R / Stata) | test |
|---|---|---|---|---|---|
ackerberg_caves_frazer |
Stata prodest, method(lp) acf valueadded; R prodest::prodestACF | Stata 18 MP; prodest SSC; R prodest 1.0.2 | StatsPAI returns an exact root of the just-identified moment conditions (criterion < 1e-25), the one nearest the stage-1 coefficients; R stops 2.4e-6 (relative) from that root with criterion 9e-16. Stata's Nelder-Mead stops at non-roots (criterion 4e-6 / 7e-6), which the test records. The parity panel has three roots, all reported in diagnostics['acf_roots']. | — / — | test_prodest_parity.py |
aft |
survival::survreg (Weibull AFT) | R 4.5.2; survival 3.8.3 | coefficients & log-scale 5e-5 abs (observed ~1e-5) | — / — | test_aft_parity.py (+1) |
augsynth |
augsynth::augsynth | R 4.5.2; augsynth 0.2.0 | rel_est<=2e-05, rel_se<=1e-06 | 7.9e-06 / 1.7e-08 | 18_augsynth.py (+2) |
blp |
pyblp 1.2.0 (Conlon & Gortmaker), identical Halton nodes via agent_data | pyblp 1.2.0; numpy 2.2.6; python 3.10.20 | beta, sigma, SEs, objective/N, own elasticities 1e-6 rel (observed <= 1.6e-8) | — / — | test_blp_pyblp_parity.py (+1) |
breakdown_m |
HonestDiD::findOptimalFLCI 0.2.8 (Rambachan & Roth), breakdown by uniroot on the bound facing zero | R 4.5.2; HonestDiD 0.2.8; CVXR 1.8.2 | vs HonestDiD with its Monte-Carlo folded-normal quantile replaced by the exact one: 1e-9 (observed 6.1e-11); vs HonestDiD as shipped: 1e-3 (observed 4.3e-4), the simulation error of .qfoldednormal (1e6 draws, seed 0; 1.96224 vs exact 1.95996 at mu = 0) | — / — | test_did_synth_R_parity.py (+1) |
cate_eval |
grf::rank_average_treatment_effect 2.6.1 (AUTOC, QINI, TOC), fed grf's nuisances; shares sp.rate's operator | R 4.5.2; grf 2.6.1 | Point estimate (AUTOC, QINI) and TOC curve on q = 0.1..1 exact at 1e-10 rel (observed 4.6e-15), including a 30-group tied-priority case against rank_average_treatment_effect.fit. SE is T3: the analytic rank-corrected influence-function SE is compared with grf's half-sample bootstrap (R = 2000) at 6% rel (observed 2.6% AUTOC, 1.7% QINI). | — / — | test_ml_causal_R_parity.py (+1) |
causal_forest |
grf::causal_forest | R 4.5.2; grf 2.6.1 | rel_est<=0.01, rel_se<=0.05 | 2.4e-03 / — | 13_causal_forest.py (+1) |
cbps |
CBPS::CBPS 0.24 (Imai & Ratkovic 2014) | — | ATE over/exact and ATT exact: rel <= 5e-3 (R's optimiser slack). ATT over is NOT pinned to R -- CBPS's ATT gradient mis-scales the balance block by n/n_1 and stops off-stationarity; StatsPAI is asserted to attain strictly better covariate balance instead. | — / — | test_matching_r_parity.py (+1) |
centrality |
R igraph degree / betweenness / closeness / page_rank / eigen_centrality | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | degree, betweenness (normalised and raw), closeness and PageRank at 1e-10 on Zachary's karate club. The eigenvector column is L2-normalised (networkx) where igraph max-scales it: the ratio is constant across nodes to 1e-14, so it is the same vector under a documented normalisation. | — / — | test_network_parity.py |
cfm_decompose |
Stata cdeco, method(logit) 1.0.2 with drprocess (Chernozhukov, Fernandez-Val & Melly) | Stata 18.0 MP; cdeco 1.0.2 01mar2023 (bmelly/Stata counterfactual, Distribution-Date 20220803) | CDFs at the thresholds 1e-8 abs (observed 2.4e-10: Stata's logit stops at its default tolerance); quantiles 1e-12 abs (observed 0) | — / — | test_decomp_qte_parity.py (+1) |
cic |
qte::CiC | R 4.5.2 | rel_est<=1e-06 | 2.4e-15 / 6.0e-03 | 74_cic.py (+2) |
cloglog |
stats::glm(binomial('cloglog')) | R 4.5.2 | coefficients 5e-5 abs (observed ~1e-5; IRLS convergence) | — / — | test_glm_ext_parity.py (+1) |
community_detection |
R igraph::cluster_louvain (T3: randomised on both sides) | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | T3, not T2: 200 seeded runs on karate have mean modularity within four combined standard errors of igraph's 200 runs (0.4157 vs 0.4145), and both reach the same maximum, 0.41979, the known optimum for this graph. | — / — | test_network_parity.py |
conditional_lr_ci |
R ivmodel::CLR | ivmodel 1.9.1 | Monte Carlo by construction and graded as such: ivmodel integrates Moreira's conditional distribution while StatsPAI simulates it, so the two cannot agree deterministically. The test asserts the error SHRINKS with n_sim rather than pinning a number -- 3.8e-3 at n_sim=5,000 and 1.7e-4 at 200,000 -- which is the only honest statement about a simulated critical value. The endpoints themselves are bisected off-grid, so the residual is the critical value and not the grid. | — / — | test_weakiv_meta_parity.py |
cox_frailty |
R survival::coxph(... + frailty(id, theta=, sparse=FALSE)); Stata stcox, shared() | R 4.5.2; survival 3.8.3; Stata 18 MP | fixed theta: beta and SE 1e-9, integrated log likelihood 1e-10 (observed 4e-14); theta maximiser 1e-6 vs R optimize (observed 5e-8) and 5e-5 vs Stata e(theta) (observed 7.5e-6) | — / — | test_survival_epi_R_parity.py (+2) |
demeaned_synth |
augsynth::augsynth(progfunc = 'None', fixedeff = TRUE) 0.2.0 (de-meaned SCM) | R 4.5.2; augsynth 0.2.0; osqp 1.0.0 | gap path and ATT 1e-7 rel (observed 1.2e-9); weights 1e-8 abs | — / — | test_did_synth_synthvar_parity.py (+1) |
dml_model_averaging |
ddml::ddml_plm 0.3.1 (shortstack = TRUE, ensemble_type = 'nnls1'), OLS candidates, shared folds | R 4.5.2; ddml 0.3.1; sandwich 3.1.1 | short-stacking weights 1e-10 (observed 4.4e-14); ddml's final lm(y_r ~ d_r) with intercept and HC1 SE rebuilt from StatsPAI's stacked residuals at 1e-10 (observed 4.8e-16) | — / — | test_ml_causal_dml_parity.py (+2) |
functional_form_test |
didFF::didFF | R 4.5.2 | rel_est<=0.001 | 1.3e-14 / — | 79_didff.py (+1) |
garch |
Stata 18 arch; rugarch::ugarchfit 1.5.6 | R 4.5.2; rugarch 1.5.6; Stata 18 | log-likelihood at the reference optimum 1e-12 (observed 8.5e-16); vs Stata: b 1e-5, SE 2e-5 (observed 8.2e-6 / 1e-5); vs rugarch: params 5e-4, SE at their parameters 1e-2 | — / — | test_timeseries_R_parity.py (+2) |
genmatch |
Matching::Match 4.10-15 (Weight = 3, Weight.matrix) | — | Deterministic kernel only: given the same diagonal W, the 1-NN assignment agrees with Matching::Match on all 163 uniquely matched treated units on MatchIt::lalonde. | — / — | test_matching_r_parity.py (+1) |
grapple |
GRAPPLE::grappleRobustEst (GitHub jingshuw/GRAPPLE 317e837) | GRAPPLE 0.2.2 | sandwich at GRAPPLE's estimates 1e-6 (l2 1e-12); fitted beta 1e-3, tau2 / SE 1e-4 | — / — | test_misc_sens_R_parity.py |
hits |
R igraph::hits_scores | igraph 2.3.3; sna 2.8; ergm 4.12.0; dyadRobust 0.0.1.0001 | Hub and authority vectors are igraph's up to normalisation: L1 here (documented), max = 1 in igraph; the ratio is constant across nodes to 1e-11. | — / — | test_network_parity.py |
lcsf |
sfaR::sfalcmcross 1.0.1 (2 classes, half-normal) | R 4.5.2; sfaR 1.0.1 | estimates 1e-6 rel (observed 1.1e-8); OIM SEs 1e-6 rel (observed 8.6e-8) | — / — | test_frontier_struct_R_parity.py (+1) |
levinsohn_petrin |
Stata prodest, method(lp) valueadded (Rovigatti & Mollisi); R prodest::prodestLP | Stata 18 MP; prodest SSC; R prodest 1.0.2 | As olley_pakes with materials as the proxy: free-input coefficient at 1e-12 against both; state coefficient at the minimiser, R 1.3e-6 and Stata up to 1.8e-3 away where their optimisers stop. | — / — | test_prodest_parity.py |
lincom |
Stata 18 lincom (after regress / ivregress / logit / poisson) | Stata 18 | 1e-6 rel (observed <= 2.3e-15 on linear fits, <= 2.3e-8 overall) | — / — | test_postestimation_stata_parity.py (+1) |
malmquist |
sfaR::sfacross 1.0.1 per period + sfaR::efficiencies (teBC / teJLMS); TC from sfaR betas | R 4.5.2; sfaR 1.0.1; metafrontier 0.3.1 | EC / TC / M 1e-6 rel (observed 2.7e-7); metafrontier::malmquist_meta(method='sfa') EC_group and 1/TC_group 1e-4 (observed 1.8e-5, reference optim BFGS with finite-difference gradients) | — / — | test_r2_frontier_parity.py (+1) |
margins |
Stata 18 margins, dydx(*) | Stata 18 | 1e-6 rel (observed 2.7e-7 AME / 7.5e-8 SE on probit and logit-with-interaction; <= 1e-10 otherwise) | — / — | test_postestimation_stata_parity.py (+1) |
melogit |
lme4::glmer | R 4.5.2; lme4 2.0.1 | rel_est<=0.0002, rel_se<=0.01 | 1.7e-04 / 1.4e-08 | 26_glmm_logit.py (+2) |
model_averaging_dml |
ddml::ddml_plm 0.3.1 (shortstack = TRUE, ensemble_type = 'nnls1'), OLS candidates, shared folds | R 4.5.2; ddml 0.3.1; sandwich 3.1.1 | short-stacking weights 1e-10 (observed 4.4e-14); ddml's final lm(y_r ~ d_r) with intercept and HC1 SE rebuilt from StatsPAI's stacked residuals at 1e-10 (observed 4.8e-16) | — / — | test_ml_causal_dml_parity.py (+2) |
mr_raps |
R mr.raps 0.4.3 (simple / overdispersed / overdispersed.robust) | MendelianRandomization 0.10.0; TwoSampleMR 0.7.9; RadialMR 1.2.4; MRPRESSO 1.0; mr.raps 0.4.3 | Simple and L2-overdispersed fits (beta, SE, tau2) at 1e-8. Robust Huber / Tukey: the sandwich reproduces R's SEs at 1e-10 when evaluated at R's own (beta, tau2) and integrate() moments; the fitted values agree to 5e-5 (beta), 5e-4 (SE), 1.5e-3 (tau2) because R stops uniroot at its default tolerance and integrate() at 1.2e-4 -- a test shows StatsPAI's root satisfies the estimating equation more tightly than R's. | — / — | test_mr_R_parity.py |
olley_pakes |
Stata prodest, method(op) valueadded (Rovigatti & Mollisi); R prodest::prodestOP | Stata 18 MP; prodest SSC; R prodest 1.0.2 | Simulated unbalanced panel with 31 calendar gaps. Free-input coefficient (stage-1 OLS) at 1e-12 against both, polynomial degree 2 and 3. State coefficient: StatsPAI solves the first-order condition and has a sum of squares no larger than at either reference estimate; the references stop their optimisers early (R BFGS 5.9e-6 away, Stata Nelder-Mead at tolerance 1e-5 up to 4.3e-3 away). | — / — | test_prodest_parity.py |
optimal_match |
optmatch::pairmatch 0.10.8 on a logit propensity score | — | Total matched distance <= optmatch's (1 + 1e-6). The matched pairs are not pinned: the assignment problem is degenerate on this data, so equally optimal solutions report different ATTs. | — / — | test_matching_r_parity.py (+1) |
pretrends_power |
pretrends::pretrends / pretrends::slope_for_power (GitHub, not CRAN) | R 4.5.2 | rel_est<=0.001 | 4.0e-05 / 1.4e-04 | 76_pretrends.py (+2) |
pretrends_slope_for_power |
pretrends::pretrends / pretrends::slope_for_power (GitHub, not CRAN) | R 4.5.2 | rel_est<=0.001 | 4.0e-05 / 1.4e-04 | 76_pretrends.py (+2) |
qdid |
qte::QDiD 1.3.1 | — | max deviation / scale < 0.08, sign agreement on the large effects and correlation > 0.999. R's quantiles come from BMisc::weighted_quantile (stats::optimize on a piecewise-linear check function, which has plateaus where every point is a minimiser) while sp.qdid interpolates the empirical inverse CDF; on this fixture the gap is at most ~152 currency units against effects running to ~8900. | — / — | test_qdid_parity.py (+1) |
qte |
qte::ci.qte / qte::ci.qtet 1.3.1 (Firpo 2007) | — | max relative deviation < 0.01 on lalonde.exp / lalonde.psid. Both sides minimise the same weighted check function; R's BMisc::weighted_quantile uses a golden-section search whose answer on a plateau is an optimiser artifact, so point-value equality is not asserted. | — / — | test_firpo_qte_parity.py (+1) |
rate |
grf::rank_average_treatment_effect and rank_average_treatment_effect.fit 2.6.1 (AUTOC, QINI, TOC), forest outputs held fixed | R 4.5.2; grf 2.6.1 | Point estimate (AUTOC, QINI) and TOC curve on q = 0.1..1 exact at 1e-10 rel (observed 4.6e-15), including a 30-group tied-priority case against rank_average_treatment_effect.fit. SE is T3: the analytic rank-corrected influence-function SE is compared with grf's half-sample bootstrap (R = 2000) at 6% rel (observed 2.6% AUTOC, 1.7% QINI). | — / — | test_ml_causal_R_parity.py (+1) |
rd_bias_aware_fuzzy |
R RDHonest 1.0.1.9000, RDHonest(y | d ~ x) | R 4.5.2; RDHonest 1.0.1.9000 | estimate, std.error, maximum.bias, conf.low/high, M, first stage: rel 1e-9 at fixed h and M (observed 1.1e-14), 1e-6 with selected h (observed 3.5e-8, optimiser) | — / — |
rd_discrete |
R RDHonest::RDHonest / RDHonestBME 1.0.1.9000 (Kolesar) | R; RDHonest 1.0.1.9000 | estimate, std.error, maximum.bias, conf.low/high rel 1e-9 at fixed h and for BME; 1e-6 when RDHonest selects h and M (optimiser) | — / — | test_rd_open_R_parity.py (+1) |
rdrbounds |
R rdlocrand::rdrbounds 2.0 | R 4.5.2; rdlocrand 2.0 | upper and lower bounds within 4 pooled binomial SEs at 4000 draws (Monte Carlo on both sides, not parity) | — / — | test_rd_iv_rd_R_parity.py (+1) |
rdsensitivity |
R rdlocrand::rdsensitivity / rdrandinf 2.0 | R 4.5.2; rdlocrand 2.0 | per-window estimates rel 1e-9 (rdrandinf observed statistic); randomization p-values within 4 pooled binomial SEs at 4000 draws (Monte Carlo on both sides, not parity) | — / — | test_rd_iv_rd_R_parity.py (+1) |
rlassologit |
R hdm::rlassologit 0.3.2 (glmnet 4.1.10 engine) | R 4.5.2; hdm 0.3.2; glmnet 4.1.10 | support exact; post-Lasso atol 1e-5 (observed rel 1.7e-10); post=False atol 1e-4 (observed 1.3e-6: glmnet coordinate-descent convergence) | — / — | test_rlassologit_parity.py (+1) |
sarar_gmm |
spatialreg::gstsls 1.4.3 (Kelejian-Prucha GS2SLS) | R 4.5.2; spatialreg 1.4.3; spdep 1.4.2 | beta, rho and SEs 1e-9 rel (observed 3.8e-11 / 3.0e-11); lambda 1e-7 rel (observed 6.5e-9) | — / — | test_spatial_survey_R_parity.py (+1) |
scest |
R scpi::scest (w.constr simplex / lasso / ridge / ols / L1-L2, V = 'separate') | scpi 4.0.1; CVXR 1.9.2; clarabel 0.11.2; osqp 1.0.0 | ols and lasso weights at 1e-9 abs; ridge Q / lambda and L1-L2 Q2 at 1e-10. simplex / ridge / L1-L2 weights are bounded by the conic solver: R's objective exceeds StatsPAI's exact optimum by <= 1e-7 relative (CLARABEL gap 1e-8), R's point is feasible, and | w_R - w_py | |
scpi |
R scpi::scpi (effect = 'unit-time', u.missp, u.sigma = HC1, u.order = e.order = 1, rho = type-2, e.method = all) | scpi 4.0.1; CVXR 1.9.2; ECOSolveR 0.6.1; Qtools 1.6.0; quantreg 6.1 | On R's weights: rho, Q.star, u.mean, Omega, Sigma, e.mean at 1e-9; out-of-sample e.var and gaussian / ls / qreg bounds at 1e-9 against R with rrq(method = 'br') (exact LP) and 5e-4 abs against the default Frisch-Newton rrq. In-sample simulation fed R's draws: per-draw median <= 1e-6 and max <= 2e-4 vs ECOS at 1e-12, quantile bounds at 1e-5; vs default ECOS (1e-8) bounds within 2e-3 abs. Average-effect CI = scdataMulti(effect = 'unit') within 5e-4. J > T0 (California, 38 donors) covered. | — / — | test_did_synth_scpi_parity.py |
shapley_inequality |
Stata shapley2 1.5 (Chavez Juarez) over regress + ineqdeco (Jenkins) | Stata 18.0 MP; shapley2 1.5 10jun15 (SSC) | value function v(S) 1e-11 rel (observed 1.4e-13); Shapley values 1e-6 rel (observed 2.7e-7, shapley2 routes v(S) through float variables); totals 1e-12 | — / — | test_decomp_qte_parity.py (+1) |
spatial_panel |
splm::spml(model = 'within') 1.6.5; Stata xsmle 1.4.5 fe type(ind) vce(oim) | R 4.5.2; splm 1.6.5; plm 2.6.7; Stata 18; xsmle version 1.4.5 5jun2017 | vs splm: estimates and SEs 1e-7 rel (observed 7.2e-8 / 1.3e-8, the splm optimize floor); vs xsmle (tightened ml tolerances): estimates and beta SEs 1e-9 (observed 2.5e-13 / 8.5e-11), spatial-parameter SE 1e-8 (observed 3.0e-9) | — / — | test_spatial_survey_R_parity.py (+2) |
survreg |
survival::survreg (Weibull AFT) | R 4.5.2; survival 3.8.3 | coefficients & log-scale 5e-5 abs (observed ~1e-5) | — / — | test_aft_parity.py (+1) |
test |
Stata 18 test (after regress / ivregress / logit / poisson) | Stata 18 | 1e-6 rel (observed <= 2.3e-15 on linear fits, <= 8e-10 on ML fits, far-tail p <= 4.7e-8) | — / — | test_postestimation_stata_parity.py (+1) |
weakrobust |
R ivmodel::CLR 1.9.1; Stata weakiv 2.4.07 (md small) | R 4.5.2; ivmodel 1.9.1; Stata 18 MP; weakiv 2.4.07 | CLR statistic and p-value vs ivmodel rel 1e-9 (observed 1.2e-11); CLR set vs ivmodel 5e-5 (observed 1.2e-5, ivmodel's uniroot default tolerance); CLR / K / AR vs Stata weakiv 1e-6 (observed 1.1e-7 / 1.1e-7 / 2.1e-8, not bisected further) | — / — | test_rd_iv_R_parity.py (+2) |
xtfrontier |
frontier::sfa | R 4.5.2; frontier 1.1.8 | rel_est<=0.001, rel_se<=0.05 | 2.8e-06 / 1.8e-06 | 29_panel_sfa.py (+2) |
zinb |
pscl::zeroinfl(dist="negbin") | R 4.5.2; pscl 1.5.9 | rel_est<=1e-05, rel_se<=0.001 | 9.5e-07 / 4.5e-11 | 64_zinb.py (+2) |
zisf |
Stata chks 1.1 (estimation(zsf) eoption(ml)); R sfa::zsfm 1.2.0 (ZISF / ZISF_Z, likelihood at its optimum) | R 4.5.2; sfa 1.2.0; numDeriv 2016.8.1.1; stata 18; chks 1.1 (chks.pkg dated 20190320) | estimates and OIM SEs 1e-6 rel (observed chks 8.6e-8 / 9.4e-8; sfa likelihood at its optimum 1.3e-8 / 5.4e-8); sfa's reported L-BFGS-B point 5e-5 / 5e-4 (observed 1.6e-5 / 1.8e-4) | — / — | test_r2_frontier_parity.py (+2) |
external-replication — 2 functions¶
Reproduces published-paper numbers; sources in tests/external_parity/PUBLISHED_REFERENCE_VALUES.md.
| function | test |
|---|---|
aggte |
test_honest_did_paper_parity.py (+1) |
parallel_trends_robustness |
test_rebel_canal_published.py |
analytical-only — 130 functions¶
Recovers a known DGP truth / closed-form identity within tolerance; no cross-package reference. See tests/reference_parity/REFERENCES.md.
unverified — 651 functions¶
These are registered public functions with no cross-language or published-reference parity evidence attached yet. This is the honest coverage gap, not a claim of incorrectness — many are frontier methods with no Stata/R sibling to align against. Query any of them with sp.parity_status(name); the closing roadmap lives in docs/dev/parity_status_roadmap.md.