Mathematical formulation
Nonparametric unit bootstrap, three causal estimators, confidence intervals, and bootstrap-assisted hypothesis testing.
Observed unit
Resample the complete subject record. Never resample Y, T, and X independently.
Multiplicity representation
Duplicates and omissions are expected. Approximately 63.2% of subjects are distinct in one large bootstrap sample.
Refit the full procedure
The algorithm 𝒜 includes refitting propensity and outcome models—not merely reusing original fitted predictions.
IPW, Horvitz–Thompson
Each bootstrap sample refits ê(X). Fitted probabilities are bounded numerically to [0.005, 0.995].
Outcome regression
Standardize both predicted treatment worlds over the same empirical X distribution.
Doubly robust / AIPW
Consistency requires at least one nuisance model to be correctly specified under standard regularity and positivity conditions.
Percentile interval
The endpoints are empirical quantiles of the successful bootstrap estimates. The percentile interval respects strictly monotone transformations, but it is generally only first-order accurate and can miscover when estimator bias or skewness is substantial.
Normal / Wald interval
This interval is symmetric around the original estimate. With the same standard error and critical value, it is algebraically dual to the corresponding two-sided Wald test.
Bootstrap-SE Wald test
Greater: p = 1−Φ(z)
Less: p = Φ(z)
The test uses normal calibration with the bootstrap-estimated standard error.
Centered, unstudentized bootstrap test
The displayed formula is for the two-sided alternative; one-sided versions use the corresponding signed inequality. This unstudentized test is approximate and is not generally dual to the percentile interval.
CI/test duality
This equivalence requires a matching dual pair. The two-sided Wald test matches the two-sided normal/Wald interval; directional Wald tests match the corresponding one-sided lower or upper confidence bound.
Independent resampling units
The unit bootstrap assumes subjects are i.i.d. If data are clustered, longitudinal, paired, or survey sampled, resample the appropriate cluster or use a design-respecting method.
Regular estimator
The ordinary bootstrap can fail for nonregular estimators, boundary parameters, extreme model selection, or some highly adaptive machine-learning procedures without additional theory.
Causal identification precedes inference
Bootstrap uncertainty does not repair unmeasured confounding, poor overlap, measurement error, nuisance-model bias, or a mismatch between the estimand and analysis.
Failed fits: summaries use only the Bs successful bootstrap analyses. Frequent failures are themselves a warning about instability or nonregularity; simply discarding many failed fits can distort the empirical bootstrap distribution.