Explore the augmented inverse probability weighted estimator for a marginal causal risk difference. Fit a treatment model and an outcome model, inspect each subject’s residual correction, and verify the “one correct nuisance model is enough” property through repeated simulations.
Core operation: outcome-regression predictions provide a starting value; inverse-probability weighted residuals correct those predictions using the observed outcomes.
AIPW • GRADUATE LEVEL • INLINE SVG
Mathematical formulation
One notation system is used throughout: φ is the uncentered AIPW estimating signal, φc is that signal centered by its empirical mean, D* is the true nonparametric efficient influence function, and IFstack is the influence value of the finite-dimensional stacked M-estimator. These objects are related, but they are not interchangeable.
Graduate-level notation
O=(X,T,Y)∼P₀
Observed data. T∈{0,1} is the observed treatment; a∈{0,1} indexes a hypothetical intervention level.
e₀(x)=P₀(T=1|X=x)
True propensity score for treatment 1.
πa,0(x)=P₀(T=a|X=x)
Treatment-level probability: πa,0(x)=e₀(x)a{1−e₀(x)}1−a. Hence π1,0=e₀ and π0,0=1−e₀.
ma,0(x)=E₀(Y|T=a,X=x)
True conditional outcome mean under observed treatment level a.
η₀=(e₀,m0,0,m1,0)
Collection of nuisance functions under the true observed-data law.
Pₙf=n⁻¹Σᵢf(Oᵢ)
Empirical average. The evaluation nuisance fit is η̂ for same-sample fitting and η̂(−k(i)) for subject i under cross-fitting.
Causal target under consistency, exchangeability, and positivity
AIPW estimating signal and estimator
Reading rule: the entire indicator–residual product is in the numerator, and πa(X) is the denominator. Thus π1(X)=e(X) and π0(X)=1−e(X); the signal requires πa(X)>0 almost surely.
Treated-world signal
The actual weighting score is ẽᵢ=Π[ε,1−ε](êᵢ). With ε=0, ẽᵢ=êᵢ apart from a machine-precision floor.
Control-world signal
The same empirical distribution of X is used for both treatment worlds.
ATE signal and estimator
φ̂Δ,i is an uncentered AIPW signal, sometimes called an AIPW pseudo-outcome. It is not itself the EIF.
Notation rule used everywhere below: φ denotes the uncentered AIPW signal; φc=φ−Pₙφ denotes its empirically centered version; D* is reserved for the true nonparametric efficient influence function; IFstack denotes the influence value from the finite-dimensional stacked estimating equations.
Nuisance evaluation rule: for same-sample fitting, φ̂a,i=φa(Oi;η̂). Under K-fold cross-fitting, if i∈Ik, then φ̂a,i=φa(Oi;η̂(−k)). Thus the observation-specific evaluation fit is not treated as a new population nuisance function.
Exact population bias identity for fixed η=(e,m₀,m₁) with 0<e(X)<1 and integrable terms
This formula uses the same φΔ defined in the estimator tab. It gives the bias of the uncentered AIPW signal relative to the true ATE. The displayed sign convention is therefore fixed and unambiguous.
Treatment-level conditional bias identity
This identity is the direct algebraic check: the conditional bias is zero if either πa=πa,0 or ma=ma,0.
Propensity nuisance correct
The outcome regressions may be misspecified, provided the untruncated propensity score is correct and positivity holds.
Outcome nuisances correct
Then each weighted residual has conditional mean zero for any positive weighting score.
Both nuisance components wrong
Double robustness is a union-model property, not protection against simultaneous misspecification.
Product-error bound
If ε≤e(X)≤1−ε, Kε may be taken proportional to 1/ε. The errors enter as a product. Here ‖h‖P₀,2={E₀[h(O)²]}1/2.
Truncation distinction: the theorem above applies to the actual function e used inside φΔ. If a correct fitted ê is replaced by a nonvanishing clipped ẽ≠e₀ on a set of positive probability, the propensity-only branch is no longer exact. The outcome-correct branch remains valid.
Simulation labels are not diagnostics: “correct” and “misspecified” are known only because the data-generating law is programmed. They cannot be verified from one observational dataset.
1. Uncentered AIPW signal
Its mean estimates the ATE. It is not mean zero and therefore is not an influence function.
2. Nonparametric EIF
This is mean zero under P₀ and attains the nonparametric efficiency bound.
3. Stacked-estimator IF
This is the dashboard’s finite-dimensional M-estimation influence value. It equals the EIF asymptotically only under the corresponding efficiency conditions.
Efficient influence function: compact and expanded forms
Critical reading check: the treated residual is multiplied by T/e₀(X), not by T·e₀(X). Likewise, the control residual is multiplied by (1−T)/{1−e₀(X)}. These are the two inverse-probability residual corrections in the nonparametric EIF.
Empirically centered AIPW signal
φ̂cΔ,i is exactly centered in the sample. It is an empirical surrogate for DΔ* only when the nuisance estimates and the target estimate converge to their true counterparts.
EIF-style variance from the centered signal
No additional centering term is needed because Pₙφ̂cΔ=0 exactly. This estimates the EIF variance only when φ̂cΔ,i consistently approximates DΔ*.
Stacked sandwich actually used
Here θ stacks the nuisance-model coefficients with ψ₁ and ψ₀, and c extracts ψ₁−ψ₀.
Crucial distinction: under one-model misspecification, the centered signal φ̂cΔ,i=φ̂Δ,i−Δ̂ is generally not the full influence function of the stacked working-model estimator. The dashboard therefore calculates and displays IF̂stackΔ,i separately.
Efficiency statement: when both nuisance functions are consistently estimated at suitable rates, no substantive truncation is active, and regularity conditions hold, IF̂stack and the centered signal φ̂cΔ both approach the nonparametric EIF. The sign convention is treatment minus control throughout: +T/e₀(X) for the treated residual and −(1−T)/{1−e₀(X)} for the control residual.
Shared-fold nuisance estimates
The propensity and both outcome-regression predictions use the same fold partition.
Cross-fitted AIPW estimator
Each observation is evaluated by nuisance fits that excluded its fold.
Product-rate condition for convergence to the nonparametric EIF
Here ẽ is the treatment probability actually used in the AIPW denominator (equal to ê when no substantive clipping is active). Together with L₂ consistency of ẽ and both outcome regressions, this product-rate condition yields convergence to the efficient influence function. Product smallness alone is not enough if one nuisance converges to a wrong limit.
Identification assumptions
Consistency: Y=Y(T). Conditional exchangeability: {Y(1),Y(0)} ⟂ T | X. Positivity: 0<e₀(X)<1 almost surely. Regularity: finite variance and stable nuisance estimates.
Point consistency versus efficient inference: one correct nuisance can make R₂ vanish, but convergence to the nonparametric EIF generally requires both nuisance functions to converge. In the finite-dimensional union model used here, the block-stacked sandwich accounts for first-order nuisance-estimation terms when one working model is misspecified.
Cross-fitting does not repair identification. It reduces own-observation overfitting and empirical-process restrictions, but it cannot fix unmeasured confounding or structural positivity failure.
1. Subjects in the observational cohort
Color is observed smoking status, ring is observed CVD, and point size can display inverse weight, augmentation, or the stacked-estimator influence value.
SmokerNon-smokerCVDlimited support
Click a subject to inspect its outcome-model contrast, augmentation, uncentered AIPW signal, centered AIPW signal, and stacked-estimator influence value.
Selected subject
The uncentered AIPW signal is assembled one person at a time. Centering creates φ̂ᶜΔ,i; stacked linearization creates IF̂stackΔ,i. Only the latter is automatically the working-model influence value under one-model misspecification.
2. Propensity-score overlap
Estimated P(smoker | X) by observed treatment group.
SmokersNon-smokerstruncation bounds
The doubly robust property does not eliminate the need for positivity. Large inverse weights can still dominate the correction term.
3. Outcome-model calibration
Observed CVD frequency versus fitted observed-treatment risk.
Prediction calibration is informative but does not establish causal exchangeability or correct counterfactual extrapolation.
4. Prediction plus augmentation correction
The final AIPW estimate is the outcome-regression contrast plus a weighted residual correction.
5. Distribution of subject-level AIPW signals
φ̂Δ,i can be negative or exceed one even though its average estimates a risk difference.
Observed smokersObserved non-smokerssample mean
Extreme AIPW signals often reflect poor overlap, large outcome residuals, or both. The signal itself is not the EIF because it is not centered at zero.
6. The double-robustness matrix
Click any cell to refit the dashboard. Green cells satisfy the theoretical union-model condition under the displayed, untruncated propensity score; amber indicates that nonvanishing truncation compromises propensity-only protection.
One cohort can still be noisy. The repeated-sampling experiment below distinguishes sampling variation from systematic misspecification bias.
7. Estimator comparison
Crude, IPW, outcome regression, and AIPW estimates under the currently selected nuisance models.
8. Repeated-sampling double-robustness experiment
Generate independent cohorts, refit all four nuisance-model combinations, and compare bias and 95% CI coverage.
No experiment run yet0%
Run the experiment to see the double-robustness pattern emerge across repeated studies.
Intervals are centered by the fixed reference-population approximation to the superpopulation ATE for the selected data-generating law. Green intervals contain zero after centering; red intervals miss it. Every repetition refits all nuisance models and recomputes the appropriate stacked sandwich.
Inspect subject-level nuisance predictions and AIPW contributions