Inverse Probability Weighted Estimator — Se Yoon Lee, Ph.D.

Explore how inverse probability weighting turns a confounded observational smoking study into a weighted pseudo-population. Move the controls, inspect individual subjects, and watch covariate distributions, balance, weights, and the estimated risk difference update together.

GRADUATE-LEVEL • SELF-CONTAINED • INLINE SVG

Mathematical formulation

The implementation uses the Horvitz–Thompson empirical mean throughout. Both potential-outcome means retain the original denominator n; there is no arm-specific self-normalization. Because this is the only IPW estimator used on the page, the superscript “HT” is omitted from the mathematical notation.

Graduate-level notation
Horvitz–Thompson estimation map
e^i=P^(Ti=1|Xi) μ^1,μ^0 Δ^ATE=μ^1μ^0
1Potential-outcome means
μ^1=1ni=1nTiYie^i, μ^0=1ni=1n(1Ti)Yi1e^i

The denominator remains n. Therefore, finite-sample treated and control weight totals are not forced to equal n.

2ATE risk difference
Δ^ATE=1ni=1n[TiYie^i(1Ti)Yi1e^i]

This is the default estimator in every estimate, confidence-interval, and repeated-sampling panel at 100% weighting.

3Crude-to-HT transition
w1i(α)=1απ^+αe^i w0i(α)=1α1π^+α1e^i

At α=0 the HT empirical means equal the crude arm means; at α=1 they equal the formal HT-IPW means.

1. Individual subjects in the observational cohort

Point size represents the current subject weight. Hover or click a subject for the exact calculation.

Smoker (T=1)Non-smoker (T=0)Had CVD (Y=1)larger point = larger current weight
Drag the Weighting transition slider: rare observed treatment assignments expand because those subjects represent more people in the weighted pseudo-population.

Selected subject

IPW is calculated subject by subject, not by naming a profile.

2. Covariate distributions

Compare the observed arms with the weighted pseudo-population.

SmokersNon-smokers
The density curves are normalized only for visual comparison. The causal estimator itself remains Horvitz–Thompson with denominator n.

3. Covariate balance

Absolute standardized mean differences (SMDs), before versus current weighting.

ObservedCurrent weightingvertical line = 0.10 guideline
A smaller SMD is better. IPW cannot repair omitted confounders, severe positivity violations, or an incorrect propensity model.

4A. Population balancing identity for a selected function

Choose a covariate and f(X), then compare the target empirical moment with its treated- and control-arm HT reconstructions.

E{Tf(X)e(X)}=E{f(X)}=E{(1T)f(X)1e(X)}
Dt(f)={M̂t(f)−M̂(f)}/SD{f(X)}. For f(X)=1, D is the HT mass error. The 0.05 and 0.10 guides are descriptive, not formal test cutoffs.

4B. Balance profile across several functions

Inspect mass, mean, higher moments, centered variance, skewness, and upper-tail balance simultaneously.

|Treated discrepancy||Control discrepancy|guides = 0.05 and 0.10
First-moment balance alone does not guarantee balance of variance, skewness, or tail probabilities. Compare fitted, uncapped, true, and current-transition weights.

5. Propensity overlap and inverse weights

The fitted propensity score is P(smoker | measured baseline covariates).

Smoker propensity distributionNon-smoker propensity distribution
A smoker with a very small fitted propensity receives 1/ê; a non-smoker with a very large fitted propensity receives 1/(1−ê). Both situations create large weights.

6. Crude versus Horvitz–Thompson causal estimate

HT treatment-specific risks use the original cohort size n as the denominator.

7. The current cohort: 95% confidence intervals

Point estimates, 95% Wald intervals, and an explicit check of whether each interval contains the known causal truth.

Point estimateKnown true ATEgray dashed line = no effect
The instant intervals use a stacked M-estimation sandwich for the fitted logistic propensity model and the two Horvitz–Thompson means. Capped weights are linearized piecewise. A 95% CI is not guaranteed to contain the truth in one dataset; 95% refers to long-run repeated-sampling coverage.

8. Repeated-sampling coverage experiment

Regenerate smoking and CVD outcomes for the same baseline covariates, refit the propensity model, and examine many 95% CIs.

No experiment run yet0%
Run the experiment to see which individual intervals contain the truth and what fraction cover it overall.
Green intervals contain the known conditional ATE; red intervals miss it. The experiment keeps the observed baseline covariates fixed, then repeatedly generates treatment and outcome, refits the propensity score, and recomputes sandwich CIs. This makes the true conditional ATE common to every repetition.
Inspect the subjects with the largest weights
SubjectAgeBMIIncome ratioFemaleCollegeHypertensionSmokerCVDRaw IPWCurrent w