Summary
Pharmacoepidemiology studies often have missing data that are nonmonotone (i.e., data missing in an arbitrary pattern). It is well known that complete case analysis (CC) reduces precision and may result in bias. Previously, inverse probability weighting (IPW) for nonmonotone missingness could not be easily implemented in standard software, but a new approach for estimating the missingness weights (unconstrained maximum likelihood estimator, UMLE) overcomes this difficulty.
Methods
We conducted a simulation study. We simulated 2000 datasets with a binary treatment, a binary outcome, and three confounders. The outcome risk was 0.1 in the untreated. We varied sample size (1500, 5000), percent treated (15%, 50%), and the true marginal risk difference (0, 0.05). We simulated missing data (missing at random) yielding 50% complete cases. We implemented four analyzes in each dataset: analysis of the full data (without missingness), CC, IPW by UMLE, and MI by chained equations. In each analysis, we used inverse probability of treatment weights to address confounding and estimated the risk difference by weighted linear-binomial model with robust standard errors. In IPW by UMLE, the propensity score model (used to estimate the treatment weights) was weighted by the missingness weights; the final weight was the product of the two weights. In MI, results from 20 imputed datasets were combined by Rubin's rule.
Results
When the true risk difference was 0.05, CC was biased in each scenario (bias > 0.02). MI had negligible bias (<0.002). IPW also had negligible bias; however, bias increased slightly at the smaller N = 1500 with 15% exposed (0.005). At N = 5000, 50% exposed, analysis of the full data yielded an empirical standard error (SE) of 0.010. IPW and MI were equally less precise (SE 0.015). As sample size and percent exposed declined, SE for analysis of the full data increased to 0.027 and MI became slightly more precise than IPW (MI SE 0.046 vs. IPW SE 0.048). MI implementation was >13 times slower than IPW in SAS. Results were similar across estimators when the true risk difference was zero.
Conclusions
IPW UMLE performed well and was easy to implement in SAS and R. Though less statistically efficient at smaller sample sizes, IPW is a viable alternative to MI for nonmonotone missingness. IPW may be more intuitive for researchers already familiar with weighting and faster computational efficiency of IPW is attractive in large datasets or when bootstrap is required.