
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
김은진 전북대학교 일반대학원 2022 국내석사
관찰연구에서 처리효과 추정치는 공변량 불균형의 영향으로 인한 편향이 발생한다. 역확률 가중치, 중복 가중치 및 매칭 가중치와 같은 성향점수의 가중치 방법은 추정치의 편향을 줄일 수 있다. 다중 처리군의 경우에는 성향점수를 확장한 일반화 성향점수 방법이 제안되었다. 주로 일반화 성향점수 추정시 다항 로지스틱 회귀 모형이 사용되며, 가중치 방법에도 적용되었다. 그러나 표본 크기에 비해 공변량이 너무 많거나 정규성을 가정할 수 없는 경우, 모수적 방법인 다항 로지스틱 회귀의 치료효과 추정치는 편향이 발생 할 수 있다. 반면 일반화 부스팅 모형은 비모수적 방법과 같은 문제에 적합하다. 또한 대용량 데이터에서도 잘 구현되어 최근에는 일반화 성향점수 추정에도 사용되고 있다. 본 연구는 모의실험을 통해 다항 로지스틱 회귀 모형과 일반화 부스팅 모형에 의해 추정된 일반화 성향점수와 이를 기반으로 하는 역확률 가중치, 중복 가중치와 매칭 가중치의 소표본 속성을 비교하였다. 모의실험 결과, 처리 또는 결과 모형이 선형인 경우에만 다항 로지스틱 회귀 방법은 편향과 MSE가 작았으며 CP는 높게 나타났다. 그러나 일반화 부스팅 방법은 모형의 선형 여부에 관계없이 편향, MSE와 CP는 일관된 결과를 보였으며, 처리 또는 결과 모형이 비선형인 경우에는 다항 로지스틱 회귀 방법보다 우수한 성능을 보였다. 가중치 방법은 전반적으로 매칭 가중치, 중복 가중치, 역확률 가중치 순으로 우수한 성능 보였다. 매칭 가중치와 중복 가중치는 일반화 성향점수 추정 방법에 따른 결과와 비슷하게 나타났으며, 역확률 가중치는 일반화 부스팅 모형으로 추정할 경우 성능이 더 좋게 나타났다. In observational studies, the treatment effect estimates are often biased by the impact of covariate imbalances. The propensity score weighting methods such as inverse probability of treatment weights(IPW), overlap weights(OW) and matching weights(MW) can be used to reduce the bias. In the case of multiple treatments, the propensity scores are extended to generalized propensity scores. They have been estimated most often by multinomial logistic regression(MLR) and used for weighting methods. However, MLR, which is a parametric methods, may cause bias in the estimate of treatment effect when there are too many covariates for the sample size or when normality cannot be assumed. The generalized boosted models(GBM) is robust to such problems as a nonparametric methods and also implemented well in very large size of data so that it is recently adopted for the estimation of GPS. In this study, we compared small sample properties of generalized propensity scores estimated by MLR and GBM and three weighting methods(IPW, OW, MW) based on these scores through simulations. The result of simulation shows; The MLR had small bias and MSE, and CP was close to the specified value only when log-odds of the treatment or outcome was linear in covariates. The GBM, however, showed consistent results regardless of the models specification in bias, MSE and CP, and performs better than the MLR as both treatment and outcome models are non-linear. The weighting methods showed good performance in the order of MW, OW and IPW. MW and OW showed similar results according to the generalization propensity score estimation methods, and IPW performed better when estimated using the GBM.
Comparison of propensity score methods flexible to model misspecification
강문종 Graduate School, Yonsei University 2020 국내석사
In an observational study, there may be differences in the distribution of pre exposure characteristics between treated and non-treated groups. This difference also affects the outcome. To address this, propensity score methods are frequently used. These methods remove the imbalance by making subgroups that have similar pre treatment characteristic distribution or creating synthetic samples by weighting. Propensity score methods are needed to specify propensity models, and misspecification of the propensity model brings notable bias to the outcome analysis. Researchers have proposed many methods which are less sensitive to the model specification to reduce the bias when the model is built to incorrect specifications. The purpose of this paper is to compare the proposed propensity estimation methods which are flexible to wrong model specification with weighting and matching methods in various scenarios. Biases, absolute bias, and MSE are used to evaluate the performance of estimating the average treatment effect (ATE), and average standardized mean difference (ASMD) is applied to check whether the covariates are balanced. From the simulated example, the scenario, the correct model specification for the generalized linear model (GLM) shows the best result for the bias, but it becomes less effective as the model complexity increases. The generalized boosted model (GBM), however, gives consistent results regardless of the model specification. Among the analysis methods incorporating propensity score, matching weight (MW) is the most effective method to reduce the bias. In terms of covariate balancing, ATE shows similar result to the bias, GLM is the worst, and GBM is the best. Matching has the smallest average standardized mean difference (ASMD). In the propensity estimation procedure, the interpretational aspect of the model is not emphasized. So GBM could be a powerful alternative to estimate propensity score. Furthermore, although matching analysis reduces ASMD better than the other methods, it also reduces the sample size. In case that sample size is small, using matching weights as an analysis method is preferable.