RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 학위유형
      • 주제분류
        펼치기
      • 수여기관
        펼치기
      • 발행연도
        펼치기
      • 작성언어
      • 지도교수
        펼치기

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • XGBoost와 Logistic Regression 기반 선택적 하이브리드 모델을 활용한 대출 부도 예측 연구

      손승연 연세대학교 정보대학원 2025 국내석사

      RANK : 2943

      Logistic Regression 모델은 전통적으로 신용 평가 분야에서 널리 활용되어 왔으며, 모델 해석이 용이하다는 장점이 있다. 그러나 복잡한 데이터 구조를 처리하거나 고위험 부도 고객을 예측하는 것은 한계를 가진다. 반면, AI 기반 모델은 비선형 관계를 효과적으로 처리하며 우수한 예측 성능을 제공한다. 그러나 복잡한 알고리즘으로 인해 해석이 어려워 '블랙박스(black-box)'로 간주되는 문제가 존재한다. 이러한 한계는 모델의 투명성과 신뢰성을 저하시킬 수 있으며, 기존 연구들은 주로 XAI(Explainable AI) 기법을 활용하여 '블랙박스' 문제를 완화하고, 모델의 해석 가능성을 제고하는 데 초점을 맞춰왔다. 본 연구는 대출 부도 예측의 정확도를 향상시키고 해석 가능성을 강화하기 위해 Logistic Regression의 해석 용이성과 AI 기반 모델의 예측 성능을 결합한 선택적 하이브리드 모델을 제안한다. 제안된 모델은 두 가지 주요 단계로 구성된다. 첫째, Logistic Regression과 AI 기반 모델의 예측이 불일치하는 영역에서 Logistic Re-gression의 예측 결과를 AI 기반 모델의 예측으로 대체하는 선택적 접근 방식을 도입하였다. 이는 기존 앙상블 모델과 달리, 두 모델 간 예측 차이를 활용하여 한 모델의 결과를 다른 모델로 보완하는 방식이다. 둘째, AI 기반 모델이 Logistic Re-gression보다 높은 부도 위험도를 예측한 고위험 고객 데이터를 별도로 추출한 후, 이를 기반으로 고위험 고객에 특화된 Logistic Regression 모델을 추가로 학습시키는 방법론을 제안하였다. 실험 결과, 선택적 하이브리드 모델은 고위험 부도 고객 세그먼트에서 단일 모델보다 우수한 성능을 나타냈다. 또한, 고위험 고객 데이터를 기반으로 학습된 Logistic Regression 모델은 Logistic Regression 및 AI 기반 단일 모델에 비해 향상된 예측 성능을 보였다. 본 연구는 해석 가능한 Logistic Regression 모델과 고성능 AI 기반 모델을 조건부로 결합함으로써 기존 단일 모델 접근법의 한계를 보완하고, 대출 부도 예측의 정확도와 해석 가능성을 동시에 제고하였다는 점에서 학술적 의의가 있다. 나아가 본 연구가 대출 부도 관리 및 리스크 평가 분야에서 실무적 통찰을 제공할 수 있을 것으로 기대한다. Logistic Regression has traditionally been widely utilized in the credit evalua-tion domain due to its interpretability. However, it faces limitations in handling complex data structures and predicting high-risk default customers. In contrast, AI-based models effectively capture non-linear relationships and exhibit supe-rior predictive performance. Nevertheless, these models are often considered "black-boxes" due to their complex algorithms, which hinder interpretability. This lack of transparency can undermine the trust and reliability of the models. Existing research has primarily focused on leveraging Explainable AI (XAI) techniques to address the "black-box" problem and enhance model interpreta-bility. This study proposes a selective hybrid model that combines the interpretabil-ity of Logistic Regression with the predictive performance of AI-based models to improve loan default prediction accuracy and interpretability. The proposed model comprises two key components. First, a selective approach replaces Lo-gistic Regression predictions with AI-based predictions in regions where the two models produce inconsistent outcomes. Unlike traditional ensemble models, this approach utilizes the differences in predictions to complement one model's results with those of the other. Second, high-risk customer data, where the AI-based model predicts a higher default risk than Logistic Regression, are sepa-rately extracted, and a specialized Logistic Regression model is trained using this data. The experimental results demonstrate that the selective hybrid model outper-forms single models in terms of AUC (Area Under the Curve) performance for high-risk default customer segments. Additionally, the specialized Logistic Re-gression model trained on high-risk customer data achieves superior predictive performance compared to the conventional Logistic Regression and AI-based models. This study contributes to the literature by addressing the limitations of sin-gle-model approaches through the conditional integration of interpretable Lo-gistic Regression and high-performance AI-based models. It simultaneously enhances the accuracy and interpretability of loan default prediction. Further-more, it is expected that this research will provide practical insights for loan default management and risk assessment.

    • Firth’s logistic regression for separation problem : Comparing between R-logistf and SAS-LOGISTIC

      Sujeong Kim 연세대학교 일반대학원 2020 국내석사

      RANK : 2911

      The phenomenon of separation or monotone likelihood commonly occurs in fitting logistic regression model. Separation problems usually occur for small sample sizes with a highly imbalanced independent variable, or for large sample sizes with rare events. In this case the logistic regression model causes serious bias problem in the maximum likelihood (Firth D 1993, King, G. and Zeng, L 2001). To solve this bias, Firth proposed a penalized maximum likelihood estimation method. It produces finite parameter estimates with a simple modification of the score function. We can use the Firth’s method through several software packages such as R and SAS. But the motivating example showed differences in the output. For example, R-logistf and SAS-LOGISTIC gave different conclusions about the significance of variables even though they referred to the same Firth’s method paper (Heinze G, Schemper M 2002). Therefore the purpose of this study is to investigate the performance of Firth’s logistic regression in R (version 3.6.0) and SAS (version 9.4). We conducted a simulation study under various scenarios to compare standardized bias of estimates, mean standard error, coverage rate and convergence issue. According to the results of simulation study, in the case of separation problem caused by continuous variables, SAS has less bias of estimates and convergence problem than R. However in the case of separation problem caused by binomial variables, R and SAS has similar performance. Therefore we recommend using SAS software when analyzing Firth’s logistic regression.

    • Logistic bagging using feature incrementation

      정진성 연세대학교 대학원 2008 국내석사

      RANK : 2911

      One of the most important issues in classification is to improve the accuracy of the prediction model. Therefore, many researches have focused on this issue recently. Ensemble method that usually referred as 'perturb and combine' strategy has gained great popularity. In this thesis, a new ensemble method called 'Logistic Bagging' is proposed, then compared with other methods such as Bagging, Double Bagging, and Logistic regression using real data set and simulations. It was found that the Logistic Bagging method outperformed Bagging and Logistic regression. It also showed better accuracy than Double Bagging for non-normal data. However, for normal data, Logistic Bagging and Double Bagging method didn't show significant difference. Finally, the bias-variance decomposition revealed that the new method reduced the variance more compared to ensemble methods such as Bagging and Double Bagging. On the other hand, it had similar variance but smaller bias when compared to the Logistic regression. 분류 모형에 있어서 주요한 이슈는 분류와 예측의 정확성을 향상시키는데 있다. 따라서 많은 연구들이 분류와 예측의 정확성을 높이기 위해서 진행되어 왔으며 대표적인 방법은 바로 변형과 통합의 전략으로 불리는 앙상블 알고리즘이다. 본 논문에서는 Logistic Bagging이라는 새로운 분류 알고리즘을 제안하고 실제 자료와 시뮬레이션 자료를 이용해 제안된 모형을 기존의 Bagging, Double Bagging, Logistic Regression과 비교하였다. 실험결과 Logistic Bagging은 Bagging과 Logistic Regression에 비해 우수한 성능을 보였으며, 자료가 정규분포를 따르지 않는 경우에 한해서 Double Bagging보다도 정확한 분류 성능을 보였다. 자료가 정규분포를 따르는 경우 Logistic Bagging은 Double Bagging보다 향상된 분류 성능을 보이지 못했다. 마지막으로 Bias-Variance Decomposition을 이용하여 Logistic Bagging이 타 분류모형에 비해 Bias와 Variance 중 어떠한 부분에서 향상된 성과를 보이는가를 분석한 결과, Logistic Bagging모형은 Bagging과 Double Bagging에 비해 모형의 Variance를 줄임으로써 분류와 예측의 정확성을 높인다는 것을 알 수 있었고, Logistic Regression에 비해서는 Variance보다는 Bias를 줄임으로 분류와 예측의 정확성을 높인다는 것을 알 수 있었다.

    • Event-related Potential analysis using elastic net logistic regression

      이우종 서울대학교 대학원 2016 국내석사

      RANK : 2907

      The objective of the thesis is to explore whether regularization techniques can be applied to ERP analysis, and which type of regularization is adequate. This thesis proposes elastic net regularization logistic regression as a good candidate of data analytic method for Event-Related Potential analysis (ERP). Specifically, regularization techniques are used to identify latency in ERP. Study 1 tested whether regularization logistic regression can classify latency using simulated ERP data. It showed that ridge and lasso could identify latency information. In study 2, the same analyses were applied to actual ERP data. Ridge regression can identify latency information wheras lasso cannot.

    • Principal weighted logistic regression for sufficient dimension reduction for binary classification

      김보영 Graduate School, Korea University 2018 국내석사

      RANK : 2893

      본 논문은 이분형 반응변수의 분류 문제에서 설명변수의 차원을 축소하는 충분 차원축소(Sufficient Dimension Reduction) 방법으로 Principal Weighted Logistic Regression을 제안하였다. 반응변수가 이분형일 때 기존에 알려져 있던 충분 차원축소 방법들이 가지는 한계점을 극복하기 위해 가중 로지스틱 회귀 분석을 활용해 축소된 변수 공간을 추정하였다. 이 때, 커널 트릭을 활용하여 비선형 충분 차원축소까지 하나의 통합된 프레임워크에서 전개함으로써 반응 변수의 결정 경계가 원점에 대해 대칭인 경우에도 적용 가능한 차원축소 방법을 제안하였다. 모의 실험과 실제 데이터 적합을 통해 그 성능과 적용 가능성을 확인함으로써 본 연구의 의의를 입증하였다. Sufficient dimension reduction(SDR) has offered an effective means to reduce the dimensionality of the data without losing the regressing information of Y onto X. In this article, we propose a new SDR tool called Principal Weighted Logistic Regression to be used in the classification problem. In order to overcome the limitations of the existing SDR methods when Y is binary, we employ weighted logistic regression to obtain the basis of the reduced predictor space. Making use of kernel trick, we developed a unified framework of linear and nonlinear SDR to be used in both situations where the true decision curve of the data is asymmetric and symmetric about the origin. Both simulation studies and empirical applications are examined in order to support the significance of the proposed method.

    • 머신러닝 기법을 활용한 유산소 운동 중 혈당 변화 예측 모형

      Oyama, Okimitsu 연세대학교 대학원 2023 국내석사

      RANK : 2889

      본 연구는 운동 중 혈당의 변화에 영향을 주는 요인을 선별하여 데이터를 수집하고 머신러닝 기법과 전통적 통계 기법을 활용하여 운동 중 혈당의 변화를 예측하는 모형을 제시하는 것에 목적이 있다. 150명(연령: 31.88±9.67세, 공복 4시간 이상 혈당: 102.09±14.09mg/dL)의 건강한 성인 남녀가 연구에 참여하였고 그 중 30명(연령: 32.67±9.91세, 공복 4시간 이상 혈당: 101.6±7.42mg/dL)을 무작위 선별하여 재측정을 하였다. 30분간의 중강도(예측 최대 심박수의 70~85%, 운동자각도: 11~13) 달리기 운동을 진행하였을 때 1차 실험에서는 83명의 혈당이 5% 이상 감소하였고, 67명의 혈당이 유지되었다. 예측 모형을 만들기 위해 측정된 대상자 특성에는 성별, 연령, 신장, 몸무게, 체지방률, 안정시 심박수, 골격근량, 내장지방량, 운동 전과 운동 중 혈당, 운동 전과 운동 중 혈액 내 젖산, 운동 중 심박수, 하루 전과 당일 식이 내역이며 예측 모형에는 성별, 연령, BMI, 체지방률, 안정시 심박수, 골격근량, 내장지방량, 4시간 이상 공복 혈당, 운동 전 젖산치, 24시간 내 섭취한 탄수화물량(g), 혈당과 체력을 반영할 수 있는 운동 후 10분 간의 심박수 회복량, 운동 중 상승한 심박수, 운동 중 변화한 젖산을 변수로 활용하였다. 운동 중 혈당 변화 예측 모형으로는 로지스틱 회귀 분류, 랜덤포레스트 분류, 다중 회귀를 응용한 분류기 3가지 방법을 사용하였으며 각각 전체 대상, 남성 대상, 여성 대상의 3가지 모델을 만들어 9가지 모델을 제시하였다. 모형의 성능은 모델 빌딩 데이터를 기준으로 랜덤 포레스트 분류 모형(전체 모형 AUC: 0.833, Youden index: 0.53, 남성 모형 AUC: 0.837, Youden index: 0.66, 여성 모형 AUC: 0.785, Youden index: 0.51), 로지스틱 회귀 분류 모형(전체 모형 AUC: 0.702, Youden index: 0.44, 남성 모형 AUC: 0.807, Youden index: 0.55, 여성 모형 AUC: 0.770, Youden index: 0.45), 다중 회귀 응용 분류 모형(전체 모형 Youden index: 0.34, 남성 모형 Youden index: 0.47, 여성 모형 Youden index: 0.31)의 순서대로 성능이 좋게 나타났다. 타당도 확인을 위한 30명의 연구 참여자의 재측정 데이터에 모형을 적용하여 보았을 때, 로지스틱 회귀 분류 모형(전체 모형 AUC: 0.817, Youden index: 0.55, 남성 모형 AUC: 0.905, Youden index: 0.38, 여성 모형 AUC: 0.735, Youden index: 0.57), 랜덤 포레스트 분류 모형(전체 모형 AUC: 0.679, Youden index: 0.21, 남성 모형 AUC: 0.730, Youden index: 0.53, 여성 모형 AUC: 0.694, Youden index: 0.14), 다중 회귀 응용 분류 모형(전체 모형 Youden index: 0.46, 남성 모형 Youden index: 0.24, 여성 모형 Youden index: 0.29)의 순으로 혈당 변화 예측 성능이 좋게 나타났다. 본 연구의 결과를 요약하면, 모델 빌딩 데이터를 기준으로 랜덤 포레스트 모형이 가장 분류 성능이 좋게 나타났으며 재측정 대상자를 기준으로는 로지스틱 회귀 분류 모형이 성능이 높은 것으로 나타났다. 재측정 대상자 수가 30명으로 매우 적은 수이기에 Youden index의 값이 낮게 나온 것으로 생각된다. 추후 연구에서는 인원을 늘려서 타당도를 검사하기를 권하며 외부 검증 데이터를 확보하여 외적 타당성의 확인도 필요할 것이다. The purposes of this study were to explore the relationship of blood glucose level change and characteristics of the body during exercise, and present three models for predicting change in blood glucose levels during exercise using machine learning and traditional statistical techniques. 150 healthy adult men and women (age: 31.88±9.67 yr, fasting blood glucose: 102.09±14.09 mg/dL) participated in the study, and 30 of them (age: 32.67±9.91 yr, fasting blood glucose: 101.6±7.42 mg/dL) were randomly selected and re-measured. In the primary experiment, blood glucose was reduced in 83 people by more than 5% and blood glucose was maintained in 63 people when running for 30 minutes of moderate intensity (70-85% of predicted maximum heart rate, RPE: 11-13). In the re-measurement experiment, blood glucose was reduced in 16 people and was maintained in 14 people. The collected factors were sex, age, height, body weight, body fat percentage, resting heart rate, skeletal muscle mass, visceral fat volume, blood glucose before and during exercise, lactate before and during exercise, heart rate before and during exercise, and the 24hour dietary history. For the prediction models of blood glucose change during exercise, the following three classification methods were used: logistic regression classification, random forest classification, and multiple regression. A total of nine models were presented since three models for all subjects, men, and women were created. Based on model building data, high performance was seen in the following order: random forest classification model (model for all AUC: 0.833, Youden index: 0.53, model for men AUC: 0.837, Youden index: 0.66, model for women model AUC: 0.785, Youden index: 0.51), logistic regression classification model (model for all AUC: 0.702, Youden index: 0.44, model for men AUC: 0.807, Youden model: 0.55, model for women AUC: 0.770, Youden index:0.45), and multiple regression classification model (model for all Youden index: 0.34, model for men Youden model: 0.47, model for women Youden index:0.31). Based on the re-measurement validation data, high performance was observed in the following order: logistic regression classification model (model for all AUC: 0.817, Youden index: 0.55, model for men AUC: 0.905, Youden index: 0.38, model for women AUC: 0.735, Youden index: 0.57), random forest classification model (model of all AUC: 0.679, Youden index: 0.21, model for men AUC: 0.730, Youden index: 0.53, model for women AUC: 0.694, Youden index: 0.14). multiple regression classification model (model for all Youden index: 0.46, model for men Youden index: 0.24, model for women Youden index: 0.29). To summarize, the random forest model showed the highest classification performance based on the model building data, and the logistic regression classification model showed the highest performance based on the remeasurement subjects. We had small values of Youden index due to the small number of re-measured subjects, and in future studies, it is warranted to use a larger dataset to check validity and also examine external verification data.

    • STOP-Bang 및 임상 지표 기반 폐쇄성수면무호흡증(OSA) 진단 예측모델 연구: 심층신경망(DNN) 제안 및 로지스틱 회귀분석과의 성능 비교

      이슬기 연세대학교 보건대학원 2025 국내석사

      RANK : 2889

      배경 및 목적: 폐쇄성수면무호흡증(Obstructive Sleep Apnea, OSA)은 수면 중 반복적인 상기도 폐쇄로 인해 호흡이 제한되는 질환으로, 고혈압, 심혈관질환, 당뇨 등 만성질환과 밀접한 연관이 있는 주요 공중보건 문제로 인식되고 있다. OSA 조기 선별을 위한 도구로 STOP-Bang 설문이 널리 사용되고 있으며, 이 설문지는 높은 민감도를 바탕으로 고위험군을 효과적으로 선별할 수 있는 장점이 있다. 그러나 낮은 특이도로 인해 개별 진단의 정밀도에는 한계가 존재하며, 단독 활용 시 진단 예측의 정확도를 확보하는 데 어려움이 있다는 지적이 제기되고 있다. 이에 본 연구는 STOP-Bang 점수의 한계를 보완하고 예측 성능을 향상시키기 위해 추가적인 임상 및 건강행태 지표를 통합한 예측모델을 구축하고자 하였으며, 전통적 통계 기법인 로지스틱 회귀분석과 머신러닝 기반 심층신경망(Deep Neural Network, DNN)의 예측 성능을 비교하여 보다 효과적인 OSA 진단 예측 모델을 제시하는 데 목적이 있다. 대상 및 방법: 본 연구는 2019년부터 2023년까지 수행된 제8–9기 국민건강영양조사(KNHANES) 자료 중 40세 이상 성인 14,939명을 대상으로 하였다. 폐쇄성수면무호흡증(OSA) 진단 여부는 자가 보고된 의사 진단을 기준으로 정의하였으며, STOP-Bang 설문 점수 및 구성 항목 외에도 혈압, 혈당, 지질 수치, 체질량지수(BMI), 음주 및 흡연 여부, 수면시간, 앉아있는 시간 등의 임상 및 건강행태 지표를 분석에 포함하였다. 통계 분석은 복합표본설계를 고려하여 로지스틱 회귀분석을 실시하였으며, STOP-Bang 점수 단일 항목의 설명력을 평가하기 위한 단변수 분석과 건강 지표를 추가로 포함한 다변수 분석을 병행하였다. 예측 성능 검증을 위해 전체 데이터를 훈련용과 검증용으로 7:3 비율로 분할하여 적용하였으며, 추가적으로 심층신경망(Deep Neural Network, DNN) 기반의 예측 모델을 구성하고, 예측 정확도 및 변수 기여도를 비교 평가하기 위해 ROC-AUC 및 SHAP 분석을 수행하였다. 연구결과: 단변수 로지스틱 회귀분석에서 STOP-Bang 점수는 OSA 진단과 유의한 양의 연관성을 보였으며(OR=2.31, 95% CI: 2.02–2.65,p<0.0001), 해당 모형의 AUC는 0.8268로 확인되었다. 다변수 로지스틱 회귀분석 결과, AUC는 0.8660으로 수면 중 무호흡 목격 여부가 가장 유의한 예측 인자로 확인되었으며 (OR=20.68, 95% CI: 9.69–44.14, p<0.0001), 코골이 여부와(OR=2.29, 95% CI: 1.09–4.79, p=0.0287) 일일 총 앉아 있는 시간(OR=1.14, 95% CI: 1.06–1.22, p=0.0003), 음주 여부(OR=0.16, 95% CI: 0.04–0.67, p=0.0130) 또한 유의한 관련성을 보였다. 한편, 심층신경망(DNN) 기반 예측모델의 경우 AUC는 0.9020으로, 세 모델 중 가장 높은 예측력을 보였다. 또한 SHAP 분석 결과 로지스틱 회귀분석에서 유의했던 변수들과 유사한 경향성을 보이며, 해석 가능성과 임상적 타당성을 동시에 확보할 수 있음을 시사하였다. 결론: STOP-Bang 설문은 OSA 고위험군 선별에 효과적인 도구로 활용될 수 있으나, 개별 진단 수준에서는 특이도의 한계가 뚜렷하여 보완이 필요하다. 본 연구는 STOP-Bang 점수에 임상 및 건강행태 정보를 결합함으로써 예측 정밀도를 향상시킬 수 있음을 확인하였으며, DNN 기반의 머신러닝 모델이 기존 통계모형보다 우수한 예측 성능을 보임을 통해 향후 OSA 조기 선별을 위한 고도화된 예측모델 개발의 가능성을 제시하였다. Background and purpose: Obstructive Sleep Apnea (OSA) is a sleep-related breathing disorder characterized by repeated upper airway obstruction during sleep. It is recognized as a major public health concern due to its strong association with chronic diseases such as hypertension, cardiovascular disease, and diabetes. The STOP-Bang questionnaire is widely used as a screening tool for early detection of OSA, offering the advantage of high sensitivity in identifying high-risk individuals. However, its low specificity limits its precision in individual diagnosis, raising concerns about its standalone predictive accuracy. This study aims to address these limitations by developing an enhanced predictive model that integrates the STOP-Bang score with additional clinical and behavioral health indicators. Furthermore, the study seeks to compare the predictive performance of a traditional statistical method—logistic regression—with that of a machine learning-based Deep Neural Network (DNN), ultimately proposing a more effective model for OSA risk prediction. Methods: This study analyzed data from 14,939 adults aged 40 years and older, derived from the 8th and 9th cycles (2019–2023) of the Korea National Health and Nutrition Examination Survey (KNHANES). Obstructive Sleep Apnea (OSA) diagnosis was defined based on self-reported physician diagnosis. In addition to the STOP-Bang score and its individual components, the analysis incorporated various clinical and behavioral health indicators, including blood pressure, blood glucose, lipid levels, body mass index (BMI), alcohol consumption, smoking status, sleep duration, and sedentary time. Logistic regression analysis was conducted considering the complex sampling design. Both univariate analysis, using only the STOP-Bang score, and multivariate analysis, incorporating additional health indicators, were performed to assess explanatory power. For model validation, the dataset was randomly split into training and validation sets in a 7:3 ratios. A Deep Neural Network (DNN)-based predictive model was also developed. Predictive performance and variable contribution were evaluated using the area under the receiver operating characteristic curve (ROC-AUC) and SHapley Additive exPlanations (SHAP) analysis. Results: In the univariate logistic regression analysis, the STOP-Bang score was significantly and positively associated with OSA diagnosis (OR = 2.31, 95% CI: 2.02–2.65, p < 0.0001), with an AUC of 0.8268. In the multivariable logistic regression model, the AUC improved to 0.8660, with witnessed apnea during sleep emerging as the most significant predictor (OR = 20.68, 95% CI: 9.69–44.14, p < 0.0001). Other significant predictors included presence of snoring (OR = 2.29, 95% CI: 1.09–4.79, p = 0.0287), total sitting time per day (OR = 1.14, 95% CI: 1.06–1.22, p = 0.0003), and alcohol consumption (OR = 0.16, 95% CI: 0.04–0.67, p = 0.0130). The Deep Neural Network (DNN)-based prediction model achieved the highest performance among the three models, with an AUC of 0.9020. SHAP analysis demonstrated variable importance trends similar to those identified in the logistic regression model, indicating both interpretability and clinical relevance of the DNN model. Conclusion: While the STOP-Bang questionnaire is an effective tool for identifying individuals at high risk for OSA, its limited specificity at the individual diagnostic level highlights the need for complementary approaches. This study demonstrates that incorporating clinical and health indicators alongside the STOP-Bang score can enhance predictive precision. Moreover, the Deep Neural Network (DNN)-based machine learning model outperformed traditional statistical models, suggesting the potential for developing more advanced and accurate prediction tools for early OSA screening in the future.

    • 외상성 늑골 골절 유형에 따른 호흡기 합병증 예측에 대한 새로운 점수 체계

      석준필 충북대학교 일반대학원 2025 국내박사

      RANK : 2889

      Purpose: Blunt chest trauma is common and is associated with post-traumatic complications and mortality. Among the various types of injury, rib fractures (RFX) are frequently observed, and many studies have attempted to identify risk factors for complications and mortality based on the number and severity of RFX. However, no clear conclusions have been reached. Therefore, this study aimed to identify new risk factors related to RFX that have not been previously reported. Using these factors, a predictive model for complications and mortality was developed and it was compared with existing chest injury scoring systems. Methods: This retrospective medical record study included patients with blunt chest trauma who presented Chungbuk National University Hospital between January 2019 and November 2024. A total of 925 patients were finally included in the study. Data were collected based on initial chest computed tomography (CT) scans and medical records. The following exclusion criteria were applied: (1) patients who died within 24 hours of arrival or had severe injuries to other organs (such as the head, abdomen, or limbs) in addition to chest trauma, making it difficult to determine the exact cause of death or complications; (2) patients who were transferred to another hospital during treatment or whose essential medical records were missing, making further investigation impossible. The number and severity of RFX were recorded both at the patient level and for each individual rib. A segmental RFX (sRFX) was defined as a single rib with two or more fracture lines. A flail segment was defined as cases containing three or more consecutive sRFX, and among patients with such a flail segment, if paradoxical respiration was observed, it was defined as flail motion. The thorax was divided into anterior, lateral, and posterior segments based on the anterior and posterior axillary lines. A “fracture line”was defined as a virtual line in which multiple RFX were aligned in a certain direction within the thorax. if multiple fracture lines existed, the single most severe line (based on the number of RFX and the degree of displacement) was designated as the primary fracture line (PFL). Depending on the location of the PFL within the thorax, it was further classified as anterior, lateral, or posterior. Because a flail segment is formed by two or more fracture lines that locate in different location within the thorax, flail segments were further subdivided according to their location, for example, anterior-lateral flail segments, anterior-posterior flail segment, and so on. Fracture severity was quantified as the percentage of displacement relative to the thickness of the rib at the fracture cross-section, and based on prior studies, fractures were graded accordingly. All results were presented as the mean and standard deviation. For nominal categorical variables or markedly skewed data, the median and interquartile range were used. The total patient cohort was randomly divided into a development set (80%) and a validation set (20%). Using the development set, significant variables were selected by using Least Absolute Shrinkage and Selection Operator (LASSO) logistic regression. To determine coefficients and construct a predictive model using a nomogram, multivariate logistic regression was conducted using these variables. Finally, the predictive performances of chest trauma scoring systems were compared in the validation set. The primary outcome was defined as the occurrence of one or more of the following events: death, pneumonia, the performance of rib fixation surgery, or complications that required surgical intervention. Results: A total of 925 patients were included in the study, of whom 148 (16.0%) experienced one or more of the primary outcomes. When randomly assigned at a 8:2 ratio, 119 of the 741 patients (16.1%) in the development set and 29 of the 184 patients (15.8%) in the validation set met the criteria for the primary outcome. Among the chest injury indicators showing significant differences in univariate analysis, LASSO logistic regression applied to the development set identified five statistically significant risk factors: age, PFR (PaO2/FiO2 ratio), an anterior-lateral flail segment, the number of completely displaced rFX, and the number of sRFX. A nomogram-based predictive model was then developed using these variables. Using the validation set, we compared the predictive power of this new model with that of five existing chest trauma scoring systems by calculating the area under the reveiver operating characteristic curve (AUROC). The new predictive model demonstrated the highest AUROC (0.820), followed by Thoracic Trauma Severity Score (0.728), Chest Trauma Score (0.701), Rib Fracture Score (0.666), and RibScore (0.635). Conclusion: In this study, new chest injury variables were identified that had not been previously reported. Using these variables, a novel chest injury scoring system were developed adverse outcome, which showed superior predictive capability compared to other existing chest trauma scoiring systems. We believe that this model may facilitate rapid patient triage and early management in the future. Key words: Rib fracture, Predictive model, Flail segment, Nomogram, Least Absolute Shrinkage and Selection Operator logistic regression (LASSO) 목적: 둔상성 흉부 외상 환자는 매우 흔하며 외상 후 합병증 및 사망률과 연 관이 있다. 외상 손상의 유형 중 늑골 골절은 가장 흔하게 관찰되며, 늑골 골 절의 개수 및 골절의 정도를 바탕으로 합병증과 사망률에 대한 위험인자를 밝 혀내기 위한 연구가 꾸준히 진행되었으나 명확한 결론은 없다. 이에 본 연구 에서는 특히 늑골 골절에 대해 동요 분절의 위치, 분절 골절의 개수, 완전히 어긋난 골절의 개수 등 예전에 보고되지 않은 새로운 위험인자를 식별하고자 하였다. 이를 토대로 합병증 및 사망률에 대한 예측 모델을 개발하여 기존에 보고된 흉부 손상 점수 체계들과 비교해 보았다. 연구 대상 및 방법: 본 후향적 의무 기록 연구는 2019년 1월부터 2024년 11월까지 충북대학교병원에 내원한 흉부 외상 환자 중 둔상 기전의 환자들을 대상으로 하였으며 925명이 최종적으로 연구에 포함되었다. 내원 초기의 흉부 전산화 단층 촬영 (chest computed tomography; chest CT) 및 의무기록지 를 토대로 하였다. 본 연구에서는 다음 기준에 따라 환자를 제외하였다: (1) 내원 초기 24시간 이내에 사망하거나 가슴 손상 외에 머리, 복부, 사지 등 타 장기에 심각한 손상이 있어서 사망 및 합병증 발생의 원인을 정확히 알 수 없 는 경우; (2) 치료 도중에 타 병원으로 전원하거나 필수적인 의무 기록이 누 락되어 더 이상의 조사가 어려운 경우. 늑골 골절의 개수, 늑골이 부러진 정도를 모든 환자에게 있어서 늑골별로 기 록하였다. 분절 골절은 단일 늑골이 두 개 이상의 골절이 있는 경우로 정의하 였다. 동요 분절 (flail segment)는 3개 이상의 연속된 분절 골절이 있는 경우 로 정의하였으며 동요 분절이 있는 환자 중 흉벽의 기이 운동 (paradoxical respiration)이 관찰되는 경우는 동요 운동 (flail motion)으로 정의하였다. 흉곽은 전방 및 후방 액와선 (anterior and posterior axillary line)을 기준 으로 전방, 측방, 그리고 후방으로 분류하였다. 다수의 늑골 골절들이 흉곽에서 일정 방향성을 가지고 형성하는 가상의 선 (line)을 골절선 (fracture line)으로 정의하였다. 한쪽 가슴에 다수의 골절선이 있는 경우에는 각 골절선을 구성하는 골절의 개수, 골절의 정도를 비교해서 가장 심한 형태 골절로 이루어진 골절선을 주 골절선 (primary fracture line; PFL)으로 정의하였고, 이 선이 있는 흉곽의 구획에 따라서 전방 주 골절선, 측방 주 골절선, 또는 후방 주 골절선으로 분류하였다. 동요 분절 역시 두 개 이상의 골절선이 모여서 하나의 분절을 만들기에 골절 선의 위치에 따라서 전방-측방 동요 분절, 전방-후방 동요 분절 및 측방-후 방 동요 분절로 세분화하여 분류하였다. 골절의 정도는 부러진 단면을 기준으로 늑골의 두께와 비교해서 어느 정도 절단면이 어긋났는지 분률 (%)로 기록하였고, 이를 기준으로 기존에 보고된 연구를 참고하여 등급 (grade)를 나누었다. 모든 결과는 평균 (average)와 표준편차 (standard deviation)을 이용하였으 며 명목형 변수 (nominal categorical variable) 또는 치우침 (skew)가 심한 값에 대해서는 중윗값 (median)과 사분위 수 (interquartile range)를 사용하 였다. 전체 환자군을 8:2 비율로 개발 그룹 (development set)과 검증 그룹 (validation set)으로 무작위 배정을 한 다음 Least Absolute Shrinkage and Selection Operator (LASSO) logistic regression을 이용해 통계적 유의성을 보이는 변수들을 찾아내었다. 이 변수들로 개발 그룹에 대해 다변량 로지스틱 회귀분석(multivariate logistic regression)을 사용해서 오즈비 (odds ratio) 등의 계수(coefficient)를 구한 다음에 nomogram을 이용한 예측 모델을 만들 었으며, 검증 그룹을 토대로 다른 흉부 손상 점수 체계와 예측력의 우위를 비 교하였다. 일차 결과 (primary outcome)은 환자의 흉부 손상 합병증으로 인한 사망, 폐렴 발생, 흉부 손상으로 인해 수술적 치료를 요구했던 합병증의 발생 및 늑 골 고정 수술의 시행 여부 중 하나 이상으로 정의하였다. 결과: 전체 925명의 환자가 본 연구에 포함되었으며 148명 (16.0%)의 환자가 하나 이상의 primary outcome을 보였다. 8:2 무작위 배정에서는 741명의 개 발 그룹 중 119명 (16.1%)가, 184명의 검증 그룹에는 29명 (15.8%)가 primary outcome에 해당하였다. 단변량 분석에서 유의한 차이를 보인 흉부 손상 지표들로 개발 그룹에 대해서 LASSO 회귀분석을 시행한 결과, 나이, PFR (PaO2/FiO2 ratio), 전 측방 (anterior-lateral)에 위치한 동요 분절, 완 전히 어긋난 골절이 있는 늑골의 개수, 그리고 분절 골절이 있는 늑골의 개수 까지 총 다섯 개의 위험인자가 통계적 유의성을 보였으며, 이 변수들을 이용 해 nomogram을 이용한 예측 모델을 만들었다. 이후 검증 그룹에 대해서 상 기 예측 모델과 기존의 다섯 종류의 흉부 손상 점수 체계의 예측도를 AUROC (Area under a receiver operating characteristic curve)를 사용해 비교하였다. 본 연구에서 제시한 새로운 예측 모델의 AUC가 0.820으로 가장 높았고, 그 뒤로 Thoracic Trauma Severity Score (0.728), Chest Trauma Score (0.701), Rib Fracture Score (0.666), 그리고 RibScore (0.635) 순서 로 예측력을 보였다. 결론: 기존의 흉부 손상 점수 체계는 늑골 골절의 개수, 나이 등 평범한 변수 들로 구성되어 있으며 사망률과 합병증에 대한 예측도는 낮은 편이다. 하지만 이번 연구에서 제시한 예측 모델은 동요 분절의 위치, 분절 골절의 개수 및 완전히 어긋난 늑골의 개수 등 구조적이고 위치적 특성을 반영한 새로운 흉부 손상 변수를 도입하였으며 기존의 예측 모델보다 우월한 예측도를 보였다. 이 를 이용하여 향후 빠른 환자 분류 및 초기 처치에 도움을 줄 수 있겠다고 생 각한다. 핵심 용어: 늑골 골절 (Rib fracture), 예측 모델 (Predictive model), 동요 분절 (Flail segment), 노모그램 (nomogram), Least Absolute Shrinkage and Selection Operator logistic regression (LASSO)

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼