RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    XGBoost와 Logistic Regression 기반 선택적 하이브리드 모델을 활용한 대출 부도 예측 연구 = Loan Default Prediction Using a Selective Hybrid Model Based on XGBoost and Logistic Regression

    한글로보기

    https://www.riss.kr/link?id=T17143835

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    Logistic Regression 모델은 전통적으로 신용 평가 분야에서 널리 활용되어 왔으며, 모델 해석이 용이하다는 장점이 있다. 그러나 복잡한 데이터 구조를 처리하거나 고위험 부도 고객을 예측하는 것은 한계를 가진다. 반면, AI 기반 모델은 비선형 관계를 효과적으로 처리하며 우수한 예측 성능을 제공한다. 그러나 복잡한 알고리즘으로 인해 해석이 어려워 '블랙박스(black-box)'로 간주되는 문제가 존재한다. 이러한 한계는 모델의 투명성과 신뢰성을 저하시킬 수 있으며, 기존 연구들은 주로 XAI(Explainable AI) 기법을 활용하여 '블랙박스' 문제를 완화하고, 모델의 해석 가능성을 제고하는 데 초점을 맞춰왔다. 본 연구는 대출 부도 예측의 정확도를 향상시키고 해석 가능성을 강화하기 위해 Logistic Regression의 해석 용이성과 AI 기반 모델의 예측 성능을 결합한 선택적 하이브리드 모델을 제안한다. 제안된 모델은 두 가지 주요 단계로 구성된다. 첫째, Logistic Regression과 AI 기반 모델의 예측이 불일치하는 영역에서 Logistic Re-gression의 예측 결과를 AI 기반 모델의 예측으로 대체하는 선택적 접근 방식을 도입하였다. 이는 기존 앙상블 모델과 달리, 두 모델 간 예측 차이를 활용하여 한 모델의 결과를 다른 모델로 보완하는 방식이다. 둘째, AI 기반 모델이 Logistic Re-gression보다 높은 부도 위험도를 예측한 고위험 고객 데이터를 별도로 추출한 후, 이를 기반으로 고위험 고객에 특화된 Logistic Regression 모델을 추가로 학습시키는 방법론을 제안하였다. 실험 결과, 선택적 하이브리드 모델은 고위험 부도 고객 세그먼트에서 단일 모델보다 우수한 성능을 나타냈다. 또한, 고위험 고객 데이터를 기반으로 학습된 Logistic Regression 모델은 Logistic Regression 및 AI 기반 단일 모델에 비해 향상된 예측 성능을 보였다. 본 연구는 해석 가능한 Logistic Regression 모델과 고성능 AI 기반 모델을 조건부로 결합함으로써 기존 단일 모델 접근법의 한계를 보완하고, 대출 부도 예측의 정확도와 해석 가능성을 동시에 제고하였다는 점에서 학술적 의의가 있다. 나아가 본 연구가 대출 부도 관리 및 리스크 평가 분야에서 실무적 통찰을 제공할 수 있을 것으로 기대한다.
    번역하기

    Logistic Regression 모델은 전통적으로 신용 평가 분야에서 널리 활용되어 왔으며, 모델 해석이 용이하다는 장점이 있다. 그러나 복잡한 데이터 구조를 처리하거나 고위험 부도 고객을 예측하는 ...

    Logistic Regression 모델은 전통적으로 신용 평가 분야에서 널리 활용되어 왔으며, 모델 해석이 용이하다는 장점이 있다. 그러나 복잡한 데이터 구조를 처리하거나 고위험 부도 고객을 예측하는 것은 한계를 가진다. 반면, AI 기반 모델은 비선형 관계를 효과적으로 처리하며 우수한 예측 성능을 제공한다. 그러나 복잡한 알고리즘으로 인해 해석이 어려워 '블랙박스(black-box)'로 간주되는 문제가 존재한다. 이러한 한계는 모델의 투명성과 신뢰성을 저하시킬 수 있으며, 기존 연구들은 주로 XAI(Explainable AI) 기법을 활용하여 '블랙박스' 문제를 완화하고, 모델의 해석 가능성을 제고하는 데 초점을 맞춰왔다. 본 연구는 대출 부도 예측의 정확도를 향상시키고 해석 가능성을 강화하기 위해 Logistic Regression의 해석 용이성과 AI 기반 모델의 예측 성능을 결합한 선택적 하이브리드 모델을 제안한다. 제안된 모델은 두 가지 주요 단계로 구성된다. 첫째, Logistic Regression과 AI 기반 모델의 예측이 불일치하는 영역에서 Logistic Re-gression의 예측 결과를 AI 기반 모델의 예측으로 대체하는 선택적 접근 방식을 도입하였다. 이는 기존 앙상블 모델과 달리, 두 모델 간 예측 차이를 활용하여 한 모델의 결과를 다른 모델로 보완하는 방식이다. 둘째, AI 기반 모델이 Logistic Re-gression보다 높은 부도 위험도를 예측한 고위험 고객 데이터를 별도로 추출한 후, 이를 기반으로 고위험 고객에 특화된 Logistic Regression 모델을 추가로 학습시키는 방법론을 제안하였다. 실험 결과, 선택적 하이브리드 모델은 고위험 부도 고객 세그먼트에서 단일 모델보다 우수한 성능을 나타냈다. 또한, 고위험 고객 데이터를 기반으로 학습된 Logistic Regression 모델은 Logistic Regression 및 AI 기반 단일 모델에 비해 향상된 예측 성능을 보였다. 본 연구는 해석 가능한 Logistic Regression 모델과 고성능 AI 기반 모델을 조건부로 결합함으로써 기존 단일 모델 접근법의 한계를 보완하고, 대출 부도 예측의 정확도와 해석 가능성을 동시에 제고하였다는 점에서 학술적 의의가 있다. 나아가 본 연구가 대출 부도 관리 및 리스크 평가 분야에서 실무적 통찰을 제공할 수 있을 것으로 기대한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Logistic Regression has traditionally been widely utilized in the credit evalua-tion domain due to its interpretability. However, it faces limitations in handling complex data structures and predicting high-risk default customers. In contrast, AI-based models effectively capture non-linear relationships and exhibit supe-rior predictive performance. Nevertheless, these models are often considered "black-boxes" due to their complex algorithms, which hinder interpretability. This lack of transparency can undermine the trust and reliability of the models. Existing research has primarily focused on leveraging Explainable AI (XAI) techniques to address the "black-box" problem and enhance model interpreta-bility. This study proposes a selective hybrid model that combines the interpretabil-ity of Logistic Regression with the predictive performance of AI-based models to improve loan default prediction accuracy and interpretability. The proposed model comprises two key components. First, a selective approach replaces Lo-gistic Regression predictions with AI-based predictions in regions where the two models produce inconsistent outcomes. Unlike traditional ensemble models, this approach utilizes the differences in predictions to complement one model's results with those of the other. Second, high-risk customer data, where the AI-based model predicts a higher default risk than Logistic Regression, are sepa-rately extracted, and a specialized Logistic Regression model is trained using this data. The experimental results demonstrate that the selective hybrid model outper-forms single models in terms of AUC (Area Under the Curve) performance for high-risk default customer segments. Additionally, the specialized Logistic Re-gression model trained on high-risk customer data achieves superior predictive performance compared to the conventional Logistic Regression and AI-based models. This study contributes to the literature by addressing the limitations of sin-gle-model approaches through the conditional integration of interpretable Lo-gistic Regression and high-performance AI-based models. It simultaneously enhances the accuracy and interpretability of loan default prediction. Further-more, it is expected that this research will provide practical insights for loan default management and risk assessment.
    번역하기

    Logistic Regression has traditionally been widely utilized in the credit evalua-tion domain due to its interpretability. However, it faces limitations in handling complex data structures and predicting high-risk default customers. In contrast, AI-base...

    Logistic Regression has traditionally been widely utilized in the credit evalua-tion domain due to its interpretability. However, it faces limitations in handling complex data structures and predicting high-risk default customers. In contrast, AI-based models effectively capture non-linear relationships and exhibit supe-rior predictive performance. Nevertheless, these models are often considered "black-boxes" due to their complex algorithms, which hinder interpretability. This lack of transparency can undermine the trust and reliability of the models. Existing research has primarily focused on leveraging Explainable AI (XAI) techniques to address the "black-box" problem and enhance model interpreta-bility. This study proposes a selective hybrid model that combines the interpretabil-ity of Logistic Regression with the predictive performance of AI-based models to improve loan default prediction accuracy and interpretability. The proposed model comprises two key components. First, a selective approach replaces Lo-gistic Regression predictions with AI-based predictions in regions where the two models produce inconsistent outcomes. Unlike traditional ensemble models, this approach utilizes the differences in predictions to complement one model's results with those of the other. Second, high-risk customer data, where the AI-based model predicts a higher default risk than Logistic Regression, are sepa-rately extracted, and a specialized Logistic Regression model is trained using this data. The experimental results demonstrate that the selective hybrid model outper-forms single models in terms of AUC (Area Under the Curve) performance for high-risk default customer segments. Additionally, the specialized Logistic Re-gression model trained on high-risk customer data achieves superior predictive performance compared to the conventional Logistic Regression and AI-based models. This study contributes to the literature by addressing the limitations of sin-gle-model approaches through the conditional integration of interpretable Lo-gistic Regression and high-performance AI-based models. It simultaneously enhances the accuracy and interpretability of loan default prediction. Further-more, it is expected that this research will provide practical insights for loan default management and risk assessment.

    더보기

    목차 (Table of Contents)

    • 제 1장 서론
    • 1.1 연구 배경 및 목적
    • 제 2장 연구 배경 지식 및 선행 연구
    • 2.1 머신러닝 알고리즘
    • 2.2 로지스틱 회귀 분석(Logistic Regression)
    • 제 1장 서론
    • 1.1 연구 배경 및 목적
    • 제 2장 연구 배경 지식 및 선행 연구
    • 2.1 머신러닝 알고리즘
    • 2.2 로지스틱 회귀 분석(Logistic Regression)
    • 2.3 의사결정나무(Decision Tree)
    • 2.4 랜덤 포레스트(Random Forest)
    • 2.5 XGBoost (Extreme Gradient Boosting)
    • 2.6 다층 퍼셉트론(MLP, Multi-Layer Perceptron)
    • 2.7 앙상블(Ensemble)
    • 2.8하이브리드 앙상블(Hybrid Ensemble)
    • 2.9 설명 가능한 인공지능
    • 2.10 SHAP(Shapley additive explanations)
    • 제3장 제안 방법론
    • 제4장 Model 1연구 방법
    • 4.1 데이터
    • 4.1.1 데이터 목록
    • 4.1.2 데이터 개요
    • 4.1.3 데이터 결측치
    • 4.1.4 데이터 이상치
    • 4.1.5 데이터 상관관계 분석
    • 4.1.6 데이터 불균형 문제
    • 4.2 데이터 전처리
    • 4.3 분석방법
    • 4.3.1 Model Training
    • 4.3.2 각 모델의 성능 비교 및 최적 모델 선정
    • 4.3.3 부도 확률 비교
    • 4.3.4 부도 예측 등급 비교
    • 제5장 Model 2연구 방법
    • 5.1 데이터
    • 5.1.1 데이터 목록
    • 5.1.2 데이터 개요
    • 5.1.3 데이터 결측치
    • 5.1.4 데이터 이상치
    • 5.1.5 데이터 상관관계 분석
    • 5.1.6 Status 데이터 분포
    • 5.2 데이터 전처리
    • 5.3 분석방법
    • 제 6장 연구 결과
    • 6.1 Model 1의 연구 결과
    • 6.1.1 선택적 대체 모델의 적용 후 AUC 변화
    • 6.1.2 선택적 대체 모델의 부도 등급별 분석
    • 6.2 Model 2의 연구 결과
    • 제 7장 시사점 및 한계점
    • 7.1 학술적 시사점
    • 7.2 실무적 시사점
    • 7.3 한계 및 향후 연구
    • 제 8장 결론
    • 참고 문헌
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼