Credit risk prediction method based on ensemble learning
By constructing and integrating LightGBM, XGBoost and CatBoost models, the integration of integrated learning strategies and weighted soft voting is adopted to solve the model limitations and inconsistencies in credit risk assessment, and more efficient credit risk prediction and assessment are achieved.
Patent Information
- Application Number
- CN202510312614.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing single machine learning model cannot fully cover all risk dimensions in credit risk assessment, and has limitations and prediction inconsistencies, making it difficult to adapt to market changes.
Three models: LightGBM, XGBoost and CatBoost are built, and the model prediction results are integrated through integrated learning strategies, and the weighted soft voting fusion model is adopted to dynamically adjust the weight allocation to adapt to changes in the market environment.
It improves the accuracy and generalization ability of credit risk prediction, can better respond to market changes and risk challenges, and helps financial institutions to more accurately evaluate borrowers' credit risks.
Smart Images

Figure CN120258959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to credit risk prediction technology, and particularly to a credit risk prediction method based on ensemble learning. Background Art
[0002] In the core area of the financial service industry, credit risk assessment is the cornerstone for ensuring the scientific nature of loan, credit card issuance, and various financing decisions. Traditionally, financial institutions rely on statistics-based methods and expert systems to evaluate borrowers' credit ratings, but these methods have limited capabilities in dealing with large datasets and capturing complex interaction effects. With the rapid development of artificial intelligence technology, especially the rise of machine learning models, it has brought innovation to credit risk prediction.
[0003] In recent years, ensemble learning methods have emerged prominently in the field of machine learning. Among them, LightGBM, XGBoost, and CatBoost, as outstanding representatives of gradient boosting decision tree models, have shown great potential in credit risk assessment with their respective excellent prediction performance, high computational speed, and good interpretability. LightGBM has greatly improved the training efficiency by introducing gradient-based unilateral sampling and histogram optimization; XGBoost has achieved high accuracy and scalability of the model with its systematic regularization term design and parallel computing framework; while CatBoost has further enhanced the model's prediction ability in dealing with categorical variables with its unique handling of categorical features.
[0004] These three models have demonstrated excellent performance in different scenarios, but at the same time, they also expose the limitations of individual models in comprehensively considering multi-dimensional risk factors. For example, LightGBM performs well on large-scale datasets, but its performance may not be as fine as other models under certain specific feature combinations;
[0005] XGBoost enhances the stability of the model through regularization, but it may consume more computing resources; CatBoost has a natural advantage in dealing with categorical variables, but it may be slightly inferior in dealing with continuous features. Therefore, relying solely on any one model for credit risk assessment may not be able to comprehensively cover all risk dimensions, thus affecting the comprehensiveness and accuracy of the prediction.
[0006] In view of this, the current urgent problem to be solved is how to integrate the respective advantages of the LightGBM, XGBoost, and CatBoost models to form a more comprehensive and accurate credit risk prediction system. Specifically, a mechanism needs to be designed that can automatically identify and emphasize the performance advantages of different models on specific risk indicators according to the specific requirements of credit risk assessment, while solving the possible inconsistencies and redundancies in the prediction results between models. In addition, how to dynamically adjust the weight allocation of each model to adapt to the changing market environment and credit behavior patterns is also a key challenge for achieving efficient credit risk control. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a credit risk prediction method based on ensemble learning, including:
[0008] S1. Model construction and initialization: For the credit risk prediction task, three basic models of LightGBM, XGBoost, and CatBoost are respectively constructed; during initialization, the learning rate and the depth of the tree are set to ensure that the model has a good learning foundation at startup and avoid overfitting or underfitting;
[0009] S2. Conduct a comprehensive preprocessing of the original data set, including: data cleaning, outlier handling, and data consistency checking;
[0010] S3. Conduct feature selection and data balancing processing on the processed data;
[0011] S4. Conduct detailed and targeted training for each of the three models of LightGBM, XGBoost, and CatBoost;
[0012] S5. Independently complete the credit risk prediction of the data through the trained LightGBM, XGBoost, and CatBoost models, and integrate the model prediction results using an ensemble learning strategy;
[0013] S6. Performance evaluation and model update decision: Regularly or when signs of model performance decline appear, conduct model performance evaluation, and compare the key indicators of prediction accuracy and recall rate of the three integrated models and single models; Set a model update mechanism to decide whether to perform model iteration and update;
[0014] The update mechanism includes: based on a time period, a data volume threshold, or a performance monitoring indicator.
[0015] Advantages of the present invention:
[0016] (1) By integrating multiple models (such as LightGBM, XGBoost, and CatBoost) and adopting a weighted soft voting fusion model, the present invention can effectively avoid the limitations of a single model, comprehensively combine the advantages of each model, improve the generalization ability of the model for new data and the accuracy of predicting the default risk of borrowers, and thus better cope with market changes and risk challenges.
[0017] (2) The model of the present invention can help financial institutions more accurately evaluate the credit risk of borrowers, thereby formulating more effective risk management measures, reducing the non-performing loan ratio, providing new technical means for the fintech field, and contributing to the digital transformation and intelligent upgrading of the financial industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of a credit risk prediction method based on integrated machine learning provided by this application;
[0019] Figure 2 It is a model structure diagram of a credit risk prediction method based on integrated machine learning provided by this application;
[0020] Figure 3 It is a schematic diagram of the steps for an enterprise to use a credit risk prediction model provided by this application;
[0021] Figure 4 It is a schematic diagram of user credit classification provided by this application;
[0022] Figure 5 It is a general loan approval flowchart of a financial institution provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] A credit risk prediction method based on ensemble learning, Figure 1 is the flowchart of the present invention, Figure 2 is the model structure diagram of the present invention, including the following detailed steps, aiming to improve the accuracy and robustness of credit risk assessment through an ensemble learning strategy:
[0025] A1. Model construction and initialization:
[0026] Build three sets of basic models: LightGBM model, XGBoost model, and CatBoost model. These models each have different algorithmic characteristics and are suitable for capturing complex patterns in the data. During initialization, the basic parameters of each model are reasonably set, such as the learning rate, tree depth, etc.
[0027] First, install the corresponding Python packages locally, and then use the sklearn-style API to simplify the process. For each library, first create a model instance, and then use the training data to fit the model; this step is LGBMRegressor() or LGBMClassifier() for LightGBM, XGBRegressor() or XGBClassifier() for XGBoost, and CatBoostRegressor() or CatBoostClassifier() for CatBoost. After training, the.predict() method can be directly used to predict new data. These three models each have unique algorithmic characteristics and advantages, jointly capturing complex patterns in the data. LightGBM adopts gradient-based one-sided sampling technology and histogram algorithms, greatly improving the training speed and efficiency, suitable for processing large-scale datasets, and automatically handling categorical features, reducing the workload of preprocessing; XGBoost enhances the model performance through an optimized gradient boosting framework, provides advanced regularization functions, effectively controls the model complexity, and reduces overfitting; while CatBoost is optimized for categorical features, automatically handles categorical features, reduces the workload of feature encoding, and introduces a unique feature combination technology, enhancing the ability to capture non-linear relationships. During the initialization phase, the basic parameters of each model need to be finely tuned to ensure that they have a good learning foundation at startup, avoiding overfitting or underfitting. Next is the process of sorting out and integrating these models to ensure that they can work together in subsequent training and prediction. All models share the same set of preprocessed datasets, including high-quality features after cleaning, feature selection, and transformation, ensuring consistent data formats, especially the encoding methods of categorical features, to facilitate the input requirements of different models.
[0028] A2. Preprocess the collected data:
[0029] Perform a comprehensive preprocessing on the original dataset, including but not limited to: Data cleaning: Fill in missing values through the mean, median, or model-predicted values to ensure data integrity. Outlier handling: Identify and reasonably handle outliers to avoid their interference with model training. Data consistency check: Unify the data format and solve data inconsistency problems, such as time format, numerical unit, etc.
[0030] A3. Feature engineering:
[0031] Use statistical methods (such as Pearson correlation coefficient, mutual information) to identify and retain features that have a significant impact on credit risk prediction. Solve the problem of class imbalance through techniques such as oversampling, undersampling, or SMOTE to ensure the fairness of model training. Create new features based on business logic and feature importance to enhance the model's ability to identify credit risk, such as debt-to-income ratio, historical default times, etc.
[0032] A4. Model training:
[0033] Use the processed dataset to train the three models separately, ensuring that each model learns on the optimized feature set.
[0034] After completing data preprocessing and feature engineering, for the optimized feature set, each model is trained independently. LightGBM is a model based on gradient boosting decision trees. In this model, LightGBM is configured with different hyperparameters. When adjusting the hyperparameters of LightGBM, this model mainly focuses on the following key parameters: num_leaves: Controls the maximum number of leaves of the tree, and this parameter directly affects the complexity of the model. max_depth: The maximum depth of the tree, which limits the growth of the tree. learning_rate: The learning rate determines the degree of weight update after each iteration. A smaller learning rate can bring a slower but more refined learning process. min_child_weight: This parameter determines the minimum sum of sample weights of the leaf nodes. Increasing this value can prevent overfitting. subsample: This parameter defines the proportion of random sampling for each tree. Increasing this value can reduce overfitting. colsample_bytree: It defines the proportion of randomly sampled features for each tree. Increasing this value also helps reduce overfitting. By adjusting these hyperparameters, the model attempts to find the optimal structural configuration to more accurately predict loan defaults. XGBoost is also a model based on gradient boosting decision trees. In this model, XGBoost is configured with different hyperparameters. When adjusting the hyperparameters of XGBoost, this model mainly focuses on the following key parameters: num_boost_round: Controls the number of iterations during training. More iterations usually can improve the accuracy of the model but may also lead to overfitting. max_depth: The maximum depth of the tree, which limits the growth of the tree. A larger value allows a more complex model but may lead to overfitting. learning_rate: The learning rate determines the degree of weight update after each iteration. A smaller learning rate can bring a slower but more refined learning process, helping to reduce overfitting. min_child_weight: This parameter determines the minimum sum of sample weights of the leaf nodes. Increasing this value can prevent overfitting. subsample: This parameter defines the proportion of random sampling for each tree. Increasing this value can reduce overfitting. colsample_bytree: It defines the proportion of randomly sampled features for each tree. Increasing this value also helps reduce overfitting. By adjusting these hyperparameters, the model attempts to find the optimal structural configuration to more accurately predict loan defaults. CatBoost is also a gradient boosting framework that is particularly good at handling categorical variables. In this model, the CatBoost model, as a component of the ensemble model, is responsible for modeling the data based on its internal logic and generating an output for loan default prediction. The CatBoost model also optimizes its performance through a series of hyperparameter adjustments.The hyperparameter tuning of the CatBoost model mainly considers the following parameters: learning_rate: Controls the speed at which the model learns. Smaller values usually result in slower learning speed but higher accuracy. iterations: The number of iterations during the training process. depth: Determines the maximum depth of the tree and controls the model complexity. l2_leaf_reg: Also known as the strength of the regularization term, used to penalize overly complex models. border_count: Determines the number of feature split points, which is particularly important for categorical features. By adjusting these hyperparameters, the model attempts to find the optimal structural configuration to more accurately predict loan defaults. This step ensures that the model fully learns on a high-quality training set, providing a strong foundation for model fusion.
[0035] A5. Model Fusion and Prediction:
[0036] For a trained model, it is first necessary to load it into memory, which can be accomplished through the load_model method of the respective model's Booster class. When there is new data that needs to be predicted, the predict method is used to make predictions on the newly added data. The new data should be in the same format as the data used during model training, including the order and type of features. If data preprocessing was performed during training (e.g., encoding categorical variables, scaling numerical features, etc.), then the same preprocessing needs to be done on the new data. After each model has completed predicting on the test set or new data, an ensemble learning strategy is adopted to integrate the prediction results. Specifically, first, the trained LightGBM model is used to predict the input data that has been preprocessed and optimized through feature engineering. By constructing multiple decision trees and using the gradient boosting algorithm to iteratively improve these trees, a risk score or probability distribution for each sample point is obtained, denoted as P1(c|x). Next, the trained XGBoost model is used to perform the same prediction task. It is based on the gradient boosting framework and introduces a regularization term to prevent overfitting. At the same time, parallel computing technology is adopted to accelerate the training process, thereby generating a corresponding risk score or probability estimate for each sample point, labeled as P2(c|x). Finally, the trained CatBoost model is applied to make predictions on the same data. This model is particularly good at handling categorical variables, can directly utilize categorical features without additional encoding, and has implemented a special mechanism to reduce the target leakage problem, ensuring the stability and accuracy of the model. It also gives the risk assessment result for each sample point, denoted as P3(c|x). Each model independently completes the credit risk prediction on the test set or new data, and soft voting is used to combine the probabilities of each class in each model, and the class with the highest probability is selected as the final prediction. Suppose there is a base model, each model makes a prediction on a given input and outputs a class probability distribution Pi(c|x), where c is the class and i represents the i-th base model. The final prediction result y of the voting is obtained as follows: where ω i is the weight of the i-th base model, which can usually be set to 1 or adjusted according to the performance of the model. The meaning of the formula is to find the class c that maximizes the sum of the probabilities that all base models predict as class c. The weight assignment is based on the performance of each model on the validation set (such as evaluation metrics like AUC-ROC, precision, recall, etc.), and the complementarity and diversity between models are considered. For example, if a certain model performs outstandingly in predicting high-risk classes, a higher weight can be assigned to it.
[0037] A6. Performance Evaluation and Model Update Decision:
[0038] Regularly or when there are signs of model performance decline, conduct model performance evaluation and compare key metrics such as prediction accuracy and recall between the ensemble model and single models.
[0039] Set a flexible model update mechanism, such as deciding whether to perform model iterative updates based on time periods (e.g., monthly updates), data volume thresholds (e.g., when the new data volume reaches a predetermined ratio), or performance monitoring metrics (e.g., accuracy drops below a certain threshold).
[0040] Regularly or according to pre-set trigger conditions (such as model performance decay, new data volume threshold reached), conduct model performance evaluation and compare the performance of the integrated model with that of the single model. Establish a flexible model update mechanism to ensure that the model is continuously optimized over time and with data changes. This mechanism combines real-time performance monitoring, regular evaluation and feedback loops, trigger conditions based on data volume growth, and an automated iterative training process. Specifically, first establish an online monitoring platform that can track the performance of the model in the production environment in real time and set warning thresholds for key performance indicators; once a significant drop in model accuracy or a predetermined data increment threshold (such as an increase of more than 20% in new samples) is detected, immediately initiate the model retraining process. In addition, conduct a comprehensive performance evaluation quarterly or semi-annually, relying not only on technical indicators but also collecting subjective evaluations from users to supplement quantitative analysis. When a model decline phenomenon is detected, an emergency update process will be triggered. Each time an update is performed, the system will prepare the latest high-quality feature set and adjust the hyperparameter configuration, and conduct a new round of training for each of the three basic models of LightGBM, XGBoost, and CatBoost, evaluate the performance using k-fold cross-validation, save the best version and decide whether to replace the existing model. Before model update, invite domain experts to review the results of the new model and make appropriate corrections to the model output according to the latest business requirements or regulatory requirements to ensure that it conforms to business logic and does not produce unexpected deviations. The entire update mechanism maintains sufficient flexibility to support small-scale trials and exploration of new technical means, thus promoting the long-term healthy development of the model and better serving the risk management needs of financial institutions.
[0041] Figure 3It is an example diagram of an enterprise using this invention for risk prediction. First, data is obtained from multiple sources, and then data cleaning is performed to remove duplicate, incorrect, or irrelevant information to ensure data quality. Next, three pre-constructed basic models, LightGBM, XGBoost, and CatBoost, are called to predict the data. Each model gives its own risk score, and these scores are aggregated to form an overall risk rating. The results output by the models are interpreted to better understand the behavior and prediction results of the models. Based on the risk scores and result interpretations, corresponding decisions are made, such as approving a loan application, rejecting a loan application, or requesting more information. The enterprise also needs to monitor the performance of the models and continuously improve the models according to the feedback to ensure the accuracy of the models. If it is found that the performance of the models has declined, model iteration is carried out to retrain and optimize the models. Throughout the process, the models continuously learn from new data and experiences, thereby continuously improving the accuracy and effectiveness of the predictions.
[0042] Figure 4 It is a schematic diagram of user credit classification provided by this invention. The user group is the input end of the system, representing all potential borrowers, who may be individuals, enterprises, or other types of organizations. The data processing part is responsible for collecting and organizing the personal information and financial data of users and converting them into structured data that can be used by the models. Data processing usually includes steps such as data cleaning, filling in missing values, and handling outliers. Then, predictions are made using three different models used in this invention. These models are all based on gradient boosting machine learning algorithms and can effectively handle complex non-linear relationships and a large number of features. The results predicted by the models are used to evaluate the credit ratings of users, which are divided into three levels: excellent, average, and poor. These results can help financial institutions decide whether to approve a loan application and the amount and interest rate of the loan.
[0043] Figure 5This is the general loan approval flow chart provided by the present invention. The process starts with the applicant submitting a loan application. After receiving the application, the financial institution will conduct a credit risk assessment of the applicant to determine whether they have the ability and willingness to repay the loan. If the applicant meets the assessment requirements, they will proceed to the next step; otherwise, the application will be rejected. If the assessment is passed, both parties will sign a loan contract, stipulating relevant terms such as the loan amount, interest rate, and term. After the contract is signed, the financial institution will disburse the loan to the applicant. After the loan is disbursed, the financial institution will conduct risk management, including monitoring the borrower's repayment situation and preventing and handling possible overdue or default behaviors. After the entire process ends, if there are no other special circumstances, the process ends. In this process, the credit risk assessment is a very crucial link, which determines whether to grant the applicant a loan. If the assessment result shows that the applicant has a high credit risk, then the financial institution may reject their loan application. On the contrary, if the assessment shows that the applicant has good credit, then the process will continue until the loan is disbursed and the subsequent risk management. Even after the loan is disbursed, the financial institution will still continuously monitor the borrower's credit status. Once a problem is discovered, it may take measures to protect its own interests, such as recalling the loan in advance.
[0044] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A credit risk prediction method based on ensemble learning, characterized in that Including: S1. Model construction and initialization: For the credit risk prediction task, construct three basic models of LightGBM, XGBoost, and CatBoost respectively; during initialization, set the learning rate and the depth of the tree to ensure that the model has a good learning foundation when starting, and avoid overfitting or underfitting. S2. Conduct a comprehensive preprocessing of the original dataset, including: data cleaning, outlier handling, and data consistency checking. S3. Conduct feature selection and data balancing processing on the processed data. S4. Conduct detailed and targeted training for the three models of LightGBM, XGBoost, and CatBoost respectively. S5. Independently complete the credit risk prediction of the data through the LightGBM, XGBoost, and CatBoost models after training, and adopt an ensemble learning strategy to integrate the model prediction results. S6. Performance evaluation and model update decision: Regularly or when there are signs of decline in model performance, conduct model performance evaluation, and compare the key indicators of prediction accuracy and recall rate of the three fusion models and single models; set a model update mechanism to decide whether to perform model iteration and update. The update mechanism includes: based on time period, data volume threshold, or performance monitoring indicators.
2. The credit risk prediction method based on ensemble learning according to claim 1, wherein The three basic models of LightGBM, XGBoost, and CatBoost include: The LightGBM adopts a gradient-based one-sided sampling technique and a histogram algorithm to accelerate the construction of decision trees, is suitable for processing large-scale datasets, and automatically processes categorical features. The XGBoost enhances the model performance through an optimized gradient boosting framework, provides a regularization function, effectively controls the model complexity, and reduces overfitting. The CatBoost is specifically optimized to process categorical features, reduces the workload of feature encoding, and introduces a unique feature combination technique to enhance the ability to capture non-linear relationships.
3. A credit risk prediction method based on ensemble learning according to claim 1, characterized in that, Conduct feature selection and data balancing processing on the processed data, including: Use statistical methods to identify and retain features that have a significant impact on credit risk prediction. Solve the class imbalance problem through oversampling, undersampling, or SMOTE technology to ensure the fairness of model training. Create new features according to business logic and feature importance to enhance the model's ability to identify credit risks.
4. The credit risk prediction method based on ensemble learning according to claim 1, wherein, Conduct detailed and targeted training for the three models of LightGBM, XGBoost, and CatBoost respectively, including: S41. Parameter Initialization: Initialize the parameters for each model. For LightGBM, set the maximum number of leaves of the tree, the maximum depth of the tree, the learning rate that determines the degree of weight update after each iteration, the sum of the minimum sample weights of the leaf nodes, the ratio of random sampling for each tree, and the ratio of randomly sampled features for each tree. For XGBoost, set the number of iterations during training, the maximum depth of the tree, the learning rate that determines the degree of weight update after each iteration, the sum of the minimum sample weights of the leaf nodes, the ratio of random sampling for each tree, the ratio of randomly sampled features for each tree. For CatBoost, set the speed of model learning, the number of iterations during training, the maximum depth of the tree, the strength of the regularization term, and the number of feature split points. S42. Hyperparameter Tuning: Use grid search, random search, or Bayesian optimization methods to finely tune the hyperparameters of each model, and evaluate the performance of different parameter combinations through cross - validation, and select the parameter configuration that makes the model perform optimally on the validation set. S43. Model Training: The LightGBM, XGBoost, and CatBoost models find the optimal structural configuration by adjusting these hyperparameters. S44. Training Process Monitoring: During the model training process, continuously monitor the learning curve, the changing trend of the validation loss function, and the performance of the model on the validation set, and promptly detect and adjust possible overfitting or underfitting problems. S45. Model Saving and Evaluation: After training is completed, save the optimal version of each model and evaluate their prediction performance on the test set.
5. The credit risk prediction method based on ensemble learning according to claim 1, wherein, The trained LightGBM, XGBoost, and CatBoost models independently complete the credit risk prediction of the data, and adopt an ensemble learning strategy to integrate the model prediction results, including: S51. Use the trained LightGBM model to predict the input data. By constructing multiple decision trees and using the gradient boosting algorithm to iteratively improve these trees, obtain the risk score or probability distribution for each sample point, denoted as P1(c|x). S52. Use the trained XGBoost model to perform the same prediction task. It is based on the gradient boosting framework and introduces a regularization term to prevent overfitting, and at the same time uses parallel computing technology to accelerate the training process, thus generating the corresponding risk score or probability estimate for each sample point, labeled as P2(c|x). S53. Make predictions on the same data through the already trained CatBoost model. This model is particularly good at handling categorical variables, can directly utilize categorical features without additional encoding, and has implemented a special mechanism to reduce the target leakage problem, ensuring the stability and accuracy of the model, and also giving the risk assessment result for each sample point, denoted as P3(c|x). S54. Each model independently completes the credit risk prediction of the test set or new data, and uses soft voting to combine the probabilities of each class in each model, and selects the class with the highest probability as the final prediction.
6. The credit risk prediction method based on ensemble learning according to claim 5, characterized in that, The soft voting combines the probabilities of each class in each model and selects the class with the highest probability as the final prediction, including: Each model makes a prediction for a given input and outputs a class probability distribution Pi(c|x). The final prediction result y of the voting is obtained as follows: where ω i represents the weight of the i-th base model, c represents the class, i represents the i-th base model, x represents the data to be predicted, and n represents the number of models.