Credit risk prediction method based on data and dynamic feature optimization and Bagging integration

The data set is balanced by the ADASYN-Tomek Links method optimized by genetic algorithm, combined with LightGBM to screen features and use the TabPFN model integrated with Bagging, which solves the problems of data imbalance and feature redundancy in credit risk prediction and improves the accuracy and adaptability of the model.

CN120823029APending Publication Date: 2025-10-21CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510910498.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing credit risk prediction suffers from problems such as data category imbalance, feature redundancy, and insufficient complexity of regression models, which leads to degraded model performance.

Method used

The dataset is balanced using the ADASYN-Tomek Links combination method based on genetic algorithm optimization. The LightGBM algorithm is used to calculate feature importance and features are selected by dynamic thresholding. The TabPFN model is constructed using the Bagging ensemble algorithm, and a dynamic weight allocation mechanism is introduced to optimize the robustness of the model.

Benefits of technology

It improves the accuracy and return rate of credit risk prediction, enhances the robustness and adaptability of the model, and enables timely response to changes in business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823029A_ABST
    Figure CN120823029A_ABST
Patent Text Reader

Abstract

The invention discloses a credit risk prediction method based on data and dynamic feature optimization and Bagging integration, and the method comprises the following steps: S1, obtaining an initial data set, carrying out the data preprocessing, balancing the data set through an ATGA algorithm, and reducing the data noise; s2, using LightGBM as an agent model, calculating importance scores of all features, gradually screening key features, adaptively adjusting feature screening standards according to data distribution through a dynamic threshold algorithm, gradually screening feature subsets, and inputting the screened features into a subsequent model; and S3, constructing a TabPFN model by using a Bagging integration algorithm, introducing a dynamic weight distribution mechanism, adaptively adjusting the weight of each base model according to sample features, and finally realizing credit risk prediction through the TabPFN model. According to the method, the credit risk prediction process is comprehensively optimized in three aspects of data set optimization, feature screening optimization and model performance optimization, so that indexes such as the accuracy rate and the rate of return are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of credit risk prediction, and in particular to a credit risk prediction method based on data and dynamic feature optimization and bagging integration. Background Art

[0002] With the current rapid economic development and rising overall consumption levels, the scale of personal credit business is showing a significant expansion trend. Curbing rising default rates and ensuring the financial security of financial institutions has become a critical issue in the current financial sector. Traditional manual risk assessment methods suffer from high costs and poor scalability in today's big data environment. Early credit assessment methods primarily relied on statistical methods, which constructed models by analyzing data distribution characteristics and patterns.

[0003] The recent rise of machine learning technology can effectively analyze financial data in depth and build risk warning models based on correlations between data features. This intelligent risk assessment system can predict pre-loan risks based on customer information, providing strong support for preventing credit defaults. With the rapid development of deep learning technology, feature selection has become a key research area in credit risk prediction. The appropriate selection and combination of different feature optimization methods can significantly improve model performance and provide higher-quality feature inputs for prediction models. Compared with traditional machine learning methods, models optimized with feature selection show significantly improved performance.

[0004] However, there are still some problems: First, the problem of data category imbalance: default samples usually account for a minority in credit data, which causes the model to tend to predict the majority class and overfit the data; second, feature redundancy: credit data usually contains a large number of redundant features, which leads to excessive model complexity and affects prediction efficiency; third, the regression model's ability to adapt to complexity is insufficient, and the internal relationship of the tabular data may be relatively complex, and may include nonlinear relationships, feature interactions, etc. Summary of the Invention

[0005] To address the above technical issues, the present invention provides a credit risk prediction method based on data and dynamic feature optimization and bagging integration (DFBI-CRP). This method comprehensively optimizes the credit risk prediction process in three aspects: dataset optimization, feature selection optimization, and model performance optimization, thereby improving indicators such as accuracy and return.

[0006] The present invention is implemented by adopting the following technical solution: a credit risk prediction method based on data and dynamic feature optimization and bagging integration, comprising the following steps: S1: Obtain the initial data set, perform data preprocessing, and balance the data set using the ADASYN-TomekLinks combined method ATGA algorithm based on genetic algorithm optimization to reduce data noise; S2: Using LightGBM as a proxy model, calculate the importance scores of all features, design a sequential forward selection strategy, gradually screen key features, and introduce confidence interval indicators. Through the dynamic threshold algorithm, adaptively adjust the feature screening criteria according to the data distribution, gradually screen out feature subsets to ensure the optimization of feature combinations, and input the screened features into subsequent models; S3: Use the Bagging ensemble algorithm to build the TabPFN model and introduce a dynamic weight allocation mechanism to adaptively adjust the weights of each base model according to sample characteristics, optimize the robustness of the ensemble model, and realize credit risk prediction through the TabPFN model.

[0007] Furthermore, step S1 includes the following sub-steps: S11: Obtain credit data through the lending platform, analyze and filter the data, and process missing values ​​to remove features that are obviously invalid in practical terms; S12: The ADASYN-Tomek Links combination method ATGA based on genetic algorithm optimization is used to dynamically generate minority class samples, remove majority class noise samples, and optimize the parameters of the combination method according to the genetic algorithm to achieve the purpose of optimizing data.

[0008] Furthermore, the combined method ATGA comprises the following processing steps: Input initial data and randomly generate multiple individuals; Combine ADASYN oversampling and Tomek Links undersampling for each population individual to generate a new training set; Train the classification model on the new training set, calculate the fitness value of the individual, and perform crossover and mutation operations based on the individual fitness; After multiple iterations, the individual with the highest fitness is selected as the optimal parameter combination [; The optimal parameter combination is used to oversample and undersample the original minority class samples to obtain the final optimized data sample set.

[0009] Furthermore, the initial data includes one or more of an input initial data sample set, a category label vector, a maximum number of iterations of a genetic algorithm, a population size, a k value in ADASYN, and an operating parameter of Tomek Links.

[0010] Furthermore, the dynamic feature screening method includes variable importance ranking and sequential forward selection.

[0011] Furthermore, the variable importance ranking adopts the LightGBM model as a proxy model to train the training dataset and target variables, and calculates the importance score of each feature after the training is completed.

[0012] Furthermore, after ranking the variable importance, a sequential forward selection strategy is adopted. Initially, the candidate feature set is empty. Subsequently, the features with the highest importance are selected from the unselected features and added to the candidate feature set, and the quality of the current candidate feature set is evaluated using K-fold cross-validation. For each candidate feature, after adding it to the current candidate feature set, the LightGBM model is used to train it on the training set and the AUC score is calculated. This process is repeated until all candidate features are evaluated.

[0013] Furthermore, a dynamic threshold adjustment mechanism based on confidence intervals is introduced to adaptively adjust the feature selection criteria according to the data distribution during the sequential forward selection process.

[0014] Furthermore, the TabPFN model is constructed by combining the TabPFN small table data basic model with the Bagging integration method.

[0015] Furthermore, the calculation method of the TabPFN model is: ; Where, P final is the final prediction result, w i It is i The weight of each sub-model, P i It is i The predicted probability of each sub-model.

[0016] The beneficial effects of the present invention are: The present invention first optimizes data based on the data optimization method ATGA of ADASYN-Tomek Links optimized by genetic algorithm; then adopts a dynamic feature screening method, uses LightGBM as a proxy model to calculate feature importance, and combines sequential forward selection with a dynamic threshold method to screen feature subsets; finally, constructs a TabPFN model pool through bagging. During the training process of each sub-model, a dynamic weight allocation mechanism is designed to record the AUC performance index of the sub-model on the training set and the validation set as the basis for initial weight allocation. The higher the AUC value of the sub-model, the greater its initial weight. The combination of dynamic weight allocation and complementary measurement for model training and prediction significantly improves the accuracy of credit risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0018] Figure 1 Flowchart of the present invention; Figure 2 Generate a flow chart for TabPFN data; Figure 3 TabPFN pre-training flowchart. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0020] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0021] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0022] See also Figure 1This credit risk prediction method, based on data and dynamic feature optimization integrated with bagging, includes three steps: data processing and optimization, dynamic feature optimization, and model integration training and prediction. Data processing and optimization involves analyzing and screening the data, handling missing values, and removing features that are clearly invalid in practical terms. To balance the dataset and reduce data noise, the ADASYN-Tomek Links combination method (ATGA) optimized based on a genetic algorithm is used to dynamically generate minority class samples and remove majority class noise samples. The genetic algorithm is then used to optimize the combination method's parameters, ensuring a relatively balanced number of positive and negative samples in the data while optimizing noise samples. This achieves the goal of optimizing data and thus improves training effectiveness.

[0023] It's important to note that ADASYN is an oversampling technique for imbalanced datasets. It generates synthetic samples for minority class samples to improve the dataset's balance. The core idea of ​​ADASYN is to dynamically generate synthetic samples of varying degrees for different minority class samples, based on the difficulty of the minority class sample (i.e., the density of surrounding majority class samples). This adjusts the data distribution while enhancing the classifier's ability to learn minority class samples, improving the model's performance and generalization capabilities on imbalanced datasets and enabling the classifier to more accurately identify minority class samples.

[0024] The Tomek Links method is an undersampling technique for processing imbalanced datasets. By identifying and removing majority class samples from sample pairs located on the boundary between categories, the boundaries between categories are made clearer, thereby improving the performance of the classifier. First, the distance between all pairs of samples in the dataset is calculated based on the Manhattan distance to construct a distance matrix. For each sample, its nearest neighbor sample is found according to the distance matrix. Then, the category labels of each sample and its nearest neighbor sample are checked. If two samples belong to different categories, they form a Tomek Link with each other. For each identified Tomek Link, the sample belonging to the majority class is removed. If both samples belong to the majority class or the minority class, no operation is performed. The samples that are not removed are retained to form a new dataset.

[0025] Genetic Algorithm (GA) is a process of finding the optimal solution by simulating the process of biological evolution in nature. The basic operation process of genetic algorithm is: 1) Randomly generate an initial population consisting of multiple individuals. Each individual represents a possible solution to the problem; 2) For each individual in the population, calculate the fitness value based on its adaptability to the problem goal; 3) Based on the fitness of each individual, some individuals are selected as parents to produce the next generation. Individuals with higher fitness have a greater probability of being selected, which allows superior individuals to have more opportunities to pass on their traits to future generations. For example, in the roulette wheel selection method, the probability of an individual being selected is proportional to its fitness. Individuals with higher fitness occupy a larger area on the wheel and have a greater chance of being selected. The formula for the roulette wheel selection method is:

[0026] ; Where i is an individual in the population, is the probability of being selected, is the fitness of individual i, and N is the population size.

[0027] 4) Pair the selected parent individuals in pairs, exchange and combine the codes of each pair of individuals with a certain crossover probability to generate new individuals; 5) Change certain gene positions of newly generated individuals with a certain mutation probability; 6) Combine the new individuals generated by the selection operation, crossover operation and mutation operation to form a new population.

[0028] Dynamic feature optimization adopts a dynamic feature screening mechanism based on LightGBM, takes LightGBM as the proxy model, calculates the importance scores of all features, designs a sequential forward selection strategy, gradually screens key features, introduces confidence interval indicators, and adds a dynamic threshold method. It adaptively adjusts the feature screening criteria according to the data distribution, gradually screens out feature subsets, ensures the optimization of feature combinations, and improves the robustness of feature subsets. Finally, the screened features are input into subsequent models.

[0029] Model ensemble training and prediction uses the TabPFN (Tabular Prior-data Fitted Network) bagging ensemble method to achieve final predictions. This method combines TabPFN's small and medium-sized tabular data base models with the bagging ensemble method to build a TabPFN model pool, enhancing model diversity and generalization. A dynamic weight allocation mechanism is also introduced to adaptively adjust the weights of each base model based on sample characteristics, optimizing the robustness of the ensemble model. TabPFN can quickly learn from data and generate models. When data is updated or business scenarios change, it can effectively handle different data distributions and feature combinations, allowing for rapid retraining and deployment of new models to promptly respond to changing business needs and ensure model timeliness and accuracy.

[0030] In this embodiment, the data is sourced from the public data of the P2P lending platform LendingClub from 2007 to 2020. To ensure the diversity of sample characteristics and avoid the influence of the time span, 887,379 relatively new data entries were selected for the research of the credit prediction model, with each data entry containing 74 features. In the data preprocessing stage, first, the features of the dataset were screened to remove factors that were obviously meaningless for model prediction, and then missing value handling was performed, retaining features with less than 20% missing values. Borrowers who repaid on time (Fully Paid) were regarded as performing customers, and borrowers with bad debts (Charged Off) were regarded as default customers. Subsequently, the data of the object type was encoded and transformed into features that the machine could understand. Finally, Z-score outlier detection was performed to limit the outliers. After data preprocessing, statistical analysis found that the ratio of the two categories of on-time repayment and bad debts was approximately 8:2, and the dataset was severely imbalanced. The processed data was output, and the number of effective data samples was 144,939, retaining 39 features.

[0031] The process of the ADASYN-Tomek Links combination method optimized based on the genetic algorithm is as follows: Input the initial data sample set D, the class label vector y, the maximum number of iterations max_gen of the genetic algorithm, the population size NP, the k value in ADASYN, and the Tomek Links operation parameters.

[0032] Randomly generate NP individuals, and each individual represents a possible parameter combination of ADASYN and Tomek Links.

[0033] When gen < max_gen: Combine the ADASYN oversampling and Tomek Links undersampling for each population individual to generate a new training set.

[0034] Train the classification model on the new training set, and calculate the fitness value (using accuracy as the fitness value) fitness of the individual.

[0035] According to the fitness of the individuals, use the roulette wheel selection method to select NP individuals as the parental population. Calculate the probability Pi = fitness_i / sum(fitness) that each individual is selected, where sum(fitness) is the sum of the fitness of all individuals.

[0036] For each pair of individuals (i, j) in the parental population, perform a crossover operation with the crossover probability Pc to generate two new offspring individuals.

[0037] For each individual in the offspring population, a mutation operation is performed with a mutation probability Pm, and a value of each parameter is randomly selected from a small integer range near it as the new k value.

[0038] The offspring population after crossover and mutation operations is used as the new population to enter the next iteration.

[0039] After max_gen iterations, the individual with the highest fitness is selected as the optimal parameter combination [k_adasyn_best, threshold_tomek_best].

[0040] The ADASYN-Tomek Links algorithm is used to oversample and undersample the original minority class samples using the optimal parameter combination to obtain the final optimized data sample set D_opt.

[0041] Output the optimized data sample set D_opt. The algorithm ends.

[0042] After the above data processing, statistical analysis of the two categories of fully paid and charged off was conducted again. The ratio of the two categories was found to be 3:2, which made the data more balanced. Most of the unreliable samples were optimized and could be used for model training. In this example, the data was divided into training and test sets with a ratio of 7:3.

[0043] To address the high complexity and poor model performance caused by redundant features in credit data, a dynamic feature screening method is proposed. Dynamic feature screening consists of two steps: variable importance ranking and sequential forward selection.

[0044] The LightGBM model is used as a proxy model for variable importance ranking. X and the target variable y The model is trained and the importance score of each feature is calculated. After training, the feature importance scores of the model are extracted and a feature importance ranking table is constructed. To simplify subsequent analysis, this example only selects the top 30 features as the candidate feature set, and displays the importance of these features through a visual bar chart.

[0045] After ranking the variables by importance, a sequential forward selection strategy is adopted. Initially, the candidate feature set is empty. Subsequently, the most important features are selected from the unselected features and added to the candidate feature set, and the quality of the current candidate feature set is evaluated using K-fold cross validation. For each candidate feature, after adding it to the current candidate feature set, the LightGBM model is used to train on the training set and the AUC score is calculated. This process is repeated until all candidate features are evaluated. The average AUC score is calculated as follows: ; Where K is the number of cross-validation folds. By averaging the AUC scores, the standard error can be calculated SE , the standard error reflects the degree of dispersion of each ROC and AUC score relative to the mean value, SE The calculation method can be expressed as: .

[0046] Next, we calculate the confidence interval indicator. This indicator allows the feature selection criteria to no longer rely solely on a single average performance indicator, but to comprehensively consider the stability and reliability of the performance. The calculation of the confidence interval can be expressed as: ; in, In statistics t The critical value of the distribution, which represents the t For a given significance level in the distribution α and degrees of freedom K −1 corresponds to the critical value of the two-tailed test. Here, the widely used significance level of 0.05 is used in statistics, and the corresponding confidence level is 95%.

[0047] To improve the robustness of feature subsets, a dynamic threshold adjustment mechanism based on confidence intervals was introduced. During the sequential forward selection process, the feature selection criteria were adaptively adjusted based on the data distribution. During feature selection, an initial feature selection threshold was set. If the average AUC score plus twice the standard error of the feature being evaluated was greater than the initial dynamic threshold, the feature was considered to have potential for inclusion in the feature subset. As the feature subset grew, the threshold was dynamically adjusted based on factors such as changes in model performance during the feature selection process. To further improve the stability of feature selection, the confidence intervals of multiple features were comprehensively considered. For an existing feature subset, when considering adding a new feature, the mean and standard error of the model AUC score of the new feature combined with each feature in the feature subset were calculated. These statistics were used to determine whether to add the new feature. A composite index, m, was defined as the ratio of the average AUC score after the new feature was added to the width of the confidence interval after the new feature was added. The threshold was dynamically adjusted based on the magnitude of this composite index, and the decision on whether to add the new feature was made. A larger composite index value indicates that the new feature has greater stability and reliability in terms of performance improvement.

[0048] The process of the dynamic feature screening method based on sequential forward selection and confidence interval is as follows: Input the training dataset D, the feature list F after variable importance sorting, the initial dynamic threshold threshold, the unselected feature list F_unused, and set the comprehensive index m.

[0049] Initialize the candidate feature set S, and the final feature subset S_final is empty.

[0050] When (F_unused ≠ empty set): Take features f from F_unused in order of importance and add them to S.

[0051] Use LightGBM to train the model on the feature S of the training set, calculate the average AUC score and standard error SE, and calculate the range and width of the confidence interval.

[0052] If (mean AUC + 2×SE) > threshold: Add feature f to the final feature subset S_final.

[0053] Remove feature f from F_unused.

[0054] Otherwise, remove feature f from F_unused.

[0055] m = mean AUC / confidence interval width Adjust the threshold based on the historical best m value.

[0056] Output the final feature subset S_final.

[0057] Based on the average AUC score and confidence interval, the feature selection threshold was dynamically adjusted to ensure the stability of the feature subset under different data distributions. Ultimately, the top 18 features with the best performance were selected as the final feature subset based on the average AUC score and confidence interval, providing efficient and robust feature input for the credit risk prediction model.

[0058] In the early stages of screening, the dynamic threshold method can set a relatively low threshold to quickly collect potential features. As long as the importance score of a feature reaches a certain basic level, it is allowed to enter the feature subset. As the feature subset gradually increases, the threshold standard is raised, and subsequently added features are screened more strictly to ensure that each newly added feature can significantly and stably improve the already relatively stable model performance, avoiding the overfitting problem that may be caused by selecting too many features at one time, and at the same time being able to more accurately determine which features are truly key features that help improve model performance.

[0059] To improve the overall prediction, effectiveness, and performance of the model for credit risk prediction, this example uses the TabPFN model for credit risk prediction and integrates the TabPFN model using the bagging method. The core concept of TabPFN is to create a large number of simulated synthetic tabular datasets and train a Transformer-based neural network to solve the simulated prediction problem. Traditional methods require manual design of solutions to deal with data challenges such as missing values. TabPFN automatically learns and generates effective strategies by solving simulated tasks that incorporate these challenges. TabPFN's performance relies on generating suitable synthetic training datasets that capture the characteristics and challenges of real-world tabular data. To generate such datasets, an approach based on structural causal models (SCMs) is used. SCMs provide a formal framework for representing the causal relationships and generative processes behind data. By using synthetic data rather than large amounts of publicly available tabular data, the underlying model avoids problems such as training data contamination by test data or limited data availability. TabPFN's underlying algorithmic process can be divided into two key steps: data generation and pre-training.

[0060] The data generation process of TabPFN (also called prior process) is as follows Figure 2 As shown, dataset hyperparameters, including dataset size, number of features, and difficulty level, are first sampled to control the overall properties of each synthetic dataset. Guided by these hyperparameters, a directed acyclic graph is constructed to represent the dataset's causal structure, where each node represents a variable and each edge denotes the causal relationship between variables. To generate each sample in the dataset, randomly generated initialization data (i.e., noise data) is propagated through the root node of the causal graph. The initialization data is sampled from a random normal or uniform distribution, with varying degrees of non-independence between samples. As this data propagates through the edges of the computational graph, a series of different computational mappings are applied, including small neural networks with linear or nonlinear activation functions (such as sigmoid, ReLU, modulo, and sine), discretization mechanisms for generating categorical features, decision tree structures for encoding local rule-based dependencies, and Gaussian noise is added to each edge to introduce uncertainty into the generated data. The intermediate data representations at each node are then saved for subsequent retrieval.

[0061] After traversing the causal graph, intermediate representations are extracted at the sampled feature nodes and target nodes, resulting in a sample consisting of feature values ​​and associated target values. Finally, a variety of post-processing techniques are applied to generate synthetic datasets. During the synthetic data generation process, missing values ​​are introduced, exposing TabPFN to synthetic datasets with different missing value patterns and proportions, allowing the model to learn effective methods for handling missing values. Feature transformations are performed using the Kumaraswamy distribution, introducing complex nonlinear distortions and simulating the quantization of discrete features to enhance data authenticity. By incorporating various data challenges and complexities into synthetic datasets, TabPFN is able to develop strategies for handling similar problems in real-world datasets. The entire generation process creates approximately 100 million synthetic datasets per model training session, each with unique causal structures, feature types, and functional characteristics.

[0062] TabPFN pre-training is to train a Transformer-based neural network architecture based on a large number of synthetic datasets generated during the data generation process, so that it can adapt to the diverse data features and prediction challenges in the synthetic datasets. The pre-training process first inputs the generated synthetic datasets into the TabPFN neural network for forward propagation. During the training process, the network is trained based on the input training dataset. X train and Y train Learn patterns and relationships in the data and make predictions on the test set X test Corresponding Y test , while calculating the training loss function, the loss function L(θ) It can be expressed as: ; in p represents the conditional probability distribution of the model, θ The parameters of TabPFN are a set of learnable parameters such as the weights and biases of the model. θ Will be continuously adjusted to optimize model performance. Then use the gradient descent algorithm to gradually update the parameters according to the gradient information of the loss function. θ , in each iteration, the loss function is calculated θ The gradient of , and then adjust according to a certain learning rate and optimization strategy θ The value of is used to reduce the value of the loss function. This process is repeated millions of times to ensure that the model can fully learn and generalize well on a large amount of synthetic data. Pre-training is performed only once in the entire model development process. The purpose is to enable the model to learn a general learning algorithm that can be used to predict any data set. The pre-training process of the TabPFN model is as follows: Figure 3After pre-training, the trained model is applied to the real dataset, and the training samples are provided to the model as context. The model predicts the labels of the unknown dataset through context learning.

[0063] The Bagging ensemble method trains the model using multiple different sub-datasets. Using the Bootstrap sampling method, samples are randomly extracted from the original dataset with replacement to generate multiple sub-datasets. Each sub-dataset is of the same size and maintains a similar distribution to the original dataset. A base learner is trained for each subset, and the prediction results of these base learners are combined to obtain the final prediction result. Since TabPFN is typically used for small to medium-sized datasets, and performs particularly well when the sample size does not exceed 10,000 and the number of features does not exceed 500, the number of sub-datasets is limited to 10,000. Each subset may contain noise or some special cases. When the prediction results of multiple base learners are combined, the effects of this noise are offset, making the final ensemble model less sensitive to noise in the data.

[0064] During the training process of each sub-model, a dynamic weight allocation mechanism is designed to record the AUC performance index of the sub-model on the training set and the validation set as the basis for initial weight allocation. The sub-model with a higher AUC value has a larger initial weight. The weight initialization formula can be expressed as: ; in, w i Indicates the i The initial weights of the sub-models, AUC i Indicates the i The AUC value of the sub-model on the validation set, n Represents the total number of sub-models. After each prediction on new data, the performance of each sub-model is re-evaluated and the weights are adjusted based on the new performance indicators. If a sub-model performs better than other models on the new data, its weight is increased; otherwise, its weight is decreased. The weight calculation formula can be expressed as: ; in, α It is a coefficient that retains the original weight, with a value between 0 and 1, used to maintain the historical performance influence of the model; β is the coefficient for updating weights, usually 1− α , used to introduce the weight of new data evaluation; AUC i It is a sub-model iThe latest AUC value on the new data validation set. After each weight adjustment, the weights of all sub-models are normalized to ensure that the sum of the weights is 1. When new credit data needs to be predicted, the data is input into the TabPFN model separately. Each sub-model will output a prediction result. Based on the dynamically adjusted weights, the prediction results of each sub-model are weighted and fused. The final prediction result is expressed as the weighted sum of the prediction probabilities of each sub-model, which can be described as: ;in, P final is the final prediction result, w i It is i The weight of each sub-model, P i It is i The predicted probability of each sub-model.

[0065] By integrating multiple TabPFN base models, we can capture diverse feature combinations and nonlinear relationships, and better fit complex data patterns. This approach improves the model's predictive accuracy and enhances its stability and diversity. By dynamically adjusting weights, we can leverage the strengths of each sub-model in specific situations, improving the generalization of the overall model and ultimately enhancing the predictive performance of the entire ensemble.

[0066] In order to further demonstrate the technical advantages of the present invention, the following experiments were conducted The experimental platform used a server consisting of an Intel Xeon Platinum 8358P 2.6GHz processor, an NVIDIA GeForce RTX 3090 GPU, and 90.0GB of RAM. JDK 1.8, Python 3, and the Pytorch framework were used as the operating environment for data screening and cleaning, as well as credit risk prediction. The experimental data was derived from LendingClub's public data from 2007 to 2020. The processed dataset contained 144,938 samples and 39 features. It was divided into training and test sets in a ratio of 7:3. The experiment selected the accuracy ( Accuracy ), recall rate (Recall), F 1 -score And AUC scores are the evaluation indicators.

[0067] Experimental results To test the performance of the ATGA method in DFBI-CRP, the most commonly used SMOTE method, ADASYN method, and BorderlineSMOTE method were selected for comparison with the ATGA method. The method was combined with TabPFN and trained on 10,000 samples randomly sampled from the full feature set. The data optimization effect was reflected by comparing the model performance. The results are shown in Table 1: Table 1 Performance comparison of ATGA and other data optimization methods Model Accuracy Recall F1-Score AUC TabPFN 0.9302 0.9336 0.9361 0.8834 SMOTE+ TabPFN 0.9412 0.9423 0.9554 0.9018 ADASYN+ TabPFN 0.9369 0.9521 0.9550 0.9067 BorderlineSMOTE+TabPFN 0.9497 0.9588 0.9567 0.9091 ATGA+ TabPFN 0.9671 0.9671 0.9683 0.9690

[0068] From the various indicators in the table, the ATGA+TabPFN method has the advantages of accuracy, recall, F The ATGA method achieved the highest values ​​in all four evaluation metrics, including 1-Score and AUC, with an accuracy of 0.9671. Compared to traditional methods, the ATGA method is an innovative data preprocessing method for addressing imbalanced datasets. First, ADASYN is used to synthetically sample minority class samples, increasing their number and improving dataset balance. Tomek Links is then used to detect and remove noise samples that occur in pairs on the boundary between different classes, further optimizing the data structure. A genetic algorithm plays a key optimization role. By simulating mechanisms such as natural selection and genetic variation in biological evolution, it searches for and optimizes key parameters in the combined method, such as determining the appropriate ratio of synthetic samples generated in ADASYN and the processing range of Tomek Links. This makes the combined method more efficient and accurate in addressing data imbalance, ultimately improving the model's overall performance and generalization ability in various application scenarios, providing a higher-quality data foundation for subsequent machine learning tasks. Experiments have demonstrated the ATGA method's advantages and applicability in addressing data optimization problems.

[0069] To test the performance of the DFBI-CRP method, we selected the benchmark models Random Forest, Decision Tree, XGBoost, KNN, Extra-Trees, and Log-CNN for comparison and tested them on the LendingClub dataset. The results are shown in Table 2: Table 2 Performance comparison of DFBI-CRP method and other methods Model Accuracy Recall F1-Score AUC Random Forest 0.9367 0.9372 0.9381 0.9353 Decision Tree 0.9257 0.9317 0.9330 0.9215 XGBoost 0.9466 0.9470 0.9470 0.9454 KNN 0.9111 0.9205 0.9297 0.8798 Extra-Trees 0.9365 0.9372 0.9379 0.9285 Log-CNN 0.9685 0.9788 0.9562 0.9673 DFBI-CRP 0.9871 0.9870 0.9883 0.9872

[0070] On the LendingClub dataset, the DFBI-CRP method achieved an accuracy of 0.9871, demonstrating superior accuracy and overall performance compared to other methods. Leveraging the advantages of data optimization, dynamic feature optimization, and integrated models, the DFBI-CRP method achieves dynamic synergy among data, features, and models. This method can adapt to external changes such as credit policy changes and economic fluctuations, while also improving prediction accuracy and efficiency through inter-module synergy and complementarity, creating a superior risk control strategy for credit risk prediction. This provides an effective approach for financial risk control in scenarios such as bank approvals and post-loan monitoring.

[0071] We conducted an ablation experiment on the LendingClub dataset for comparison. We removed the ATGA part of the model, the dynamic feature optimization part, and the model integration part. We used the TabPFN model for training and prediction and compared the performance to reflect the contribution of each part to the DFBI-CRP method. The results are shown in Table 3. Table 3 Ablation experiment Model Accuracy Recall F1-Score AUC -ATGA 0.9437 0.9406 0.9455 0.9426 -Dynamic feature optimization 0.9508 0.9475 0.9564 0.9568 -Bagging integration 0.9538 0.9593 0.9580 0.9586 DFBI-CRP 0.9871 0.9870 0.9883 0.9872

[0072] Ablation experiments show that the full version of the DFBI-CRP method performs best across all metrics. Removing ATGA, the dynamic feature optimization method, and the bagging ensemble method reduces accuracy by 4.34%, 3.63%, and 3.33%, respectively, compared to the full model. This demonstrates that all three components contribute significantly to the overall model performance. Overall, the DFBI-CRP method significantly improves its effectiveness in credit risk prediction through these optimizations.

[0073] For the sake of simplicity, the aforementioned embodiments are described as a series of actions. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions involved are not necessarily required by this application.

[0074] The above embodiments describe the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Without departing from the spirit and scope of the present invention, modifications and variations made by those skilled in the art without departing from the spirit and scope of the present invention should be within the scope of protection of the appended claims.

Claims

1. A credit risk prediction method based on data and dynamic feature optimization integrated with bagging, characterized by: The steps include: S1: Obtain the initial data set, perform data preprocessing, and balance the data set using the ADASYN-TomekLinks combined method ATGA algorithm based on genetic algorithm optimization to reduce data noise; S2: Using LightGBM as a proxy model, calculate the importance scores of all features, design a sequential forward selection strategy, gradually screen key features, and introduce confidence interval indicators. Through the dynamic threshold algorithm, adaptively adjust the feature screening criteria according to the data distribution, gradually screen out feature subsets to ensure the optimization of feature combinations, and input the screened features into subsequent models; S3: Use the Bagging ensemble algorithm to build the TabPFN model and introduce a dynamic weight allocation mechanism to adaptively adjust the weights of each base model according to sample characteristics, optimize the robustness of the ensemble model, and realize credit risk prediction through the TabPFN model.

2. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 1, characterized in that: Step S1 includes the following sub-steps: S11: Obtain credit data through the lending platform, analyze and filter the data, and process missing values ​​to remove features that are obviously invalid in practical terms; S12: The ADASYN-Tomek Links combination method ATGA based on genetic algorithm optimization is used to dynamically generate minority class samples, remove majority class noise samples, and optimize the parameters of the combination method according to the genetic algorithm to achieve the purpose of optimizing data.

3. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 2, characterized in that: The combined method ATGA comprises the following processing steps: Input initial data and randomly generate multiple individuals; Combine ADASYN oversampling and Tomek Links undersampling for each population individual to generate a new training set; Train the classification model on the new training set, calculate the fitness value of the individual, and perform crossover and mutation operations based on the individual fitness; After multiple iterations, the individual with the highest fitness is selected as the optimal parameter combination [; The optimal parameter combination is used to oversample and undersample the original minority class samples to obtain the final optimized data sample set.

4. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 3, characterized in that: The initial data includes one or more of an input initial data sample set, a category label vector, a maximum number of iterations of a genetic algorithm, a population size, a k value in ADASYN, and an operating parameter of Tomek Links.

5. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 1, characterized in that: The dynamic feature screening method includes variable importance ranking and sequential forward selection.

6. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 5, characterized in that: The variable importance ranking adopts the LightGBM model as a proxy model to train the training dataset and target variables, and calculates the importance score of each feature after the training is completed.

7. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 6, characterized in that: After ranking the variable importance, a sequential forward selection strategy is adopted. Initially, the candidate feature set is empty. Subsequently, the features with the highest importance are selected from the unselected features and added to the candidate feature set. The quality of the current candidate feature set is evaluated using K-fold cross-validation. For each candidate feature, after adding it to the current candidate feature set, the LightGBM model is used to train it on the training set and the AUC score is calculated. This process is repeated until all candidate features are evaluated.

8. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 7, characterized in that: A dynamic threshold adjustment mechanism based on confidence intervals is introduced to adaptively adjust the feature selection criteria according to the data distribution during the sequential forward selection process.

9. The credit risk prediction method based on data and dynamic feature optimization and bagging integration as claimed in claim 1, characterized in that: The TabPFN model is constructed by combining the TabPFN small table data basic model with the Bagging integration method.

10. The credit risk prediction method based on data and dynamic feature optimization and bagging integration according to claim 1, characterized in that: The calculation method of the TabPFN model is: ; Where, P final is the final prediction result, w i It is i The weight of each sub-model, P i It is i The predicted probability of each sub-model.

Citation Information

Cited By

  • Foundation pit deformation control method based on TabPFN model and supporting axial force servo system

    CN120995802A

  • Classification and grading early warning method for traffic potential safety hazards of key vehicles on expressway

    CN121686785A