Forest gradient tree-based manufacturing industry personal credit investigation credit assessment method

By adopting the forest gradient tree method in personal credit credit assessment, using a random forest and XGBoost model combined with a logistic regression meta-learner to dynamically adjust the model weight and optimize residual modeling, the problems of insufficient model fusion strategy, limited residual optimization and lack of adaptability in the existing technology are solved, and more efficient and accurate credit evaluation is achieved.

CN120163644APending Publication Date: 2025-06-17LONGYING ZHIDA (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340190.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, there are problems such as insufficient model fusion strategy, limited residual optimization and lack of adaptability in personal credit credit assessment, and it is difficult to dynamically adjust model weights and optimize residual modeling.

Method used

The personal credit credit evaluation method for manufacturing industry using forest gradient trees is used to make preliminary predictions through a random forest model, residuals are calculated and residuals are fitted using XGBoost regression model, and the model weight is dynamically adjusted with logistic regression as a meta-learner, and an adaptive learning rate strategy is adopted.

Benefits of technology

By dynamically adjusting the model weights and optimizing residual modeling, the prediction performance and adaptability of the credit evaluation model are improved, and the model's fitting ability and accuracy of complex data are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163644A_ABST
    Figure CN120163644A_ABST
Patent Text Reader

Abstract

The invention discloses a manufacturing industry personal credit investigation credit assessment method based on a forest gradient tree. The credit assessment method comprises the following steps: data preprocessing; performing preliminary prediction on the data by using a random forest model to obtain a prediction probability of each sample; calculating the residual error of the random forest model, fitting the residual error by using an XGBoost regression model, and compensating the error of the random forest model; logistic regression is used as a meta-learner, prediction results of the random forest and the XGBoost are used as input features, and the weight of the model is dynamically adjusted by training the meta-learner; dynamically adjusting the weight of each basic model in the integration process by adopting an adaptive learning rate strategy to obtain a prediction model; and evaluating the personal credit investigation credit of the manufacturing industry according to the prediction model. The part which cannot be fitted by the random forest can be captured, and the expression ability of the model is enhanced. Through XGBoost fitting residual errors, complex modes in the data can be better captured, and the accuracy of the model is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of personal credit assessment, and particularly to a method for manufacturing personal credit assessment using forest gradient trees. Background Art

[0002] The disadvantages of the prior art include:

[0003] Insufficient model fusion strategy: Weighted fusion and stacking methods fail to effectively combine the outputs of multiple models and are difficult to dynamically adjust weights.

[0004] Limited residual optimization: Lack of an adaptive fusion framework for multi-model residuals, affecting prediction performance.

[0005] Lack of adaptability: The model is not optimized in real time for the dynamic data characteristics during the training process.

[0006] Although certain achievements have been made in the fields of ensemble learning and residual modeling in the prior art, there are still significant deficiencies in how to more efficiently and adaptively fuse the outputs of multiple models, dynamically adjust model weights, and optimize residual modeling.

[0007] Traditional ensemble learning methods (such as stacking, weighted voting, etc.) usually adopt static weight settings during model fusion, which causes the model to be unable to adaptively adjust the importance of each model when facing different datasets or tasks.

[0008] In many ensemble learning methods, residual modeling usually relies on the correction of a single model and ignores the fusion of residuals between different models. The residual correction effect of a single model is limited, and in complex tasks, the errors of a single model may not be fully compensated, resulting in a still large error in the final prediction. Summary of the Invention

[0009] In view of the above problems, the present invention is proposed to provide a method for manufacturing personal credit assessment using forest gradient trees that overcomes or at least partially solves the above problems.

[0010] According to one aspect of the present invention, there is provided a method for manufacturing personal credit assessment using forest gradient trees, the credit assessment method comprising:

[0011] Data preprocessing;

[0012] Using a random forest model to make a preliminary prediction on the data to obtain the prediction probability of each sample;

[0013] Calculating the residuals of the random forest model and using an XGBoost regression model to fit the residuals to compensate for the errors of the random forest model;

[0014] Using logistic regression as the meta-learner, taking the prediction results of random forest and XGBoost as input features, and training the meta-learner to dynamically adjust the weights of the model;

[0015] Adopting an adaptive learning rate strategy to dynamically adjust the weights of each base model during the integration process to obtain a prediction model;

[0016] Evaluating the personal credit of manufacturing enterprises according to the prediction model.

[0017] Optionally, calculating the residuals of the random forest model specifically includes: for each prediction sample, calculating the difference between the predicted probability of the random forest and the actual value as the residual.

[0018] Optionally, the data preprocessing specifically includes:

[0019] Data extraction, extracting the personal credit and credit-related tables of legal persons of manufacturing enterprises from the data warehouse;

[0020] Feature processing, processing feature indicators according to the logic mapping defined by the business and handling missing values;

[0021] Data annotation, generating target variables based on business requirements;

[0022] Data splitting, dividing it into a training set and a test set.

[0023] Optionally, using the random forest model to make a preliminary prediction on the data specifically includes:

[0024] Training the random forest model, including: using the random forest model to train the training data to obtain the predicted probability of each sample

[0025] Optionally, calculating the residuals of the random forest model and using the XGBoost regression model to fit the residuals to compensate for the errors of the random forest model specifically includes:

[0026] Calculating the residuals, including: calculating the true label y i and the prediction result of the random forest to obtain the difference between them as the residual

[0027] Using the XGBoost model to train the residuals and further improving the prediction ability of the model by minimizing the residuals;

[0028] Stacking the random forest prediction probability and the prediction result of XGBoost for the residuals into a new feature matrix:

[0029]

[0030] Optionally, the random forest model algorithm specifically includes:

[0031] Training phase: The random forest generates each decision tree T through the following steps i :

[0032] T i = Train(Subset(X, y, B i ))

[0033] where X is the training data, y is the target variable, and B i is a subset randomly sampled from the training set (X, y), and sampling allows repetition;

[0034] Prediction phase: The prediction result of each tree T i for the input data X is The final prediction result is obtained through majority voting or averaging:

[0035]

[0036] where N is the number of decision trees in the random forest, and T i (X) is the predicted value of each decision tree.

[0037] Optionally, the XGBoost regression model algorithm specifically includes:

[0038] Loss function: The goal of XGBoost is to minimize the loss function L, which is usually the difference between the model prediction and the true label on the training set;

[0039] Assume the loss function is:

[0040]

[0041] where is the loss function, and Ω(f k ) is the regularization term to prevent overfitting;

[0042] At each step t, the model is adjusted by calculating the residuals and using gradient boosting:

[0043]

[0044] where is the prediction result of the t-th iteration, f t (X) is the base model learned in the t-th iteration, and η is the learning rate.

[0045] A method for manufacturing personal credit assessment of forest gradient trees provided by the present invention, the credit assessment method comprising: data preprocessing; using a random forest model to make a preliminary prediction on the data to obtain the prediction probability of each sample; calculating the residual of the random forest model, and using an XGBoost regression model to fit the residual to compensate for the error of the random forest model; using logistic regression as a meta-learner, taking the prediction results of the random forest and XGBoost as input features, and dynamically adjusting the weights of the model by training the meta-learner; adopting an adaptive learning rate strategy to dynamically adjust the weights of each base model during the integration process to obtain a prediction model; evaluating the manufacturing personal credit according to the prediction model. It can capture the part that the random forest fails to fit, enhancing the performance ability of the model. By fitting the residual with XGBoost, it can better capture the complex patterns in the data and optimize the accuracy of the model.

[0046] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 It is a flowchart of a method for manufacturing personal credit assessment of forest gradient trees provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0050] The terms "including" and "having" and any variations thereof in the description of the embodiments, claims and drawings of the present invention are intended to cover non-exclusive inclusion. For example, a series of steps or units are included.

[0051] The technical solutions of the present invention will be further described in detail below in conjunction with the drawings and embodiments.

[0052] As Figure 1 shown, a method for manufacturing personal credit assessment of forest gradient trees includes:

[0053] Data preprocessing;

[0054] Initial training of random forest;

[0055] XGBoost residual fitting;

[0056] Meta-learner integration;

[0057] Model prediction;

[0058] Model evaluation.

[0059] The adaptive fusion forest-gradient boosting tree algorithm (AHF-XGB) of the present invention mainly adopts the idea of ensemble learning, combines the advantages of the random forest (RandomForest, RF) and XGBoost models, and introduces the residual modeling and meta-learning mechanisms to improve the prediction performance through the combination of multiple models. First, the random forest model is used to generate prediction probabilities, then the XGBoost model is used to model the residuals of the random forest, and finally a meta-learner (logistic regression) is used to fuse the prediction results from the random forest and XGBoost to improve the final classification performance.

[0060] The algorithm of the random forest model (RandomForest) includes:

[0061] Random forest is an ensemble learning method that makes predictions based on the combination of multiple decision trees. It trains multiple decision trees and uses the voting method (for classification) or the averaging method (for regression) to make the final prediction. Its training process is expressed as:

[0062] Training stage:

[0063] The random forest generates each decision tree T through the following steps i :

[0064] T i = Train(Subset(X, y, B i ))

[0065] where X is the training data, y is the target variable, and B i is a subset randomly sampled from the training set (X, y), and sampling allows repetition.

[0066] Prediction stage:

[0067] The prediction result of each tree T i for the input data X is The final prediction result is obtained through majority voting or averaging:

[0068]

[0069] where N is the number of decision trees in the random forest, and T i (X) is the predicted value of each decision tree.

[0070] The algorithm principle of the XGBoost model (gradient boosting tree)

[0071] XGBoost (Extreme Gradient Boosting) is an optimized algorithm based on the gradient boosting framework. It improves the prediction ability by gradually reducing the residuals of the model prediction.

[0072] The basic idea of XGBoost is to optimize the prediction result by modeling the residuals, and the specific process is expressed by the following formula.

[0073] Loss function:

[0074] The goal of XGBoost is to minimize the loss function L, which is usually the difference between the model prediction and the true label on the training set.

[0075] Suppose the loss function is:

[0076]

[0077] where is the loss function (such as mean squared error), and Ω(f k ) is the regularization term to prevent overfitting.

[0078] Update formula:

[0079] At each step t, the model is adjusted by calculating the residuals and using gradient boosting:

[0080]

[0081] where is the prediction result of the t-th iteration, and f t (X) is the base model learned in the t-th iteration (usually a decision tree), and η is the learning rate.

[0082] The algorithm of the Adaptive Hybrid Forest - eXtreme Gradient Boosting (AHF-XGB) includes:

[0083] Residual modeling includes:

[0084] Typically, the prediction error (residual) of a model contains pattern information that has not been captured. By using the prediction residual of a random forest as a new target variable and modeling it with XGBoost, the part that the random forest cannot effectively fit can be better captured.

[0085] For each sample x i , assuming the true value is y i , the prediction of the random forest is The residual r i is defined as:

[0086]

[0087] Using XGBoost to fit the residual, assuming is close to the true value y after optimization i , then the residual r i captures the part that the random forest fails to fit. XGBoost can further improve the performance of the model by optimizing the residual:

[0088]

[0089] XGBoost further optimizes the model by learning the residual to make up for the deficiencies of the random forest model.

[0090] The present invention uses a random forest as the basic model, which can better fit most of the data, but there are model errors (residuals). The errors contain patterns that the model fails to capture. By modeling the prediction error (residual) as the target of XGBoost, the part that the random forest fails to fit can be captured, enhancing the performance of the model. By fitting the residual with XGBoost, complex patterns in the data can be better captured, optimizing the accuracy of the model.

[0091] Model weighting and feature stacking include:

[0092] By stacking the output features of two models to form a new feature matrix and using a meta-learner (logistic regression) to further fuse these features. The meta-learner generates the final prediction by learning the outputs of the random forest and XGBoost.

[0093] Let be the prediction result of the random forest, be the prediction result of XGBoost for the residual, then the stacked feature matrix is:

[0094]

[0095] The meta-learner f meta Then train on these stacked features to generate the final prediction:

[0096]

[0097] By stacking the output features of different models, the meta-learner is jointly trained based on different models to optimize the final prediction results. Through this feature fusion and weighting strategy, the advantages of the two models can be combined to improve the accuracy.

[0098] Adaptive training process

[0099] In the training stage, first, the prediction probability is trained and calculated through a random forest, then the residuals are calculated and used as the input for training the XGBoost. This process is adaptive because each training dynamically adjusts the training objective of the XGBoost model according to the output of the random forest.

[0100] Data preprocessing and feature engineering:

[0101] The predicted target variable includes: Definition of the predicted label (1 for yes, 0 for no): At the observation time point, this legal person has a loan balance and is not overdue, and the number of overdue days of the enterprise in the next 90 days is greater than or equal to 30. That is: Label 1 (yes): The enterprise has a loan balance and is not overdue at the observation time point, but the number of overdue days in the next 90 days is greater than or equal to 30 days; Label 0 (no): The enterprise has a loan balance and is not overdue at the observation time point, and there is no overdue or the number of overdue days is less than 30 days in the next 90 days.

[0102] Data cleaning and processing, including: Removing features with a null value rate greater than 75%; Removing features with a correlation greater than 0.9.

[0103] Dataset division and data augmentation include:

[0104] Pre-training and test sets extracted: There are 2,572 records in total, and the good-bad ratio is 0.932:0.069.

[0105] Post-extraction training set: 290 records, and the good-bad ratio is 17:12.

[0106] Post-extraction test set: 165 records, and the good-bad ratio is 2:1.

[0107] The total number of post-extraction training set and test set records is 496, and the number of real samples meeting the project requirements does not exceed 455.

[0108] The data time points of the training set and the test set are: 202301 - 202404.

[0109] Data augmentation

[0110] Enhance the training set using the GAN algorithm. After enhancement, the training set has 4,290 records, and the ratio of good to bad is 0.902:0.098, meeting the project requirements.

[0111] The test set is used to evaluate the true performance of the model. No data enhancement is performed on it, and its true distribution is retained.

[0112] Feature selection

[0113] Use the XGBoost and LightGBM algorithms to rank the feature importance, and finally select the modeling features.

[0114] Table 1 Feature index names and descriptions

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121] The following is the business explanation of each feature variable in the personal credit model for small and medium-sized enterprise legal persons in the manufacturing industry:

[0122] Degree

[0123] This indicator reflects the highest degree held by the legal person, which is usually related to their personal educational background and professional qualities. Legal persons with higher education may possess stronger decision-making and management abilities, thus affecting their personal credit performance.

[0124] 2. Number of credit card accounts

[0125] This indicator represents the number of credit cards held by the legal person. As a credit tool, a larger number of credit cards may mean that the legal person has a certain credit history, but it may also indicate a larger debt burden, which needs to be considered comprehensively.

[0126] 3. Number of quasi-credit card accounts

[0127] This indicator reflects the type of credit limit used by the legal person and has reference value for their debt and consumption behaviors.

[0128] 4. Age

[0129] The age of a legal person can reflect its professional experience and stability. Generally speaking, older legal persons may have more management experience and financial stability, but they may also face risks in aspects such as health.

[0130] 5. Average Overdraft Limit of Quasi-credit Card in the Past 6 Months

[0131] This indicator measures the credit limit of a legal person's use of a quasi-credit card. A higher overdraft limit may mean that the legal person has a greater reliance on credit and may have higher financial risks.

[0132] 6. Average Usage Limit of Credit Card in the Past 6 Months

[0133] This indicator reflects the credit card usage of a legal person in the past 6 months. A higher average usage limit may mean that the legal person has a higher consumption and debt level, which may thus affect its credit score.

[0134] 7. Average Usage Limit of Quasi-credit Card in the Past 6 Months

[0135] This indicator reflects the usage of the credit limit by a legal person. A higher usage limit may indicate that the legal person has higher financial activities.

[0136] Number of Credit Card Approval Queries

[0137] This indicator represents the number of times a legal person has recently queried for credit card approval. Frequent queries may indicate that the legal person is seeking to increase its credit limit or that its financial situation is unstable.

[0138] 9. Number of Credit Card Approval Query Institutions

[0139] This indicator shows how many different institutions a legal person has applied to for credit card approval. More institutional queries may indicate that the legal person is seeking multiple loans or has frequent credit applications, which may pose credit risks.

[0140] 10. Credit Limit of Credit Card

[0141] This indicator shows the total credit limit of a legal person's credit card. A higher credit limit can reflect the legal person's credit status in the banking system, but if overused, it may also increase debt risks.

[0142] 11. Credit Limit of Quasi-credit Card

[0143] This indicator measures the credit limit of a legal person's quasi-credit card. A high limit may represent that the legal person has a good credit record with multiple financial institutions.

[0144] 12. Credit Score of the People's Bank of China

[0145] This indicator is the credit score of legal persons by the Credit Reference Center of the People's Bank of China and is an important reference for assessing the credit risk of legal persons. A higher score usually means lower credit risk.

[0146] 13. Number of work units

[0147] This indicator reflects the number of units where the legal person has worked. A larger number of work units may indicate greater job mobility of the legal person.

[0148] 14. Gender

[0149] Gender may indirectly affect the credit score of legal persons. Although gender itself has no direct relation to credit, it may be used as an auxiliary feature in some models for assessment.

[0150] Number of credit card issuers

[0151] This indicator shows the number of issuers of credit cards held by the legal person. A larger number of issuers may indicate that the legal person has an extensive credit network, but it may also mean that they have more debts.

[0152] 16. Number of quasi-credit card issuers

[0153] This indicator reflects the number of issuers of quasi-credit cards held by the legal person. More issuers may indicate that the legal person has more credit products.

[0154] 17. Number of loan approval inquiries

[0155] It represents the number of inquiries when the legal person applied for a loan recently. A higher number of inquiries may mean that the legal person is seeking a loan or has had multiple loan applications rejected, which may reflect its unstable financial situation.

[0156] 18. Number of institutions for loan approval inquiries

[0157] It indicates the number of different institutions to which the legal person has applied for loan approval. Frequent inquiries may mean that the legal person is facing financing difficulties or has a high debt pressure.

[0158] 19. Highest credit limit per single bank for credit cards

[0159] This indicator reflects the highest credit limit obtained by the legal person from a single bank. A higher credit limit may indicate that the legal person has a good credit record with that bank.

[0160] 20. Highest credit limit per single institution for quasi-credit cards

[0161] It reflects the credit limit of the quasi-credit card obtained by the legal person from a single institution. A higher limit may represent a good credit record of the legal person.

[0162] 21. Lowest credit limit per single bank for credit cards

[0163] Indicates the minimum credit limit obtained by a legal entity from a single bank. A lower credit limit may reflect a lower credit level of the legal entity or more cautious credit granting by the bank.

[0164] 22. Minimum Credit Limit of Quasi-credit Card per Single Bank

[0165] This indicator reflects the minimum credit limit of a legal entity for quasi-credit cards. A lower limit may indicate that the bank has a more conservative credit assessment of it.

[0166] 23. Number of Lending Legal Entities

[0167] This indicator shows how many different institutions the legal entity has applied for loans from. A larger number of lending institutions may indicate that the legal entity has multiple sources of debt or a frequent need for loans.

[0168] 24. Loan Balance

[0169] This indicator reflects the total amount of loans currently owed by the legal entity. A higher loan balance may mean that the legal entity has a heavy debt burden, affecting its credit risk.

[0170] 25. Number of Loan Transactions

[0171] This indicator reflects the number of loan transactions the legal entity has obtained. A larger number of loan transactions may indicate that the legal entity has a frequent need for financing, which may increase the credit risk.

[0172] 26. Number of Lending Institutions

[0173] Indicates how many different institutions the legal entity has obtained loans from. A larger number of lending institutions may reflect a high credit demand of the legal entity, but it may also mean that its credit status is unstable.

[0174] 27. Loan Contract Amount

[0175] Shows the total amount of each loan contract of the legal entity. A higher loan amount may indicate that the legal entity faces a large capital demand or has a high credit risk.

[0176] 28. Number of Approval Queries for Other Reasons

[0177] This indicator reflects the number of approval queries made by the legal entity for other reasons. Frequent queries may be due to poor credit management of the legal entity.

[0178] 29. Number of Institutions for Approval Queries for Other Reasons

[0179] Indicates the number of times the legal entity has queried different institutions. Frequent queries may be related to the legal entity's attempt to change its credit status.

[0180] 30. Overdraft Balance of Quasi-credit Card

[0181] This indicator reflects the overdraft situation of legal persons on quasi-credit cards. A higher overdraft balance may indicate that the current financial situation of the legal person is relatively tight.

[0182] 31. Latest living status

[0183] Reflects the current living situation of legal persons, such as whether they own their own properties, etc. A stable living status is usually related to a better credit status.

[0184] 32. Number of residential addresses

[0185] Indicates the number of times the legal person has changed their place of residence. A larger number of residential addresses may indicate that the legal person lacks a stable living situation, which may affect their credit score.

[0186] 33. Professional title

[0187] The professional title of a legal person reflects their position and responsibilities in the company, and is usually related to their economic situation and credit performance. Legal persons with higher professional titles may have higher incomes and stronger financial stability.

[0188] 34. Used credit card limit

[0189] Shows the credit card limit currently used by the legal person. A higher used limit may mean that the legal person is currently under a higher debt burden, affecting their credit status.

[0190] The characteristic variables comprehensively reflect multiple dimensions in aspects such as credit, financial situation, credit card and loan usage, and accurately evaluate the credit risks of legal persons of small and medium-sized enterprises in the manufacturing industry.

[0191] Training the random forest model includes: training the training data using the random forest model to obtain the prediction probability of each sample

[0192] Calculating the residuals includes:

[0193] Calculating the true label y i and the random forest prediction result The difference between them to obtain the residuals

[0194] Training the XGBoost model includes: training the residuals using the XGBoost model to further improve the prediction ability of the model by minimizing the residuals.

[0195] Stacking features includes: stacking the random forest prediction probabilities and the XGBoost prediction results for the residuals into a new feature matrix:

[0196]

[0197] Train the meta - learner, including: using logistic regression (or other classifiers) to train the stacked features to obtain the final prediction model.

[0198] Model prediction specifically includes:

[0199] Generate random forest prediction: Use the trained random forest model to predict the test data to obtain the prediction probability

[0200] Generate XGBoost residual prediction: Use the trained XGBoost model to predict the test data to obtain the residuals

[0201] Stack the features: Stack the prediction results of random forest and XGBoost into a new feature matrix:

[0202]

[0203] Generate the final prediction: Use the trained meta - learner to predict the stacked features to obtain the final prediction result

[0204] In the AHF - XGB algorithm, the advantages of RandomForestClassifier and XGBoostClassifier are combined.

[0205] The specific selection is based on the following considerations:

[0206] 1. RandomForestClassifier: Good at dealing with non - linear relationships of features and high - dimensional data, improving the robustness and stability of the model through random sampling and feature selection.

[0207] 2. XGBoostClassifier: Based on the idea of gradient - boosting trees, effectively dealing with complex non - linear problems and having high prediction accuracy.

[0208] 3. Ensemble idea: Through an adaptive weighting strategy, dynamically adjust the contribution weights of the two models, making the algorithm have both robustness and accuracy.

[0209] The modeling process includes: data pre - processing → initial random forest training → XGBoost residual fitting → meta - learner integration → model prediction → model evaluation.

[0210] Step 1: Data pre - processing

[0211] Data extraction: Extract the personal credit investigation and credit - related tables of the legal persons of manufacturing enterprises from the data warehouse.

[0212] Feature processing: Process the feature indicators according to the logical mapping defined by the business and handle the missing values.

[0213] Data annotation: Generate the target variable (whether overdue) based on business requirements.

[0214] Data splitting: Divide it into a training set and a test set.

[0215] Step 2: Initial prediction by random forest: Use the RandomForest training model to calculate the residuals.

[0216] Step 3: Residual fitting by XGBoost: Use the XGBoost model to fit the prediction residuals of the random forest with the prediction residuals of the random forest as the target variable.

[0217] Step 4: Ensemble of the meta - learner (logistic regression), then stack the prediction results of these two models as new features, and use logistic regression as the meta - learner for final training.

[0218] Step 5: Model prediction. First, perform probability prediction through the random forest model, and then predict the residuals through XGBoost; stack the prediction results of the two models together, and then hand them over to logistic regression for the final prediction.

[0219] Step 6: Model evaluation. The evaluation metrics include: AUC and KS.

[0220] Model evaluation and result analysis: By comparing the performance of different models on the test set, the present invention focuses on two key evaluation metrics, AUC and KS:

[0221] Table 2 Model evaluation results

[0222]

[0223]

[0224] The present invention combines residual modeling and meta - learning and applies it to personal credit investigation in the manufacturing industry, effectively enhancing the accuracy and interpretability of the credit scoring model, especially when dealing with complex, multi - dimensional, and highly dynamic data.

[0225] In the present invention, the residuals are used as important input features and modeled through the XGBoost model, forming an adaptive residual optimization strategy. This strategy can not only enhance the model's fitting ability for complex data but also improve the accuracy of prediction results. Among them, stacked feature fusion stacks the prediction results from different models (such as RF and XGBoost) as input features and uses a meta-learner (such as logistic regression) for final learning. This process realizes cross-model fusion of features, making the model's prediction more accurate. AHF-XGB (Adaptive Hybrid Forest-XGBoost algorithm) is an innovative ensemble learning method that combines Random Forest (RF) and XGBoost and adopts residual modeling and meta-learning strategies, which are mainly reflected in the following aspects:

[0226] Residual Modeling

[0227] Taking the prediction residuals of the random forest model as the new target, XGBoost is used to model these residuals. The advantage of this strategy is that although the random forest model performs well in dealing with complex data, there are still certain errors in its prediction results. By modeling these errors, the defects of the model can be compensated and the prediction accuracy can be improved.

[0228] The present invention further optimizes by taking the model error as the target, enabling the model to better handle the parts that the random forest fails to effectively fit.

[0229] Meta-Learning

[0230] The present invention introduces a meta-learner (such as logistic regression) to fuse the prediction results from the random forest and XGBoost. The role of the meta-learner is to generate the final prediction by weighting the outputs of different models. This fusion method effectively utilizes the advantages of different models, improving the overall prediction performance. Different from the traditional simple weighting strategy, this algorithm combines the weighted outputs of multiple models, improving the generalization ability while reducing the risk of overfitting.

[0231] Dynamic Adaptive Learning Process

[0232] The present invention is not just a static model combination method but has a dynamic adaptive training process. During the training process, first, probability predictions are obtained through the random forest, then the residuals are calculated, and XGBoost is used to fit these residuals, thereby adaptively improving the performance of the model. Compared with traditional ensemble methods, this algorithm is more flexible and effective.

[0233] Beneficial Effects:

[0234] 1. Introduction of Adaptive Weight Adjustment Strategy

[0235] In traditional ensemble learning methods, model fusion usually relies on fixed weights or simple weighted averaging, and cannot dynamically adapt to the characteristics of different data distributions or tasks. By introducing an adaptive weight adjustment strategy based on meta-learning, the present invention can dynamically adjust the fusion weights of the base models according to data characteristics, thereby improving the prediction accuracy and generalization ability of the model under complex tasks and changing data distributions.

[0236] 2. Improving Accuracy by Fusing Residual Modeling

[0237] The present invention breaks through the limitations of existing residual modeling methods. Instead of only optimizing the residuals for a single model, it combines the residuals of multiple base models and corrects the fused residuals through adaptive optimization techniques, further improving the prediction accuracy of the model, especially showing more excellent performance in complex non-linear tasks.

[0238] 3. Dynamic Learning Rate Optimization

[0239] Traditional ensemble learning methods usually adopt fixed training strategies, ignoring the dynamic optimization requirements of the model during training. By introducing a dynamic learning rate mechanism, the present invention can track the error changes during model training in real time, adaptively adjust the learning rate, and significantly improve the model training efficiency and convergence performance.

[0240] 4. An Efficient Integration Framework Compatible with Multiple Base Models

[0241] The AHF-XGB algorithm designs a highly compatible fusion framework that can efficiently integrate different types of base models such as RandomForest and XGBoost, and at the same time realizes the efficient combination of model outputs through stacking and weighted fusion strategies. This framework significantly reduces the limitations of a single model and enhances the model's ability to handle high-dimensional data and complex tasks.

[0242] 5. Strong Practicality in the Fields of Manufacturing and Personal Credit Investigation

[0243] The algorithm of the present invention is specifically optimized for the characteristics of complex data scenarios in the fields of manufacturing and personal credit investigation, and can better adapt to the dynamic changes of data characteristics in these scenarios, providing a higher-precision prediction tool for financial institutions, manufacturing enterprises, and credit assessment services. This can not only reduce the financial default risk, but also improve the scientific nature of corporate loans and credit decisions.

[0244] The present invention proposes an innovative algorithm integrating multi-model fusion, residual modeling, adaptive weight adjustment, and dynamic learning rate optimization, which solves the deficiencies of traditional ensemble learning methods in multi-model fusion, residual optimization, and dynamic adjustment strategies, improves the accuracy, robustness, and practicality of the model, and has significant advantages and broad application prospects in multiple actual scenarios.

[0245] The above specific implementation manners further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A forest gradient tree-based credit assessment method for manufacturing industry personal credit, characterized in that: The credit assessment method comprises: Data preprocessing; Use the random forest model to make preliminary predictions on the data and obtain the predicted probability of each sample; Calculate the residuals of the random forest model and use the XGBoost regression model to fit the residuals to compensate for the errors of the random forest model; Use logistic regression as a meta-learner, take the prediction results of random forest and XGBoost as input features, and dynamically adjust the weights of the model by training the meta-learner; Adopting adaptive learning rate strategy, dynamically adjusting the weights of each basic model in the integration process, and obtaining the prediction model; The credit of individuals in the manufacturing industry is evaluated based on the prediction model.

2. According to claim 1, a forest gradient tree manufacturing personal credit assessment method is characterized in that: The calculating the residual of the random forest model specifically includes: for each predicted sample, calculating the difference between the predicted probability of the random forest and the actual value as the residual.

3. The method for evaluating personal credit in manufacturing industry based on forest gradient tree according to claim 1 is characterized in that: The data preprocessing specifically includes: Data extraction: extracting personal credit and credit-related tables of manufacturing enterprise legal persons from the data warehouse; Feature processing: processing feature indicators according to the logical mapping defined by the business and handling missing values; Data annotation, generating target variables based on business needs; Data is split into training and testing sets.

4. The method for evaluating personal credit in manufacturing industry based on forest gradient tree according to claim 1 is characterized in that: The use of the random forest model to make preliminary predictions on the data specifically includes: Training the random forest model, including: using the random forest model to train the training data to obtain the predicted probability of each sample 5. The method for evaluating personal credit in the manufacturing industry based on a forest gradient tree according to claim 4 is characterized in that: The calculation of the residual of the random forest model and the use of the XGBoost regression model to fit the residual and compensate for the error of the random forest model specifically include: Calculate the residual, including: Calculate the true label y i And the random forest prediction results The difference between Use the XGBoost model to train the residuals and further improve the model's predictive ability by minimizing the residuals; Stack the random forest prediction probability and XGBoost's prediction results for the residual into a new feature matrix:

6. The method for evaluating personal credit in manufacturing industry based on forest gradient tree according to claim 1, characterized in that: The random forest model algorithm specifically includes: Training phase: Random forest generates each decision tree T through the following steps i : T i =Train(Subset(X,y,B i )) Among them, X is the training data, y is the target variable, and B i It is a subset randomly sampled from the training set (X, y), and repetition is allowed during sampling; Prediction stage: Each tree T i The prediction result for the input data X is The final prediction result is obtained by majority voting or averaging: Where N is the number of decision trees in the random forest, T i (x) is the predicted value of each decision tree.

7. The method for evaluating personal credit in manufacturing industry based on forest gradient tree according to claim 1 is characterized in that: The XGBoost regression model algorithm specifically includes: Loss function: The goal of XGBoost is to minimize the loss function L, which is usually the difference between the model predictions and the true labels on the training set; Assume the loss function is: in, is the loss function, Ω(f k ) is a regularization term to prevent overfitting; At each step t, the model is adjusted by computing the residual and using gradient boosting: in, is the prediction result of the tth iteration, f t (X) is the base model learned in the tth iteration, and η is the learning rate.