Method for constructing retail credit risk prediction model and retail credit Scoreslogplus model

By building a retail credit risk prediction model based on large-scale bank data, the problems of financial institutions' insufficient data diversity and modeling skills have been solved, and more accurate and stable credit risk assessment has been achieved.

CN120707263APending Publication Date: 2025-09-26CCB FINTECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410306191.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When building retail credit risk scoring systems, financial institutions face problems such as insufficient data diversity and coverage, and insufficient modeling skills, which leads to insufficient stability and accuracy of the scoring systems. In particular, small and medium-sized financial institutions find it difficult to effectively use internal data during the digital transformation process.

Method used

Build a retail credit risk prediction model based on massive data from large banks. By acquiring, processing and handling multi-dimensional credit sample data, using statistical principles and decision tree methods, combined with logistic regression models, calculate the probability of credit default and generate a credit score.

Benefits of technology

It improves the accuracy and stability of credit risk assessment, can better cover the risk characteristics in different business scenarios, and provide more accurate credit risk prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004745922460000061
    Figure BDA0004745922460000061
  • Figure BDA0004745922460000071
    Figure BDA0004745922460000071
  • Figure BDA0004745922460000072
    Figure BDA0004745922460000072
Patent Text Reader

Abstract

The invention relates to a method for calculating retail credit risks, and the method comprises the steps: data collection: obtaining retail credit prediction data of a to-be-predicted sample; a data processing step of processing the obtained retail credit prediction data to obtain second generation derivative retail credit prediction data; and a credit default probability calculation step: substituting the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the to-be-predicted sample. According to the application, during model construction, the credit scores constructed based on different credit business scenes of the financial institution are used as the characteristic variables, and the repayment performance of all credit products of the institution by the customer is used as the performance variable to construct the model, so that behaviors and risk trends of the customer in different credit scenes are integrated; therefore, the final model can cover the risk characteristics of different customer groups under different businesses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a credit risk management system and method. The method and system of this application can assist financial institutions in making more accurate risk decisions and accelerate their digital transformation. Specifically, this application relates to a method for constructing a retail credit risk prediction model and a retail credit Scoresigmaplus model. Background Art

[0002] In the current environment of booming consumer lending, some financial institutions' manual approval mechanisms are no longer able to cope with the increasing demand for credit. Consequently, there is an urgent need to enhance their intelligent risk control capabilities. Financial institutions hope to establish risk mitigation mechanisms that cover the entire credit process, from customer pre-screening, pre-loan review, mid-loan approval, post-loan management, and early collection.

[0003] If a scoring system can be developed based on the principles of early identification, early warning, early detection and early disposal, and credit business can be monitored and managed quickly and conveniently, the financial institutions' own business volume, competitive advantage and asset quality can be improved while risks are controllable.

[0004] However, building a scoring system is highly dependent on data and technology. The diversity and coverage of data dimensions, as well as modeling techniques and methodologies, directly impact the ultimate stability and ranking of the scoring system. Some financial institutions lack experience with intelligent risk control for retail businesses and therefore have weak risk control capabilities. In actual applications, factors such as insufficient data mining and analysis capabilities and weak risk modeling techniques hinder financial institutions from fully leveraging the value of internal data and effectively improving model accuracy and stability, leading to technical control challenges. This is also one of the main obstacles facing small and medium-sized financial institutions in their digital transformation. Summary of the Invention

[0005] To address the shortcomings of the aforementioned prior art, this application aims to provide a credit risk management system and method that can provide effective risk management for financial institutions. The credit risk prediction method and system of this application is based on massive amounts of data from large banks and utilizes statistical principles to extract risk patterns, which is inherently valuable for promotion.

[0006] The construction of other popular scoring models currently on the market is often hampered by factors such as small sample sizes, relatively limited data sources, and high homogeneity in data dimensions. Furthermore, because most current risk assessment models utilize data with weak financial attributes—that is, they are based on non-credit transaction data such as smart terminal device data, social platform data, and online shopping mall data—and non-overdue prediction targets, their predictions often deviate significantly from actual credit overdue situations.

[0007] This application first provides a method that can be used to construct a retail credit risk prediction model, and the model constructed based on this method provides financial institutions with a set of accurate methods for calculating the retail credit risk of samples to be predicted.

[0008] The method and system of this application are developed based on highly stable and high-coverage data samples, and systematically innovate the relatively mature credit risk control system. They use credit samples covering various business forms to predict the potential retail credit risks of financial institutions.

[0009] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0010] This application involves the following technical solutions:

[0011] 1. A method for calculating retail credit risk, comprising:

[0012] Data collection step, which obtains retail credit prediction data of the sample to be predicted;

[0013] a data processing step, which processes the acquired retail credit forecast data to obtain second-generation derived retail credit forecast data;

[0014] The credit default probability calculation step is to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

[0015] 2. The method according to item 1 further includes:

[0016] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.

[0017] 3. The method according to item 1 or 2, wherein

[0018] The retail credit prediction data includes original retail credit prediction data of the sample to be predicted and derived retail credit prediction data processed based on the original retail credit prediction data;

[0019] Preferably, the original retail credit prediction data includes:

[0020] Credit card basic data, which is based on all available data during the sample user's credit card creation and usage process.

[0021] Basic data on personal loans, which is all available data based on the loan application and usage behavior of sample users.

[0022] Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions.

[0023] Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

[0024] 4. The method according to any one of items 1 to 3, wherein

[0025] The derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension;

[0026] Preferably, the derived retail credit forecast data includes but is not limited to:

[0027] Derivative retail credit forecast data obtained by processing based on sample relationship length,

[0028] Derivative retail credit forecast data obtained by processing time interval variables,

[0029] Derivative retail credit forecast data obtained based on the frequency of sample behavior.

[0030] Derivative retail credit forecast data obtained by processing the sample at the current time point,

[0031] Derivative retail credit forecast data obtained based on sample continuous behavior processing,

[0032] Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

[0033] 5. The method according to any one of items 1 to 4, wherein

[0034] In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data.

[0035] The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

[0036] 6. The method according to claim 5, wherein

[0037] The comprehensive business credit score includes a comprehensive business credit total score and 12 seed scores; wherein the comprehensive business credit total score and the 12 seed scores are calculated based on the retail credit prediction data using the 12 seed models of the calculation model; wherein, based on the 12 seed models, the 12 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score;

[0038] The credit card business credit score includes four credit card business credit total scores and 20 seed scores; wherein, the four credit card business credit total scores and 20 seed scores are calculated based on the retail credit prediction data using the 20 seed models obtained by the calculation model; wherein, based on the 20 seed models, the 20 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method using the sub-model and the score calculated by the sub-model is used as the four credit card business credit total scores;

[0039] The personal loan business credit score includes a total personal loan business credit score and five sub-scores; wherein the total personal loan business credit score and the five sub-scores are calculated based on the retail credit prediction data using the five sub-models of the calculation model; wherein, based on the five sub-models, the five sub-scores of a sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal loan business credit score;

[0040] The personal consumer loan business credit score includes three personal consumer loan business credit total scores and 19 seed scores; wherein, the three personal consumer loan business credit total scores and 19 seed scores are calculated based on the retail credit prediction data using the 19 seed models of the calculation model; wherein, based on the 19 seed models, the 19 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the three personal consumer loan business credit total scores;

[0041] The personal mortgage business credit score includes a total personal mortgage business credit score and five seed scores; wherein the total personal mortgage business credit score and the five seed scores are calculated based on the retail credit prediction data using the five seed models of the calculation model; wherein, based on the five seed models, the five seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal mortgage business credit score;

[0042] The personal quick loan business credit score includes a personal quick loan business credit total score and 6 seed scores; wherein, the personal quick loan business credit total score and 6 seed scores are calculated based on the retail credit prediction data using the 6 seed models of the calculation model; wherein, based on the 6 seed models, the 6 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the personal quick loan business credit total score.

[0043] 7. The method according to claim 6, wherein

[0044] The 12 sub-models of the comprehensive service class are shown in formulas (1-1) to (1-12).

[0045] The 20 sub-models of the credit card business class are respectively shown as formulas (2-1) to (2-5), formulas (3-1) to (3-5), formulas (4-1) to (4-5), and formulas (5-1) to (5-5);

[0046] The six sub-models of the personal loan business are shown in formulas (6-1) to (6-6).

[0047] The 19 seed models of the personal consumption loan business are shown in Formulas (7-1) to (7-7), (8-1) to (8-6), and (9-1) to (9-6).

[0048] The five sub-models of the personal mortgage business are shown in formulas (10-1) to (10-5).

[0049] The six seed models of the personal quick loan business are shown in formulas (11-1) to (11-6).

[0050] 8. The method according to claim 6, wherein

[0051] The second-generation derivative retail credit prediction data is subjected to feature conversion and then substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes:

[0052] Based on the feature type of the second-generation derivative retail credit prediction data that needs to be substituted into the credit default probability model, the WOE method, dummy feature method or continuous method is selected for feature conversion.

[0053] 9. The method according to claim 8, wherein

[0054] The continuous method for feature conversion includes the following methods: calculating the square root of the second-generation derivative retail credit forecast data, calculating the natural logarithm of the second-generation derivative retail credit forecast data, or calculating the cube root of the second-generation derivative retail credit forecast data.

[0055] 10. The method according to claim 9, wherein

[0056] The credit default probability model is a model constructed based on the second-generation derivative retail credit forecast data and credit default probability using logistic regression based on the existing user population.

[0057] 11. The method according to claim 10, wherein

[0058] The second-generation derivative retail credit prediction data is selected from one or more of the comprehensive business credit score, the credit card business credit score, the personal mortgage business credit score, the personal loan business credit score and the personal consumption loan business credit score.

[0059] 12. The method according to claim 11, wherein

[0060] The credit default probability calculation step includes:

[0061] The total credit score of comprehensive business, credit card business, personal mortgage business, personal loan business and personal consumption loan business are converted into features.

[0062] It is preferred to use a continuous conversion method to convert the total credit score of comprehensive business; a continuous conversion method to convert the total credit score of credit card business; a continuous conversion method to convert the total credit score of personal mortgage business; a dummy variable conversion method to convert the total credit score of personal loan business; and a continuous conversion method to convert the total credit score of personal consumer loan business.

[0063] Further preferably, the total credit score of the comprehensive business type is converted into a calculation method of taking the cube root of the total credit score of the comprehensive business type by a continuous conversion method; the total credit score of the credit card business type is converted into a calculation method of taking the square root of the total credit score of the credit card business type by a continuous conversion method; the total credit score of the personal mortgage business type is converted into a calculation method of taking the natural logarithm of the total credit score of the personal mortgage business type by a continuous conversion method; the total credit score of the personal consumption loan business type is converted into a calculation method of taking the natural logarithm of the total credit score of the personal consumption loan business type by a continuous conversion method.

[0064] 13. The method according to claim 12, wherein

[0065] The converted values ​​of the five features, namely the comprehensive business credit score, credit card business credit score, personal mortgage business credit score, personal loan business credit score and personal consumption loan business credit score, are substituted into the second-generation derivative retail credit prediction data and the credit default probability, and the credit default probability model constructed using logistic regression is used to calculate the credit default probability of the sample to be predicted.

[0066] 14. The method according to claim 13, wherein

[0067] The credit default probability model is shown in the following formula 12-1:

[0068]

[0069] Where k is the number of features entering the model, and in Formula 1, k is 5;

[0070] α is the intercept term, with a value range of (1.3714, 1.2146) and an optimal value of 1.293;

[0071] β1 is the coefficient corresponding to the total credit score of credit card business, with a value range of (-0.725, -0.921) and an optimal value of -0.823;

[0072] β2 is the corresponding coefficient of the total credit score of personal mortgage business, with a value range of (-0.4084, -0.4476) and an optimal value of -0.428;

[0073] β3 is the coefficient corresponding to the total credit score of comprehensive business, with a value range of (-0.7294, -0.9646) and an optimal value of -0.847;

[0074] β4 is the coefficient corresponding to the total credit score of personal loan business, with a value range of (-0.6156, -0.9684) and an optimal value of -0.792;

[0075] β5 is the coefficient corresponding to the total credit score of personal consumer loan business, with a range of values ​​(-0.5424, -0.5816), and the optimal value is -0.562 (Note: the numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0076] x1 is the square root transformation of the total credit score of the credit card business generated in the feature conversion step;

[0077] x2 is the natural logarithm transformation value of the total credit score of personal mortgage business generated in the feature conversion step;

[0078] x3 is the cube root conversion value of the comprehensive business credit score generated in the feature conversion step;

[0079] x4 is the Woe conversion value of the total credit score of the personal loan business generated in the feature conversion step;

[0080] x5 is the natural logarithm transformation of the total credit score of the personal consumption loan business generated in the feature transformation step.

[0081] 15. The method according to claim 14, wherein

[0082] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower:

[0083]

[0084]

[0085] Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0086] 16. A device for calculating retail credit risk, comprising:

[0087] A data collection module, which is used to obtain retail credit prediction data of the sample to be predicted;

[0088] A module for processing samples to be predicted, which is used to process the acquired retail credit prediction data to obtain second-generation derived retail credit prediction data;

[0089] A credit default probability calculation module is used to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

[0090] 17. The apparatus according to item 16, wherein the apparatus executes the steps of the method for calculating retail credit risk according to any one of items 1 to 15.

[0091] 18. A system for calculating retail credit risk, characterized in that the system for calculating retail credit risk comprises: a memory, a processor, and a program for the method for calculating retail credit risk stored in the memory and executable on the processor, wherein the program for calculating retail credit risk, when executed by the processor, implements the steps of the method for calculating retail credit risk as described in any one of items 1 to 15.

[0092] 19. A computer storage medium, characterized in that a program for calculating retail credit risk is stored on the computer storage medium, and when the program for calculating retail credit risk is executed by a processor, the steps of the method for calculating retail credit risk as described in any one of items 1 to 15 are implemented.

[0093] 20. A method for constructing a retail credit risk prediction model, comprising:

[0094] A data collection step, which obtains the original retail credit prediction data of the sample used to build the model;

[0095] A data derivation step, which processes the original retail credit forecast data into derived retail credit forecast data;

[0096] a data processing step, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data;

[0097] a feature initial screening step, which performs a preliminary screening of all categories, that is, all features, comprising the second-generation derivative retail credit forecast data to obtain features after the preliminary screening;

[0098] The initial screening data conversion step determines the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method, and continuous conversion method for feature conversion, and uses the optimal method for feature conversion for each feature after the initial screening;

[0099] Feature fine screening step, in which the features initially screened after feature conversion are deeply screened to obtain finely screened features;

[0100] The credit default probability modeling step selects the logistic regression method to build the model based on the carefully screened features and the probability relationship between them and credit default, and confirms the method used to calculate the credit default probability.

[0101] 21. The method according to claim 20, wherein

[0102] In the data collection step, the original retail credit prediction data of the samples obtained for building the model include:

[0103] Credit card basic data, which is based on all available data during the sample user's credit card creation and usage process.

[0104] Basic data on personal loans, which is all available data based on the loan application and usage behavior of sample users.

[0105] Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions.

[0106] Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

[0107] 22. The method according to claim 20, wherein

[0108] In the data derivation step, the derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension;

[0109] Preferably, the derived retail credit forecast data includes but is not limited to:

[0110] Derivative retail credit forecast data obtained by processing based on sample relationship length,

[0111] Derivative retail credit forecast data obtained by processing time interval variables,

[0112] Derivative retail credit forecast data obtained based on the frequency of sample behavior.

[0113] Derivative retail credit forecast data obtained by processing the sample at the current time point,

[0114] Derivative retail credit forecast data obtained based on sample continuous behavior processing,

[0115] Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

[0116] 23. The method according to claim 20, wherein

[0117] In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data.

[0118] The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

[0119] 24. The method according to any one of items 20 to 23, wherein

[0120] The feature initial screening step includes the following steps:

[0121] The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model.

[0122] The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high.

[0123] The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features;

[0124] The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order,

[0125] The fourth preliminary screening step uses a stepwise discriminant algorithm to perform preliminary screening of features after the first to third preliminary screenings;

[0126] The fifth preliminary screening step is to perform preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

[0127] 25. A method according to any one of items 20 to 23, wherein, in the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.

[0128] 26. The method according to claim 25, wherein

[0129] The initial screening data conversion steps include the following steps based on the judgment of concentration and data type:

[0130] Classify the data type of each feature into character variables and numeric variables.

[0131] For character variables, dummy feature conversion is used to convert the initial screening data.

[0132] The process of further classifying numerical variables includes the following sub-steps:

[0133] If the value of the numerical variable is less than n, the WOE conversion method is used to convert the initial screening data.

[0134] If the value of the numerical variable is more than n, further judge if the conversion to a continuous variable has more values ​​and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used.

[0135] Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.

[0136] 27. The method according to claim 26, further comprising:

[0137] For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature.

[0138] Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

[0139] 28. The method according to any one of items 20 to 27, wherein the feature fine screening step comprises:

[0140] The first fine screening step is based on the stepwise regression algorithm, and the significance of the features is screened based on the F test and T test.

[0141] The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate features with higher variance inflation factors to screen features.

[0142] The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default in order to further perform feature screening.

[0143] 29. A method according to any one of items 20 to 28, wherein the credit default probability modeling step substitutes the features screened in the feature fine screening step into a Sigmoid function to perform logistic regression to calculate a model for the credit default probability.

[0144] 30. A device for constructing a retail credit risk prediction model, characterized in that the device comprises:

[0145] A data acquisition module, which is used to obtain raw retail credit prediction data of samples used to build a model;

[0146] A data derivation module, which is used to process derived retail credit forecast data based on the original retail credit forecast data;

[0147] A data processing module, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data;

[0148] A feature initial screening module is used to perform initial screening on all categories, i.e., all features, of the original retail credit forecast data and the derived retail credit forecast data to obtain features after initial screening;

[0149] The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening;

[0150] A feature fine-screening module is used to perform in-depth screening on the initially screened features after feature conversion to obtain fine-screened features;

[0151] The credit default probability modeling module is used to select the logistic regression method to construct a model based on the probability relationship between the carefully screened features and the credit default, and confirm the model used to calculate the credit default probability.

[0152] 31. The apparatus according to item 30, wherein the apparatus executes the steps of the method for constructing a retail credit risk prediction model according to any one of items 20 to 29.

[0153] 32. A system for constructing a retail credit risk prediction model, characterized in that the system includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor, wherein the program for constructing a retail credit risk prediction model method, when executed by the processor, implements the steps of constructing a retail credit risk prediction model method as described in any one of items 20 to 29.

[0154] Effects of the Invention

[0155] The method and system for constructing a retail credit risk prediction model in this application combines a large number of samples from a large financial institution when constructing the model, and deeply processes and derives the original data obtained from the samples. It uses advanced statistical analysis methods and explainable machine learning technology to construct a general retail credit scoring model based on the inherent characteristics of the original data and derived data, as well as the strong financial attribute information contained in the data that is difficult to obtain in the market.

[0156] When constructing the model, this application uses the credit score constructed based on the different credit business scenarios of the financial institution as the characteristic variable, and the customer's repayment performance of all credit products of the institution as the performance variable to construct the model. It integrates the customer's behavior and risk trends in different credit scenarios, thereby ensuring that the final model can cover the risk characteristics of different customer groups under different businesses. BRIEF DESCRIPTION OF THE DRAWINGS

[0157] Figure 1 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0158] Figure 2 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0159] Figure 3 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0160] Figure 4-1 Shows a typical flow chart for classifying loan customers based on a decision tree.

[0161] Figure 4-2 A typical flow chart for classifying non-loan customers based on a decision tree is shown.

[0162] Figure 5 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0163] Figure 6 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0164] Figure 7 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0165] Figure 8 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0166] Figure 9 A typical flow chart for classifying all sample users based on a decision tree is shown.

[0167] Figure 10 This is a typical flow chart for classifying all sample users based on a decision tree.

[0168] Figure 11 A typical flow chart for classifying all sample users based on a decision tree is shown. DETAILED DESCRIPTION

[0169] Credit risk is the risk that arises when a borrower's financial capacity changes, resulting in a reduced willingness to repay or an inability to fulfill the loan contract, rather than the risk of default due to deliberate fraud on the part of the borrower. Credit default occurs in all types of retail credit business scenarios and is particularly related to the borrower's personal financial situation. The reasons for borrowers' credit defaults can be divided into four main categories: 1. Short credit history, where such borrowers have little experience in managing their financial situation; 2. Borrowers temporarily forget to repay; 3. Overborrowing, where such borrowers have relatively low repayment capacity due to large debts; and 4. Major negative factors, where such borrowers experience long-term impacts on their repayment capacity due to factors such as reduced income, unemployment, or divorce. Each of these different reasons will, to varying degrees, lead to a borrower's default, potentially leading to a more serious default. Credit risk scoring aims to uncover the inherent mathematical relationship between various historical customer information and the probability of future default, and convert it into a score to quantify the probability of default.

[0170] Currently, credit risk scores in existing technologies are primarily developed using historical credit data and multi-credit data (multi-credit data refers to the statistical data of borrowers' loan requests from multiple financial institutions. It is generally believed that a greater number of multi-credit requests in a short period of time indicates a greater probability of future defaults). These data reflect their payment behavior and willingness to pay. The sample data used to build the model in this application not only includes transaction information on credit business but also adds asset data that is typically difficult to obtain. This not only reflects the borrower's payment behavior and willingness to pay, but also provides a more comprehensive assessment of their debt repayment ability, personal qualifications, and other aspects, thereby providing more accurate prediction results.

[0171] <Overall description of model building method>

[0172] Specifically, the present application relates to a method for constructing a retail credit risk prediction model, which includes: a data collection step, which obtains original retail credit prediction data of samples used to construct the model; a data derivation step, which processes derived retail credit prediction data based on the original retail credit prediction data; a data processing step, which processes the acquired retail credit prediction data and derived retail credit prediction data using a calculation model of second-generation derived retail credit prediction data to obtain the second-generation derived retail credit prediction data; a feature initial screening step, which performs preliminary screening on all categories, that is, all features, including the second-generation derived retail credit prediction data to obtain features after preliminary screening; a preliminary screening data conversion step, which judges the conversion method of the features after preliminary screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and uses the best method judged for each feature after preliminary screening to perform feature conversion; a feature fine screening step, which performs in-depth screening on the features after preliminary screening to obtain fine-screened features; a credit default probability modeling step, which selects the logistic regression method for model construction based on the fine-screened features in combination with the probability relationship between the features and credit default, and confirms the method for calculating the credit default probability.

[0173] <Raw Retail Credit Forecast Data>

[0174] When constructing the model for this application, we first obtain approximately original retail credit forecast data based on the basic data of financial institutions, and then split or process the retail credit risk points as much as possible according to different dimensions under the premise of optimal effect, generating more than 3,000 derivative features in total.

[0175] When building the model, first of all, based on all historical data of large financial institutions, the credit card application and behavior data, personal loan application and behavior data, personal financial asset transaction data and personal customer information data of all retail customers over the past four years were initially collected when building the model of this application, totaling 780 million people, of which each person has corresponding data to be processed every month. It can be seen that the data system used to build the model of this application is comprehensive and the data volume is very large. When building a model based on such a data system, it is necessary to consider the modeling methodology, otherwise it will be trapped in the huge data, resulting in the special sample groups that need to be paid attention to being covered in the huge data volume and unable to be effectively identified, causing the computer program to run slowly or even unable to run, so that it is impossible to accurately build the most suitable prediction model.

[0176] In the data collection step, the original retail credit prediction data of the sample obtained for building the model include: basic credit card data, which is all available data based on the sample's credit card creation process and usage process (i.e., retail customers' credit card application and behavior data); basic personal loan data, which is all available data based on the sample's loan application and usage behavior (i.e., personal loan application and behavior data); basic customer information data, which is based on the attributes of the sample itself but is not directly related to the behavior in the financial institution (i.e., personal customer information data); or basic personal financial asset data, which is all other financial assets and financial transaction data of the sample in the financial institution that are not related to credit cards and loans (personal financial asset transaction data).

[0177] In a specific embodiment, basic credit card data includes, but is not limited to, information on accounts, overdue payments, balances, limits, due payments, and actual payments under different dimensions (e.g., time, space, and frequency) such as credit card bills, credit card withdrawals, credit card installments, and interest accrued. The above basic data is not limited to the specific categories listed. As credit card services evolve, those skilled in the art can further encompass newly emerging data types during implementation. In other words, all types of data available during the user's credit card creation and usage process can serve as basic credit card data.

[0178] In a specific embodiment, basic personal loan data includes, but is not limited to, personal loan accounts, overdue personal loan status, personal loan balances, total personal loan amounts, and personal loan repayments due and actual. The above basic data is not limited to the specific categories listed. As the personal loan business evolves, those skilled in the art can further encompass emerging data types during implementation. Specifically, all available data types based on a user's loan application and behavior can serve as basic personal loan data.

[0179] In a specific implementation, in accordance with the requirements of the Personal Privacy Protection Act, the basic data for basic customer information only includes basic customer information, including gender, age, and the administrative region where the business is located. The above basic data is not limited to the specific categories listed. As customer situations and social relationships evolve, those skilled in the art can further include other or emerging data types within the scope of the relevant business application scenarios. In other words, all types of data based on the attributes of the user sample itself but not directly related to the behavior of the financial institution can serve as basic customer information data.

[0180] In one specific embodiment, basic data on personal financial assets includes, but is not limited to, AUM (assets under management), deposits, wealth management, and payroll information. This basic data is not limited to the specific categories listed above. As financial assets evolve, those skilled in the art can further encompass emerging data types during implementation. Specifically, all other financial assets and financial transaction data unrelated to credit cards and loans at financial institutions, based on the sample, can serve as basic data on personal financial assets.

[0181] In a specific embodiment, the original retail credit prediction data of the samples obtained for building the model are based on the data types (i.e., basic variables or basic features) obtained from 780 million people, including but not limited to: credit card bills, credit card cash withdrawals, credit card installments, interest generated by credit cards, and other information on accounts, overdue payments, balances, amounts, repayments due, actual repayments, etc. in different dimensions (e.g., time dimension, space dimension, frequency dimension); personal loan accounts, overdue status of personal loans, balances of personal loans, total amounts of personal loans, repayments due, actual repayments, etc.; basic customer information including gender, age, and administrative region to which the business belongs; basic information such as AUM (i.e., assets under management), deposits, wealth management, and payroll.

[0182] <Derivative Retail Credit Forecast Data>

[0183] In this application, in the data derivation step, the derived retail credit prediction data processed based on the original retail credit prediction data refers to the data obtained by processing the collected original retail credit prediction data based on the time dimension, space dimension, frequency dimension, and statistical information dimension.

[0184] In one specific embodiment, derived retail credit prediction data includes, but is not limited to, derived retail credit prediction data processed based on sample relationship length, derived retail credit prediction data processed based on time interval variables, derived retail credit prediction data processed based on the frequency of sample behavior, derived retail credit prediction data processed based on the current time point of the sample, derived retail credit prediction data processed based on the continuous behavior of the sample, or derived retail credit prediction data processed based on statistical information dimensions. For example, monthly customer data can be obtained and processed based on this monthly data. In this application, processing based on statistical information dimensions includes obtaining the maximum, minimum, and average values ​​of the data to describe the data situation.

[0185] In a specific implementation, for example, starting from the time dimension, for example, customer relationship length variables: for example, the customer account opening time, the customer's maximum account age, etc. are used as types of derived retail credit prediction data, that is, as derived features or derived variables.

[0186] In a specific implementation, for example, starting from the time dimension, time intervals are considered, such as the number of months from the customer's most recent repayment to the current time point, the number of months from the customer's most recent overdue payment to the current time point, etc. as derived features or derived variables.

[0187] In a specific implementation, for example, starting from the frequency level, behavioral frequency level variables are considered: for example, the number of times a customer has made repayments greater than N in the last X months, the number of times a customer's credit limit utilization rate greater than N in the last X months, etc. are used as derived features or derived variables. There is no limit on X and N, and they can be any positive integer greater than 0 as long as the business logic is reasonable.

[0188] In a specific implementation, for example, starting from the time dimension, current point variables are considered: the customer's current monthly credit limit, the customer's current monthly balance, etc. as derived features or derived variables.

[0189] In a specific approach, starting from the time dimension, continuous behavioral variables are considered, such as the maximum number of consecutive overdue payments of a customer > N in the last X months, the number of consecutive repayment rates of a customer > N in the last X months, etc. as derived features or derived variables.

[0190] In a specific approach, starting from the dimension of statistical information, statistical variables are considered, such as the maximum number of overdue payments of a customer in the last X months, the average credit limit utilization rate of a customer in the last X months, etc. as derived features or derived variables.

[0191] It is clear to those skilled in the art that the above-mentioned methods for processing derived variables are merely examples and can be selected arbitrarily. In this application, data, data types, data types, variables or features are sometimes mixed, and those skilled in the art can understand them based on common sense in statistics.

[0192] In this application, derived data can be derived from the original retail credit forecast data through simple processing or complex processing. Simple processed derived data, such as current variables such as monthly salary and monthly balance, can be used directly after data aggregation. Complex processed derived data requires time slicing and logical processing based on the current class variables, and can generate derived variables such as the maximum salary in the past three months and the minimum balance in the past 12 months.

[0193] <Second-generation derivative retail credit forecast data>

[0194] The second-generation derived retail credit prediction data refers to data generated by processing the original retail credit prediction data and the derived retail credit prediction data using a calculation model based on the second-generation derived retail credit prediction data. The second-generation derived retail credit prediction data is selected from the following categories: comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores, and personal mortgage business credit scores.

[0195] The comprehensive business credit score includes a comprehensive business credit total score and 12 seed scores; wherein the comprehensive business credit total score and the 12 seed scores are calculated based on the retail credit prediction data using the 12 seed models of the calculation model; wherein, based on the 12 seed models, the 12 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score;

[0196] The credit card business credit score includes four credit card business credit total scores and 20 seed scores; wherein, the four credit card business credit total scores and 20 seed scores are calculated based on the retail credit prediction data using the 20 seed models obtained by the calculation model; wherein, based on the 20 seed models, the 20 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method using the sub-model and the score calculated by the sub-model is used as the four credit card business credit total scores;

[0197] The personal loan business credit score includes a total personal loan business credit score and five sub-scores; wherein the total personal loan business credit score and the five sub-scores are calculated based on the retail credit prediction data using the five sub-models of the calculation model; wherein, based on the five sub-models, the five sub-scores of a sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal loan business credit score;

[0198] The personal consumer loan business credit score includes three personal consumer loan business credit total scores and 19 seed scores; wherein, the three personal consumer loan business credit total scores and 19 seed scores are calculated based on the retail credit prediction data using the 19 seed models of the calculation model; wherein, based on the 19 seed models, the 19 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the three personal consumer loan business credit total scores;

[0199] The personal mortgage business credit score includes a total personal mortgage business credit score and five seed scores; wherein the total personal mortgage business credit score and the five seed scores are calculated based on the retail credit prediction data using the five seed models of the calculation model; wherein, based on the five seed models, the five seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal mortgage business credit score;

[0200] The personal quick loan business credit score includes a personal quick loan business credit total score and 6 seed scores; wherein, the personal quick loan business credit total score and 6 seed scores are calculated based on the retail credit prediction data using the 6 seed models of the calculation model; wherein, based on the 6 seed models, the 6 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the personal quick loan business credit total score.

[0201] The 12 sub-models of the integrated service class are respectively shown in Formula (1-1) to Formula (1-12). For details, please refer to Example 1-1 to Example 1-17.

[0202] The 20 seed models of the credit card business category are respectively shown as Formula (2-1) to Formula (2-5), Formula (3-1) to Formula (3-5), Formula (4-1) to Formula (4-5), and Formula (5-1) to Formula (5-5). For details, please refer to Example 2-1 to Example 2-10, Example 3-1 to Example 3-10, Example 4-1 to Example 4-10, and Example 5-1 to Example 5-10.

[0203] The six sub-models of the personal loan business are shown in Formula (6-1) to Formula (6-6), respectively. For details, please refer to Example 6-1 to Example 6-11.

[0204] The 19 seed models of the personal consumption loan business are respectively shown as formulas (7-1) to (7-7), formulas (8-1) to (8-6), and formulas (9-1) to (9-6). For details, please refer to Examples 7-1 to 7-12, Examples 8-1 to 8-11, and Examples 9-1 and 9-11.

[0205] The five seed models of the personal housing loan business are shown in Formula (10-1) to Formula (10-5), respectively. For details, please refer to Example 10-1 to Example 10-17.

[0206] The six seed models for the personal quick loan business are shown in Formulas (11-1) to (11-6). For details, please refer to Examples 11-1 to 11-11.

[0207] <Preliminary feature screening>

[0208] In the initial screening data conversion step of the present application, the judgment of the conversion method of the features after the initial screening is based on the concentration and data type of the features after the initial screening. The initial screening data conversion step based on the judgment of concentration and data type includes the following steps: classifying the data type of each feature into character variables and numerical variables, using dummy feature conversion to perform initial screening data conversion on the character variables, and further classifying the numerical variables includes the following sub-steps: if the value of the numerical variable is less than n, using the WOE conversion method to perform initial screening data conversion, if the value of the numerical variable is more than n, further judging if the value of the continuous variable is large and the concentration of the single value is greater than m%, then using the WOE conversion method, if the concentration of the single value is less than or equal to m%, then using the continuous conversion method, preferably, n and m are both positive integers, where n = 5 to 10, m = 90 to 99.

[0209] For example, n is 5, 6, 7, 8, 9 or 10, and m is 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.

[0210] Specifically, in this application, a variety of different feature initial screening methods are used to screen a large number of features, so that the screening can be performed effectively from the largest dimension of features. Existing credit risk scores often use missing rate, concentration and information value IV for feature initial screening.

[0211] In a specific embodiment, features are screened based on the data missingness of each feature of the sample used to build the model. When screening the missing rate, it is generally considered to delete variables with a missing rate greater than 90%, 91%, 92%, 93%, 94%, or 95%. For example, features with a data missing rate exceeding 95% can be eliminated, and features with a data missing rate exceeding 90% can also be eliminated.

[0212] In a specific method, features are screened based on the fact that a single value of a feature sample is too high. When screening for concentration, it is generally considered to delete variables whose single values ​​account for more than 99%, 98%, 97%, 96%, or 95%. For example, features whose single values ​​exceed 99% can be eliminated, and features whose single values ​​exceed 95% can also be eliminated.

[0213] In a specific method, the IV value of each feature is calculated to perform a preliminary screening of the features. During IV screening, the IV value can be used to measure the predictive power of the feature. The larger the IV value, the stronger the predictive power of the feature. The calculation method of the IV value of a single feature is as follows:

[0214]

[0215] Where k is the number of groups after this feature is discretized; y i is the number of non-defaulting customers in group i; y s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s The IV quantitative indicators have the following meanings: when the calculated IV value is less than 0.02, it indicates that the predictive power of this feature is very weak; when the calculated IV value is above 0.02 but less than 0.1, it indicates that the predictive power of this feature is weak; when the calculated IV value is above 0.1 but less than 0.3, it indicates that the predictive power of this feature is good; and when the calculated IV value is above 0.3, it indicates that the predictive power of this feature is strong.

[0216] Of course, the IV calculation value deletion threshold can also be set to 0.03, 0.04, 0.05, etc.

[0217] In the model building method of the present application, the steps of performing preliminary feature screening using missing rate, concentration, and information value IV can be performed in any order. For example, screening can be performed first based on missing rate, then based on concentration, and finally based on information value IV. Screening can also be performed first based on concentration, then based on missing rate, and finally based on information value IV. Screening can also be performed first based on missing rate, then based on information value IV, and finally based on concentration. Screening can also be performed first based on information value IV, then based on missing rate, and finally based on concentration. Screening can also be performed first based on information value IV, then based on concentration, and finally based on missing rate. Screening can also be performed first based on information value IV, then based on concentration, and finally based on missing rate. Screening can also be performed first based on concentration, then based on information value IV, and finally based on missing rate. Those skilled in the art can make their choice based on the sample data. Therefore, based on these three methods, features with obvious defects in certain aspects can be effectively removed, which can effectively reduce the data dimension and improve the effect of model building.

[0218] In one specific method, after the features have been screened using missing rate, concentration and information value IV respectively, a stepwise discriminant algorithm is used to perform preliminary screening of the features, and then the features are screened based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

[0219] On this basis, the technical solution of the present application introduces a step-by-step discriminant method to make the initial screening of variables more efficient and accurate in order to improve the overall efficiency of model development. In actual data, there may be a situation where the distribution of good and bad samples on a certain variable is close, and the ability to distinguish between good and bad is weak. There may also be a class of variables, each of which can distinguish good and bad samples well, but if all are included in the model, they will be redundant due to the duplication of data dimensions covered by the variables. In this regard, the step-by-step discriminant analysis method is used in this application, and the WILKS'S LAMBDA value is used as the statistical criterion for entry and removal, and variables with weak or redundant discriminant effects in the data are deleted.

[0220] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0221] In this application, the step of screening using the stepwise discriminant method is a more critical step. In order to capture as many credit risk points as possible, the model involved in this application uses a huge amount of data and many data dimensions, so variable screening must be performed to further reduce the time cost of model development. In existing financial risk model evaluations, variable importance methods are generally used to reduce variables, such as calculating the importance of variables through algorithms such as the Gini index and information entropy, and selecting variables with high importance. Modeling methods that use stepwise discriminant methods for screening are rarely used. Compared with the variable importance screening schemes commonly used in the industry, the methodology used in this application can retain a large number of variables with relatively weak importance but relatively independent information dimensions.

[0222] In one specific approach, over 3,000 derived data points can be generated from 79 categories of basic data. After three rounds of screening based on missing rate, concentration, and information value (IV), approximately 20-30% of features with poor data quality can be removed. However, over 2,000 features will still remain. Subsequent variable refinement based on all of these features would severely impact development efficiency. Therefore, after comparing various variable reduction schemes, stepwise discrimination was ultimately determined to be the optimal approach.

[0223] For the important features that have been screened by the step-by-step discrimination algorithm, further screening of features is performed based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample. Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and eliminate the features whose actual bad debt rate distribution does not conform to the business trend. Specifically, (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of the values ​​of each box and the corresponding bad debt rate. (2) Use the median of the values ​​of each box and the previous box and the corresponding bad debt rate to calculate the rate of change (slope). (3) Count the number of boxes greater than 0 and non-0 boxes (non-greater than 0 boxes) in the rate of change between two adjacent boxes, and calculate the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change. (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend. The criteria for whether they are approximately consistent are as follows: First, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate increases (such as the credit limit utilization rate, etc.). In this module, the features whose number of boxes greater than 0 in the above-calculated slope accounts for less than 70% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated. Second, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate decreases (such as the deposit amount, etc.). In this module, the features whose number of boxes greater than 0 in the above-calculated slope accounts for more than 30% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated.

[0224] For example, the table below provides an example of binning. In this example, the feature values ​​are divided into 10 bins, and the value ranges of the bins are also summarized in the table below. Using this method, we can further filter features based on the risk characteristics of each risk point (preset by the system) and the actual bad debt rate of the sample. Using this binning method to further filter features can effectively select the features that best match business development trends, thereby further obtaining features suitable for modeling.

[0225]

[0226] <Conversion of Eigenvalues>

[0227] The initial feature processing of existing credit risk scoring methods primarily uses word of essence (WOE) conversion (multi-classification) and dummy feature conversion (binary classification) to discretize continuous features (such as age, account age, etc.). The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes continuous features according to the optimal cut point and embeds the data of good and bad samples into the WOE value. Therefore, it performs better during model construction. However, the binning results may overfit the modeling sample, resulting in a serious decline in model effectiveness when applied to the population (poor generalization ability). At the same time, due to the normalization operation used in binning, the original features falling into different bins are converted into a single value corresponding to each bin, thus losing the ability to distinguish the risks of people falling into the same interval. The dummy feature conversion method is mainly applied to grouping features. Its advantage is that it can eliminate the distinction between good and bad values ​​of different feature values, but it becomes extremely complex and redundant when processing continuous variables.

[0228] In the model building method of the present application, the conversion method of the features after preliminary screening is judged to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and the optimal method is used to perform feature conversion for each feature after preliminary screening.

[0229] The initial screening data conversion step includes the following steps based on the judgment of concentration and data type: classifying the data type of each feature into character variables and numerical variables, and performing initial screening data conversion on character variables using dummy feature conversion. The process of further classifying numerical variables includes the following sub-steps: if the value of the numerical variable is less than n, the WOE conversion method is used for initial screening data conversion. If the value of the numerical variable is more than n, further judging if the continuous variable has more values ​​and the concentration of a single value is greater than m%, the WOE conversion method is used. If the concentration of a single value is less than or equal to m%, the continuous conversion method is used. Preferably, n and m are both positive integers, where n = 5 to 10 and m = 90 to 99.

[0230] Specifically, taking the character variable "education level" as an example, the values ​​of this feature variable can be elementary school, middle school, college, graduate school, etc. For a numeric variable, if the numeric variable is the number of overdue months in the past three months, the values ​​are 0, 1, 2, and 3.

[0231] In the present application, the above n can be 5, 6, 7, 8, 9 or 10, and m can be 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.

[0232] In a specific embodiment, m is selected as 5 and n is selected as 95.

[0233] Specifically, the WOE conversion method is to find the optimal cut point of the feature, divide the value range of the original feature into multiple bins, and then calculate the WOE conversion value corresponding to each bin based on the good and bad performance of each bin and output it. The original feature is divided according to the bin results and the WOE conversion value is output. For each bin, the WOE value is calculated as follows:

[0234]

[0235] where y i is the number of non-defaulting customers in group i; y s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s is the total number of defaulting customers.

[0236] Let's use age as an example to explain the WOE binning process. The original age feature contains values ​​ranging from 18 to 50 years old. After binning, we obtain five bins: 18-24, 25-30, 31-35, 36-42, and 43-50. We then calculate the WOE conversion value for each bin based on the number of good and bad customers in each bin. Finally, each user data entry in each bin is output according to the corresponding conversion value. For example, for a 23-year-old, the WOE conversion value corresponding to bins 18 to 24 is output, and for a 46-year-old, the WOE conversion value corresponding to bins 43 to 50 is output.

[0237] The conversion method of dummy features is: convert a single classification feature into an equal number of dummy features according to the number of values ​​it contains. If a customer belongs to the corresponding value of the generated dummy feature, the corresponding dummy feature value is 1, and the remaining dummy feature values ​​are 0.

[0238] Let's use gender as an example to explain how to convert dummy features. The original features include "male" and "female." After dummy feature conversion, two dummy features are generated: "Gender-Male" and "Gender-Female." If the customer's gender is male, "Gender-Male" is recorded as 1, and "Gender-Female" is recorded as 0. For example, if the original features include "junior college or below," "bachelor's degree," and "master's degree or above," three dummy features are generated: "Education-Junior College or Below," "Education-Bachelor's degree," and "Education-Master's degree or above." If the customer's education level is only a bachelor's degree, "Education-Junior College or Below" is recorded as 0, "Education-Bachelor's degree" is recorded as 1, and "Education-Master's degree or above" is recorded as 0.

[0239] The continuous conversion method is to perform various continuous conversions on the original features (continuous conversion methods include but are not limited to: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, and calculating the natural logarithm of the original data). Calculate the correlation coefficient r (Correlation Coefficient) between the feature value after continuous conversion and the overdue label, select the conversion method with the largest absolute value of the correlation coefficient, and output the conversion method corresponding to the original feature. The calculation formula for the correlation coefficient is as follows:

[0240]

[0241] Where Σ is the summation symbol in mathematics; n is the total number of observations; x i is the conversion value of the original feature of the i-th observation after continuous conversion; This is the mean of the transformed values; where y i is a binary feature indicating whether the i-th observation is in default; The average value of this binary feature. The closer the absolute value of the correlation coefficient is to 1, the more closely the converted value correlates with the default situation, and the more effective the conversion method is. The quantitative meaning of the correlation coefficient r is as follows: when the absolute value of the correlation coefficient is greater than 0 and less than 0.3, it indicates a low correlation; when the absolute value of the correlation coefficient is greater than 0.3 and less than 0.8, it indicates a moderate correlation; and when the absolute value of the correlation coefficient is greater than 0.8 and less than 1, it indicates a high correlation.

[0242] In the model construction of the present application, for the features that are confirmed to adopt a continuous conversion method, the optimal conversion method is selected to perform the continuous feature conversion of the feature based on the correlation between the feature and credit default under different continuous conversion methods. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

[0243] Let's use age as an example to explain continuous transformation methods. The original feature "age" contains values ​​ranging from 18 to 50. Continuous transformation yields the original value, square, square root, cube root, and natural logarithm of age. The word of essence (WOE) value and the absolute value of the correlation coefficient between each transformed value and the good / bad label (prediction result) are then calculated. The transformation method with the largest absolute value of the correlation coefficient is selected for conversion and output. If the cube root of the age feature has the largest absolute value of the correlation coefficient with the good / bad label compared to other transformation methods, the cube root of age is output as the transformed value.

[0244] In the existing technical solutions, a feature conversion method between WOE and dummy features is mainly selected for model construction. For example, in the Chinese patent CN112686749B, the WOE method is used to perform feature conversion on the eigenvalues. The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes continuous features according to the optimal cutting point, so it performs better in model construction, but the binning results may overfit the modeling sample, resulting in a serious decline in the model effect when applied to the overall population (poor generalization ability). At the same time, due to the normalization operation used in the binning, the original features falling into different bins are converted into a single value corresponding to each bin, thereby losing the ability to distinguish the risks of people falling into the same interval.

[0245] The dummy feature conversion method is primarily used for grouping features. Its advantage is that it can eliminate differences between good and bad feature values. For example, in the customer's industry, because there is no clear advantage or disadvantage between retail and wholesale, dummy feature processing is more suitable. There is a hierarchical difference between associate and bachelor's degrees, so although dummy features can be used, the WOE processing method is actually more appropriate.

[0246] On the other hand, the continuous conversion method avoids the overfitting of modeling samples that occurs with the WOE conversion method. It has strong generalization capabilities for the overall sample, and because it does not perform interval mapping, it is less likely that most customer groups fall into a single value. However, it cannot be applied to some features with poor monotonicity or discrete characteristics (such as occupation and position).

[0247] As described above, the technical solution of the present application takes a different approach and creatively combines three methods: continuous conversion, WOE conversion, and dumb feature conversion. It reprocesses some carefully screened features, combines the advantages and disadvantages of the three conversion methods, and creatively designs a conversion judgment method. It selects the optimal feature conversion method based on parameters such as feature data attributes, missing rate, concentration, and supplemented by business logic judgment.

[0248] <Logistic regression and deep feature screening based on logistic regression>

[0249] In existing credit risk scoring, due to the requirement of model interpretability, the logistic regression model is mainly used for model development, and the software that can be used are generally: SAS, R, Python, etc.

[0250] In a specific embodiment, the present application develops the model based on SAS software.

[0251] Specifically, the Sigmoid function is used in logistic regression to fit the probability of predicted default. The Sigmoid function is:

[0252]

[0253] Where Z is a linear combination of the model coefficients and the feature transformation values, and Z is defined as follows:

[0254] Z=α+β1x1+β2x2+...+β k-1 x k-1 +β k x k

[0255] The predicted probability of default is:

[0256] P=P(Y=1|x1,x2,x3,...,x k-1 , x k )

[0257] The fitted prediction for the probability of default is:

[0258]

[0259] From the above formula we can further deduce:

[0260] Substituting the Z value into the above formula can calculate the probability P of predicted default.

[0261] The core of logistic regression model construction is feature screening. The steps of feature screening are as follows: First, batch screen the features based on missing rate, concentration and information value IV. Second, screen all remaining features one by one based on whether the features are consistent with business trends, and retain the features with correct business trends. For example: if it is found that the bad debt rate of the customer group decreases with the increase of the feature loan balance, then this feature is considered to be inconsistent with the business trend. In the understanding of credit business, the higher the loan balance, the higher the customer's default risk exposure (EAD) level, and the greater the risk. At this time, this feature will be removed from the feature list. Third, use the stepwise regression function of logistic regression to eliminate features that are less important and highly correlated with other features. Fourth, filter the coefficients based on the positive and negative signs of the training coefficients and the business trend of the feature conversion values, and retain features whose feature coefficient signs are consistent with the business logic. Defining the Y label as 0 for good customers and 1 for bad customers, for a feature whose bad debt rate increases monotonically with the feature value (e.g., loan balance), the training coefficient should be positive. Conversely, for a feature whose bad debt rate decreases monotonically with the feature value (e.g., deposit balance), the training coefficient should be negative. Features that do not meet these criteria should be removed. Fifth, use the variance inflation factor (VIF) and correlation coefficient to further eliminate highly correlated features: For the variance inflation factor, eliminate features with the highest VIF greater than 4 one by one. For highly correlated features, eliminate features with low IV values ​​within the feature group with the highest correlation coefficient greater than 0.80 one by one. Sixth, use the population stability index (PSI) to eliminate features with significant distribution differences at different time points, thereby causing instability. Features with a PSI greater than 0.25 should be eliminated directly. Features with a PSI greater than 0.25 and greater than 0.1 should be carefully removed based on the impact of eliminating these features on the model's discriminatory ability.

[0262] This application strictly adheres to the above rules to screen features, ensuring the model's interpretability, stability, and ability to distinguish between good and bad customers.

[0263] The feature fine-screening steps of the present application include: a first fine-screening step, based on the stepwise regression algorithm, screening the features based on the F test and the T test for the significance of the features; a second fine-screening step, calculating the variance inflation factor for each feature and eliminating features with higher variance inflation factors to screen the features; a third fine-screening step, based on logistic regression, analyzing whether the feature coefficients of the features after the first fine-screening step and the second fine-screening step are consistent with the trend of the prediction results for credit default in order to further screen the features.

[0264] Feature fine-screening step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the previously introduced features are individually tested. If a previously introduced feature becomes less significant due to the introduction of subsequent features, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation.

[0265] Feature fine-screening step 2 is based on the method of eliminating features with high variance inflation factors to further reduce multicollinearity in the model.

[0266] Feature fine screening step 3 is based on the comparison of the risk characteristics of each risk point itself (system preset) with the positive and negative signs of the model training coefficients, and determines whether the feature coefficients of the remaining features in the feature fine screening step 3 in the model are consistent with the business trend, and the features whose model coefficients do not conform to the business trend are eliminated, and the iteration is repeated. The specific implementation plan of the feature fine screening module 3 is as follows: 1. For features whose feature conversion method is WOE type, the corresponding model training coefficient should be negative, and WOE conversion type features with positive training coefficients should be eliminated. 2. For continuous conversion methods, if the bad debt rate should increase as the value of this feature increases in business logic (such as the credit limit utilization rate, etc.), the corresponding model training coefficient should be positive, and such continuous conversion type features with negative training coefficients should be eliminated; if the bad debt rate should decrease as the value of this feature increases in business logic (such as the deposit amount, etc.), the corresponding model training coefficient should be negative, and such continuous conversion type features with positive training coefficients should be eliminated. 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0267] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0268] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model, and the iteration is stopped to obtain the final feature list and its conversion value. After the above steps, the input variables for constructing the model of this application can be obtained.

[0269] The credit default probability modeling step of the present application substitutes the features screened in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.

[0270] In the prior art, the core of all scoring models constructed lies in whether the data they use is representative. For the scoring models that already exist in the prior art, due to information security and cost reasons, the sample size and bad labels of most models are small, and the stability of the model cannot be guaranteed, let alone independent modeling of a certain type of customer. At the same time, because the information dimensions that can be obtained by the scoring models on the market during the modeling process are mostly multiple loan data and weak financial attribute data (such as smart terminal device data, social platform data, online shopping mall data and other non-credit transaction data), they cannot accurately reflect the customer's asset status and repayment ability. The data source used in this application is the full business data of large banks, and the amount of modeling samples and bad labels is very large. This application designs modeling samples based on different sample groups, which can more finely distinguish the risk differences between such customers, and at the same time ensure the stability of data and models based on cross-time verification and PSI verification. This application uses personal credit history data and asset data for model development, so it has a good reflection of the borrower's repayment willingness and repayment ability.

[0271] For some small and medium-sized financial institutions, the internally built scoring models are relatively small or started relatively late in their retail credit business, and the accumulated historical data is insufficient to develop a credit risk model with stable data and strong differentiation capabilities. Therefore, they rely heavily on manual approval for credit review. The efficiency limitations of manual approval have restricted the development of their retail credit business. At the same time, the subjectivity of manual approval has increased the operational risk in the credit review process. The model constructed in this application can assist such financial institutions in making digital decisions, enhance their approval accuracy and speed, and reduce the above-mentioned adverse effects.

[0272] In terms of feature conversion, compared to the traditional WOE conversion method, which requires coarse binning and discretization of continuous features based on data performance and the modeler's experience, the results of coarse binning are significantly affected by the modeler's subjective factors. Furthermore, the discretization process of continuous features may result in a large number of single-valued scores due to too many customers falling into the same interval. This application combines the continuous conversion method, the WOE conversion method, and the dummy feature conversion method to encode the original features, reducing the impact of human factors and single values ​​while enhancing the discriminative ability of the scores.

[0273] <Method for calculating retail credit risk> Please refer to the description of Example 12-1 and Example 12-6.

[0274] <Device, system, and computer storage medium for calculating retail credit risk>

[0275] The present application relates to an apparatus for constructing a retail credit risk prediction model, the apparatus comprising: a data acquisition module for acquiring original retail credit prediction data of a sample for constructing a model; a data derivation module for processing derived retail credit prediction data based on the original retail credit prediction data; a data processing module for processing second-generation derived retail credit prediction data based on the original retail credit prediction data and the derived retail credit prediction data; a feature initial screening module for performing preliminary screening of all categories, i.e., all features, including the original retail credit prediction data and the derived retail credit prediction data, to obtain features after preliminary screening; a preliminary screening data conversion module for determining a conversion method for the features after preliminary screening to confirm whether to use one of a WOE conversion method, a dummy feature conversion method, and a continuous conversion method for feature conversion, and for each feature after preliminary screening, using the determined optimal method for feature conversion; a feature fine screening module for performing in-depth screening of the features after preliminary screening to obtain fine-screened features; and a credit default probability modeling module for selecting a logistic regression method for model construction based on the relationship between the fine-screened features and the probability of credit default, and for determining a method for calculating the credit default probability.

[0276] The present application relates to a system for constructing a retail credit risk prediction model, which includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor. When the program for constructing a retail credit risk prediction model method is executed by the processor, the steps of constructing a retail credit risk prediction model method as described above are implemented.

[0277] All the contents described above for the method for calculating retail credit risk are fully applicable to the device, system and computer storage medium for calculating retail credit risk.

[0278] The method for calculating retail credit risk in this application can avoid the technical shortcomings of overfitting when building a model with complete WOE variables and the inability to adapt well to categorical variables when building a model with complete continuous variables. Therefore, when this model is used for retail credit risk calculation, it can better cover the retail credit risk prediction needs of various types of customer groups and meet the demand for more accurate prediction results when predicting credit risk as a general score.

[0279] Example

[0280] 1) Comprehensive business credit score

[0281] Example 1-1 Collection of computational modeling samples

[0282] During the development of this example, we collected credit card application and behavior data, personal loan application and behavior data, personal financial asset transaction data, and personal customer information data from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample, selecting 620 million data points from 2018 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 120 million and 500 million, respectively.

[0283] When designing the specific model, analysis and design will be conducted on the 120 million customers who have already applied. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0284] The model design includes 1) exclusion rules: excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 65.74 million customers; 2) time window setting: using 2018 data as the modeling sample and the next 24 months as the performance period of Y; 3) sample sampling: using a good / bad sample ratio of 4:1 for sampling modeling. Figure 1 The decision tree shown in the figure is used to design different segmentation schemes, and the parent-child model comparison method is used to confirm the final model segmentation scheme, and then each sub-model is modeled separately. The model segmentation scheme of this application is designed based on whether the customer has applied, whether he is a new customer, whether he is overdue, whether he has a mortgage, whether he has a consumer loan, and the customer's region. This fully reflects the control variables closely related to risk characteristics during business development and can more comprehensively match various customer groups in the market.

[0285] In this embodiment, the modeling samples are first divided into the following sub-model modeling samples based on the decision tree for the subsequent construction of sub-models. The first layer of the decision tree when determining the modeling samples is the account age, which is used to distinguish the length of the customer's credit history. In this embodiment, for customers with an account age of less than 3 (who have applied for registration for credit business but the application has not exceeded 3 months) (i.e., defined as new customers), the regions are divided according to the administrative divisions to which they belong. The division of administrative regions is, for example, derived from data released by authoritative statistical departments, or can be based on the division made by a rating agency, or can be based on another constructed financial model. It is fully understood by those skilled in the art that after selecting the division criteria, the user population can be divided into three different regions without duplication. At the same time, it is further combined with the three different situations of low, medium and high bad debt rates to divide them into three sub-models, Sigma1, Sigma2 and Sigma3, where the bad debt rate refers to the ratio of bad customers in a certain type of sample to the total number of samples in this type.

[0286] For customers with an account age greater than or equal to 3, the judgment is made based on their current overdue situation: if the current overdue period (i.e., the number of months) is greater than 3 periods, the Sigma4 sub-model is entered; if the current overdue period (i.e., the number of months) is less than or equal to 3 periods, the Sigma5 sub-model is entered.

[0287] For customers whose account age is greater than or equal to 3 and who are not currently overdue, a judgment is made based on whether they currently have a mortgage: if they do have a mortgage, the Sigma6 sub-model is entered.

[0288] For customers whose account age is greater than or equal to 3, who are not currently overdue and have no mortgage, the judgment is made based on whether they currently have consumer loans (including consumer loans, business loans and special installments): if they have the above loans, they enter the Sigma7 sub-model.

[0289] For customers whose account age is greater than or equal to 3, who are not currently overdue, have no mortgage, and do not have consumer loans / business loans / special installments, a judgment is made based on whether they use a revolving credit line (for example, a credit card, after repayment, the remaining balance is restored to the maximum balance, which is revolving): if a revolving credit line is used, the Sigma8 sub-model is entered; if a revolving credit line is not used, the Sigma9 sub-model is entered.

[0290] In addition, in this embodiment, in order to build a model for the sample group that has not applied for credit business registration, the first layer of the decision tree is also used as the account age, and only customers with an account age of less than 3 months are selected to enter the model. For customers with an account age of less than 3 months, the area is divided according to the administrative division to which they belong. The division of administrative areas is, for example, based on data released by an authoritative statistical department, or it can be based on the division made by a rating agency, or it can be divided based on another constructed financial model. Those skilled in the art can fully understand that after selecting the division criteria, the user population can be divided into three different areas without duplication. At the same time, it is further combined with the three different situations of low, medium and high bad debt rates to divide them into three sub-models, Sigma10, Sigma11, and Sigma12.

[0291] Specifically, the sample size used to construct the Sigma 1 sub-model is about 220,000. Its customer base is mainly new customers who have applied for credit business and are from economically underdeveloped areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0292] Specifically, the sample size used to construct the Sigma 2 sub-model is approximately 230,000. Its customer base is mainly new customers who have applied for credit business and are from moderately developed economic regions. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0293] Specifically, the sample size used to construct the Sigma 3 sub-model is about 460,000. Its customer base is mainly new customers who have applied for credit business and are from economically developed areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0294] Specifically, the sample size used to construct the Sigma 4 sub-model is approximately 1.19 million. Its customer base is mainly non-new customers who have applied for credit services and have already experienced serious overdue payments. The model is used to predict the probability of this group of people experiencing credit overdue payments of more than 90 days.

[0295] Specifically, the sample size used to construct the Sigma 5 sub-model is about 730,000. Its customer base is mainly non-new customers who have applied for credit business and have already experienced mild to moderate overdue payments. The model is used to predict the probability of this group of people experiencing credit overdue payments of more than 90 days.

[0296] Specifically, the sample size used to construct the Sigma 6 sub-model is about 14.25 million. Its customer base is mainly non-new customers who have applied for credit business and have not had any overdue payments but hold housing loans. The model is used to predict the probability of this group of people having credit overdue payments of more than 90 days.

[0297] Specifically, the sample size used to construct the Sigma 7 sub-model is approximately 5.77 million. Its customer base is mainly non-new customers who have applied for credit business and have not had any overdue payments but hold consumer loans. The model is used to predict the probability of this group of people having credit overdue payments of more than 90 days.

[0298] Specifically, the sample size used to construct the Sigma 8 sub-model is approximately 2.56 million. Its customer base is mainly non-new customers who have applied for credit business and have not had any overdue payments but only have credit cards and use them in a revolving manner. The model is built to predict the probability of this group of people having credit overdue payments of more than 90 days.

[0299] Specifically, the sample size used to construct the Sigma 9 sub-model is approximately 40.33 million. Its customer base is mainly non-new customers who have applied for credit business and have not had any overdue payments but only have credit cards and do not use them in a revolving manner. The model is built to predict the probability of this group of people having credit overdue payments of more than 90 days.

[0300] Specifically, the sample size used to construct the Sigma 10 sub-model is about 220,000. Its customer base is mainly potential customers who have not applied for credit business and are from economically underdeveloped areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0301] Specifically, the sample size used to construct the Sigma 11 sub-model is approximately 230,000. Its customer base is mainly potential customers who have not applied for credit business and are from moderately developed economic areas. The model is used to predict the probability of credit delinquency of more than 90 days in this group of people.

[0302] Specifically, the sample size used to construct the Sigma 12 sub-model is approximately 460,000. Its customer base is mainly potential customers who have not applied for credit business and are from economically developed areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0303] In this embodiment, whether a customer is a new customer refers to a customer with an account age of less than 3 years, that is, the time since the customer applied for credit business is less than 3 months. A serious overdue payment means that the customer has been overdue for more than 3 periods (i.e., the number of months). A moderate or light overdue payment means that the customer is currently overdue, and the number of overdue periods (i.e., the number of months) is less than or equal to 3 periods.

[0304] Based on the different customer samples of each sub-model confirmed above, historical data information of borrowers is obtained: 1) Credit card category, basic fields include account, overdue, balance, credit limit, repayment due, actual repayment and other information under different dimensions such as bills, cash withdrawals, installments, and interest; 2) Personal loan category, including account, overdue, balance, credit limit, repayment due, actual repayment and other information; 3) Customer basic information, including gender, age and administrative region information of the business; 4) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0305] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 3,067 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount,' 'Average Number of Overdue Payments in the Last 6 Months,' and 'Number of Credit Card Installments in the Last 12 Months.' In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0306] Table 1-1 Summary of basic variables and derived variables used in this example

[0307]

[0308] Example 1-2 Characteristic Screening

[0309] A preliminary screening was conducted on the 3067 features (variables) collected in Example 1-1 that have potential predictive power for customer overdue payments.

[0310] The preliminary screening of the nine sub-model features, Sigma 1 to Sigma 9, divided in Example 1-1, was performed as follows:

[0311] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 1-1 were first eliminated, resulting in a total of 154 variables being deleted, leaving 2912 variables.

[0312] In the second round of preliminary screening, for the 2912 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 87 variables were eliminated, leaving 2825 variables.

[0313] In the third round of preliminary screening, the 2825 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box, and if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 706 features were eliminated, and 2119 feature variables remained.

[0314] In the fourth round of preliminary screening, the 2119 features after the third round of screening were further screened based on the stepwise discrimination algorithm. After this round of screening, 221 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, this embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0315] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0316] The fifth round of preliminary screening targets the 221 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0317] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0318] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0319] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0320] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0321] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0322] The criteria for approximate consistency are as follows:

[0323] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0324] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0325] After the fifth round of screening, 66 features were eliminated, leaving 155 features.

[0326] For the three sub-models Sigma 10 to Sigma 12 divided in Example 1-1, the selected variables are customer basic information and personal financial asset variables (a total of 316 feature variables) used to approximately predict the probability of credit defaults in the future for customers who have not registered for credit applications. The initial feature screening step and subsequent steps are the same as those of other sub-models.

[0327] Example 1-3 Conversion of features after initial screening

[0328] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0329] First, the remaining 155 features in Example 1-2 are judged for their conversion methods, and the following three methods are selected based on the concentration of the features, data type, etc.

[0330] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0331] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0332] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0333] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0334] Through this embodiment, 155 features were converted, of which 32 features were converted to WOE, 10 features were converted to dummy features (expanded to 23 variables), and 113 features were converted to continuous features. The conversion of the post-primary screening features from the Sigma 10 to Sigma 12 submodels was also basically similar.

[0335] Example 1-4 Feature Depth Screening (Feature Fine Screening Step)

[0336] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0337] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the total number of features from 168 to 56.

[0338] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 56 features to 48.

[0339] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0340] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0341] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0342] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0343] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0344] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0345] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 48 ​​features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0346] Example 1-Construction of 5Sigma 1 Model

[0347] Generally speaking, the stronger the correlation between the feature variables and the target variable, the more accurate the final model. In this example, Sigma 1 is used as a submodel, employing the aforementioned decision tree method. The sample size obtained for classification is approximately 220,000. This sample primarily consists of new customers who have applied for credit and are located in economically underdeveloped areas. The model is constructed to predict the probability (target variable) of this group experiencing a credit delinquency of more than 90 days.

[0348] In the present embodiment 1-5, it is finally confirmed that the six features of the average utilization rate of the credit revolving line in the past three months, the remaining available credit card limit, the average number of cash withdrawals from the credit card in the past three months, the minimum asset size in the past six months, the deposit situation at the current time point, and the average interest of the credit card in the past three months are used as examples to describe the final modeling results. In the feature conversion step, the six modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0349] The average utilization rate of the revolving credit line over the past three months reflects the borrower's credit demand. In economically underdeveloped regions, large, short-term credit demands can lead to excessive debt and ultimately delinquency. The characteristic is that the lower the average utilization rate over the past three months, the lower the borrower's credit demand and the lower the likelihood of a serious default. Conversely, the higher the average utilization rate over the past three months, the higher the borrower's credit demand and the higher the likelihood of a serious default. The conversion method is WOE.

[0350] The current available credit limit on a credit card reflects the borrower's credit need. In economically underdeveloped regions, excessive spending caused by an imbalance between income and expenditure can easily lead to credit card installments or even delinquencies. Therefore, the remaining credit limit on a credit card can be a good indicator of a customer's spending behavior. The higher the current available credit limit on a credit card, the lower the borrower's credit need and the lower the likelihood of a serious default. Conversely, the lower the current available credit limit on a credit card, the higher the borrower's credit need and the higher the likelihood of a serious default. The conversion method is the cube root conversion within the continuous conversion method.

[0351] The average number of credit card cash withdrawals over the past three months reflects the borrower's credit need. In economically underdeveloped regions, short-term financial pressures may lead customers to withdraw cash from their credit cards at high interest rates, further exacerbating financial pressure and ultimately leading to delinquency. The lower the average number of credit card cash withdrawals over the past three months, the lower the borrower's credit need and the lower the likelihood of a serious default. Conversely, the higher the average number of credit card cash withdrawals over the past three months, the higher the borrower's credit need and the higher the likelihood of a serious default. The conversion method is WOE.

[0352] The minimum asset size over the past six months reflects the borrower's worst-case asset situation. For economically underdeveloped regions, the worst-case medium- and long-term asset size helps gauge repayment pressure. The larger the minimum asset size over the past six months, the higher the borrower's asset quality and the lower the likelihood of a major default. Conversely, the smaller the minimum asset size over the past six months, the lower the borrower's asset quality and the higher the likelihood of a major default. The conversion method is WOE.

[0353] The current balance in the deposit account reflects the borrower's recent asset level. For economically underdeveloped regions, the current deposit situation directly reflects the borrower's qualifications. The lower the current deposit account balance, the lower the borrower's solvency and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the stronger the borrower's solvency and the lower the likelihood of a serious default. The conversion method is the cube root conversion of the continuous conversion method.

[0354] The average credit card interest rate over the past three months reflects a borrower's credit usage. In economically underdeveloped regions, high credit card interest rates indicate a decline in repayment capacity. The lower the average credit card interest rate over the past three months, the stronger the borrower's repayment capacity and the lower the likelihood of a serious default. Conversely, the higher the average credit card interest rate over the past three months, the weaker the borrower's repayment capacity and the higher the likelihood of a serious default. The conversion method is WOE.

[0355] In the calculation model acquisition step, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0356]

[0357] Where k is the number of features entering the model, and in Formula 1-1, k is 6.

[0358] α is the intercept term, with a range of values ​​from (-0.7434, -0.4926), and an optimal value of -0.618; β1 is the coefficient corresponding to the average utilization rate of the credit revolving limit in the past three months, with a range of values ​​from (-0.6703, -0.6037), and an optimal value of -0.637; β2 is the coefficient corresponding to the current credit card balance, with a range of values ​​from (-0.0649, -0.0531), and an optimal value of -0.059; β3 is the coefficient corresponding to the average number of credit card withdrawals in the past three months, with a range of values ​​from (-0.52 36, -0.3864), with an optimal value of -0.455; β4 is the coefficient corresponding to the minimum asset size in the past six months, with a range of values ​​(-0.0743, -0.1057), and an optimal value of -0.090; β5 is the coefficient corresponding to the current deposit account balance, with a range of values ​​(-0.3466, -0.5034), and an optimal value of -0.425; β6 is the coefficient corresponding to the average credit card interest rate in the past three months, with a range of values ​​(-0.2373, -0.3627), and an optimal value of -0.300. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0359] x1 is the Word of Exchange Rate (WOE)-converted value of the average revolving credit limit utilization rate over the past three months, generated by the feature transformation step; x2 is the cube root-converted value of the current remaining available credit limit on the credit card, generated by the feature transformation step; x3 is the Word of Exchange Rate (WOE)-converted value of the average number of credit card withdrawals over the past three months, generated by the feature transformation step; x4 is the Word of Exchange Rate (WOE)-converted value of the minimum asset size over the past six months, generated by the feature transformation step; x5 is the cube root-converted value of the current remaining balance in the deposit account, generated by the feature transformation step; and x6 is the Word of Exchange Rate (WOE)-converted value of the average credit card interest rate over the past three months, generated by the feature transformation step. The model performance of some features is shown in Table 1-2 below.

[0360] Table 1-2

[0361]

[0362] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0363] Example 1-6 Construction of Sigma 2 Model

[0364] Taking Sigma 2 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 230,000. Its customer base was mainly new customers who had applied for credit business and were from moderately developed economic areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0365] In the present embodiment 1-6, it is finally confirmed that the final modeling result is described by taking seven features as an example: the current total consumption loan utilization rate, the average utilization rate of the revolving loan limit in the past three months, the current credit card remaining available limit, the average monthly asset size in the past three months, the average number of cash withdrawals from the credit card in the past three months, the average monthly salary in the past six months, and the current deposit account balance. In the feature conversion step, the seven modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the seven features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0366] The current total consumer loan utilization rate reflects the borrower's credit demand. For regions with medium economic development, the repayment pressure on short-term consumer credit loans may lead to overdue payments due to excessive borrowing. If the feature "Current Total Consumer Loan Utilization Rate" is missing, meaning there is no history of consumer loans, the lower the borrower's credit demand, the lower the likelihood of a serious default. Conversely, if the feature "Current Total Consumer Loan Utilization Rate" is present, meaning there is a history of consumer loans, the higher the borrower's credit demand, the higher the likelihood of a serious default. The conversion method is WOE.

[0367] The average utilization rate of the revolving credit line over the past three months reflects the borrower's credit need. For regions with medium economic development, the revolving credit line utilization rate reflects the health of the borrower's credit habits. The lower the average utilization rate over the past three months, the lower the borrower's credit need and the lower the likelihood of a serious default. Conversely, the higher the average utilization rate over the past three months, the higher the borrower's credit need and the higher the likelihood of a serious default. The conversion method is WOE.

[0368] The current available credit limit on a credit card reflects the borrower's credit demand. For regions with medium economic development, the current available credit limit reflects the borrower's reliance on credit. The higher the current available credit limit, the lower the borrower's credit demand and the lower the likelihood of a serious default. Conversely, the lower the current available credit limit, the higher the borrower's credit demand and the higher the likelihood of a serious default. The conversion method is the square root conversion of the continuous conversion method.

[0369] The average monthly asset size over the past three months reflects the borrower's debt repayment capacity. For regions with medium economic development, the average monthly asset size over the past three months reflects the borrower's recent financial strength. The lower the average monthly asset size over the past three months, the lower the borrower's debt repayment capacity and the higher the likelihood of a serious default. Conversely, the higher the average monthly asset size over the past three months, the stronger the borrower's debt repayment capacity and the lower the likelihood of a serious default. The conversion method is the square root conversion in the continuous conversion method.

[0370] The average number of credit card cash withdrawals over the past three months reflects a borrower's credit usage. For regions with medium economic development, this average number of withdrawals over the past three months reflects a borrower's recent reliance on high-interest credit cards. The lower the average number of credit card withdrawals over the past three months, the more prudent the borrower's credit use and the lower the likelihood of a serious default. Conversely, the higher the average number of credit card withdrawals over the past three months, the more aggressive the borrower's credit use and the higher the likelihood of a serious default. The conversion method is WOE.

[0371] The average monthly salary over the past six months reflects the borrower's ability to acquire assets. For regions with medium economic development, the average monthly salary over the past six months provides an early indicator of the borrower's medium- and long-term debt repayment capacity. The lower the average monthly salary over the past six months, the weaker the borrower's debt repayment capacity and the higher the likelihood of a major default. Conversely, the higher the average monthly salary over the past six months, the stronger the borrower's debt repayment capacity and the lower the likelihood of a major default. The conversion method is WOE.

[0372] The current deposit account balance reflects the borrower's deposit level. For regions with medium economic development, current deposits are a key indicator of near-term repayment ability. The lower the current deposit account balance, the weaker the borrower's near-term repayment ability and the higher the likelihood of a serious default. Conversely, the higher the past-current deposit account balance, the stronger the borrower's near-term repayment ability and the lower the likelihood of a serious default. The transformation method is the natural logarithm of the continuous transformation.

[0373] In the calculation model acquisition step, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0374]

[0375] Where k is the number of features entering the model, and in Formula 1-2, k is 7.

[0376] α is the intercept term, with a range of values ​​from (-0.05096, -0.43904), and an optimal value of -0.245; β1 is the coefficient corresponding to the current utilization rate of the total consumer loan amount, with a range of values ​​from (-1.82132, -2.04868), and an optimal value of -1.935; β2 is the coefficient corresponding to the average utilization rate of the revolving loan amount in the past three months, with a range of values ​​from (-0.48984, -0.57216), and an optimal value of -0.531; β3 is the coefficient corresponding to the current remaining available credit card amount, with a range of values ​​from (-0.04112, -0.05288), and an optimal value of -0.047; β4 is The coefficient corresponding to the average monthly asset size over the past three months ranges from (-0.20864 to -0.27136), with an optimal value of -0.24. β5 is the coefficient corresponding to the average number of credit card cash withdrawals over the past three months, with a range from (-0.41744 to -0.55856), and an optimal value of -0.488. β6 is the coefficient corresponding to the average monthly salary over the past six months, with a range from (-1.02372 to -1.48628), and an optimal value of -1.255. β7 is the coefficient corresponding to the current deposit account balance, with a range from (-0.06136 to -0.09664), and an optimal value of -0.079. (Note: The numerical ranges are derived from the 95% confidence intervals, i.e., the 95% CI in the table below.)

[0377] x1 is the Word of Exchange (WOE)-converted value of the current total consumer loan utilization rate, generated in the feature conversion step; x2 is the Word of Exchange (WOE)-converted value of the average revolving loan utilization rate over the past three months, generated in the feature conversion step; x3 is the square root-converted value of the current remaining available credit limit on the credit card, generated in the feature conversion step; x4 is the square root-converted value of the average monthly asset size over the past three months, generated in the feature conversion step; x5 is the Word of Exchange (WOE)-converted value of the average number of credit card cash withdrawals over the past three months, generated in the feature conversion step; x6 is the Word of Exchange (WOE)-converted value of the average monthly salary over the past six months, generated in the feature conversion step; and x7 is the natural logarithm-converted value of the current deposit account balance, generated in the feature conversion step. The model performance of some features is shown in Tables 1-3 below:

[0378] Table 1-3

[0379]

[0380] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0381] Example 1-7 Construction of Sigma3 Model

[0382] Taking Sigma 3 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 460,000. Its customer base was mainly new customers who had applied for credit business and were from economically developed areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0383] In the present embodiment 1-7, it is finally confirmed that the final modeling result is described by taking six features as an example: the proportion of months with a credit utilization rate greater than 10% in the past three months, the current total amount of personal loan revolving loans, the number of months with a credit utilization rate greater than 90% in the past three months, the average monthly asset size in the past 12 months, the number of months with a repayment rate greater than or equal to 100% in the past three months, and the current number of credit card cash withdrawals. In the feature conversion step, the six modeling variables have been converted in a corresponding form according to the correlation between the features and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0384] The percentage of months in the past three months with a credit limit utilization rate exceeding 10% reflects the borrower's credit demand and usage habits. In economically developed regions, credit limits for credit loans (including credit cards) are generally higher, so their revolving credit limit utilization rate reflects their willingness to use credit. The lower the percentage of months in the past three months with a revolving credit limit utilization rate greater than 10%, the lower the borrower's credit demand and the lower the likelihood of a serious default; conversely, the higher the percentage of months in the past three months with a revolving credit limit utilization rate greater than 10%, the higher the borrower's credit demand and the higher the likelihood of a serious default. The conversion method is WOE.

[0385] The current total revolving loan limit for individual loans reflects the borrower's creditworthiness. For economically developed regions, the mere fact of being approved for a revolving loan reflects a borrower's creditworthiness, and the ability to obtain a high-value loan indicates a higher level of assets. The higher the current total revolving loan limit for individual loans, the stronger the borrower's creditworthiness and the lower the likelihood of a serious default. Conversely, the lower the current total revolving loan limit for individual loans, the weaker the borrower's creditworthiness and the higher the likelihood of a serious default. The conversion method is WOE.

[0386] The number of months in the past three months with a credit limit utilization rate exceeding 90% reflects the borrower's credit need. For economically developed regions, recent frequent and substantial credit limit utilization indicates significant financial pressure on the client, potentially leading to overdue payments. The higher the number of months in the past three months with a credit limit utilization rate exceeding 90%, the higher the borrower's credit need and the higher the likelihood of a serious default. Conversely, the lower the number of months in the past three months with a credit limit utilization rate exceeding 90%, the lower the borrower's credit need and the lower the likelihood of a serious default. The conversion method is WOE.

[0387] The average monthly asset size over the past 12 months reflects a borrower's debt repayment capacity. For economically developed regions, the long-term average asset size reflects a client's ability to accumulate assets over time. The lower the average monthly asset size over the past 12 months, the lower the borrower's debt repayment capacity and the higher the likelihood of a major default. Conversely, the higher the average monthly asset size over the past 12 months, the stronger the borrower's debt repayment capacity and the lower the likelihood of a major default. The transformation method is the natural logarithm transformation within the continuous transformation.

[0388] The number of months in the past three months with a credit repayment ratio greater than or equal to 100% reflects the borrower's credit repayment performance. In economically developed regions, where incomes are generally higher, installment payments or overdue payments may indicate a decline in the borrower's repayment ability. The fewer months in the past three months with a credit repayment ratio greater than or equal to 100%, the better the borrower's credit repayment performance and the lower the likelihood of a serious default. Conversely, the more months in the past three months with a credit repayment ratio greater than or equal to 100%, the worse the borrower's credit repayment performance and the higher the likelihood of a serious default. The conversion method is WOE.

[0389] The current number of credit card cash withdrawals reflects the borrower's credit repayment status. In economically developed regions, where incomes are generally higher, credit card cash withdrawals may indicate increasing financial pressure on borrowers. The lower the number of credit card cash withdrawals, the less financial pressure on the borrower and the higher the likelihood of a serious default. Conversely, the higher the number of credit card cash withdrawals, the greater the credit pressure on the borrower and the higher the likelihood of a serious default. The conversion method is WOE.

[0390] In the calculation model acquisition step, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0391]

[0392] Where k is the number of features entering the model, and in formula 1-3, k is 6.

[0393] α is the intercept term, with a range of values ​​from (-2.32928, -2.45472), and the optimal value is -2.392; β1 is the corresponding coefficient for the number of months in the past three months when the credit limit utilization rate is greater than 10%, with a range of values ​​from (-0.60156, -0.65644), and the optimal value is -0.629; β2 is the corresponding coefficient for the total amount of the current personal loan revolving loan, with a range of values ​​from (-0.84424, -0.96576), and the optimal value is -0.905; β3 is the corresponding coefficient for the number of months in the past three months when the credit limit utilization rate is greater than 90%, with a range of values ​​from (-0.499 76, -0.57424), with an optimal value of -0.537; β4 is the coefficient corresponding to the average monthly asset size over the past 12 months, with a range of values ​​from (-0.05612 to -0.06788), and an optimal value of -0.062; β5 is the coefficient corresponding to the number of months with a credit repayment rate greater than or equal to 100% over the past three months, with a range of values ​​from (-0.74304 to -0.94296), and an optimal value of -0.843; β6 is the coefficient corresponding to the current number of credit card withdrawals, with a range of values ​​from (-0.35724 to -0.47876), and an optimal value of -0.418. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0394] x1 is the WOE conversion value of the percentage of months in the past three months with a credit utilization rate greater than 10% generated in the feature conversion step; x2 is the WOE conversion value of the current total amount of personal revolving loans generated in the feature conversion step; x3 is the WOE conversion value of the number of months in the past three months with a credit utilization rate greater than 90% generated in the feature conversion step; x4 is the natural logarithm conversion value of the average monthly asset size in the past 12 months generated in the feature conversion step; x5 is the WOE conversion value of the number of months in the past three months with a credit repayment rate greater than or equal to 100% generated in the feature conversion step; x6 is the WOE conversion value of the current number of credit card cash withdrawals generated in the feature conversion step.

[0395] The model performance of some features is shown in Table 1-4 below:

[0396] Table 1-4

[0397]

[0398] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0399] Example 1-8 Sigma 4 model construction

[0400] Taking Sigma 4 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 1.19 million. Its customer base was mainly non-new customers who had applied for credit business and had serious overdue customers. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0401] In Example 1-8, it is finally confirmed that the final modeling results are described by taking five features as an example: the number of months from the time when the credit card interest amount in the past 12 months is greater than 0 to the observation point, the maximum number of overdue periods for the current credit customer, the current credit card installment balance, the proportion of the longest continuous increase in the deposit account balance in the past 3 months, and the number of months from the last credit card repayment amount in the past 12 months to the observation point, which is greater than the minimum repayment amount of the previous period. In the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0402] The number of months in the past 12 months during which credit card interest payments were greater than 0 up to the observation point reflects the borrower's credit demand. For borrowers currently experiencing significant overdue payments, overdue payments may occur across all their products, as overdue payments are transmitted across products. Overdue status on credit card repayments often deteriorates from installment payments or minimum payments. Since both installment payments and minimum payments accrue interest, determining when credit card interest first accrued can estimate the likelihood of improvement in repayment status. The fewer the number of months in the past 12 months during which credit card interest payments were greater than 0 up to the observation point, the greater the likelihood that the borrower will improve and the lower the likelihood that they will remain in default. Conversely, the more months in the past 12 months during which credit card interest payments were greater than 0 up to the observation point, the lower the likelihood that the borrower will improve and the higher the likelihood that they will remain in default. The conversion method is WOE.

[0403] The current maximum overdue period for a credit customer reflects the borrower's delinquency status. For credit customers currently experiencing severe overdue payments, their worst-case overdue status reflects the likelihood of returning to normal repayments. The higher the current maximum overdue period for a credit customer, the more severe the borrower's delinquency and the higher the likelihood of continued default. Conversely, the lower the current maximum overdue period for a credit customer, the less severe the borrower's delinquency and the higher the likelihood of escaping default. The conversion method is WOE.

[0404] The current credit card installment balance reflects the borrower's repayment pressure. For credit customers currently experiencing significant overdue payments, their credit card installment balance reflects the difficulty of repaying their loans. The higher the current credit card installment balance, the greater the repayment pressure on the borrower and the higher the likelihood of continued default. Conversely, the lower the current credit card installment balance, the less repayment pressure on the borrower and the higher the likelihood of escaping default. The conversion method is WOE.

[0405] The percentage of months with the longest consecutive increase in deposit account balances over the past three months reflects the borrower's deposit balance changes. For borrowers with serious overdue payments, this percentage reflects their earning capacity. The lower the percentage of months with the longest consecutive increase in deposit account balances over the past three months, the fewer months of deposit increases the borrower has experienced, and the higher the likelihood of continued default. Conversely, the higher the percentage of months with the longest consecutive increase in deposit account balances over the past three months, the more months of deposit increases the borrower has experienced, and the higher the likelihood of escaping default. The conversion method is WOE.

[0406] The number of months since the last credit card payment in the past 12 months exceeded the previous minimum payment reflects the borrower's repayment ability. For seriously overdue borrowers, the month since the change in repayment behavior indicates the likelihood of returning to normal repayment status. The longer the number of months since the last credit card payment in the past 12 months exceeded the previous minimum payment, the weaker the borrower's recent credit repayment ability and the higher the likelihood of continued default. Conversely, the shorter the number of months since the last credit card payment in the past 12 months exceeded the previous minimum payment, the stronger the borrower's recent repayment ability and the lower the likelihood of default. The conversion method is WOE.

[0407] In the step of obtaining the calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0408]

[0409] Where k is the number of features entering the model. In formula 1-4, k is 5.

[0410] α is the intercept term, with a range of values ​​from (3.49848, 3.44752) and an optimal value of 3.473; β1 is the coefficient corresponding to the number of months from the past 12 months when the credit card interest amount was greater than 0 to the observation point, with a range of values ​​from (-1.7838, -1.8622) and an optimal value of -1.823; β2 is the coefficient corresponding to the maximum number of overdue periods for credit customers, with a range of values ​​from (-0.59456, -0.64944) and an optimal value of -0.622; β3 is the coefficient corresponding to the current installment balance of the credit card, The values ​​range from (-0.71372, -0.78428), with an optimal value of -0.749. β4 is the coefficient corresponding to the proportion of consecutive months with the longest increase in deposit account balances over the past three months, with a range from (-0.3624, -0.4016), and an optimal value of -0.382. β5 is the coefficient corresponding to the number of months from the observation point when the last credit card payment in the past 12 months was greater than the previous minimum payment, with a range from (-0.60592, -0.69608), and an optimal value of -0.651. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0411] x1 is the WOE conversion value of the number of months from the past 12 months when the interest amount of the credit card is greater than 0 to the observation point, generated by the feature conversion step; x2 is the WOE conversion value of the maximum number of overdue periods of the current credit customer, generated by the feature conversion step; x3 is the WOE conversion value of the current credit card installment balance, generated by the feature conversion step; x4 is the WOE conversion value of the proportion of the longest consecutive increase in the deposit account balance in the past three months, generated by the feature conversion step; x5 is the WOE conversion value of the number of months from the observation point when the last credit card repayment amount in the past 12 months is greater than the minimum repayment amount of the previous period, generated by the feature conversion step.

[0412] The model performance of some features is shown in Table 1-5 below:

[0413] Table 1-5

[0414]

[0415] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0416] Example 1-9 Sigma 5 model construction

[0417] Taking Sigma 5 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 730,000. Its customer base was mainly non-new customers who had applied for credit business and had mild to moderate overdue customers. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0418] In Example 1-9, it is finally confirmed that the final modeling results are described by taking five features as examples: the number of months from the time the credit card interest amount is greater than 0 to the observation point in the past 12 months, the remaining available credit limit of the current credit card, the longest consecutive number of months in which the credit customer’s overdue payment is greater than 0 in the past 12 months, the longest consecutive number of months in which the credit card limit utilization rate is greater than 90% in the past 12 months, and the proportion of the number of months in which the credit card overdue period is greater than or equal to 1 in the past 12 months. In the selection of the conversion method in the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0419] The number of months in the past 12 months during which credit card interest payments were greater than zero up to the observation point reflects the borrower's credit demand. For borrowers currently experiencing mild to moderate delinquency, delinquency may occur across all products, as delinquency is contagious across products. Credit card delinquency often deteriorates from installment payments or minimum payments. Since both installment payments and minimum payments accrue interest, determining when interest first accrues can help distinguish between delinquency caused by excessive borrowing and accidental forgetfulness. The fewer the number of months in the past 12 months during which credit card interest payments were greater than zero up to the observation point, the closer the borrower's default occurred, and the lower the likelihood that the delinquency will escalate into a major default. Conversely, the longer the number of months in the past 12 months during which credit card interest payments were greater than zero up to the observation point, the further the borrower's default occurred, and the higher the likelihood that the delinquency will escalate into a major default. The conversion method is WOE.

[0420] The current remaining available credit limit on a credit card reflects the borrower's credit need. For borrowers currently experiencing minor to moderate overdue payments, their remaining credit limit reflects the difficulty of repaying in full. The lower the current remaining available credit limit, the lower the borrower's credit need and the lower the likelihood that an overdue payment will escalate into a serious default. Conversely, the higher the current remaining available credit limit, the higher the borrower's credit need and the higher the likelihood that an overdue payment will escalate into a serious default. The conversion method is the cube root conversion of the continuous conversion method.

[0421] The longest consecutive number of months with more than zero overdue payments over the past 12 months reflects the borrower's overdue status. For borrowers currently experiencing mild to moderate overdue payments, the longest consecutive number of months with more than zero overdue payments over the past 12 months reflects the frequency of their overdue payments. The higher the number of consecutive months with more than zero overdue payments over the past 12 months, the more frequent the overdue behavior and the higher the likelihood that the overdue payments will become a serious default. Conversely, the lower the number of consecutive months with more than zero overdue payments over the past 12 months, the more occasional the overdue behavior and the lower the likelihood that the overdue payments will become a serious default. The conversion method is WOE.

[0422] The longest consecutive number of months with a credit card limit utilization rate exceeding 90% over the past 12 months reflects changes in the borrower's deposit balance. For borrowers currently experiencing minor to moderate delinquencies, recent changes in income indicate changes in their debt repayment ability. The lower the number of consecutive months with a credit card limit utilization rate exceeding 90% over the past 12 months, the fewer months the borrower's deposit increased and the higher the likelihood that the delinquency will become a major default. Conversely, the longer the longest consecutive number of months with a credit card limit utilization rate exceeding 90% over the past 12 months, the more months the borrower's deposit increased and the lower the likelihood that the delinquency will become a major default. The conversion method is WOE.

[0423] The percentage of months with one or more credit card delinquencies over the past 12 months reflects the borrower's repayment status. For borrowers with mild to moderate delinquencies, the percentage of months with one or more credit card delinquencies over the past 12 months reflects the limits of their repayment ability. The smaller the percentage of months with one or more credit card delinquencies over the past 12 months, the shorter the period of deterioration in the borrower's repayment ability and the lower the likelihood that the delinquency will turn into a serious default. Conversely, the larger the percentage of months with one or more credit card delinquencies over the past 12 months, the longer the period of deterioration in the borrower's repayment ability and the higher the likelihood that the delinquency will turn into a serious default. The conversion method is WOE.

[0424] In the step of obtaining the calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0425]

[0426] Where k is the number of features entering the model. In formula 1-5, k is 5.

[0427] α is the intercept term, with a range of values ​​of (1.2853, 1.3265) and an optimal value of 1.3059; β1 is the coefficient corresponding to the number of months from the time the interest amount of the credit card in the past 12 months was greater than 0 to the observation point, with a range of values ​​of (-0.5117, -0.4925) and an optimal value of -0.5021; β2 is the coefficient corresponding to the current remaining available credit limit of the credit card, with a range of values ​​of (-0.0631, -0.0603) and an optimal value of -0.0617; β3 is the coefficient corresponding to the number of months from the time the credit customer’s overdue payment was greater than 0 in the past 12 months. The coefficient for the longest consecutive number of months ranges from (-0.3677 to -0.3445), with an optimal value of -0.3561. β4 is the coefficient for the longest consecutive number of months with a credit card limit utilization rate > 90% over the past 12 months, with a range from (-0.3344 to -0.3124), with an optimal value of -0.3234. β5 is the coefficient for the proportion of months with one or more credit card overdue payments over the past 12 months, with a range from (-0.3369 to -0.3141), with an optimal value of -0.3255. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0428] x1 is the WOE conversion value of the number of months from the past 12 months when the interest amount of the credit card is greater than 0 to the observation point, generated by the feature conversion step; x2 is the continuous cube root conversion value of the remaining available credit limit of the current credit card, generated by the feature conversion step; x3 is the WOE conversion value of the longest consecutive number of months with overdue payments of credit customers greater than 0 in the past 12 months, generated by the feature conversion step; x4 is the WOE conversion value of the longest consecutive number of months with a credit card limit utilization rate greater than 90% in the past 12 months, generated by the feature conversion step; x5 is the WOE conversion value of the proportion of months with overdue payments of credit cards greater than or equal to 90% in the past 12 months, generated by the feature conversion step.

[0429] The model performance of some features is shown in Table 1-6 below:

[0430] Table 1-6

[0431]

[0432] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0433] Example 1-10Sigma 6 model construction

[0434] Taking Sigma 6 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 14.25 million. Its customer base was mainly non-new customers who had applied for credit business and had no overdue customers with mortgage loans. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0435] In Example 1-10, it is finally confirmed that the final modeling results are described by taking eight features as an example: the maximum number of overdue periods in the past 12 months, the ratio of credit card interest to bill balance in the past 3 months, the current credit card remaining available limit, the minimum asset size in the past 6 months, the deposit account balance at the current time point, the current credit card limit utilization rate, whether it is a payroll customer, and the number of months with a credit limit utilization rate greater than 50% in the past 12 months. In the feature conversion step, the eight modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the eight features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0436] The maximum number of overdue payments over the past 12 months reflects the borrower's historical credit performance. For borrowers holding housing loans, overdue payments may indicate a deterioration in repayment ability. The fewer the maximum overdue payments over the past 12 months, the better the borrower's credit performance and the lower the likelihood of a serious default. Conversely, the more overdue payments over the past 12 months, the worse the borrower's credit performance and the higher the likelihood of a serious default. The conversion method is WOE.

[0437] The ratio of credit card interest to bill balance over the past three months reflects a borrower's credit demand. For borrowers holding mortgages, excessive spending due to an imbalance between income and expenditure is likely to lead to credit card installment payments. Therefore, the remaining balance on a credit card is a good indicator of a customer's credit behavior. The lower the current available credit limit, the higher the borrower's credit demand and the greater the likelihood of a serious default. Conversely, the higher the current available credit limit, the lower the borrower's credit demand and the lower the likelihood of a serious default. The conversion method is the square root conversion within the continuous conversion method.

[0438] The current available credit limit on a credit card reflects the borrower's credit needs. For borrowers holding mortgages, credit card interest typically accrues from installments, minimum payments, or overdue payments. Therefore, the incurring of credit card interest indicates financial constraints. The lower the current available credit limit, the greater the borrower's financial pressure and the lower the likelihood of a serious default. Conversely, the higher the current available credit limit, the less financial pressure the borrower faces and the higher the likelihood of a serious default. The conversion method is WOE.

[0439] The minimum asset size over the past six months reflects the medium-term volatility of a borrower's assets. For borrowers holding mortgages, the worst asset level in the short to medium term provides a more objective assessment of their asset strength. The lower the minimum asset size over the past six months, the lower the borrower's asset level and the higher the likelihood of a serious default. Conversely, the higher the minimum asset size over the past six months, the more aggressive the borrower's credit use and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0440] The current deposit account balance reflects the borrower's asset status. For borrowers holding housing loans, the current deposit situation directly reflects the borrower's ability to make continuous repayments. The lower the current deposit account balance, the lower the borrower's solvency and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the stronger the borrower's solvency and the lower the likelihood of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0441] The current credit card limit utilization rate reflects the borrower's current credit needs. For borrowers holding mortgages, this rate reflects their current financial pressures. The lower the current credit card limit utilization rate, the less repayment pressure the borrower faces and the lower the likelihood of a serious default. Conversely, the higher the current credit card limit utilization rate, the greater the repayment pressure and the higher the likelihood of a serious default. The conversion method is WOE.

[0442] Whether or not a borrower is a payroll agent reflects the stability of their current funding sources. For borrowers holding housing loans, the repayment pressure of a payroll agent can be assessed based on their income. Payroll agents have a more stable funding source and are less likely to default. Conversely, non-payroll agents have a higher likelihood of default due to the unmonitored funding source. The conversion method is WOE.

[0443] The number of months in the past 12 months with a credit limit utilization ratio exceeding 50% reflects the borrower's recent credit needs. For borrowers with mortgages, this variable reflects their repayment pressure. The more months in the past 12 months with a credit limit utilization ratio exceeding 50%, the greater the repayment pressure and the higher the likelihood of a serious default. Conversely, the fewer months in the past 12 months with a credit limit utilization ratio exceeding 50%, the lower the likelihood of a serious default. The conversion method is WOE.

[0444] In the step of obtaining the calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0445]

[0446] Where k is the number of features entering the model. In formula 1-6, k is 8.

[0447] α is the intercept term, with a range of values ​​from (-3.14952, -3.20048), and an optimal value of -3.175; β1 is the coefficient corresponding to the maximum number of overdue periods in the past 12 months, with a range of values ​​from (-0.48512, -0.49688), and an optimal value of -0.491; β2 is the coefficient corresponding to the ratio of credit card interest to bill balance in the past three months, with a range of values ​​from (-0.3742, -0.3938), and an optimal value of -0.384; β3 is the coefficient corresponding to the current remaining available credit card balance, with a range of values ​​from (-0.0008, -0.0012), and an optimal value of -0.001; β4 is the coefficient corresponding to the minimum asset size in the past six months, with a range of values ​​from (-0.0800 8, -0.08792), with an optimal value of -0.084; β5 is the coefficient corresponding to the current deposit account balance, with a range of values ​​(-0.05208, -0.05992), and an optimal value of -0.056; β6 is the coefficient corresponding to the current credit card limit utilization rate, with a range of values ​​(-0.1542, -0.1738), and an optimal value of -0.164; β7 is the coefficient corresponding to whether the customer is a payroll agent, with a range of values ​​(-0.65276, -0.72724), and an optimal value of -0.69; β8 is the coefficient corresponding to the number of months in the past 12 months with a credit limit utilization rate greater than 0%, with a range of values ​​(-0.45264, -0.51536), and an optimal value of -0.484. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0448] x1 is the WOE conversion value of the maximum number of overdue periods in the past 12 months generated by the feature conversion step; x2 is the square root conversion value of the ratio of credit card interest to bill balance in the past 3 months generated by the feature conversion step; x3 is the WOE conversion value of the current remaining available credit card limit generated by the feature conversion step; x4 is the natural logarithm conversion value of the minimum asset size in the past 6 months generated by the feature conversion step; x5 is the natural logarithm conversion value of the deposit account balance at the current time point generated by the feature conversion step; x6 is the WOE conversion value of the current credit card limit utilization rate generated by the feature conversion step; x7 is the WOE conversion value of whether the company is a payroll customer generated by the feature conversion step; x8 is the WOE conversion value of the number of months in the past 12 months with a credit limit utilization rate greater than 50% generated by the feature conversion step.

[0449] The model performance of some features is shown in Table 1-7 below:

[0450] Table 1-7

[0451]

[0452] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0453] Example 1-11 Sigma7 model construction

[0454] Taking Sigma 7 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 5.77 million. Its customer base was mainly non-new customers who had applied for credit business and had no overdue loans but held consumer loans. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0455] In Example 1-11, it is finally confirmed that the final modeling results are described by taking the average credit card limit utilization rate in the past 6 months, the number of overdue months of credit customers in the past 6 months, the current credit card remaining available limit, the maximum increase in the balance of investment and financial management accounts in the past 3 months, the current total consumer loan limit utilization rate, the number of months of credit card cycle or installment use in the past 12 months, and the difference between the minimum repayment amount of the average monthly credit in the past 3 months and the average monthly asset size as an example. In the feature conversion step conversion method selection, the 7 model variables have been converted in the corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the 7 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0456] The average credit card utilization rate over the past six months reflects the borrower's credit demand. For customers holding consumer loans, an excessively high credit card utilization rate may indicate overspending and may ultimately lead to significant delinquencies. The lower the average credit card utilization rate over the past six months, the lower the borrower's credit demand and the lower the likelihood of a serious default. Conversely, the higher the average credit card utilization rate over the past six months, the higher the borrower's credit demand and the higher the likelihood of a serious default. The conversion method is WOE.

[0457] The current remaining available credit limit on a credit card reflects the borrower's credit need. For customers holding consumer loans, high additional credit card spending can lead to overspending and consequently, an inability to repay their loans. The higher the current remaining available credit limit, the lower the borrower's credit need and the lower the likelihood of a serious default. Conversely, the lower the current remaining available credit limit, the higher the borrower's credit need and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation within the continuous transformation model.

[0458] The number of months a credit customer has been overdue over the past six months reflects the borrower's historical performance. For consumers holding consumer loans, overdue payments often indicate a deterioration in their balance sheet. The more months a credit customer has been overdue over the past six months, the higher the likelihood of a subsequent serious default. Conversely, if a credit customer has had no overdue payments over the past six months, the likelihood of a serious default is relatively low. The conversion method is WOE.

[0459] The maximum increase in investment and wealth management account balances over the past three months reflects the borrower's investment behavior and asset growth. For customers holding consumer loans, the maximum increase in investment and wealth management account balances over the past three months reflects the borrower's assets and income. The lower the maximum increase in investment and wealth management account balances over the past three months, the stronger the borrower's debt repayment ability and the lower the likelihood of a major default; conversely, the higher the maximum increase, the higher the likelihood of a major default. The transformation method is the natural logarithm transformation within the continuous transformation.

[0460] The current consumer loan utilization rate reflects the borrower's credit usage attitude and repayment status. For customers holding consumer loans, the utilization rate reflects both usage attitude and repayment progress. The lower the current consumer loan utilization rate, the more prudent the borrower's credit use or the closer they are to the end of repayment, and the lower the likelihood of a serious default. Conversely, the higher the current consumer loan utilization rate, the more aggressive the borrower's credit use or the earlier they are in the repayment process, and the higher the likelihood of a serious default. The conversion method is WOE.

[0461] The number of months of credit card revolving or installment use over the past 12 months reflects a customer's credit need. The lower the number of months of credit card revolving or installment use over the past 12 months, the weaker the borrower's credit need and the lower the likelihood of a serious default. Conversely, the more months of credit card revolving or installment use over the past 12 months, the stronger the borrower's credit need and the higher the likelihood of a serious default. The conversion method is WOE.

[0462] The difference between the monthly minimum loan repayment amount and the monthly average asset size over the past three months reflects the client's income and expenditure. The greater the difference between the monthly minimum loan repayment amount and the monthly average asset size over the past three months, the greater the repayment pressure on the borrower and the higher the likelihood of a serious default. Conversely, the smaller the difference between the monthly minimum loan repayment amount and the monthly average asset size over the past three months, the less repayment pressure on the borrower and the lower the likelihood of a serious default. The conversion method is the cube root conversion in the continuous conversion method.

[0463] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0464]

[0465] Where k is the number of features entering the model. In formula 1-7, k is 7.

[0466] α is the intercept term, with a range of values ​​from (-0.56984, -0.65216), and the optimal value is -0.611; β1 is the coefficient corresponding to the average credit card limit utilization rate in the past 6 months, with a range of values ​​from (-0.51112, -0.52288), and the optimal value is -0.517; β2 is the coefficient corresponding to the number of overdue months for credit customers in the past 6 months, with a range of values ​​from (-0.74224, -0.76576), and the optimal value is -0.754; β3 is the coefficient corresponding to the current remaining available credit limit on the credit card, with a range of values ​​from (-0.17808, -0.18592), and the optimal value is -0.182; β4 is the coefficient corresponding to the investment and wealth management account in the past 3 months The coefficient for the maximum balance increase ranges from (-0.13008 to -0.13792), with an optimal value of -0.134. β5 represents the coefficient for the current total consumer loan utilization rate, with a range of (-0.2092 to -0.2288), and an optimal value of -0.219. β6 represents the coefficient for the number of months of credit card revolving or installment use over the past 12 months, with a range of (-0.15616 to -0.17184), and an optimal value of -0.164. β7 represents the coefficient for the difference between the average minimum monthly loan repayment and the average monthly asset size over the past three months, with a range of (0.0112 to 0.0108), and an optimal value of 0.011. (Note: The numerical ranges are derived from the 95% confidence intervals (95% CIs) in the table below.)

[0467] x1 is the Word of Exchange (WOE)-converted value of the average credit card credit limit utilization rate over the past six months, generated by the feature conversion step; x2 is the Word of Exchange (WOE)-converted value of the number of months overdue for credit customers over the past six months, generated by the feature conversion step; x3 is the continuous natural logarithm-converted value of the current remaining credit limit on the credit card, generated by the feature conversion step; x4 is the continuous natural logarithm-converted value of the maximum increase in the investment and wealth management account balance over the past three months, generated by the feature conversion step; x5 is the Word of Exchange (WOE)-converted value of the current total consumer loan limit utilization rate, generated by the feature conversion step; x6 is the Word of Exchange (WOE)-converted value of the number of months credit card revolving or installment use over the past 12 months, generated by the feature conversion step; and x7 is the continuous cube root-converted value of the difference between the monthly average minimum credit repayment amount and the monthly average asset size over the past three months, generated by the feature conversion step. The model performance of some features is shown in Tables 1-8 below:

[0468] Table 1-8

[0469]

[0470] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0471] Example 1-12 Sigma 8 model construction

[0472] Taking Sigma 8 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 2.56 million. Its customer base was mainly non-new customers who had applied for credit business and had no overdue payments but only had credit cards and used them in a revolving manner. The model was constructed to predict the probability (target variable) of this group of people having credit overdue payments of more than 90 days.

[0473] In Example 1-12, it is finally confirmed that the final modeling results are described using six features: the average credit card usage rate in the past 12 months, the number of overdue periods of credit cards in the past 12 months greater than 0 months, the current time point deposit account balance, the minimum credit repayment rate in the past 12 months, the number of consecutive months of increased credit card bill balances in the past 3 months, and the current credit card remaining available credit limit. In the feature conversion step, the six model variables have been converted in a corresponding form according to the correlation between the features and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0474] The average credit card limit utilization rate over the past 12 months reflects a borrower's credit demand. For credit card installment customers, an excessively high average credit card limit utilization rate may indicate long-term over-indebtedness and ultimately delinquency. The lower the average credit card limit utilization rate over the past 12 months, the lower the borrower's long-term repayment pressure and the lower the likelihood of a serious default. Conversely, the higher the average credit card limit utilization rate over the past 12 months, the higher the borrower's long-term repayments and the higher the likelihood of a serious default. The conversion method is the original value conversion in the continuous conversion method.

[0475] The number of credit card delinquencies exceeding 0 in the past 12 months reflects the borrower's repayment behavior. For installment credit card customers, past delinquencies often indicate a decline in repayment ability. The higher the number of credit card delinquencies exceeding 0 in the past 12 months, the higher the likelihood of a serious default. Conversely, the likelihood of a serious default is lower if there are no delinquencies in the past 12 months. The conversion method is WOE conversion.

[0476] The minimum credit repayment rate over the past 12 months reflects the borrower's repayment ability. For credit card installment customers, the repayment rate is a key criterion for assessing their repayment ability. The lower the minimum credit repayment rate over the past 12 months, the weaker the borrower's repayment ability and the higher the likelihood of a serious default. Conversely, the higher the minimum credit repayment rate over the past 12 months, the stronger the borrower's repayment ability and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0477] The current deposit account balance reflects the borrower's asset level. For credit card installment customers, the current deposit situation directly reflects the borrower's repayment ability. The lower the current deposit account balance, the lower the borrower's debt repayment ability and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the stronger the borrower's debt repayment ability and the lower the likelihood of a serious default. The conversion method is the cube root conversion of the continuous conversion method.

[0478] The number of consecutive months of credit card balance increases over the past three months reflects a borrower's credit usage. For installment credit card customers, a continuously increasing credit card balance may indicate a deterioration in their financial situation. The lower the number of consecutive months of credit card balance increases over the past three months, the stronger the borrower's repayment ability and the lower the likelihood of a serious default. Conversely, the higher the number of consecutive months of credit card balance increases over the past three months, the weaker the borrower's repayment ability and the higher the likelihood of a serious default. The conversion method is WOE conversion.

[0479] The current available credit limit on a credit card reflects the borrower's credit usage. For installment credit card customers, the higher the available credit limit, the lower the borrower's capital needs and the lower the likelihood of a serious default. Conversely, the lower the available credit limit, the greater the borrower's capital needs and the higher the likelihood of a serious default. The conversion method is the natural logarithm transformation of the continuous transformation.

[0480] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0481]

[0482] Where k is the number of features entering the model. In formula 1-8, k is 6.

[0483] α is the intercept term, with a range of values ​​from (0.66824, 0.59376), and an optimal value of 0.631; β1 is the coefficient corresponding to the average credit card credit limit utilization rate in the past 12 months, with a range of values ​​from (0.00125, 0.00075), and an optimal value of 0.001; β2 is the coefficient corresponding to the number of credit card overdue periods greater than 0 months in the past 12 months, with a range of values ​​from (-0.6112, -0.6308), and an optimal value of -0.621; β3 is the coefficient corresponding to the deposit account balance at the current time point, with a range of values ​​from (-0.08604, -0. 08996), with an optimal value of -0.088; β4 is the coefficient corresponding to the minimum credit repayment rate in the past 12 months, with a range of values ​​(-0.44932, -0.48068), and an optimal value of -0.465; β5 is the coefficient corresponding to the number of consecutive months of increase in credit card bill balances in the past three months, with a range of values ​​(-0.18944, -0.23256), and an optimal value of -0.211; β6 is the coefficient corresponding to the current remaining available credit card balance, with a range of values ​​(-0.10604, -0.10996), and an optimal value of -0.108. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0484] x1 is the raw value-converted value of the average credit card credit limit utilization rate over the past 12 months, generated by the feature conversion step; x2 is the word-of-mouth (WOE)-converted value of the number of credit card overdue periods greater than 0 over the past 12 months, generated by the feature conversion step; x3 is the cube root-converted value of the current deposit account balance, generated by the feature conversion step; x4 is the word-of-mouth (WOE)-converted value of the minimum credit repayment rate over the past 12 months, generated by the feature conversion step; x5 is the word-of-mouth (WOE)-converted value of the number of consecutive months of credit card bill balance increases over the past three months, generated by the feature conversion step; and x6 is the natural logarithm-converted value of the current credit card remaining available balance, generated by the feature conversion step. The model performance of some features is shown in Tables 1-9 below:

[0485] Table 1-9

[0486]

[0487] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0488] Example 1-13 Sigma 9 model construction

[0489] Taking Sigma 9 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 40.33 million. Its customer base was mainly non-new customers who had applied for credit business and had no overdue payments but only had credit cards and did not use them in a revolving manner. The model was constructed to predict the probability (target variable) of this group of people having credit overdue payments of more than 90 days.

[0490] In Example 1-13, the final modeling results are described by taking five features as an example: the length of time it takes to open a credit card account, the balance of the deposit account at the current time point, the utilization rate of the maximum cash withdrawal limit of the credit card in the past 12 months, the number of months in which the credit card is fully repaid in the past 12 months, and the difference between the average monthly minimum repayment amount of the credit card and the average monthly asset size in the past 6 months. In the selection of the conversion method in the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0491] The length of time a credit card has been open reflects the borrower's credit history. For credit card users who haven't used installment plans, overdue payments are more closely related to the length of their credit history. The longer the card has been open, the longer and more stable the borrower's credit history, and the lower the likelihood of a serious default. Conversely, the shorter the card has been open, the more ambiguous the borrower's credit behavior, and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation within the continuous transformation model.

[0492] The utilization rate of the maximum cash withdrawal limit on a credit card over the past 12 months reflects the borrower's credit demand. For credit card users who do not have installment plans, credit card cash withdrawals may indicate recent financial constraints. The lower the utilization rate of the maximum cash withdrawal limit over the past 12 months, the less financially strained the borrower is and the lower the likelihood of a serious default. Conversely, the higher the utilization rate of the maximum cash withdrawal limit over the past 12 months, the greater the borrower's financial strain and the higher the likelihood of a serious default. The conversion method is WOE conversion.

[0493] The number of months with full credit card payments over the past 12 months reflects a borrower's repayment ability. For customers who don't have installment credit card loans, a history of overdue payments or installment payments reflects their long-term solvency. The higher the number of months with full credit card payments over the past 12 months, the stronger the borrower's solvency and the lower the likelihood of a serious default. Conversely, the lower the number of months with full credit card payments over the past 12 months, the weaker the borrower's solvency and the higher the likelihood of a serious default. The conversion method is WOE conversion.

[0494] The current deposit account balance reflects the borrower's asset status. For credit card users who haven't paid installments, the current deposit balance directly reflects the borrower's qualifications. The lower the current deposit account balance, the lower the borrower's solvency and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the stronger the borrower's solvency and the lower the likelihood of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0495] The difference between the average monthly minimum credit repayment amount and the average monthly asset size over the past six months reflects a borrower's credit repayment capacity. For credit card users who do not have installment plans, the difference between the minimum repayment amount and the average monthly asset size is a more clear indicator of repayment capacity. When the difference between the average monthly minimum credit repayment amount and the average monthly asset size over the past six months is positive, the borrower's liabilities exceed their assets, increasing the likelihood of a serious default. Conversely, when the difference between the average monthly minimum credit repayment amount and the average monthly asset size over the past six months is negative, the borrower's assets exceed their liabilities, decreasing the likelihood of a serious default. The conversion method is the cube root conversion in the continuous conversion method.

[0496] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0497]

[0498] Where k is the number of features entering the model. In formula 1-9, k is 5.

[0499] α is the intercept term, with a range of values ​​from (0.64448, 0.59352) and an optimal value of 0.619; β1 is the coefficient corresponding to the credit card account opening time, with a range of values ​​from (-0.38812, -0.39988) and an optimal value of -0.394; β2 is the coefficient corresponding to the deposit account balance at the current time point, with a range of values ​​from (-0.08204, -0.08596) and an optimal value of -0.084; β3 is the coefficient corresponding to the maximum cash withdrawal rate of the credit card in the past 12 months. The corresponding coefficient is (-0.22912, -0.24088), with an optimal value of -0.235; β4 is the coefficient corresponding to the number of months with full credit card repayments in the past 12 months, with a range of (-0.17112, -0.18288), and an optimal value of -0.177; β5 is the coefficient corresponding to the difference between the average monthly minimum credit repayment amount and the average monthly asset size in the past six months, with a range of (0.00602, 0.00598), and an optimal value of 0.006. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0500] x1 is the natural logarithm-transformed value of the credit card account opening time, generated in the feature transformation step; x2 is the natural logarithm-transformed value of the deposit account balance at the current time, generated in the feature transformation step; x3 is the Word of Equity (WOE)-transformed value of the credit card maximum cash withdrawal limit utilization rate over the past 12 months, generated in the feature transformation step; x4 is the WOE-transformed value of the number of months with full credit card payments over the past 12 months, generated in the feature transformation step; and x5 is the cube root-transformed value of the difference between the average minimum monthly credit repayment amount and the average monthly asset size over the past six months, generated in the feature transformation step. The model performance of some features is shown in Tables 1-10 below:

[0501] Table 1-10

[0502]

[0503] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0504] Example 1-14 Construction of Sigma 10 Model

[0505] Taking Sigma 10 as a sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 220,000. Its customer base was mainly potential customers who had not applied for credit business and were from economically underdeveloped areas. The model was constructed to predict the probability (target variable) of this group of people being overdue for credit for more than 90 days.

[0506] In Example 1-14, it is finally confirmed that the final modeling results are described by taking the five features of the current deposit account balance, the maximum balance of the investment and financial management account in the past three months, the number of consecutive months of asset size reduction in the past three months, the percentage of the average monthly salary in the past three months to the average monthly asset size, and the quantile number of the region where the current salary is paid as an example. In the feature conversion step, the five model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0507] The current deposit account balance reflects the borrower's asset level. For new applicants from less developed regions, the higher the asset level, the lower the likelihood of a serious default; conversely, the lower the asset level, the higher the likelihood of a serious default. The conversion method is a logarithmic transformation within the continuous transformation.

[0508] The maximum balance in the investment and wealth management account over the past three months reflects the borrower's investment habits. For new applicants in economically underdeveloped regions, good investment habits guarantee repayment ability. The lower the maximum balance in the investment and wealth management account over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the maximum balance in the investment and wealth management account over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0509] The number of consecutive months of asset decline over the past three months reflects the borrower's asset fluctuations. For new applicants from less developed regions, a continuous decline in asset levels indicates a persistent deficit. The greater the number of consecutive months of asset decline over the past three months, the more severe the borrower's deficit and the higher the likelihood of a major default. Conversely, the fewer consecutive months of asset decline over the past three months, the more stable the borrower's financial situation and the lower the likelihood of a major default. The conversion method is WOE conversion.

[0510] The percentage of the average monthly salary over the past three months to the average monthly asset size reflects the borrower's savings status. For new applicants in less developed regions, long-term savings habits determine their long-term repayment ability. The lower the percentage of the average monthly salary over the past three months to the average monthly asset size, the smaller the proportion of the borrower's monthly salary to their total assets, and therefore, the better their long-term savings habits and the lower the likelihood of a serious default. Conversely, the higher the percentage of the average monthly salary over the past three months to the average monthly asset size, the greater the proportion of the average monthly salary to their total assets, and therefore, the worse their long-term savings habits and the higher the likelihood of a serious default. The conversion method is WOE conversion.

[0511] The regional quantile for current payroll payment reflects the borrower's income level. For new applicants from less developed regions, the lower the regional quantile for current payroll payment, the lower the borrower's income level and the higher the likelihood of a serious default. Conversely, the higher the regional quantile for current payroll payment, the higher the borrower's income level and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0512] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0513]

[0514] Where k is the number of features entering the model. In formula 1-10, k is 5.

[0515] α is the intercept term, with a range of values ​​from (-1.65504, -1.85496) and an optimal value of -1.755; β1 is the coefficient corresponding to the current deposit account balance, with a range of values ​​from (-0.12828, -0.15572) and an optimal value of -0.142; β2 is the coefficient corresponding to the maximum balance of the investment and wealth management account in the past three months, with a range of values ​​from (-0.52164, -0.78036) and an optimal value of -0.651; β3 is the coefficient corresponding to the asset size in the past three months. The coefficient corresponding to the number of consecutive months of reduction ranges from (-1.25768 to -2.20632), with an optimal value of -1.732. β4 is the coefficient corresponding to the percentage of average monthly wages to average monthly assets over the past three months, with a range from (-0.20416 to -0.41584), and an optimal value of -0.31. β5 is the coefficient corresponding to the quantile of the region where the current payroll is distributed, with a range from (-0.10288 to -0.48312), and an optimal value of -0.293. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0516] x1 is the continuous logarithmic transformation value of the deposit account balance at the current time point generated by the feature conversion step; x2 is the WOE conversion value of the maximum balance of the investment and wealth management account in the past three months generated by the feature conversion step; x3 is the WOE conversion value of the number of consecutive months of asset size reduction in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the percentage of average monthly salary to average monthly asset size in the past three months generated by the feature conversion step; x5 is the WOE conversion value of the percentile of the current payroll area generated by the feature conversion step.

[0517] The model performance of some features is shown in Table 1-11 below:

[0518] Table 1-11

[0519]

[0520] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0521] Example 1-15 Construction of Sigma 11 Model

[0522] Taking Sigma 11 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 230,000. Its customer base was mainly potential customers who had not applied for credit business and were from moderately developed economic areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0523] In Example 1-15, it is finally confirmed that the final modeling result is described by taking the current deposit account balance, the average monthly asset size in the past three months, the average balance of the investment and financial management account in the past three months, the longest continuous reduction in the number of months of the asset size value in the past three months, the maximum salary in the past 12 months, and the number of months from the maximum balance of the deposit account in the past six months to the observation point as an example. In the feature conversion step, the six model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0524] The current deposit account balance reflects the borrower's asset level. For new applicants in moderately developed regions, the higher the asset level, the lower the likelihood of a serious default; conversely, the lower the asset level, the higher the likelihood of a serious default. The conversion method is the natural logarithm transformation of the continuous transformation.

[0525] The average monthly asset size over the past three months reflects the borrower's asset level. For new applicants in moderately developed regions, the absolute value of their asset level is a better indicator of asset quality. The higher the average monthly asset size over the past three months, the more cash-rich the borrower is and the lower the likelihood of a serious default. Conversely, the lower the average monthly asset size over the past three months, the more cash-strapped the borrower is and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0526] The average balance of investment and wealth management accounts over the past three months reflects the borrower's investment habits. For new applicants in moderately developed regions, good investment habits guarantee their repayment ability. The lower the average balance of investment and wealth management accounts over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the average balance of investment and wealth management accounts over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0527] The longest consecutive number of months of asset size decline over the past three months reflects the borrower's asset fluctuations. For new applicants in moderately developed regions, a continuous decline in their asset levels indicates a persistent deficit. The longer the number of consecutive months of asset size decline over the past three months, the more severe the borrower's deficit and the higher the likelihood of a major default. Conversely, the shorter the number of consecutive months of asset size decline over the past three months, the more stable the borrower's financial situation and the lower the likelihood of a major default. The conversion method is WOE conversion.

[0528] The maximum wage over the past 12 months reflects the borrower's income level. For new applicants in moderately developed regions, this figure reflects their overall income level. The lower the maximum wage over the past 12 months, the lower the borrower's overall income level and the higher the likelihood of a major default. Conversely, the higher the maximum wage over the past 12 months, the higher the borrower's overall income level and the lower the likelihood of a major default. The conversion method is WOE.

[0529] The number of months between the maximum deposit account balance at a point in time over the past six months and the observation point reflects the borrower's deposit changes. The closer the maximum deposit account balance at a point in time over the past six months is to the observation point, the greater the likelihood that the borrower's assets will gradually increase and the lower the probability of a serious default. Conversely, the closer or further the maximum deposit account balance is to the observation point, the greater the likelihood that the borrower's assets will gradually decrease and the higher the probability of a serious default. The conversion method is WOE conversion.

[0530] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0531]

[0532] Where k is the number of features entering the model. In formula 1-11, k is 6.

[0533] α is the intercept term, with a range of values ​​from (-0.50444, -0.84156), and an optimal value of -0.673; β1 is the coefficient corresponding to the current deposit account balance, with a range of values ​​from (-0.13036, -0.16564), and an optimal value of -0.148; β2 is the coefficient corresponding to the average monthly asset size in the past three months, with a range of values ​​from (-0.19664, -0.25936), and an optimal value of -0.228; β3 is the coefficient corresponding to the average balance of the investment and wealth management account in the past three months, with a range of values ​​from (-0.26772, -0.632 28), with an optimal value of -0.45; β4 is the coefficient corresponding to the longest consecutive decrease in asset size over the past three months, with a range of (-0.78496, -1.17304), and an optimal value of -0.979; β5 is the coefficient corresponding to the maximum salary over the past 12 months, with a range of (-0.35304, -0.74896), and an optimal value of -0.551; β6 is the coefficient corresponding to the number of months from the maximum deposit account balance over the past six months to the observation point, with a range of (-0.29792, -0.78008), and an optimal value of -0.539. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0534] x1 is the natural logarithm conversion value of the deposit account balance at the current time point generated by the feature conversion step; x2 is the natural logarithm conversion value of the average monthly asset size in the past three months generated by the feature conversion step; x3 is the WOE conversion value of the average balance of the investment and wealth management account in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the longest consecutive number of months of decrease in asset size value in the past three months generated by the feature conversion step; x5 is the WOE conversion value of the maximum salary in the past 12 months generated by the feature conversion step; x6 is the WOE conversion value of the number of months from the observation point to the maximum balance of the deposit account in the past six months generated by the feature conversion step.

[0535] The model performance of some features is shown in Table 1-12 below:

[0536] Table 1-12

[0537]

[0538] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0539] Example 1-1 Construction of 6Sigma 12 Model

[0540] Taking Sigma 12 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 460,000. Its customer base was mainly potential customers who had not applied for credit business and were from economically developed areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0541] In Example 1-16, it is finally confirmed that the final modeling results are described by taking the maximum asset size in the past 3 months, the minimum balance of the deposit account in the past 3 months, the maximum balance of the investment and financial management account in the past 3 months, the unit quantile of the current salary payment unit, and the number of months from the maximum balance of the deposit account in the past 12 months to the observation point as an example. In the feature conversion step conversion mode selection, the 5 model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the 5 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0542] The maximum asset size over the past three months reflects the borrower's asset level. For new applicants in economically developed regions, asset levels fluctuate significantly. Therefore, the short-term maximum asset value better reflects asset quality. Higher asset levels indicate a lower likelihood of a major default; conversely, lower asset levels indicate a higher likelihood of a major default. The conversion method is the cube root conversion in the continuous conversion method.

[0543] The minimum deposit balance over the past three months reflects the borrower's deposit status. For new applicants in economically developed regions, the minimum deposit is a better indicator of asset quality. The higher the minimum deposit balance over the past three months, the more abundant the borrower's funds are and the lower the likelihood of a serious default. Conversely, the lower the minimum deposit balance over the past three months, the less cash the borrower has and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation within the continuous transformation model.

[0544] The maximum balance in the investment and wealth management account over the past three months reflects the borrower's investment habits. For new applicants in economically developed regions, good investment habits guarantee their repayment ability. The lower the maximum balance in the investment and wealth management account over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the maximum balance in the investment and wealth management account over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE.

[0545] The current payroll unit's percentile reflects the borrower's income capacity and job stability. For new applicants in economically developed regions, the higher the unit's percentile, the greater the borrower's income capacity and stability, and the lower the likelihood of a serious default. Conversely, the lower the unit's percentile, the weaker the borrower's income capacity and stability, and the higher the likelihood of a serious default. The conversion method is WOE.

[0546] The number of months between the maximum deposit balance in the past 12 months and the observation point reflects changes in the borrower's deposits. For new applicants in economically developed regions, the closer the maximum deposit balance in the past 12 months was to the observation point, the more recent the borrower's deposits are and the lower the likelihood of a serious default. Conversely, the further the maximum deposit balance in the past 12 months was from the observation point, the lower the borrower's deposits are and the higher the likelihood of a serious default. The conversion method is WOE.

[0547] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0548]

[0549] Where k is the number of features entering the model. In formula 1-12, k is 5.

[0550] α is the intercept term, with a range of values ​​from (-2.18528, -2.31072) and an optimal value of -2.248; β1 is the coefficient corresponding to the maximum asset size in the past three months, with a range of values ​​from (-0.02008, -0.02792) and an optimal value of -0.024; β2 is the coefficient corresponding to the minimum balance of the deposit account in the past three months, with a range of values ​​from (-0.08928, -0.11672) and an optimal value of -0.103; β3 is the investment and financial management in the past three months. The coefficient corresponding to the maximum account balance ranges from (-0.45036 to -0.68164), with an optimal value of -0.566. β4 is the coefficient corresponding to the unit quantile of the current payroll payment agency, with a range of (-0.11148 to -0.64852), and an optimal value of -0.38. β5 is the coefficient corresponding to the number of months from the observation point to the maximum deposit account balance in the past 12 months, with a range of (-0.35424 to -0.57376), and an optimal value of -0.464. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0551] x1 is the cube root conversion value of the maximum asset size in the past three months generated by the feature conversion step in a continuous manner; x2 is the natural logarithm conversion value of the minimum balance of the deposit account in the past three months generated by the feature conversion step in a continuous manner; x3 is the WOE conversion value of the maximum balance of the investment and wealth management account in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the unit percentile of the current payroll payment generated by the feature conversion step; x5 is the WOE conversion value of the number of months from the observation point to the maximum balance of the deposit account in the past 12 months generated by the feature conversion step.

[0552] The model performance of some features is shown in Table 1-13 below:

[0553] Table 1-13

[0554]

[0555] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0556] Examples 1-17

[0557] The P values ​​calculated by the above formulas can be used to further calculate the sub-scores and total scores of any customer.

[0558] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0559] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0560]

[0561]

[0562] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0563] In this embodiment, the P value calculated in the above embodiments can be substituted to calculate the score.

[0564] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0565] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0566] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0567] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Sigma model is 75.25. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0568] The following table shows the KS values ​​for the development and validation samples when the model was applied to the development and validation sets. Table 1-14 shows that the model performs well in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample set. This means that all sub-models have excellent discrimination.

[0569] Table 1-14

[0570]

[0571] 2) Credit card business credit score (Alpha model credit score)

[0572] Example 2-1 Collection of modeling samples

[0573] During the construction of this example, we collected credit card application and behavior data, personal loan application and behavior data, personal financial asset transaction data, and personal customer information data from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample, selecting 630 million data points from 2019 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two parts: those with registered credit applications and those without registered credit applications, with sample sizes of 130 million and 500 million, respectively.

[0574] The specific Alpha model is a scenario scoring for credit card business. When designing the model, analysis and design are carried out on the 130 million customers who have applied for registered credit business. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0575] Model design includes: 1) Exclusion rules: For example, exclude customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.5 million customers (this group of people are those who have applied for credit card business); 2) Time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15 months as the performance period of Y; 3) Sample sampling: use a good / bad sample number of 5:1 for sampling modeling. Figure 2 Different segmentation schemes are designed using the decision tree shown in the figure. The final model segmentation scheme is confirmed by comparing parent-child models, and each sub-model is then modeled separately. The model segmentation scheme in this application is designed based on whether the customer's account age is less than 3 years, whether the current payment is overdue, and whether the current repayment rate is less than 100%. This fully reflects the control variables closely related to risk characteristics during business operations and can more comprehensively match various customer groups in the market.

[0576] In this embodiment, the modeling samples are first divided into the following sub-model modeling samples based on the decision tree to facilitate the construction of subsequent sub-models. When determining the modeling samples, the first layer of the decision tree is the account age, which is used to distinguish the length of the customer's credit history. In this embodiment, for customers with an account age greater than or equal to 3 (i.e., defined as new applications), the user population is non-repeatedly divided into three models: Alpha2, Alpha3, and Alpha4 based on their current number of overdue periods and credit repayment rate; customers with a credit account age less than 3 are divided into Alpha1, where the bad debt rate refers to the ratio of bad customers in a certain type of sample to the total number of samples in this type.

[0577] In addition, in this embodiment, in order to build a model for the sample group that has not applied for credit business registration, the user group with an account age of less than 3 months is selected as the approximate customer group and enters the Alpha5 segmentation for model training.

[0578] Specifically, the sample size used to construct the Alpha1 sub-model is about 300,000. Its customer base mainly has an account age of less than 3 years. The model is used to predict the probability of credit overdue for more than 60 days in this group of people.

[0579] Specifically, the sample size used to construct the Alpha2 sub-model is about 2 million. Its customer base mainly consists of accounts with an age greater than or equal to 3 years and are currently overdue. The model is used to predict the probability of this group of people being overdue for credit for more than 60 days.

[0580] Specifically, the sample size used to construct the Alpha3 sub-model is approximately 3.8 million. Its customer base mainly consists of those with an account age greater than 3 years, who are not currently overdue and whose current repayment rate is less than 100%. The model is used to predict the probability of this group of people being overdue for credit for more than 60 days.

[0581] Specifically, the sample size used to construct the Alpha4 sub-model is approximately 46 million. Its customer base is mainly those with an account age greater than 3 years who are not currently overdue and whose current repayment rate is greater than or equal to 100%. The model is used to predict the probability of this group of people being overdue for credit for more than 60 days.

[0582] Specifically, the sample size used to construct the Alpha5 sub-model is about 300,000. Its customer base is mainly those who have applied for credit services and have an account age of less than 3 years. The model is used to predict the probability of this group of people being overdue for credit for more than 60 days.

[0583] After this model is successfully constructed, it can be used to predict the probability of credit being overdue for 60 days or more for samples that have not applied for registered credit business.

[0584] Based on the different customer samples of each sub-model confirmed above, historical data information of borrowers is obtained: 1) Credit card category, basic fields include account, overdue, balance, credit limit, repayment due, actual repayment and other information under different dimensions such as bills, cash withdrawals, installments, and interest; 2) Personal loan category, including account, overdue, balance, credit limit, repayment due, actual repayment and other information; 3) Customer basic information, including gender, age and administrative region information of the business; 4) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0585] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods, including the following: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of their account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months, the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods in the last X months and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 3,067 features with potential predictive power for customer delinquency were derived, including "current repayment amount," "average number of overdue payments in the past six months," and "number of credit card installments in the past 12 months." In this application, "observation point" refers to the point in time at which samples were collected before modeling. "Current" also refers to the sampling cutoff time. These two terms have the same meaning.

[0586] Table 2-1 Summary of basic variables and derived variables used in this example

[0587]

[0588]

[0589] Example 2-2 Feature Screening

[0590] A preliminary screening was conducted on the 3067 features (variables) collected in Example 2-1 that have potential predictive power for customer overdue payments.

[0591] The initial screening of the features of the four sub-models Alpha1 to Alpha4 divided in Example 2-1 is performed as follows:

[0592] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 2-1 were first eliminated, and a total of 589 variables were deleted, leaving 2478 variables.

[0593] In the second round of preliminary screening, for the 2478 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 189 variables were eliminated, leaving 2309 variables.

[0594] In the third round of preliminary screening, the 2,309 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is as follows: if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 976 features were eliminated, and 1,333 feature variables remained.

[0595] In the fourth round of preliminary screening, the features are further screened based on the stepwise discrimination algorithm for the 1,333 features after the third round of screening. After this round of screening, 276 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment adopts a pioneering stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0596] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0597] The fifth round of preliminary screening targets the 276 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0598] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0599] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0600] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0601] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0602] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0603] The criteria for approximate consistency are as follows:

[0604] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0605] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0606] After the fifth round of screening, 78 features were eliminated, leaving 198 features.

[0607] For the Alpha5 sub-model divided in Example 2-1, the selected variables are customer basic information and personal financial asset variables (a total of 316 feature variables) used to approximately predict the probability of credit defaults in the future for customers who have not registered for credit applications. The initial feature screening step and subsequent steps are the same as other sub-models.

[0608] Example 2-3 Conversion of features after initial screening

[0609] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the feature initial screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numerical variables. For character variables, the dummy feature (dummy variable) conversion method is generally used for variable conversion; for numerical variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. The features using different conversion methods are grouped and divided into different data sets. Specifically,

[0610] First, the remaining 198 features in Example 2-2 are judged for their conversion methods, and the following three methods are selected based on the concentration of the features, data type, etc.

[0611] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0612] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0613] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0614] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0615] In this example, 198 features were converted, of which 20 were converted to WOE, 25 were converted to dummy features (expanded to 15 variables), and 143 were converted to continuous features. The conversion of the features after the initial screening of the Alpha5 sub-model is basically similar.

[0616] Example 2-4 Feature Depth Screening (Feature Fine Screening Step)

[0617] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0618] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the total number of features from 198 to 87.

[0619] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 87 features to 53.

[0620] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0621] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0622] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0623] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0624] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0625] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0626] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 48 ​​features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0627] Example 2-5 Construction of Alpha1 model

[0628] Generally speaking, the stronger the correlation between the feature variables and the target variable, the more accurate the final model. In this example, Alpha1 is used as a sub-model, employing the aforementioned decision tree method. The sample size obtained for classification is approximately 300,000. Its customer base primarily consists of customers with an account age of less than 3 years. The model is constructed to predict the probability (target variable) of this group experiencing a credit delinquency of more than 60 days.

[0629] In this embodiment 2-5, it is finally confirmed that the final modeling results are described by taking the difference between the minimum monthly repayment amount of the average credit in the past three months and the average monthly asset size, the average interest in the past three months, the maximum balance of the investment and financial management account in the past 12 months, the unit level (quantile) of the salary payment agency, and the minimum balance of the deposit account in the past six months as an example. In the feature conversion step, the five model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0630] The difference between the monthly minimum loan repayment amount and the monthly average asset size over the past three months reflects a borrower's credit repayment capacity. For new credit applicants, the difference between the monthly minimum loan repayment amount and the monthly average asset size is a more accurate indicator of repayment capacity. A larger difference between the monthly minimum loan repayment amount and the monthly average asset size over the past three months indicates a higher likelihood of a major default; conversely, a smaller difference indicates a lower likelihood of a major default. The transformation is a natural logarithm.

[0631] The average interest rate over the past three months reflects the borrower's repayment history. For new credit applicants, higher average interest rates indicate more underpaid loans and a greater likelihood of defaults. Conversely, lower average interest rates indicate a lower likelihood of defaults. The conversion method is raw value conversion.

[0632] The maximum balance in the investment and wealth management account over the past 12 months reflects the client's financial management skills. For new applicants, the higher the balance in the wealth management account, the stronger their financial management skills, the better their overall asset condition, and the lower the likelihood of default. Conversion is done using cube root conversion.

[0633] The payroll agency level (quantile) reflects the client's income level. For new credit applicants, the more payroll agencies their employers have, the more income they can use to repay their loans and the lower the likelihood of defaults. Conversely, the higher the likelihood of defaults, the higher the likelihood. The conversion method is WOE conversion.

[0634] The minimum balance in the deposit account over the past six months reflects the borrower's asset level. For new applicants, it directly reflects the borrower's repayment ability. The lower the minimum balance in the deposit account over the past six months, the lower the borrower's repayment ability and the higher the likelihood of a serious default; conversely, the lower the minimum balance, the lower the likelihood of a serious default. The conversion method is square root conversion.

[0635] In the step of obtaining the calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0636]

[0637] Where k is the number of features entering the model, and in Formula 2-1, k is 5.

[0638] α is the intercept term, with a range of values ​​from (-0.64252, -0.91692), and the optimal value is -0.77972; β1 is the coefficient corresponding to the difference between the minimum monthly repayment amount of the average credit in the past three months and the average monthly asset size, with a range of values ​​from (-0.7582, -0.8758), and the optimal value is -0.817; β2 is the coefficient corresponding to the average interest in the past three months, with a range of values ​​from (-0.66666, -0.74506), and the optimal value is -0.70586; β3 is the coefficient corresponding to the maximum balance of the investment and wealth management account in the past 12 months. The value range is (-0.32009, -0.67289), and the optimal value is -0.49649; β4 is the corresponding coefficient of the unit level (quantile) where the payroll is paid, and the value range is (-0.70276, -0.82036), and the optimal value is -0.76156; β5 is the corresponding coefficient of the minimum balance of the deposit account in the past 6 months, and the value range is (-0.09393, -0.25073), and the optimal value is -0.17233 (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below).

[0639] x1 is the natural logarithm-transformed value of the difference between the minimum monthly loan repayment and the average monthly asset size over the past three months, generated by the feature transformation step; x2 is the original value of the average interest over the past three months, generated by the feature transformation step; x3 is the cube root-transformed value of the maximum balance in the investment and wealth management account over the past 12 months, generated by the feature transformation step; x4 is the Word of Equity (WOE)-transformed value at the payroll unit level (quantile), generated by the feature transformation step; and x5 is the square root-transformed value of the minimum balance in the deposit account over the past six months, generated by the feature transformation step. The model performance of some features is shown in Table 2-2 below.

[0640] Table 2-2

[0641]

[0642] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0643] Example 2-6 Construction of Alpha2 Model

[0644] For Alpha 2 as a sub-model, the above-mentioned decision tree method is adopted, and the sample size obtained through classification is about 2 million. Its customer base is mainly customers who have applied for credit services and are currently overdue. The model is constructed to predict the probability (target variable) of this group of people being overdue for credit for more than 60 days.

[0645] In this embodiment 2-6, it is finally confirmed that the final modeling results are described by taking five features as examples: the proportion of the maximum consecutive months of decrease in the balance of the deposit account in the past 6 months, the maximum number of overdue periods of credit customers in the past 12 months, the average balance of investment and financial management accounts in the past 3 months, the maximum consecutive months when the credit utilization rate in the past 12 months is greater than 70%, and the average repayment rate of credit in the past 3 months. In the selection of the conversion method in the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0646] The percentage of months with the largest consecutive month-over-month decrease in deposit account balances over the past six months reflects the fluctuations in a borrower's asset balances. For currently overdue customers, a higher percentage of consecutive months with a decrease in deposit account balances indicates a longer period of continuous asset decline, which increases the likelihood of a serious default. Conversely, a lower percentage of months with the largest consecutive month-over-month decrease in deposit account balances over the past six months indicates a shorter period of continuous asset decline, indicating a more stable financial situation and a lower likelihood of a serious default. This conversion is done using the cube root method.

[0647] The maximum number of overdue payments for a credit customer over the past 12 months reflects the borrower's historical credit behavior. For customers currently in overdue payments, the smaller the maximum number of overdue payments over the past 12 months, the lower the likelihood of a serious default. Conversely, a larger number of overdue payments over the past 12 months indicates poorer credit behavior and a higher likelihood of a serious default. The conversion method is raw value conversion.

[0648] The average balance of investment and wealth management accounts over the past three months reflects a client's financial management skills. For clients currently experiencing overdue payments, higher balances in wealth management accounts indicate stronger financial management skills, better overall asset health, and a lower likelihood of default. The conversion method is natural logarithm transformation.

[0649] The maximum number of consecutive months with a credit limit utilization ratio exceeding 70% over the past 12 months reflects the borrower's long-term credit needs. For currently overdue customers, the greater the number of consecutive months with a credit limit utilization ratio exceeding 70% over the past 12 months, the greater the repayment pressure and the higher the likelihood of a serious default. Conversely, the fewer consecutive months with a credit limit utilization ratio exceeding 70% over the past 12 months, the lower the likelihood of a serious default. The conversion method is raw value conversion.

[0650] The average repayment rate over the past three months reflects the borrower's recent repayment performance. For currently overdue customers, the higher the recent repayment rate, the better their asset condition and the lower the likelihood of continued overdue payments. Conversely, the higher the rate, the higher the likelihood of continued overdue payments. The conversion method is square root conversion.

[0651] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0652]

[0653] Where k is the number of features entering the model, and in Formula 2-2, k is 5.

[0654] α is the intercept term, with a range of values ​​from (-0.21277, -0.25627), and the optimal value is -0.23452; β1 is the coefficient corresponding to the proportion of the maximum consecutive months of decrease in the deposit account balance over the past 6 months, with a range of values ​​from (-0.00355, -0.0805), and the optimal value is -0.04202; β2 is the coefficient corresponding to the maximum number of overdue periods for credit customers in the past 12 months, with a range of values ​​from (-0.7076, -0.78119), and the optimal value is -0.7444; β3 is the coefficient corresponding to the maximum number of overdue periods for credit customers in the past 3 The coefficient for the average monthly investment and wealth management account balance ranges from (-0.73382 to -0.74012), with an optimal value of -0.73697. β4 represents the maximum number of consecutive months with a credit limit utilization rate exceeding 70% over the past 12 months, with a range from (-0.46684 to -0.50273), and an optimal value of -0.48479. β5 represents the average credit repayment rate over the past three months, with a range from (-0.58221 to -0.87047), and an optimal value of -0.72634. (Note: The numerical ranges are derived from the 95% confidence intervals, i.e., the 95% CI in the table below.)

[0655] x1 is the cube root-transformed value of the percentage of months with the largest consecutive month-on-month decrease in deposit account balances over the past six months, generated by the feature conversion step; x2 is the original value of the maximum number of overdue periods for credit customers over the past 12 months, generated by the feature conversion step; x3 is the natural logarithm-transformed value of the average balance in investment and wealth management accounts over the past three months, generated by the feature conversion step; x4 is the original value of the maximum consecutive months with a credit limit utilization rate greater than 70% over the past 12 months, generated by the feature conversion step; and x5 is the square root-transformed value of the average credit repayment rate over the past three months, generated by the feature conversion step. The model performance of some features is shown in Table 2-3 below:

[0656] Table 2-3

[0657]

[0658] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0659] Example 2-7 Construction of Alpha 3 Model

[0660] For Alpha3 as a sub-model, the above-mentioned decision tree method is adopted, and the sample size obtained by classification is about 3.8 million. Its customer base is mainly customers who have applied for credit business and are not currently overdue and have a repayment rate of less than 100%. The model is constructed to predict the probability of credit overdue for more than 60 days in this group (target variable).

[0661] In this embodiment 2-7, it is finally confirmed that the final modeling results are described by taking seven features as an example: the maximum salary of the revolving loan in the past 12 months, the minimum utilization rate of the revolving loan in the past 6 months, the minimum repayment rate of the credit in the past 6 months, the current remaining available credit limit of the credit card, the average monthly minimum repayment amount of the credit in the past 3 months / the average monthly asset size, the current deposit account balance, and the maximum number of consecutive months with credit customers overdue>0 in the past 12 months. In the feature conversion step, the 7 modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the 7 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0662] The maximum salary over the past 12 months reflects the borrower's income level. For clients with no overdue payments and a repayment rate less than 100%, the lower the maximum salary over the past 12 months, the lower the borrower's overall income level, and the higher the likelihood of a serious default. Conversely, the higher the maximum salary over the past 12 months, the higher the borrower's overall income level, and the lower the likelihood of a serious default. The transformation method is natural logarithm.

[0663] The minimum utilization rate of a revolving loan over the past six months reflects the borrower's recent loan demand. For customers with no current overdue payments and a repayment rate below 100%, the lower the minimum utilization rate over the past six months, the less repayment pressure the borrower faces and the lower the likelihood of a serious default. Conversely, the higher the minimum utilization rate over the past six months, the greater the repayment pressure and the higher the likelihood of a serious default. The conversion method is square root conversion.

[0664] The minimum repayment rate over the past six months reflects the borrower's recent repayment performance. For customers who are not currently overdue and have a repayment rate below 100%, a higher recent repayment rate indicates a relatively stronger asset position and a lower likelihood of persistent overdue payments. Conversely, a lower repayment rate indicates a higher likelihood of persistent overdue payments. The conversion method is square root conversion.

[0665] The current available credit limit on a credit card reflects the borrower's credit needs. For customers with no overdue payments and a repayment rate below 100%, the lower the current available credit limit, the greater the borrower's financial pressure and the lower the likelihood of a serious default. Conversely, the higher the current available credit limit, the less financial pressure the borrower faces and the higher the likelihood of a serious default. The natural logarithm transformation is used.

[0666] The average monthly minimum credit repayment amount over the past three months divided by the average monthly asset size reflects the borrower's credit repayment ability. For customers with no current overdue payments and a repayment rate less than 100%, the larger the ratio (average monthly minimum credit repayment amount over the past three months) is, the higher the borrower's debt service ratio, the greater the debt repayment pressure, the higher the debt-to-asset ratio, and the higher the likelihood of a major default. Conversely, the smaller the average monthly minimum credit repayment amount over the past three months is (less than 1), the lower the debt repayment pressure, the higher the debt-to-asset ratio, and the lower the likelihood of a major default. This conversion is done using the cube root method.

[0667] The current deposit account balance reflects the borrower's deposit level. For customers with no overdue payments and a repayment rate less than 100%, the current deposit is an important indicator of near-term repayment ability. The lower the current deposit account balance, the weaker the borrower's near-term debt repayment ability and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the stronger the borrower's near-term debt repayment ability and the lower the likelihood of a serious default. The transformation method is natural logarithm.

[0668] The maximum number of consecutive months with >0 delinquencies over the past 12 months indicates the severity of a customer's overdue status. For customers with no current overdue payments and a repayment rate below 100%, the greater the number of consecutive months with >0 delinquencies over the past 12 months, the more severe the customer's consecutive overdue status, the greater the debt repayment risk, and the higher the likelihood of future overdue payments. Conversely, the lower the likelihood of future overdue payments. The conversion method is WOE conversion.

[0669] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0670]

[0671] Where k is the number of features entering the model, and k is preferably 7;

[0672] α is the intercept term, with a range of values ​​from (-0.13202, -0.41) and an optimal value of -0.27101; β1 is the coefficient corresponding to the maximum salary in the past 12 months, with a range of values ​​from (-0.22917, -0.36318) and an optimal value of -0.29617; β2 is the coefficient corresponding to the minimum utilization rate of the revolving loan in the past 6 months, with a range of values ​​from (-0.34922, -0.63599) and an optimal value of -0.49261; β3 is the coefficient corresponding to the minimum repayment rate of the credit in the past 6 months, with a range of values ​​from (-0.04513, -0.3209) and an optimal value of -0.18301; β4 is the current remaining available balance on the credit card The corresponding coefficient of degree is (-0.82704, -1.05304), and the optimal value is -0.94004; β5 is the corresponding coefficient of the average monthly minimum credit repayment amount / average monthly asset size in the past three months, with a value range of (-0.20947, -0.567) and an optimal value of -0.38823; β6 is the corresponding coefficient of the deposit account balance at the current time point, with a value range of (-0.26127, -0.54833) and an optimal value of -0.4048; β7 is the corresponding coefficient of the maximum number of consecutive months with overdue payments of credit customers > 0 in the past 12 months, with a value range of (-0.06634, -0.43891) and an optimal value of -0.25262;

[0673] x1 is the natural logarithm conversion value of the maximum salary in the past 12 months generated by the feature conversion step; x2 is the square root conversion value of the minimum utilization rate of the revolving loan in the past 6 months generated by the feature conversion step; x3 is the square root conversion value of the minimum repayment rate of the credit in the past 6 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the current remaining available credit limit of the credit card generated by the feature conversion step; x5 is the cube root conversion value of the average monthly minimum credit repayment amount / average monthly asset size in the past 3 months generated by the feature conversion step; x6 is the natural logarithm conversion value of the deposit account balance at the current time point generated by the feature conversion step; x7 is the WOE conversion value of the maximum number of consecutive months with overdue payments of credit customers > 0 in the past 12 months generated by the feature conversion step.

[0674] The model performance of some features is shown in Table 2-4 below:

[0675] Table 2-4

[0676]

[0677] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0678] Example 2-8 Alpha4 model construction

[0679] For Alpha4 as a sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 45.5 million. Its customer base was mainly customers who had applied for credit services and were not currently overdue and whose credit repayment rate was greater than or equal to 100. The model was constructed to predict the probability (target variable) of this group of customers being overdue for credit for more than 60 days.

[0680] In Example 2-8, it is finally confirmed that the final modeling results are described using five features as examples: the average asset size in the past 12 months, the maximum repayment amount in the past 6 months, the current account age of the credit customer, the minimum utilization rate of the revolving loan in the past 6 months, and the percentage of the overdue balance of the credit customer in the loan amount in the past 12 months. In the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0681] The average asset size over the past 12 months reflects the borrower's debt repayment ability. For customers with no overdue payments and a credit repayment ratio greater than or equal to 100%, the lower the average asset size over the past 12 months, the lower the borrower's debt repayment ability and the higher the likelihood of a serious default. Conversely, the higher the average asset size over the past 12 months, the stronger the borrower's debt repayment ability and the lower the likelihood of a serious default. The conversion method is raw value conversion.

[0682] The maximum repayment amount over the past six months reflects the borrower's recent repayment ability. For customers with no current overdue payments and a credit repayment ratio greater than or equal to 100, the maximum repayment amount helps measure the borrower's asset level and repayment ability. The higher the maximum repayment amount over the past three months, the greater the borrower's repayment ability and the lower the likelihood of persistent default. Conversely, the lower the maximum repayment amount over the past 12 months, the lower the borrower's recent repayment ability and the higher the likelihood of default. The conversion method is raw value conversion.

[0683] The current age of a credit customer reflects the length of the borrower's current credit history. For customers with no overdue payments and a credit repayment rate greater than or equal to 100%, the longer the age, the stronger the credit product management capabilities and the lower the likelihood of bad debts. Conversely, the shorter the age, the weaker the credit product management capabilities and the higher the likelihood of default. This conversion method uses dummy variable conversion.

[0684] The minimum utilization rate of a revolving loan over the past six months reflects the borrower's recent loan demand. For customers with no current overdue payments and a credit repayment ratio greater than or equal to 100%, the lower the minimum utilization rate over the past six months, the lower the borrower's repayment pressure and the lower the likelihood of a serious default. Conversely, the higher the minimum utilization rate over the past six months, the greater the borrower's repayment pressure and the higher the likelihood of a serious default. This conversion method uses a dummy variable conversion.

[0685] The percentage of overdue balances as a percentage of the loan amount over the past 12 months reflects the borrower's medium- to long-term delinquency situation. For customers who are not currently overdue and have a credit repayment ratio greater than or equal to 100%, the higher the percentage of overdue balances as a percentage of the loan amount over the past six months, the higher the borrower's level of delinquency and the higher the likelihood of continued default. Conversely, the lower the percentage of overdue balances as a percentage of the loan amount over the past six months, the less delinquent the borrower is and the higher the likelihood of escaping default. This conversion is done using the square root method.

[0686] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0687]

[0688] Where k is the number of features entering the model. In formula 2-4, k is 5.

[0689] α is the intercept term, with a range of values ​​from (-0.17311, -0.25645) and an optimal value of -0.21478; β1 is the coefficient corresponding to the average asset size in the past 12 months, with a range of values ​​from (-0.59164, -0.6956) and an optimal value of -0.64362; β2 is the coefficient corresponding to the maximum repayment amount in the past 6 months, with a range of values ​​from (-0.26779, -0.65612) and an optimal value of -0.46196; β3 is the coefficient corresponding to the current aging of the credit customer The corresponding coefficient is (-0.59414, -0.72344), with the optimal value being -0.65879; β4 is the corresponding coefficient for the minimum utilization rate of the revolving loan in the past 6 months, with the numerical range being (-0.77339, -1.01385), and the optimal value being -0.89362; β5 is the corresponding coefficient for the percentage of overdue balances of credit customers in the loan amount in the past 12 months, with the numerical range being (-0.02998, -0.08368), and the optimal value being -0.05683;

[0690] x1 is the original value of the average asset size for 12 months generated by the feature conversion step; x2 is the original value of the maximum repayment amount in the past 6 months generated by the feature conversion step; x3 is the dummy variable conversion value of the current account age of the credit customer generated by the feature conversion step; x4 is the dummy variable conversion value of the minimum utilization rate of the revolving loan in the past 6 months generated by the feature conversion step; x5 is the square root conversion value of the percentage of the credit customer's overdue balance in the loan amount in the past 12 months generated by the feature conversion step.

[0691] The model performance of some features is shown in Table 2-5 below:

[0692] Table 2-5

[0693]

[0694] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0695] Example 2-9 Alpha5 model construction

[0696] For Alpha5 as a sub-model, the above-mentioned decision tree method is adopted, and the sample size obtained by classification is about 300,000. Its customer base is mainly new credit applicants. The model is constructed to predict the probability of credit overdue for more than 60 days in this group (target variable).

[0697] In Example 2-9, it is finally confirmed that the final modeling results are described by taking the four features of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months, the minimum balance of the deposit account in the past 3 months, the number of months from the minimum balance of the financial management account in the past 3 months, and the number of salary payments in the past 12 months as an example. In the feature conversion step, the four model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the four features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0698] The average monthly assets over the past six months divided by the average monthly assets over the past twelve months reflects changes in the borrower's asset level. For recently registered borrowers, an increase in their asset level indicates a continued increase in their asset size. A larger value (greater than 1) indicates a stronger and more stable borrower's financial situation, and a lower likelihood of default. A smaller value (less than 1) indicates a weaker and more unstable borrower's financial situation, and a higher likelihood of default. The conversion method is raw value conversion.

[0699] The minimum deposit balance over the past three months reflects the borrower's recent deposit status. For newly registered borrowers, the minimum deposit is a better indicator of asset quality. The higher the minimum deposit balance over the past three months, the more abundant the borrower's funds are and the lower the likelihood of a major default. Conversely, the lower the minimum deposit balance over the past three months, the less cash the borrower has and the higher the likelihood of a major default. Natural logarithm transformation is used.

[0700] The number of months from the minimum balance in the wealth management account over the past three months reflects recent changes in the borrower's wealth management assets. For newly registered clients, the higher the number of months from the minimum balance in the wealth management account over the past three months and the higher the recent wealth management assets, the lower the likelihood of a serious default; conversely, the higher the likelihood of a serious default. This conversion method uses dummy variable conversion.

[0701] The number of payroll payments made by agents in the past 12 months reflects the borrower's job stability. For recent new applicants, the lower the number of payroll payments made by agents in the past 12 months, the lower the borrower's stability and the higher the likelihood of a serious default. Conversely, the higher the number of payroll payments made by agents in the past 12 months, the greater the borrower's stability and the lower the likelihood of a serious default. The transformation method is natural logarithm.

[0702] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0703]

[0704] Where k is the number of features entering the model. In formula 2-5, k is 4.

[0705] α is the intercept term, with a range of values ​​from (-2.44505 to -2.74473), and an optimal value of -2.59489. β1 is the coefficient corresponding to the average monthly asset size in the past six months / the average monthly asset size in the past 12 months, with a range of values ​​from (-0.25311 to -0.41086), and an optimal value of -0.33199. β2 is the coefficient corresponding to the minimum balance of the deposit account in the past three months, with a range of values ​​from (-0.07435 to -0.46009), and an optimal value of -0.26722. β3 is the coefficient corresponding to the number of months from the minimum balance of the wealth management account in the past three months, with a range of values ​​from (-0.14379 to -0.14841), and an optimal value of -0.1461. β4 is the coefficient corresponding to the number of payroll payments in the past 12 months, with a range of values ​​from (-0.8194 to -1.10539), and an optimal value of -0.9624.

[0706] x1 is the original value of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months generated by the feature conversion step; x2 is the natural logarithm conversion value of the minimum balance of the deposit account in the past 3 months generated by the feature conversion step; x3 is the dummy variable conversion value of the number of months from the minimum balance of the wealth management account in the past 3 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the number of payroll payments in the past 12 months generated by the feature conversion step.

[0707] The model performance of some features is shown in Table 2-6 below:

[0708] Table 2-6

[0709]

[0710] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0711] Example 2-10

[0712] The P values ​​calculated by the above formulas can be used to further calculate the sub-scores and total scores of any customer.

[0713] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0714] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0715]

[0716]

[0717] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0718] In this embodiment, the P value calculated in the above embodiments can be substituted to calculate the score.

[0719] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0720] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0721] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0722] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Alpha model is 74. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0723] The following table shows the KS values ​​for the development and validation samples when the model was applied to the development and validation sets. Table 2-7 shows that the model performs well in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample set. This means that all sub-models have excellent discrimination.

[0724] Table 2-7

[0725]

[0726] 3) Credit card business credit score (Alpha2 model credit score)

[0727] Example 3-1 Collection of modeling samples

[0728] In the process of constructing this embodiment, credit card application and behavior data, personal loan application and behavior data, personal financial asset transaction data, and personal customer information data of all retail customers of a large bank between 2017 and 2021 were collected, totaling 730 million people. The modeling sample was confirmed through a professional model design solution, and 620 million data from 2018 and 650 million data from 2019 were selected as analysis samples. Because performance variables are required for modeling and the performance period of credit cards and special installment services is relatively short, 2019 was selected as the observation period, and the 650 million analysis samples were divided into two parts: those who have applied for registered credit business and those who have not applied for registered credit business, with sample sizes of 130 million and 520 million people, respectively.

[0729] The specific Alpha2 model is a scenario scoring for credit cards and special installment services. When designing the model, 110 million customers who have applied for credit cards and special installment services are analyzed and designed. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0730] The model design includes 1) exclusion rules: excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.59 million customers; 2) time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15 months as the performance period of Y; 3) sample sampling: use a good / bad sample number of 10:1 for sampling modeling. Figure 3 Different segmentation schemes are designed using the decision tree method shown in the figure. The final model segmentation scheme is confirmed by comparing the parent-child model, and then each sub-model is modeled separately. The model segmentation scheme of this application is designed based on whether the customer has applied, whether they are new customers, whether they are overdue, and the customer's repayment rate level. This fully reflects the control variables closely related to risk characteristics during business development and can more comprehensively match various customer groups in the market.

[0731] In this embodiment, the modeling sample is first divided into the following sub-samples based on the decision tree to facilitate the construction of subsequent sub-models. When determining the modeling sample, the first level of the decision tree makes a judgment based on whether the customer's account age is less than 3 (the customer has applied for credit registration but the application has not exceeded 3 months). If the account age is less than 3, the customer is classified as a new customer in the Alpha2 seg1 model. If the account age is greater than or equal to 3, further subdivision is performed. In this embodiment, customers with an account age greater than or equal to 3 will be divided based on whether the customer's current credit business is overdue. Customers with overdue credit business will be classified into the Alpha2 seg2 model, and customers with no overdue credit business will be further subdivided.

[0732] For existing customers whose account age is greater than or equal to 3 and whose current credit business has no overdue payments, they will be divided according to the repayment rate of their credit cards in the current month. If the repayment rate of the credit card in the current month is <90%, this customer group will be classified as Alpha2seg3, otherwise this customer group will be classified as Alpha2seg4.

[0733] In addition, in this embodiment, in order to build a model for the sample group that has not applied for credit business registration, new customers with an account age of less than 3 months are selected from the overall sample as an approximate customer group and enter the Alpha2 seg5 segment for model training.

[0734] Specifically, the customer sample size used to construct the Alpha2 seg1 sub-model is approximately 430,000. The customer base is mainly new customers who have applied for credit registration but the application period is less than 3 months. The model is used to predict the probability of their credit being overdue for more than 60 days.

[0735] Specifically, the customer sample size used to construct the Alpha2 seg2 sub-model is about 400,000. Its customer base mainly consists of existing customers who have applied for credit business registration for more than 3 months and whose current credit business is overdue. The model is used to predict the probability of their credit being overdue for more than 60 days.

[0736] Specifically, the customer sample size used to construct the Alpha2 seg3 sub-model is approximately 4.47 million. Its customer base mainly consists of existing customers who have applied for and registered for credit business for more than 3 months, have no overdue credit business, and have a current monthly credit card repayment rate of less than 90%. The model is built to predict the probability of this group of people having credit overdue for more than 60 days.

[0737] Specifically, the customer sample size used to construct the Alpha2 seg4 sub-model is approximately 53.29 million. Its customer base mainly consists of existing customers who have applied for and registered for credit business for more than 3 months, have no overdue credit business, and have a current monthly credit card repayment rate of more than 90%. The model is used to predict the probability of this group of people having credit overdue for more than 60 days.

[0738] Specifically, the customer sample size used to construct the Alpha2 seg5 sub-model is about 430,000. Its customer base is mainly new customers who have applied for credit business and have an account age of less than 3 months. The model is used to predict the probability of credit overdue for more than 60 days in the future for the sample group that has not applied for registered credit business.

[0739] Based on the different customer samples of each sub-model confirmed above, historical data information of borrowers is obtained: 1) Credit card category, basic fields include account, overdue, balance, credit limit, repayment due, actual repayment and other information under different dimensions such as bills, cash withdrawals, installments, and interest; 2) Personal loan category, including account, overdue, balance, credit limit, repayment due, actual repayment and other information; 3) Customer basic information, including gender, age and administrative region information of the business; 4) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0740] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 3,067 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount,' 'Average Number of Overdue Payments in the Last 6 Months,' and 'Number of Credit Card Installments in the Last 12 Months.' In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0741] Table 3-1 Summary of basic variables and derived variables used in this example

[0742]

[0743]

[0744] Example 3-2 Feature Screening

[0745] A preliminary screening was conducted on the 3067 features (variables) collected in Example 3-1 that have potential predictive power for customer overdue payments.

[0746] The preliminary screening of the features of the four sub-models Alpha2 seg1 to Alpha2 seg4 divided in Example 3-1 is performed as follows:

[0747] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 3-1 were first eliminated, resulting in a total of 206 variables being deleted, leaving 2861 variables.

[0748] In the second round of preliminary screening, for the 2861 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% of the remaining features were eliminated, and a total of 350 variables were eliminated, leaving 2511 variables.

[0749] In the third round of preliminary screening, the 2511 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box, and if the feature is a numeric variable, it is sorted according to the value from small to large), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 1209 features were eliminated, and 1302 feature variables remained.

[0750] In the fourth round of preliminary screening, the features are further screened based on the stepwise discrimination algorithm for the 1,302 features after the third round of screening. After this round of screening, 268 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0751] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0752] The fifth round of preliminary screening targets the 268 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0753] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0754] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0755] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0756] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0757] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0758] The criteria for approximate consistency are as follows:

[0759] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0760] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0761] After the fifth round of screening, 121 features were eliminated, leaving 147 features.

[0762] For the Alpha2 seg5 sub-model divided in Example 3-1, the selected variables are customer basic information and personal financial asset variables (a total of 316 feature variables) used to approximately predict the probability of credit defaults in the future for customers who have not registered for credit applications. The initial feature screening step and subsequent steps are the same as other sub-models.

[0763] Example 3-3 Conversion of features after initial screening

[0764] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0765] First, the remaining 147 features in Example 3-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0766] Featu...

Claims

1. A method for calculating retail credit risk, comprising: Data collection step, which obtains retail credit prediction data of the sample to be predicted; a data processing step, which processes the acquired retail credit forecast data to obtain second-generation derived retail credit forecast data; The credit default probability calculation step is to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

2. The method according to claim 1, further comprising: After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.

3. The method according to claim 1 or 2, wherein: The retail credit prediction data includes original retail credit prediction data of the sample to be predicted and derived retail credit prediction data processed based on the original retail credit prediction data; Preferably, the original retail credit prediction data includes: Credit card basic data, which is based on all available data during the sample user's credit card creation and usage process. Basic data on personal loans, which is all available data based on the loan application and usage behavior of sample users. Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions. Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

4. The method according to any one of claims 1 to 3, wherein The derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension; Preferably, the derived retail credit forecast data includes but is not limited to: Derivative retail credit forecast data obtained by processing based on sample relationship length, Derivative retail credit forecast data obtained by processing time interval variables, Derivative retail credit forecast data obtained based on the frequency of sample behavior. Derivative retail credit forecast data obtained by processing the sample at the current time point, Derivative retail credit forecast data obtained based on sample continuous behavior processing, Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

5. The method according to any one of claims 1 to 4, wherein In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data. The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

6. The method according to claim 5, wherein: The comprehensive business credit score includes a comprehensive business credit total score and 12 seed scores; wherein the comprehensive business credit total score and the 12 seed scores are calculated based on the retail credit prediction data using the 12 seed models of the calculation model; wherein, based on the 12 seed models, the 12 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score; The credit card business credit score includes four credit card business credit total scores and 20 seed scores; wherein, the four credit card business credit total scores and 20 seed scores are calculated based on the retail credit prediction data using the 20 seed models obtained by the calculation model; wherein, based on the 20 seed models, the 20 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method using the sub-model and the score calculated by the sub-model is used as the four credit card business credit total scores; The personal loan business credit score includes a total personal loan business credit score and five sub-scores; wherein the total personal loan business credit score and the five sub-scores are calculated based on the retail credit prediction data using the five sub-models of the calculation model; wherein, based on the five sub-models, the five sub-scores of a sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal loan business credit score; The personal consumer loan business credit score includes three personal consumer loan business credit total scores and 19 seed scores; wherein, the three personal consumer loan business credit total scores and 19 seed scores are calculated based on the retail credit prediction data using the 19 seed models of the calculation model; wherein, based on the 19 seed models, the 19 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the three personal consumer loan business credit total scores; The personal mortgage business credit score includes a total personal mortgage business credit score and five seed scores; wherein the total personal mortgage business credit score and the five seed scores are calculated based on the retail credit prediction data using the five seed models of the calculation model; wherein, based on the five seed models, the five seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine the sub-model and the score calculated by the sub-model is used as the total personal mortgage business credit score; The personal quick loan business credit score includes a personal quick loan business credit total score and 6 seed scores; wherein, the personal quick loan business credit total score and 6 seed scores are calculated based on the retail credit prediction data using the 6 seed models of the calculation model; wherein, based on the 6 seed models, the 6 seed scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on the decision tree method to determine the sub-model and the score calculated by the sub-model is used as the personal quick loan business credit total score.

7. The method according to claim 6, wherein: The 12 sub-models of the comprehensive service class are shown in formulas (1-1) to (1-12). The 20 sub-models of the credit card business class are respectively shown as formulas (2-1) to (2-5), formulas (3-1) to (3-5), formulas (4-1) to (4-5), and formulas (5-1) to (5-5); The six sub-models of the personal loan business are shown in formulas (6-1) to (6-6). The 19 seed models of the personal consumption loan business are shown in Formulas (7-1) to (7-7), (8-1) to (8-6), and (9-1) to (9-6). The five sub-models of the personal mortgage business are shown in formulas (10-1) to (10-5). The six seed models of the personal quick loan business are shown in formulas (11-1) to (11-6).

8. The method according to claim 6, wherein: The second-generation derivative retail credit prediction data is subjected to feature conversion and then substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes: Based on the feature type of the second-generation derivative retail credit prediction data that needs to be substituted into the credit default probability model, the WOE method, dummy feature method or continuous method is selected for feature conversion.

9. The method according to claim 8, wherein The continuous method for feature conversion includes the following methods: calculating the square root of the second-generation derivative retail credit forecast data, calculating the natural logarithm of the second-generation derivative retail credit forecast data, or calculating the cube root of the second-generation derivative retail credit forecast data.

10. The method according to claim 9, wherein: The credit default probability model is a model constructed based on the second-generation derivative retail credit forecast data and credit default probability using logistic regression based on the existing user population.

11. The method according to claim 10, wherein: The second-generation derivative retail credit prediction data is selected from one or more of the comprehensive business credit score, the credit card business credit score, the personal mortgage business credit score, the personal loan business credit score and the personal consumption loan business credit score.

12. The method according to claim 11, wherein The credit default probability calculation step includes: The total credit score of comprehensive business, credit card business, personal mortgage business, personal loan business and personal consumption loan business are converted into features. It is preferred to use a continuous conversion method to convert the total credit score of comprehensive business; a continuous conversion method to convert the total credit score of credit card business; a continuous conversion method to convert the total credit score of personal mortgage business; a dummy variable conversion method to convert the total credit score of personal loan business; and a continuous conversion method to convert the total credit score of personal consumer loan business. Further preferably, the total credit score of the comprehensive business type is converted into a calculation method of taking the cube root of the total credit score of the comprehensive business type by a continuous conversion method; the total credit score of the credit card business type is converted into a calculation method of taking the square root of the total credit score of the credit card business type by a continuous conversion method; the total credit score of the personal mortgage business type is converted into a calculation method of taking the natural logarithm of the total credit score of the personal mortgage business type by a continuous conversion method; the total credit score of the personal consumption loan business type is converted into a calculation method of taking the natural logarithm of the total credit score of the personal consumption loan business type by a continuous conversion method.

13. The method according to claim 12, wherein: The converted values ​​of the five features, namely the comprehensive business credit score, credit card business credit score, personal mortgage business credit score, personal loan business credit score and personal consumption loan business credit score, are substituted into the second-generation derivative retail credit prediction data and the credit default probability, and the credit default probability model constructed using logistic regression is used to calculate the credit default probability of the sample to be predicted.

14. The method according to claim 13, wherein The credit default probability model is shown in the following formula 12-1: Where k is the number of features entering the model, and in Formula 12-1, k is 5; α is the intercept term, with a value range of (1.3714, 1.2146) and an optimal value of 1.293; β1 is the coefficient corresponding to the total credit score of credit card business, with a value range of (-0.725, -0.921) and an optimal value of -0.823; β2 is the corresponding coefficient of the total credit score of personal mortgage business, with a value range of (-0.4084, -0.4476) and an optimal value of -0.428; β3 is the coefficient corresponding to the total credit score of comprehensive business, with a value range of (-0.7294, -0.9646) and an optimal value of -0.847; β4 is the coefficient corresponding to the total credit score of personal loan business, with a value range of (-0.6156, -0.9684) and an optimal value of -0.792; β5 is the coefficient corresponding to the total credit score of personal consumer loan business, with a value range of (-0.5424, -0.5816), and the optimal value is -0.562 (Note: the value range comes from the 95% confidence interval, i.e., the 95% CI in the table below); x1 is the square root transformation of the total credit score of the credit card business generated in the feature conversion step; x2 is the natural logarithm transformation value of the total credit score of personal mortgage business generated in the feature conversion step; x3 is the cube root conversion value of the comprehensive business credit score generated in the feature conversion step; x4 is the Woe conversion value of the total credit score of the personal loan business generated in the feature conversion step; x5 is the natural logarithm transformation of the total credit score of the personal consumption loan business generated in the feature transformation step.

15. The method according to claim 14, wherein After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower: Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

16. A device for calculating retail credit risk, comprising: A data collection module, which is used to obtain retail credit prediction data of the sample to be predicted; A module for processing samples to be predicted, which is used to process the acquired retail credit prediction data to obtain second-generation derived retail credit prediction data; A credit default probability calculation module is used to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

17. The device according to claim 16, wherein The device executes the steps of the method for calculating retail credit risk according to any one of claims 1 to 15.

18. A system for calculating retail credit risk, characterized in that: The system for calculating retail credit risk includes: a memory, a processor, and a program for the method for calculating retail credit risk stored in the memory and executable on the processor. When the program for calculating retail credit risk is executed by the processor, the steps of the method for calculating retail credit risk according to any one of claims 1 to 15 are implemented.

19. A computer storage medium, characterized in that The computer storage medium stores a program for calculating the retail credit risk method. When the program for calculating the retail credit risk method is executed by a processor, the steps of the method for calculating the retail credit risk according to any one of claims 1 to 15 are implemented.

20. A method for constructing a retail credit risk prediction model, comprising: A data collection step, which obtains the original retail credit prediction data of the sample used to build the model; A data derivation step, which processes the original retail credit forecast data into derived retail credit forecast data; a data processing step, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data; a feature initial screening step, which performs a preliminary screening of all categories, that is, all features, comprising the second-generation derivative retail credit forecast data to obtain features after the preliminary screening; The initial screening data conversion step determines the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method, and continuous conversion method for feature conversion, and uses the optimal method for feature conversion for each feature after the initial screening; Feature fine screening step, in which the features initially screened after feature conversion are deeply screened to obtain finely screened features; The credit default probability modeling step selects the logistic regression method to build the model based on the carefully screened features and the probability relationship between them and credit default, and confirms the method used to calculate the credit default probability.

21. The method according to claim 20, wherein In the data collection step, the original retail credit prediction data of the samples obtained for building the model include: Credit card basic data, which is based on all available data during the sample user's credit card creation and usage process. Basic data on personal loans, which is all available data based on the loan application and usage behavior of sample users. Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions. Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

22. The method according to claim 20, wherein In the data derivation step, the derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension; Preferably, the derived retail credit forecast data includes but is not limited to: Derivative retail credit forecast data obtained by processing based on sample relationship length, Derivative retail credit forecast data obtained by processing time interval variables, Derivative retail credit forecast data obtained based on the frequency of sample behavior. Derivative retail credit forecast data obtained by processing the sample at the current time point, Derivative retail credit forecast data obtained based on sample continuous behavior processing, Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

23. The method according to claim 20, wherein In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data. The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

24. The method according to any one of claims 20 to 23, wherein The feature initial screening step includes the following steps: The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model. The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high. The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features; The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order, The fourth preliminary screening step uses a stepwise discriminant algorithm to perform preliminary screening of features after the first to third preliminary screenings; The fifth preliminary screening step is to perform preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

25. The method according to any one of claims 20 to 23, wherein In the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.

26. The method according to claim 25, wherein The initial screening data conversion steps include the following steps based on the judgment of concentration and data type: Classify the data type of each feature into character variables and numeric variables. For character variables, dummy feature conversion is used to convert the initial screening data. The process of further classifying numerical variables includes the following sub-steps: If the value of the numerical variable is less than n, the WOE conversion method is used to convert the initial screening data. If the value of the numerical variable is more than n, further judge if the conversion to a continuous variable has more values ​​and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used. Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.

27. The method according to claim 26, wherein Also includes: For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

28. The method according to any one of claims 20 to 27, wherein The feature fine screening steps include: The first fine screening step is based on the stepwise regression algorithm, and the significance of the features is screened based on the F test and T test. The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate features with higher variance inflation factors to screen features. The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default in order to further perform feature screening.

29. The method according to any one of claims 20 to 28, wherein The credit default probability modeling step substitutes the features selected in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.

30. A device for constructing a retail credit risk prediction model, characterized in that: The device comprises: A data acquisition module, which is used to obtain raw retail credit prediction data of samples used to build a model; A data derivation module, which is used to process derived retail credit forecast data based on the original retail credit forecast data; A data processing module, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data; A feature initial screening module is used to perform initial screening on all categories, i.e., all features, of the original retail credit forecast data and the derived retail credit forecast data to obtain features after initial screening; The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening; A feature fine-screening module is used to perform in-depth screening on the preliminarily screened features after feature conversion to obtain fine-screened features; The credit default probability modeling module is used to select the logistic regression method to construct a model based on the probability relationship between the carefully screened features and the credit default, and confirm the model used to calculate the credit default probability.

31. The device according to claim 30, wherein The device executes the steps of the method for constructing a retail credit risk prediction model according to any one of claims 20 to 29.

32. A system for constructing a retail credit risk prediction model, characterized in that: The system includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor. When the program for constructing a retail credit risk prediction model method is executed by the processor, the steps of the method for constructing a retail credit risk prediction model method as described in any one of claims 20 to 29 are implemented.

Citation Information

Patent Citations

  • A credit risk assessment method and apparatus based on logistic regression technology

    CN112686749B