Method for constructing retail credit risk prediction model and retail credit Scoreassetplus model

By building a retail credit risk prediction model based on large bank data, the problem of single data sample quantity and source faced by financial institutions has been solved, achieving more accurate credit risk assessment and improved risk control capabilities.

CN120707264APending Publication Date: 2025-09-26CCB FINTECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410306193.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When building retail credit risk prediction models, financial institutions face problems such as small data sample size, single data source, and high homogeneity of data dimensions. These problems lead to insufficient model accuracy and stability, making it difficult to effectively improve risk control capabilities.

Method used

Adopting statistical principles based on massive data from large banks, we build a retail credit risk prediction model. By utilizing highly stable and high-coverage data samples, combined with various credit business scenarios and customer asset information, we calculate the probability of credit default and generate a credit score through data processing, feature conversion, and logistic regression models.

Benefits of technology

It improves the accuracy and stability of credit risk prediction, can more comprehensively measure the borrower's debt repayment ability and willingness to pay, and provide more accurate credit risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004745922830000051
    Figure BDA0004745922830000051
  • Figure BDA0004745922830000061
    Figure BDA0004745922830000061
  • Figure BDA0004745922830000062
    Figure BDA0004745922830000062
Patent Text Reader

Abstract

The invention relates to a method for calculating retail credit risks, and the method comprises the steps: data collection: obtaining retail credit prediction data of a to-be-predicted sample; a data processing step of processing the obtained retail credit prediction data to obtain second generation derivative retail credit prediction data; and a credit default probability calculation step: substituting the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the to-be-predicted sample. According to the application, during model construction, the credit scores constructed based on different credit business scenes of the financial institution are used as the characteristic variables, and the repayment performance of all credit products of the institution by the customer is used as the performance variable to construct the model, so that behaviors and risk trends of the customer in different credit scenes are integrated; therefore, the final model can cover the risk characteristics of different customer groups under different businesses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a credit risk management system and method. The method and system of this application can assist financial institutions in making more accurate risk decisions and accelerate their digital transformation. Specifically, this application relates to a method for constructing a retail credit risk prediction model and a retail credit Scoreassetplus model. Background Art

[0002] In the current environment of booming consumer lending, some financial institutions' manual approval mechanisms are no longer able to cope with the increasing demand for credit. Consequently, there is an urgent need to enhance their intelligent risk control capabilities. Financial institutions hope to establish risk mitigation mechanisms that cover the entire credit process, from customer pre-screening, pre-loan review, mid-loan approval, post-loan management, and early collection.

[0003] If a scoring system can be developed based on the principles of early identification, early warning, early detection and early disposal, and credit business can be monitored and managed quickly and conveniently, the financial institutions' own business volume, competitive advantage and asset quality can be improved while risks are controllable.

[0004] However, building a scoring system is highly dependent on data and technology. The diversity and coverage of data dimensions, as well as modeling techniques and methodologies, directly impact the ultimate stability and ranking of the scoring system. Some financial institutions lack experience with intelligent risk control for retail businesses and therefore have weak risk control capabilities. In actual applications, factors such as insufficient data mining and analysis capabilities and weak risk modeling techniques hinder financial institutions from fully leveraging the value of internal data and effectively improving model accuracy and stability, leading to technical control challenges. This is also one of the main obstacles facing small and medium-sized financial institutions in their digital transformation. Summary of the Invention

[0005] To address the shortcomings of the aforementioned prior art, this application aims to provide a credit risk management system and method that can provide effective risk management for financial institutions. The credit risk prediction method and system of this application is based on massive amounts of data from large banks and utilizes statistical principles to extract risk patterns, which is inherently valuable for promotion.

[0006] The construction of other popular scoring models currently on the market is often hampered by factors such as small sample sizes, relatively limited data sources, and high homogeneity in data dimensions. Furthermore, because most current risk assessment models utilize data with weak financial attributes—that is, they are based on non-credit transaction data such as smart terminal device data, social platform data, and online shopping mall data—and non-overdue prediction targets, their predictions often deviate significantly from actual credit overdue situations.

[0007] This application first provides a method that can be used to construct a retail credit risk prediction model, and the model constructed based on this method provides financial institutions with a set of accurate methods for calculating the retail credit risk of samples to be predicted.

[0008] The method and system of this application are developed based on highly stable and high-coverage data samples, and systematically innovate the relatively mature credit risk control system. They use credit samples covering various business forms to predict the potential retail credit risks of financial institutions.

[0009] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0010] This application involves the following technical solutions:

[0011] 1. A method for calculating retail credit risk, comprising:

[0012] Data collection step, which obtains retail credit prediction data of the sample to be predicted;

[0013] a data processing step, which processes the acquired retail credit forecast data to obtain second-generation derived retail credit forecast data;

[0014] The credit default probability calculation step is to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

[0015] 2. The method according to claim 1, further comprising:

[0016] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.

[0017] 3. The method according to claim 1 or 2, wherein

[0018] The retail credit prediction data includes original retail credit prediction data of the sample to be predicted and derived retail credit prediction data processed based on the original retail credit prediction data;

[0019] Preferably, the original retail credit prediction data includes:

[0020] Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions.

[0021] Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

[0022] 4. The method according to any one of items 1 to 3, wherein

[0023] The derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension;

[0024] Preferably, the derived retail credit forecast data includes but is not limited to:

[0025] Derivative retail credit forecast data obtained by processing based on sample relationship length,

[0026] Derivative retail credit forecast data obtained by processing time interval variables,

[0027] Derivative retail credit forecast data obtained based on the frequency of sample behavior.

[0028] Derivative retail credit forecast data obtained by processing the sample at the current time point,

[0029] Derivative retail credit forecast data obtained based on sample continuous behavior processing,

[0030] Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

[0031] 5. The method according to any one of items 1 to 4, wherein

[0032] In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data.

[0033] The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

[0034] 6. The method according to claim 5, wherein

[0035] The comprehensive business credit score includes a comprehensive business credit total score and three sub-scores; wherein the comprehensive business credit total score and the three sub-scores are calculated based on the retail credit prediction data using the three sub-models of the calculation model; wherein, based on the three sub-models, the three sub-scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score;

[0036] The credit card business credit score includes four credit card business credit scores; wherein the four credit card business credit scores are calculated based on the retail credit prediction data using the calculation model;

[0037] The personal loan business credit score is calculated based on the retail credit prediction data using the calculation model;

[0038] The personal consumer loan business credit score includes three types of personal consumer loan business credit scores; wherein the three types of personal consumer loan business credit scores are calculated based on the retail credit prediction data using the calculation model;

[0039] The personal mortgage business credit score is calculated based on the retail credit prediction data using the calculation model;

[0040] The personal quick loan business credit score is calculated based on the retail credit prediction data using the calculation model.

[0041] 7. The method according to claim 6, wherein

[0042] The three sub-models of the integrated service class are shown in formulas (1-1) to (1-3).

[0043] The four calculation models of the credit card business class are shown in Formula 2, Formula 3, Formula 4 and Formula 5 respectively;

[0044] The calculation models of the personal loan business are shown in Formula 6 respectively;

[0045] The three calculation models for the personal consumption loan business are shown in Formula 7, Formula 8, and Formula 9 respectively;

[0046] The calculation model of the personal mortgage business is shown in Formula 10;

[0047] The calculation models for the personal quick loan business are shown in Formula 11.

[0048] 8. The method according to claim 6, wherein

[0049] The second-generation derivative retail credit prediction data is subjected to feature conversion and then substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes:

[0050] Based on the feature type of the second-generation derivative retail credit prediction data that needs to be substituted into the credit default probability model, the WOE method, dummy feature method or continuous method is selected for feature conversion.

[0051] 9. The method according to claim 8, wherein

[0052] The continuous method for feature conversion includes the following methods: performing continuous feature conversion by calculating the cube root of the second-generation derivative retail credit forecast data or calculating the natural logarithm of the second-generation derivative retail credit forecast data.

[0053] 10. The method according to claim 9, wherein

[0054] The credit default probability model is a model constructed based on the second-generation derivative retail credit forecast data and credit default probability using logistic regression based on the existing user population.

[0055] 11. The method according to claim 10, wherein

[0056] The second-generation derivative retail credit prediction data is selected from one or more of the comprehensive business credit score, the personal mortgage business credit score, and the personal loan business credit score.

[0057] 12. The method according to claim 11, wherein

[0058] The credit default probability calculation step includes:

[0059] Perform feature conversion on the comprehensive business credit score, personal mortgage business credit score and personal loan business credit score.

[0060] It is preferred to use the continuous conversion method for the comprehensive business credit score; the continuous conversion method for the credit card business credit score; the continuous conversion method for the personal mortgage business credit score; and the WOE conversion method for the personal loan business credit score;

[0061] Further preferred is to use a continuous conversion method to convert the total comprehensive business credit score into a calculation method of taking the cube root of the total comprehensive business credit score; and to use a continuous conversion method to convert the personal mortgage business credit score into a calculation method of taking the natural logarithm of the personal mortgage business credit score.

[0062] 13. The method according to claim 12, wherein

[0063] The converted values ​​of the three features, namely the comprehensive business credit score, the personal mortgage business credit score and the personal loan business credit score, are substituted into the second-generation derivative retail credit prediction data and the credit default probability, and the credit default probability model constructed using logistic regression is used to calculate the credit default probability of the sample to be predicted.

[0064] 14. The method according to claim 13, wherein

[0065] The credit default probability model is shown in the following formula 12-1:

[0066]

[0067] Where k is the number of features entering the model, and in Formula 12-1 k is 3;

[0068] α is the intercept term, the value range is (5.5636, 5.4068), and the optimal value is 5.458;

[0069] The value range of β1 is (0.10676, -0.12844), and the optimal value is -0.011;

[0070] The β2 value range is (0.10474, -0.24806), and the optimal value is -0.072;

[0071] The value range of β3 is (0.00984, -0.02936), and the optimal value is -0.010;

[0072] x1 is the cube root transformation of the comprehensive business credit score generated in the feature transformation step;

[0073] x2 is the WOE conversion value of the total credit score of personal loan business generated in the feature conversion step;

[0074] x3 is the logarithmic transformation value of the total credit score of the personal mortgage business generated in the feature conversion step.

[0075] 15. The method according to claim 14, wherein

[0076] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower:

[0077]

[0078]

[0079] Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0080] 16. A device for calculating retail credit risk, comprising:

[0081] A data collection module, which is used to obtain retail credit prediction data of the sample to be predicted;

[0082] A module for processing samples to be predicted, which is used to process the acquired retail credit prediction data to obtain second-generation derived retail credit prediction data;

[0083] A credit default probability calculation module is used to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

[0084] 17. The apparatus according to item 16, wherein the apparatus executes the steps of the method for calculating retail credit risk according to any one of items 1 to 15.

[0085] 18. A system for calculating retail credit risk, characterized in that the system for calculating retail credit risk comprises: a memory, a processor, and a program for the method for calculating retail credit risk stored in the memory and executable on the processor, wherein the program for calculating retail credit risk, when executed by the processor, implements the steps of the method for calculating retail credit risk as described in any one of items 1 to 15.

[0086] 19. A computer storage medium, characterized in that a program for calculating retail credit risk is stored on the computer storage medium, and when the program for calculating retail credit risk is executed by a processor, the steps of the method for calculating retail credit risk as described in any one of items 1 to 15 are implemented.

[0087] 20. A method for constructing a retail credit risk prediction model, comprising:

[0088] A data collection step, which obtains the original retail credit prediction data of the sample used to build the model;

[0089] A data derivation step, which processes the original retail credit forecast data into derived retail credit forecast data;

[0090] a data processing step, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data;

[0091] a feature initial screening step, which performs a preliminary screening of all categories, that is, all features, comprising the second-generation derivative retail credit forecast data to obtain features after the preliminary screening;

[0092] The initial screening data conversion step determines the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method, and continuous conversion method for feature conversion, and uses the optimal method for feature conversion for each feature after the initial screening;

[0093] Feature fine screening step, in which the features initially screened after feature conversion are deeply screened to obtain finely screened features;

[0094] The credit default probability modeling step selects the logistic regression method to build the model based on the carefully screened features and the probability relationship between them and credit default, and confirms the method used to calculate the credit default probability.

[0095] 21. The method according to claim 20, wherein

[0096] In the data collection step, the original retail credit prediction data of the samples obtained for building the model include:

[0097] Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions.

[0098] Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

[0099] 22. The method according to claim 20, wherein

[0100] In the data derivation step, the derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension;

[0101] Preferably, the derived retail credit forecast data includes but is not limited to:

[0102] Derivative retail credit forecast data obtained by processing based on sample relationship length,

[0103] Derivative retail credit forecast data obtained by processing time interval variables,

[0104] Derivative retail credit forecast data obtained based on the frequency of sample behavior.

[0105] Derivative retail credit forecast data obtained by processing the sample at the current time point,

[0106] Derivative retail credit forecast data obtained based on sample continuous behavior processing,

[0107] Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

[0108] 23. The method according to claim 20, wherein

[0109] In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data.

[0110] The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

[0111] 24. The method according to any one of items 20 to 23, wherein

[0112] The feature initial screening step includes the following steps:

[0113] The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model.

[0114] The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high.

[0115] The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features;

[0116] The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order,

[0117] The fourth preliminary screening step uses a stepwise discriminant algorithm to perform preliminary screening of features after the first to third preliminary screenings;

[0118] The fifth preliminary screening step is to perform preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

[0119] 25. A method according to any one of items 20 to 23, wherein, in the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.

[0120] 26. The method according to claim 25, wherein

[0121] The initial screening data conversion steps include the following steps based on the judgment of concentration and data type:

[0122] Classify the data type of each feature into character variables and numeric variables.

[0123] For character variables, dummy feature conversion is used to convert the initial screening data.

[0124] The process of further classifying numerical variables includes the following sub-steps:

[0125] If the value of the numerical variable is less than n, the WOE conversion method is used to convert the initial screening data.

[0126] If the value of the numerical variable is more than n, further judge if the conversion to a continuous variable has more values ​​and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used.

[0127] Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.

[0128] 27. The method according to claim 26, further comprising:

[0129] For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature.

[0130] Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

[0131] 28. The method according to any one of items 20 to 27, wherein the feature fine screening step comprises:

[0132] The first fine screening step is based on the stepwise regression algorithm, and the significance of the features is screened based on the F test and T test.

[0133] The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate features with higher variance inflation factors to screen features.

[0134] The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default in order to further perform feature screening.

[0135] 29. A method according to any one of items 20 to 28, wherein the credit default probability modeling step substitutes the features screened in the feature fine screening step into a Sigmoid function to perform logistic regression to calculate a model for the credit default probability.

[0136] 30. A device for constructing a retail credit risk prediction model, characterized in that the device comprises:

[0137] A data acquisition module, which is used to obtain raw retail credit prediction data of samples used to build a model;

[0138] A data derivation module, which is used to process derived retail credit forecast data based on the original retail credit forecast data;

[0139] A data processing module, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data;

[0140] A feature initial screening module is used to perform initial screening on all categories, i.e., all features, of the original retail credit forecast data and the derived retail credit forecast data to obtain features after initial screening;

[0141] The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening;

[0142] A feature fine-screening module is used to perform in-depth screening on the preliminarily screened features after feature conversion to obtain fine-screened features;

[0143] The credit default probability modeling module is used to select the logistic regression method to construct a model based on the probability relationship between the carefully screened features and the credit default, and confirm the model used to calculate the credit default probability.

[0144] 31. The apparatus according to item 30, wherein the apparatus executes the steps of the method for constructing a retail credit risk prediction model according to any one of items 20 to 29.

[0145] 32. A system for constructing a retail credit risk prediction model, characterized in that the system includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor, wherein the program for constructing a retail credit risk prediction model method, when executed by the processor, implements the steps of constructing a retail credit risk prediction model method as described in any one of items 20 to 29.

[0146] Effects of the Invention

[0147] The method and system for constructing a retail credit risk prediction model in this application combines a large number of samples from a large financial institution when constructing the model, and deeply processes and derives the original data obtained from the samples. It uses advanced statistical analysis methods and explainable machine learning technology to construct a general retail credit scoring model based on the inherent characteristics of the original data and derived data, as well as the strong financial attribute information contained in the data that is difficult to obtain in the market.

[0148] When constructing the model, this application uses the credit score constructed based on the different credit business scenarios of the financial institution as the characteristic variable, and all the asset information and demographic characteristics of the customer in the institution as the performance variables to construct the model. It integrates the customer's behavior and risk trends in different credit scenarios, thereby ensuring that the final model can cover the risk characteristics of different customer groups under different businesses. DETAILED DESCRIPTION

[0149] Credit risk is the risk that arises when a borrower's financial capacity changes, resulting in a reduced willingness to repay or an inability to fulfill the loan contract, rather than the risk of default due to deliberate fraud on the part of the borrower. Credit default occurs in all types of retail credit business scenarios and is particularly related to the borrower's personal financial situation. The reasons for borrowers' credit defaults can be divided into four main categories: 1. Short credit history, where such borrowers have little experience in managing their financial situation; 2. Borrowers temporarily forget to repay; 3. Overborrowing, where such borrowers have relatively low repayment capacity due to large debts; and 4. Major negative factors, where such borrowers experience long-term impacts on their repayment capacity due to factors such as reduced income, unemployment, or divorce. Each of these different reasons will, to varying degrees, lead to a borrower's default, potentially leading to a more serious default. Credit risk scoring aims to uncover the inherent mathematical relationship between various historical customer information and the probability of future default, and convert it into a score to quantify the probability of default.

[0150] Currently, credit risk scores in existing technologies are primarily developed using historical credit data and multi-credit data (multi-credit data refers to the statistical data of borrowers' loan requests from multiple financial institutions. It is generally believed that a greater number of multi-credit requests in a short period of time indicates a greater probability of future defaults). These data reflect their payment behavior and willingness to pay. The sample data used to build the model in this application not only includes transaction information on credit business but also adds asset data that is typically difficult to obtain. This not only reflects the borrower's payment behavior and willingness to pay, but also provides a more comprehensive assessment of their debt repayment ability, personal qualifications, and other aspects, thereby providing more accurate prediction results.

[0151] <Overall description of model building method>

[0152] Specifically, the present application relates to a method for constructing a retail credit risk prediction model, which includes: a data collection step, which obtains original retail credit prediction data of samples used to construct the model; a data derivation step, which processes derived retail credit prediction data based on the original retail credit prediction data; a data processing step, which processes the acquired retail credit prediction data and derived retail credit prediction data using a calculation model of second-generation derived retail credit prediction data to obtain the second-generation derived retail credit prediction data; a feature initial screening step, which performs preliminary screening on all categories, that is, all features, including the second-generation derived retail credit prediction data to obtain features after preliminary screening; a preliminary screening data conversion step, which judges the conversion method of the features after preliminary screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and uses the best method judged for each feature after preliminary screening to perform feature conversion; a feature fine screening step, which performs in-depth screening on the features after preliminary screening to obtain fine-screened features; a credit default probability modeling step, which selects the logistic regression method for model construction based on the fine-screened features in combination with the probability relationship between the features and credit default, and confirms the method for calculating the credit default probability.

[0153] <Raw Retail Credit Forecast Data>

[0154] When constructing the model for this application, we first obtain approximately original retail credit forecast data based on the basic data of financial institutions, and split or process the retail credit risk points as much as possible according to different dimensions under the premise of optimal effect, generating more than one hundred derivative features in total.

[0155] When building the model, first of all, based on all historical data of large financial institutions, the personal financial asset transaction data and personal customer information data of all retail customers over the past four years were initially collected when building the model of this application, totaling 780 million people, of which each person has corresponding data to be processed every month. It can be seen that the data system used to build the model of this application is comprehensive and the data volume is very large. When building a model based on such a data system, it is necessary to consider the modeling methodology, otherwise it will be trapped in the huge data, resulting in the special sample groups that need to be paid attention to being covered in the huge data volume and unable to be effectively identified, causing the computer program to run slowly or even unable to run, so that it is impossible to accurately build the most suitable prediction model.

[0156] In the data collection step, the original retail credit prediction data of the samples obtained for building the model include: basic customer information data, which is based on the attributes of the sample itself but is not directly related to the behavior in the financial institution (i.e., personal customer information data), or basic personal financial asset data, which is all other financial assets and financial transaction data of the sample in the financial institution that are not related to credit cards and loans (personal financial asset transaction data).

[0157] In a specific implementation, in accordance with the requirements of the Personal Privacy Protection Act, the basic data for basic customer information only includes basic customer information, including gender, age, and the administrative region where the business is located. The above basic data is not limited to the specific categories listed. As customer situations and social relationships evolve, those skilled in the art can further include other or emerging data types within the scope of the relevant business application scenarios. In other words, all types of data based on the attributes of the user sample itself but not directly related to the behavior of the financial institution can serve as basic customer information data.

[0158] In one specific embodiment, basic data on personal financial assets includes, but is not limited to, AUM (assets under management), deposits, wealth management, and payroll information. This basic data is not limited to the specific categories listed above. As financial assets evolve, those skilled in the art can further encompass emerging data types during implementation. Specifically, all other financial assets and financial transaction data unrelated to credit cards and loans at financial institutions, based on the sample, can serve as basic data on personal financial assets.

[0159] In a specific embodiment, the original retail credit prediction data of the sample obtained for building the model is based on the data types (i.e., basic variables or basic features) obtained from 780 million people, including but not limited to: AUM (i.e., asset management scale), deposits, wealth management, payroll payment and other basic information.

[0160] <Derivative Retail Credit Forecast Data>

[0161] In this application, in the data derivation step, the derived retail credit prediction data processed based on the original retail credit prediction data refers to the data obtained by processing the collected original retail credit prediction data based on the time dimension, space dimension, frequency dimension, and statistical information dimension.

[0162] In one specific embodiment, derived retail credit prediction data includes, but is not limited to, derived retail credit prediction data processed based on sample relationship length, derived retail credit prediction data processed based on time interval variables, derived retail credit prediction data processed based on the frequency of sample behavior, derived retail credit prediction data processed based on the current time point of the sample, derived retail credit prediction data processed based on the continuous behavior of the sample, or derived retail credit prediction data processed based on statistical information dimensions. For example, monthly customer data can be obtained and processed based on this monthly data. In this application, processing based on statistical information dimensions includes obtaining the maximum, minimum, and average values ​​of the data to describe the data situation.

[0163] In a specific implementation, for example, starting from the time dimension, for example, customer relationship length variables: for example, the customer account opening time, the customer's maximum account age, etc. are used as types of derived retail credit prediction data, that is, as derived features or derived variables.

[0164] In a specific implementation, for example, starting from the time dimension, time intervals are considered, such as the number of months from the customer's most recent repayment to the current time point, the number of months from the customer's most recent overdue payment to the current time point, etc. as derived features or derived variables.

[0165] In a specific implementation, for example, starting from the frequency level, behavioral frequency level variables are considered: for example, the number of times a customer has made repayments greater than N in the last X months, the number of times a customer's credit limit utilization rate greater than N in the last X months, etc. are used as derived features or derived variables. There is no limit on X and N, and they can be any positive integer greater than 0 as long as the business logic is reasonable.

[0166] In a specific implementation, for example, starting from the time dimension, current point variables are considered: the customer's current monthly credit limit, the customer's current monthly balance, etc. as derived features or derived variables.

[0167] In a specific approach, starting from the time dimension, continuous behavioral variables are considered, such as the maximum number of consecutive overdue payments of a customer > N in the last X months, the number of consecutive repayment rates of a customer > N in the last X months, etc. as derived features or derived variables.

[0168] In a specific approach, starting from the dimension of statistical information, statistical variables are considered, such as the maximum number of overdue payments of a customer in the last X months, the average credit limit utilization rate of a customer in the last X months, etc. as derived features or derived variables.

[0169] It is clear to those skilled in the art that the above-mentioned methods for processing derived variables are merely examples and can be selected arbitrarily. In this application, data, data types, data types, variables or features are sometimes mixed, and those skilled in the art can understand them based on common sense in statistics.

[0170] In this application, derived data can be derived from the original retail credit forecast data through simple processing or complex processing. Simple processed derived data, such as current variables such as monthly salary and monthly balance, can be used directly after data aggregation. Complex processed derived data requires time slicing and logical processing based on the current class variables, and can generate derived variables such as the maximum salary in the past three months and the minimum balance in the past 12 months.

[0171] <Second-generation derivative retail credit forecast data>

[0172] The second-generation derived retail credit prediction data refers to data generated by processing the original retail credit prediction data and the derived retail credit prediction data using a calculation model based on the second-generation derived retail credit prediction data. The second-generation derived retail credit prediction data is selected from the following categories: comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores, and personal mortgage business credit scores.

[0173] The comprehensive business credit score includes a comprehensive business credit total score and three sub-scores; wherein the comprehensive business credit total score and the three sub-scores are calculated based on the retail credit prediction data using the three sub-models of the calculation model; wherein, based on the three sub-models, the three sub-scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score;

[0174] The credit card business credit score includes four credit card business credit scores; wherein the four credit card business credit scores are calculated based on the retail credit prediction data using the calculation model;

[0175] The personal loan business credit score is calculated based on the retail credit prediction data using the calculation model;

[0176] The personal consumer loan business credit score includes three types of personal consumer loan business credit scores; wherein the three types of personal consumer loan business credit scores are calculated based on the retail credit prediction data using the calculation model;

[0177] The personal mortgage business credit score is calculated based on the retail credit prediction data using the calculation model;

[0178] The personal quick loan business credit score is calculated based on the retail credit prediction data using the calculation model.

[0179] The three sub-models of the integrated service class are shown in Formula (1-1) to Formula (1-3) respectively; for details, please refer to Example 1-5 to Example 1-7.

[0180] The four calculation models of the credit card business class are shown in Formula 2, Formula 3, Formula 4 and Formula 5 respectively; for details, please refer to Example 2-5, Example 3-5, Example 4-5 and Example 5-5.

[0181] The calculation models for the personal loan business are shown in Formula 6; for details, please refer to Example 6-5.

[0182] The three calculation models for the personal consumption loan business are shown in Formula 7, Formula 8, and Formula 9, respectively. For details, please refer to Example 7-5, Example 8-5, and Example 9-5.

[0183] The calculation model for the personal mortgage business is shown in Formula 10. For details, please refer to Example 10-5.

[0184] The calculation models for the personal quick loan business are shown in Formula 11. For details, please refer to Example 11-5.

[0185] <Preliminary feature screening>

[0186] In the initial screening data conversion step of the present application, the judgment of the conversion method of the features after the initial screening is based on the concentration and data type of the features after the initial screening. The initial screening data conversion step based on the judgment of concentration and data type includes the following steps: classifying the data type of each feature into character variables and numerical variables, using dummy feature conversion to perform initial screening data conversion on the character variables, and further classifying the numerical variables includes the following sub-steps: if the value of the numerical variable is less than n, using the WOE conversion method to perform initial screening data conversion, if the value of the numerical variable is more than n, further judging if the value of the continuous variable is large and the concentration of the single value is greater than m%, then using the WOE conversion method, if the concentration of the single value is less than or equal to m%, then using the continuous conversion method, preferably, n and m are both positive integers, where n = 5 to 10, m = 90 to 99.

[0187] For example, n is 5, 6, 7, 8, 9 or 10, and m is 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.

[0188] Specifically, in this application, a variety of different feature initial screening methods are used to screen a large number of features, so that the screening can be performed effectively from the largest dimension of features. Existing credit risk scores often use missing rate, concentration and information value IV for feature initial screening.

[0189] In a specific embodiment, features are screened based on the data missingness of each feature of the sample used to build the model. When screening the missing rate, it is generally considered to delete variables with a missing rate greater than 90%, 91%, 92%, 93%, 94%, or 95%. For example, features with a data missing rate exceeding 95% can be eliminated, and features with a data missing rate exceeding 90% can also be eliminated.

[0190] In a specific method, features are screened based on the fact that a single value of a feature sample is too high. When screening for concentration, it is generally considered to delete variables whose single values ​​account for more than 99%, 98%, 97%, 96%, or 95%. For example, features whose single values ​​exceed 99% can be eliminated, and features whose single values ​​exceed 95% can also be eliminated.

[0191] In a specific method, the IV value of each feature is calculated to perform a preliminary screening of the features. During IV screening, the IV value can be used to measure the predictive power of the feature. The larger the IV value, the stronger the predictive power of the feature. The calculation method of the IV value of a single feature is as follows:

[0192]

[0193] Where k is the number of groups after this feature is discretized; y i is the number of non-defaulting customers in group i; y s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s The IV quantitative indicators have the following meanings: when the calculated IV value is less than 0.02, it indicates that the predictive power of this feature is very weak; when the calculated IV value is above 0.02 but less than 0.1, it indicates that the predictive power of this feature is weak; when the calculated IV value is above 0.1 but less than 0.3, it indicates that the predictive power of this feature is good; and when the calculated IV value is above 0.3, it indicates that the predictive power of this feature is strong.

[0194] Of course, the IV calculation value deletion threshold can also be set to 0.03, 0.04, 0.05, etc.

[0195] In the model building method of the present application, the steps of performing preliminary feature screening using missing rate, concentration, and information value IV can be performed in any order. For example, screening can be performed first based on missing rate, then based on concentration, and finally based on information value IV. Screening can also be performed first based on concentration, then based on missing rate, and finally based on information value IV. Screening can also be performed first based on missing rate, then based on information value IV, and finally based on concentration. Screening can also be performed first based on information value IV, then based on missing rate, and finally based on concentration. Screening can also be performed first based on information value IV, then based on concentration, and finally based on missing rate. Screening can also be performed first based on information value IV, then based on concentration, and finally based on missing rate. Screening can also be performed first based on concentration, then based on information value IV, and finally based on missing rate. Those skilled in the art can make their choice based on the sample data. Therefore, based on these three methods, features with obvious defects in certain aspects can be effectively removed, which can effectively reduce the data dimension and improve the effect of model building.

[0196] In one specific method, after the features have been screened using missing rate, concentration and information value IV respectively, a stepwise discriminant algorithm is used to perform preliminary screening of the features, and then the features are screened based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

[0197] On this basis, the technical solution of the present application introduces a step-by-step discriminant method to make the initial screening of variables more efficient and accurate in order to improve the overall efficiency of model development. In actual data, there may be a situation where the distribution of good and bad samples on a certain variable is close, and the ability to distinguish between good and bad is weak. There may also be a class of variables, each of which can distinguish good and bad samples well, but if all are included in the model, they will be redundant due to the duplication of data dimensions covered by the variables. In this regard, the step-by-step discriminant analysis method is used in this application, and the WILKS'S LAMBDA value is used as the statistical criterion for entry and removal, and variables with weak or redundant discriminant effects in the data are deleted.

[0198] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0199] In this application, the step of screening using the stepwise discriminant method is a more critical step. In order to capture as many credit risk points as possible, the model involved in this application uses a huge amount of data and many data dimensions, so variable screening must be performed to further reduce the time cost of model development. In existing financial risk model evaluations, variable importance methods are generally used to reduce variables, such as calculating the importance of variables through algorithms such as the Gini index and information entropy, and selecting variables with high importance. Modeling methods that use stepwise discriminant methods for screening are rarely used. Compared with the variable importance screening schemes commonly used in the industry, the methodology used in this application can retain a large number of variables with relatively weak importance but relatively independent information dimensions.

[0200] In one specific approach, over 100 derived data sets can be generated from, for example, 14 categories of basic data. After three rounds of screening based on missingness rate, concentration, and information value (IV), approximately 20-30% of features with poor data quality can be removed. However, over 2,000 features remain. Subsequent variable refinement based on all of these features would severely impact development efficiency. Therefore, after comparing various variable reduction schemes, stepwise discrimination was ultimately determined to be the optimal approach.

[0201] For the important features that have been screened by the step-by-step discrimination algorithm, further screening of features is performed based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample. Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and eliminate the features whose actual bad debt rate distribution does not conform to the business trend. Specifically, (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of the values ​​of each box and the corresponding bad debt rate. (2) Use the median of the values ​​of each box and the previous box and the corresponding bad debt rate to calculate the rate of change (slope). (3) Count the number of boxes greater than 0 and non-0 boxes (non-greater than 0 boxes) in the rate of change between two adjacent boxes, and calculate the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change. (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend. The criteria for whether they are approximately consistent are as follows: First, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate increases (such as the credit limit utilization rate, etc.). In this module, the features whose number of boxes greater than 0 in the above-calculated slope accounts for less than 70% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated. Second, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate decreases (such as the deposit amount, etc.). In this module, the features whose number of boxes greater than 0 in the above-calculated slope accounts for more than 30% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated.

[0202] For example, the table below provides an example of binning. In this example, the feature values ​​are divided into 10 bins, and the value ranges of the bins are also summarized in the table below. Using this method, we can further filter features based on the risk characteristics of each risk point (preset by the system) and the actual bad debt rate of the sample. Using this binning method to further filter features can effectively select the features that best match business development trends, thereby further obtaining features suitable for modeling.

[0203]

[0204] <Conversion of Eigenvalues>

[0205] The initial feature processing of existing credit risk scoring methods primarily uses word of essence (WOE) conversion (multi-classification) and dummy feature conversion (binary classification) to discretize continuous features (such as age, account age, etc.). The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes continuous features according to the optimal cut point and embeds the data of good and bad samples into the WOE value. Therefore, it performs better during model construction. However, the binning results may overfit the modeling sample, resulting in a serious decline in model effectiveness when applied to the population (poor generalization ability). At the same time, due to the normalization operation used in binning, the original features falling into different bins are converted into a single value corresponding to each bin, thus losing the ability to distinguish the risks of people falling into the same interval. The dummy feature conversion method is mainly applied to grouping features. Its advantage is that it can eliminate the distinction between good and bad values ​​of different feature values, but it becomes extremely complex and redundant when processing continuous variables.

[0206] In the model building method of the present application, the conversion method of the features after preliminary screening is judged to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and the optimal method is used to perform feature conversion for each feature after preliminary screening.

[0207] The initial screening data conversion step includes the following steps based on the judgment of concentration and data type: classifying the data type of each feature into character variables and numerical variables, and performing initial screening data conversion on character variables using dummy feature conversion. The process of further classifying numerical variables includes the following sub-steps: if the value of the numerical variable is less than n, the WOE conversion method is used for initial screening data conversion. If the value of the numerical variable is more than n, further judging if the continuous variable has more values ​​and the concentration of a single value is greater than m%, the WOE conversion method is used. If the concentration of a single value is less than or equal to m%, the continuous conversion method is used. Preferably, n and m are both positive integers, where n = 5 to 10 and m = 90 to 99.

[0208] Specifically, taking the character variable "education level" as an example, the values ​​of this feature variable can be elementary school, middle school, college, graduate school, etc. For a numeric variable, if the numeric variable is the number of overdue months in the past three months, the values ​​are 0, 1, 2, and 3.

[0209] In the present application, the above n can be 5, 6, 7, 8, 9 or 10, and m can be 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.

[0210] In a specific embodiment, m is selected as 5 and n is selected as 95.

[0211] Specifically, the WOE conversion method is to find the optimal cut point of the feature, divide the value range of the original feature into multiple bins, and then calculate the WOE conversion value corresponding to each bin based on the good and bad performance of each bin and output it. The original feature is divided according to the bin results and the WOE conversion value is output. For each bin, the WOE value is calculated as follows:

[0212]

[0213] where y i is the number of non-defaulting customers in group i; y s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s is the total number of defaulting customers.

[0214] Let's use age as an example to explain the WOE binning process. The original age feature contains values ​​ranging from 18 to 50 years old. After binning, we obtain five bins: 18-24, 25-30, 31-35, 36-42, and 43-50. We then calculate the WOE conversion value for each bin based on the number of good and bad customers in each bin. Finally, each user data entry in each bin is output according to the corresponding conversion value. For example, for a 23-year-old, the WOE conversion value corresponding to bins 18 to 24 is output, and for a 46-year-old, the WOE conversion value corresponding to bins 43 to 50 is output.

[0215] The conversion method of dummy features is: convert a single classification feature into an equal number of dummy features according to the number of values ​​it contains. If a customer belongs to the corresponding value of the generated dummy feature, the corresponding dummy feature value is 1, and the remaining dummy feature values ​​are 0.

[0216] Let's use gender as an example to explain how to convert dummy features. The original features include "male" and "female." After dummy feature conversion, two dummy features are generated: "Gender-Male" and "Gender-Female." If the customer's gender is male, "Gender-Male" is recorded as 1, and "Gender-Female" is recorded as 0. For example, if the original features include "junior college or below," "bachelor's degree," and "master's degree or above," three dummy features are generated: "Education-Junior College or Below," "Education-Bachelor's degree," and "Education-Master's degree or above." If the customer's education level is only a bachelor's degree, "Education-Junior College or Below" is recorded as 0, "Education-Bachelor's degree" is recorded as 1, and "Education-Master's degree or above" is recorded as 0.

[0217] The continuous conversion method is to perform various continuous conversions on the original features (continuous conversion methods include but are not limited to: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, and calculating the natural logarithm of the original data). Calculate the correlation coefficient r (Correlation Coefficient) between the feature value after continuous conversion and the overdue label, select the conversion method with the largest absolute value of the correlation coefficient, and output the conversion method corresponding to the original feature. The calculation formula for the correlation coefficient is as follows:

[0218]

[0219] Where Σ is the summation symbol in mathematics; n is the total number of observations; x i is the conversion value of the original feature of the i-th observation after continuous conversion; This is the mean of the transformed values; where y i is a binary feature indicating whether the i-th observation is in default; The average value of this binary feature. The closer the absolute value of the correlation coefficient is to 1, the more closely the converted value correlates with the default situation, and the more effective the conversion method is. The quantitative meaning of the correlation coefficient r is as follows: when the absolute value of the correlation coefficient is greater than 0 and less than 0.3, it indicates a low correlation; when the absolute value of the correlation coefficient is greater than 0.3 and less than 0.8, it indicates a moderate correlation; and when the absolute value of the correlation coefficient is greater than 0.8 and less than 1, it indicates a high correlation.

[0220] In the model construction of the present application, for the features that are confirmed to adopt a continuous conversion method, the optimal conversion method is selected to perform the continuous feature conversion of the feature based on the correlation between the feature and credit default under different continuous conversion methods. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

[0221] Let's use age as an example to explain continuous transformation methods. The original feature "age" contains values ​​ranging from 18 to 50. Continuous transformation yields the original value, square, square root, cube root, and natural logarithm of age. The word of essence (WOE) value and the absolute value of the correlation coefficient between each transformed value and the good / bad label (prediction result) are then calculated. The transformation method with the largest absolute value of the correlation coefficient is selected for conversion and output. If the cube root of the age feature has the largest absolute value of the correlation coefficient with the good / bad label compared to other transformation methods, the cube root of age is output as the transformed value.

[0222] In the existing technical solutions, a feature conversion method between WOE and dummy features is mainly selected for model construction. For example, in the Chinese patent CN112686749B, the WOE method is used to perform feature conversion on the eigenvalues. The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes continuous features according to the optimal cutting point, so it performs better in model construction, but the binning results may overfit the modeling sample, resulting in a serious decline in the model effect when applied to the overall population (poor generalization ability). At the same time, due to the normalization operation used in the binning, the original features falling into different bins are converted into a single value corresponding to each bin, thereby losing the ability to distinguish the risks of people falling into the same interval.

[0223] The dummy feature conversion method is primarily used for grouping features. Its advantage is that it can eliminate differences between good and bad feature values. For example, in the customer's industry, because there is no clear advantage or disadvantage between retail and wholesale, dummy feature processing is more suitable. There is a hierarchical difference between associate and bachelor's degrees, so although dummy features can be used, the WOE processing method is actually more appropriate.

[0224] On the other hand, the continuous conversion method avoids the overfitting of modeling samples that occurs with the WOE conversion method. It has strong generalization capabilities for the overall sample, and because it does not perform interval mapping, it is less likely that most customer groups fall into a single value. However, it cannot be applied to some features with poor monotonicity or discrete characteristics (such as occupation and position).

[0225] As described above, the technical solution of the present application takes a different approach and creatively combines three methods: continuous conversion, WOE conversion, and dumb feature conversion. It reprocesses some carefully screened features, combines the advantages and disadvantages of the three conversion methods, and creatively designs a conversion judgment method. It selects the optimal feature conversion method based on parameters such as feature data attributes, missing rate, concentration, and supplemented by business logic judgment.

[0226] <Logistic regression and deep feature screening based on logistic regression>

[0227] In existing credit risk scoring, due to the requirement of model interpretability, the logistic regression model is mainly used for model development, and the software that can be used are generally: SAS, R, Python, etc.

[0228] In a specific embodiment, the present application develops the model based on SAS software.

[0229] Specifically, the Sigmoid function is used in logistic regression to fit the probability of predicted default. The Sigmoid function is:

[0230]

[0231] Where Z is a linear combination of the model coefficients and the feature transformation values, and Z is defined as follows:

[0232] Z=α+β1x1+β2x2+...+β k-1 x k-1 +β k x k

[0233] The predicted probability of default is:

[0234] P=P(Y=1|x1,x2,x3,…,x k-1 , x k )

[0235] The fitted prediction for the probability of default is:

[0236]

[0237] From the above formula we can further deduce:

[0238] Substituting the Z value into the above formula can calculate the probability P of predicted default.

[0239] The core of logistic regression model construction is feature screening. The steps of feature screening are as follows: First, batch screen the features based on missing rate, concentration and information value IV. Second, screen all remaining features one by one based on whether the features are consistent with business trends, and retain the features with correct business trends. For example: if it is found that the bad debt rate of the customer group decreases with the increase of the feature loan balance, then this feature is considered to be inconsistent with the business trend. In the understanding of credit business, the higher the loan balance, the higher the customer's default risk exposure (EAD) level, and the greater the risk. At this time, this feature will be removed from the feature list. Third, use the stepwise regression function of logistic regression to eliminate features that are less important and highly correlated with other features. Fourth, filter the coefficients based on the positive and negative signs of the training coefficients and the business trend of the feature conversion values, and retain features whose feature coefficient signs are consistent with the business logic. Defining the Y label as 0 for good customers and 1 for bad customers, for a feature whose bad debt rate increases monotonically with the feature value (e.g., loan balance), the training coefficient should be positive. Conversely, for a feature whose bad debt rate decreases monotonically with the feature value (e.g., deposit balance), the training coefficient should be negative. Features that do not meet these criteria should be removed. Fifth, use the variance inflation factor (VIF) and correlation coefficient to further eliminate highly correlated features: For the variance inflation factor, eliminate features with the highest VIF greater than 4 one by one. For highly correlated features, eliminate features with low IV values ​​within the feature group with the highest correlation coefficient greater than 0.80 one by one. Sixth, use the population stability index (PSI) to eliminate features with significant distribution differences at different time points, thereby causing instability. Features with a PSI greater than 0.25 should be eliminated directly. Features with a PSI greater than 0.25 and greater than 0.1 should be carefully removed based on the impact of eliminating these features on the model's discriminatory ability.

[0240] This application strictly adheres to the above rules to screen features, ensuring the model's interpretability, stability, and ability to distinguish between good and bad customers.

[0241] The feature fine-screening steps of the present application include: a first fine-screening step, based on the stepwise regression algorithm, screening the features based on the F test and the T test for the significance of the features; a second fine-screening step, calculating the variance inflation factor for each feature and eliminating features with higher variance inflation factors to screen the features; a third fine-screening step, based on logistic regression, analyzing whether the feature coefficients of the features after the first fine-screening step and the second fine-screening step are consistent with the trend of the prediction results for credit default in order to further screen the features.

[0242] Feature fine-screening step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the previously introduced features are individually tested. If a previously introduced feature becomes less significant due to the introduction of subsequent features, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation.

[0243] Feature fine-screening step 2 is based on the method of eliminating features with high variance inflation factors to further reduce multicollinearity in the model.

[0244] Feature fine screening step 3 is based on the comparison of the risk characteristics of each risk point itself (system preset) with the positive and negative signs of the model training coefficients, and determines whether the feature coefficients of the remaining features in the feature fine screening step 3 in the model are consistent with the business trend, and the features whose model coefficients do not conform to the business trend are eliminated, and the iteration is repeated. The specific implementation plan of the feature fine screening module 3 is as follows: 1. For features whose feature conversion method is WOE type, the corresponding model training coefficient should be negative, and WOE conversion type features with positive training coefficients should be eliminated. 2. For continuous conversion methods, if the bad debt rate should increase as the value of this feature increases in business logic (such as the credit limit utilization rate, etc.), the corresponding model training coefficient should be positive, and such continuous conversion type features with negative training coefficients should be eliminated; if the bad debt rate should decrease as the value of this feature increases in business logic (such as the deposit amount, etc.), the corresponding model training coefficient should be negative, and such continuous conversion type features with positive training coefficients should be eliminated. 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0245] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0246] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model, and the iteration is stopped to obtain the final feature list and its conversion value. After the above steps, the input variables for constructing the model of this application can be obtained.

[0247] The credit default probability modeling step of the present application substitutes the features screened in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.

[0248] In the prior art, the core of all scoring models constructed lies in whether the data they use is representative. For the scoring models that already exist in the prior art, due to information security and cost reasons, the sample size and bad labels of most models are small, and the stability of the model cannot be guaranteed, let alone independent modeling of a certain type of customer. At the same time, because the information dimensions that can be obtained by the scoring models on the market during the modeling process are mostly multiple loan data and weak financial attribute data (such as smart terminal device data, social platform data, online shopping mall data and other non-credit transaction data), they cannot accurately reflect the customer's asset status and repayment ability. The data source used in this application is the full business data of large banks, and the amount of modeling samples and bad labels is very large. This application designs modeling samples based on different sample groups, which can more finely distinguish the risk differences between such customers, and at the same time ensure the stability of data and models based on cross-time verification and PSI verification. This application uses personal credit history data and asset data for model development, so it has a good reflection of the borrower's repayment willingness and repayment ability.

[0249] For some small and medium-sized financial institutions, the internally built scoring models are relatively small or started relatively late in their retail credit business, and the accumulated historical data is insufficient to develop a credit risk model with stable data and strong differentiation capabilities. Therefore, they rely heavily on manual approval for credit review. The efficiency limitations of manual approval have restricted the development of their retail credit business. At the same time, the subjectivity of manual approval has increased the operational risk in the credit review process. The model constructed in this application can assist such financial institutions in making digital decisions, enhance their approval accuracy and speed, and reduce the above-mentioned adverse effects.

[0250] In terms of feature conversion, compared to the traditional WOE conversion method, which requires coarse binning and discretization of continuous features based on data performance and the modeler's experience, the results of coarse binning are significantly affected by the modeler's subjective factors. Furthermore, the discretization process of continuous features may result in a large number of single-valued scores due to too many customers falling into the same interval. This application combines the continuous conversion method, the WOE conversion method, and the dummy feature conversion method to encode the original features, reducing the impact of human factors and single values ​​while enhancing the discriminative ability of the scores.

[0251] <Method for Calculating Retail Credit Risk> Please refer to the descriptions of Examples 12-1 to 12-6.

[0252] <Device, system, and computer storage medium for calculating retail credit risk>

[0253] The present application relates to an apparatus for constructing a retail credit risk prediction model, the apparatus comprising: a data acquisition module for acquiring original retail credit prediction data of a sample for constructing a model; a data derivation module for processing derived retail credit prediction data based on the original retail credit prediction data; a data processing module for processing second-generation derived retail credit prediction data based on the original retail credit prediction data and the derived retail credit prediction data; a feature initial screening module for performing preliminary screening of all categories, i.e., all features, including the original retail credit prediction data and the derived retail credit prediction data, to obtain features after preliminary screening; a preliminary screening data conversion module for determining a conversion method for the features after preliminary screening to confirm whether to use one of a WOE conversion method, a dummy feature conversion method, and a continuous conversion method for feature conversion, and for each feature after preliminary screening, using the determined optimal method for feature conversion; a feature fine screening module for performing in-depth screening of the features after preliminary screening to obtain fine-screened features; and a credit default probability modeling module for selecting a logistic regression method for model construction based on the relationship between the fine-screened features and the probability of credit default, and for determining a method for calculating the credit default probability.

[0254] The present application relates to a system for constructing a retail credit risk prediction model, which includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor. When the program for constructing a retail credit risk prediction model method is executed by the processor, the steps of constructing a retail credit risk prediction model method as described above are implemented.

[0255] All the contents described above for the method for calculating retail credit risk are fully applicable to the device, system and computer storage medium for calculating retail credit risk.

[0256] The method for calculating retail credit risk in this application can avoid the technical shortcomings of overfitting when building a model with complete WOE variables and the inability to adapt well to categorical variables when building a model with complete continuous variables. Therefore, when this model is used for retail credit risk calculation, it can better cover the retail credit risk prediction needs of various types of customer groups and meet the demand for more accurate prediction results when predicting credit risk as a general score.

[0257] Example

[0258] 1) Comprehensive business credit score (Sigma_a model credit score)

[0259] Example 1-1 Collection of computational modeling samples

[0260] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample and selected 620 million data points from 2018 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 120 million and 500 million, respectively.

[0261] When designing the specific model, analysis and design will be conducted on the 120 million customers who have already applied. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0262] Model design included 1) exclusion rules: excluding data from customers with settled accounts, closed accounts, no performance, or special circumstances. The modeled population was 65.74 million customers. 2) time window setting: 2018 data was used as the modeling sample, with the next 24 months as the performance period for Y. 3) sample selection: a 4:1 ratio of good to bad samples was used for modeling. Different segmentation schemes were designed using a decision tree approach, and the final model was confirmed by comparing parent-child models. Each sub-model was then modeled separately.

[0263] In this embodiment, in order to build a model for the sample group that has not applied for credit business registration, the first layer of the decision tree is the account age, and only customers with an account age of less than 3 months are selected to enter the model. For customers with an account age of less than 3 months, the area is divided according to the administrative district to which they belong. The division of administrative areas can be based on data released by authoritative statistical departments, or based on the division conducted by a rating agency, or based on another constructed financial model. It is fully understood by those skilled in the art that after selecting the division criteria, the user population can be divided into three different areas without duplication. At the same time, it is further combined with the three different situations of low, medium and high bad debt rates to divide them into three sub-models, Sigma10, Sigma11, and Sigma12.

[0264] Specifically, the sample size used to construct the Sigma 10 sub-model is about 220,000. Its customer base is mainly potential customers who have not applied for credit business and are from economically underdeveloped areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0265] Specifically, the sample size used to construct the Sigma 11 sub-model is approximately 230,000. Its customer base is mainly potential customers who have not applied for credit business and are from moderately developed economic areas. The model is used to predict the probability of credit delinquency of more than 90 days in this group of people.

[0266] Specifically, the sample size used to construct the Sigma 12 sub-model is approximately 460,000. Its customer base is mainly potential customers who have not applied for credit business and are from economically developed areas. The model is used to predict the probability of credit overdue for more than 90 days in this group of people.

[0267] In this embodiment, whether a customer is a new customer refers to a customer with an account age of less than 3 years, that is, the time since the customer applied for credit business is less than 3 months. A serious overdue payment means that the customer has been overdue for more than 3 periods (i.e., the number of months). A moderate or light overdue payment means that the customer is currently overdue, and the number of overdue periods (i.e., the number of months) is less than or equal to 3 periods.

[0268] Based on the different customer samples of each sub-model confirmed above, obtain the borrower's historical data information: personal financial assets, including AUM, deposits, wealth management and payroll information.

[0269] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 141 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount,' 'Average Number of Overdue Payments in the Last 6 Months,' and 'Number of Credit Card Installments in the Last 12 Months.' In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0270] Table 1-1 Summary of basic variables and derived variables used in this example

[0271]

[0272] Example 1-2 Characteristic Screening

[0273] A preliminary screening was conducted on the 141 features (variables) collected in Example 1-1 that have potential predictive power for customer overdue payments.

[0274] The preliminary screening of the three sub-model features, Sigma 10 to Sigma 12, divided in Example 1-1, was performed as follows:

[0275] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 1-1 were first eliminated, resulting in a total of 13 variables being deleted, leaving 128 variables.

[0276] In the second round of preliminary screening, for the 128 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 23 variables were eliminated, leaving 105 variables.

[0277] In the third round of preliminary screening, the 105 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box, and if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 16 features were eliminated, and 89 feature variables remained.

[0278] In the fourth round of preliminary screening, the 89 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 65 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, this embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0279] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0280] The fifth round of preliminary screening targets the 65 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0281] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0282] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0283] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0284] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0285] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0286] The criteria for approximate consistency are as follows:

[0287] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0288] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0289] After the fifth round of screening, 25 features were eliminated, leaving 40 features.

[0290] Example 1-3 Conversion of features after initial screening

[0291] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0292] First, the remaining 40 features in Example 1-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0293] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0294] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0295] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0296] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0297] Through this embodiment, 40 features are converted, of which 32 features are converted into WOE, 8 features are converted into dummy features (expanded into 32 variables), and 64 features are converted into continuous features.

[0298] Example 1-4 Feature Depth Screening (Feature Fine Screening Step)

[0299] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0300] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the 64 features to 43.

[0301] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 43 features to 32.

[0302] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0303] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0304] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0305] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0306] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0307] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0308] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The iterations are stopped, resulting in the final feature list and its conversion values. After these steps, 32 features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0309] Example 1 - Construction of 5Sigma 10 Model

[0310] Taking Sigma 10 as a sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 220,000. Its customer base was mainly potential customers who had not applied for credit business and were from economically underdeveloped areas. The model was constructed to predict the probability (target variable) of this group of people being overdue for credit for more than 90 days.

[0311] In the present embodiment 1-5, it is finally confirmed that the final modeling result is described by taking the five features of the current deposit account balance, the maximum balance of the investment and financial management account in the past three months, the number of consecutive months of asset size reduction in the past three months, the percentage of the average monthly salary in the past three months to the average monthly asset size, and the quantile number of the region where the current salary is paid as an example. In the feature conversion step, the five model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0312] The current deposit account balance reflects the borrower's asset level. For new applicants from less developed regions, the higher the asset level, the lower the likelihood of a serious default; conversely, the lower the asset level, the higher the likelihood of a serious default. The conversion method is a logarithmic transformation within the continuous transformation.

[0313] The maximum balance in the investment and wealth management account over the past three months reflects the borrower's investment habits. For new applicants in economically underdeveloped regions, good investment habits guarantee repayment ability. The lower the maximum balance in the investment and wealth management account over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the maximum balance in the investment and wealth management account over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0314] The number of consecutive months of asset decline over the past three months reflects the borrower's asset fluctuations. For new applicants from less developed regions, a continuous decline in asset levels indicates a persistent deficit. The greater the number of consecutive months of asset decline over the past three months, the more severe the borrower's deficit and the higher the likelihood of a major default. Conversely, the fewer consecutive months of asset decline over the past three months, the more stable the borrower's financial situation and the lower the likelihood of a major default. The conversion method is WOE conversion.

[0315] The percentage of the average monthly salary over the past three months to the average monthly asset size reflects the borrower's savings status. For new applicants in less developed regions, long-term savings habits determine their long-term repayment ability. The lower the percentage of the average monthly salary over the past three months to the average monthly asset size, the smaller the proportion of the borrower's monthly salary to their total assets, and therefore, the better their long-term savings habits and the lower the likelihood of a serious default. Conversely, the higher the percentage of the average monthly salary over the past three months to the average monthly asset size, the greater the proportion of the average monthly salary to their total assets, and therefore, the worse their long-term savings habits and the higher the likelihood of a serious default. The conversion method is WOE conversion.

[0316] The regional quantile for current payroll payment reflects the borrower's income level. For new applicants from less developed regions, the lower the regional quantile for current payroll payment, the lower the borrower's income level and the higher the likelihood of a serious default. Conversely, the higher the regional quantile for current payroll payment, the higher the borrower's income level and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0317] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0318]

[0319] Where k is the number of features entering the model. In formula 1-1, k is 5.

[0320] α is the intercept term, with a range of values ​​from (-1.65504, -1.85496) and an optimal value of -1.755; β1 is the coefficient corresponding to the current deposit account balance, with a range of values ​​from (-0.12828, -0.15572) and an optimal value of -0.142; β2 is the coefficient corresponding to the maximum balance of the investment and wealth management account in the past three months, with a range of values ​​from (-0.52164, -0.78036) and an optimal value of -0.651; β3 is the coefficient corresponding to the asset size in the past three months. The coefficient corresponding to the number of consecutive months of reduction ranges from (-1.25768 to -2.20632), with an optimal value of -1.732. β4 is the coefficient corresponding to the percentage of average monthly wages to average monthly assets over the past three months, with a range from (-0.20416 to -0.41584), and an optimal value of -0.31. β5 is the coefficient corresponding to the quantile of the region where the current payroll is distributed, with a range from (-0.10288 to -0.48312), and an optimal value of -0.293. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0321] x1 is the continuous logarithmic transformation value of the deposit account balance at the current time point generated by the feature conversion step; x2 is the WOE conversion value of the maximum balance of the investment and wealth management account in the past three months generated by the feature conversion step; x3 is the WOE conversion value of the number of consecutive months of asset size reduction in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the percentage of average monthly salary to average monthly asset size in the past three months generated by the feature conversion step; x5 is the WOE conversion value of the percentile of the current payroll area generated by the feature conversion step.

[0322] The model performance of some features is shown in Table 1-1 below:

[0323] Table 1-1

[0324]

[0325] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0326] Example 1-6 Construction of Sigma 11 Model

[0327] Taking Sigma 11 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 230,000. Its customer base was mainly potential customers who had not applied for credit business and were from moderately developed economic areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0328] In Example 1-6, it is finally confirmed that the final modeling results are described using six features: the current deposit account balance, the average monthly asset size in the past three months, the average balance of the investment and financial management account in the past three months, the longest continuous reduction in the number of months of the asset size value in the past three months, the maximum salary in the past 12 months, and the number of months from the maximum balance of the deposit account in the past six months to the observation point. In the feature conversion step, the six model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0329] The current deposit account balance reflects the borrower's asset level. For new applicants in moderately developed regions, the higher the asset level, the lower the likelihood of a serious default; conversely, the lower the asset level, the higher the likelihood of a serious default. The conversion method is the natural logarithm transformation of the continuous transformation.

[0330] The average monthly asset size over the past three months reflects the borrower's asset level. For new applicants in moderately developed regions, the absolute value of their asset level is a better indicator of asset quality. The higher the average monthly asset size over the past three months, the more cash-rich the borrower is and the lower the likelihood of a serious default. Conversely, the lower the average monthly asset size over the past three months, the more cash-strapped the borrower is and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0331] The average balance of investment and wealth management accounts over the past three months reflects the borrower's investment habits. For new applicants in moderately developed regions, good investment habits guarantee their repayment ability. The lower the average balance of investment and wealth management accounts over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the average balance of investment and wealth management accounts over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0332] The longest consecutive number of months of asset size decline over the past three months reflects the borrower's asset fluctuations. For new applicants in moderately developed regions, a continuous decline in their asset levels indicates a persistent deficit. The longer the number of consecutive months of asset size decline over the past three months, the more severe the borrower's deficit and the higher the likelihood of a major default. Conversely, the shorter the number of consecutive months of asset size decline over the past three months, the more stable the borrower's financial situation and the lower the likelihood of a major default. The conversion method is WOE conversion.

[0333] The maximum wage over the past 12 months reflects the borrower's income level. For new applicants in moderately developed regions, this figure reflects their overall income level. The lower the maximum wage over the past 12 months, the lower the borrower's overall income level and the higher the likelihood of a major default. Conversely, the higher the maximum wage over the past 12 months, the higher the borrower's overall income level and the lower the likelihood of a major default. The conversion method is WOE.

[0334] The number of months between the maximum deposit account balance at a point in time over the past six months and the observation point reflects the borrower's deposit changes. The closer the maximum deposit account balance at a point in time over the past six months is to the observation point, the greater the likelihood that the borrower's assets will gradually increase and the lower the probability of a serious default. Conversely, the closer or further the maximum deposit account balance is to the observation point, the greater the likelihood that the borrower's assets will gradually decrease and the higher the probability of a serious default. The conversion method is WOE conversion.

[0335] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0336]

[0337] Where k is the number of features entering the model. In formula 1-2, k is 6.

[0338] α is the intercept term, with a range of values ​​from (-0.50444, -0.84156), and an optimal value of -0.673; β1 is the coefficient corresponding to the current deposit account balance, with a range of values ​​from (-0.13036, -0.16564), and an optimal value of -0.148; β2 is the coefficient corresponding to the average monthly asset size in the past three months, with a range of values ​​from (-0.19664, -0.25936), and an optimal value of -0.228; β3 is the coefficient corresponding to the average balance of the investment and wealth management account in the past three months, with a range of values ​​from (-0.26772, -0.632 28), with an optimal value of -0.45; β4 is the coefficient corresponding to the longest consecutive decrease in asset size over the past three months, with a range of (-0.78496, -1.17304), and an optimal value of -0.979; β5 is the coefficient corresponding to the maximum salary over the past 12 months, with a range of (-0.35304, -0.74896), and an optimal value of -0.551; β6 is the coefficient corresponding to the number of months from the maximum deposit account balance over the past six months to the observation point, with a range of (-0.29792, -0.78008), and an optimal value of -0.539. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0339] x1 is the natural logarithm conversion value of the deposit account balance at the current time point generated by the feature conversion step; x2 is the natural logarithm conversion value of the average monthly asset size in the past three months generated by the feature conversion step; x3 is the WOE conversion value of the average balance of the investment and wealth management account in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the longest consecutive number of months of decrease in asset size value in the past three months generated by the feature conversion step; x5 is the WOE conversion value of the maximum salary in the past 12 months generated by the feature conversion step; x6 is the WOE conversion value of the number of months from the observation point to the maximum balance of the deposit account in the past six months generated by the feature conversion step.

[0340] The model performance of some features is shown in Table 1-2 below:

[0341] Table 1-2

[0342]

[0343] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0344] Example 1-7 Construction of Sigma 12 Model

[0345] Taking Sigma 12 as the sub-model, the above-mentioned decision tree method was adopted, and the sample size obtained by classification was about 460,000. Its customer base was mainly potential customers who had not applied for credit business and were from economically developed areas. The model was constructed to predict the probability of credit overdue for more than 90 days in this group (target variable).

[0346] In Example 1-7, it is finally confirmed that the final modeling results are described by taking the maximum asset size in the past 3 months, the minimum balance of the deposit account in the past 3 months, the maximum balance of the investment and financial management account in the past 3 months, the unit quantile of the current salary payment unit, and the number of months from the maximum balance of the deposit account in the past 12 months to the observation point as an example. In the feature conversion step conversion mode selection, the 5 model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the 5 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0347] The maximum asset size over the past three months reflects the borrower's asset level. For new applicants in economically developed regions, asset levels fluctuate significantly. Therefore, the short-term maximum asset value better reflects asset quality. Higher asset levels indicate a lower likelihood of a major default; conversely, lower asset levels indicate a higher likelihood of a major default. The conversion method is the cube root conversion in the continuous conversion method.

[0348] The minimum deposit balance over the past three months reflects the borrower's deposit status. For new applicants in economically developed regions, the minimum deposit is a better indicator of asset quality. The higher the minimum deposit balance over the past three months, the more abundant the borrower's funds are and the lower the likelihood of a serious default. Conversely, the lower the minimum deposit balance over the past three months, the less cash the borrower has and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation within the continuous transformation model.

[0349] The maximum balance in the investment and wealth management account over the past three months reflects the borrower's investment habits. For new applicants in economically developed regions, good investment habits guarantee their repayment ability. The lower the maximum balance in the investment and wealth management account over the past three months, the worse the borrower's investment habits and the higher the likelihood of a serious default. Conversely, the higher the maximum balance in the investment and wealth management account over the past three months, the better the borrower's investment habits and the lower the likelihood of a serious default. The conversion method is WOE.

[0350] The current payroll unit's percentile reflects the borrower's income capacity and job stability. For new applicants in economically developed regions, the higher the unit's percentile, the greater the borrower's income capacity and stability, and the lower the likelihood of a serious default. Conversely, the lower the unit's percentile, the weaker the borrower's income capacity and stability, and the higher the likelihood of a serious default. The conversion method is WOE.

[0351] The number of months between the maximum deposit balance in the past 12 months and the observation point reflects changes in the borrower's deposits. For new applicants in economically developed regions, the closer the maximum deposit balance in the past 12 months was to the observation point, the more recent the borrower's deposits are and the lower the likelihood of a serious default. Conversely, the further the maximum deposit balance in the past 12 months was from the observation point, the lower the borrower's deposits are and the higher the likelihood of a serious default. The conversion method is WOE.

[0352] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0353]

[0354] Where k is the number of features entering the model. In formula 1-3, k is 5.

[0355] α is the intercept term, with a range of values ​​from (-2.18528, -2.31072) and an optimal value of -2.248; β1 is the coefficient corresponding to the maximum asset size in the past three months, with a range of values ​​from (-0.02008, -0.02792) and an optimal value of -0.024; β2 is the coefficient corresponding to the minimum balance of the deposit account in the past three months, with a range of values ​​from (-0.08928, -0.11672) and an optimal value of -0.103; β3 is the investment and financial management in the past three months. The coefficient corresponding to the maximum account balance ranges from (-0.45036 to -0.68164), with an optimal value of -0.566. β4 is the coefficient corresponding to the unit quantile of the current payroll payment agency, with a range of (-0.11148 to -0.64852), and an optimal value of -0.38. β5 is the coefficient corresponding to the number of months from the observation point to the maximum deposit account balance in the past 12 months, with a range of (-0.35424 to -0.57376), and an optimal value of -0.464. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0356] x1 is the cube root conversion value of the maximum asset size in the past three months generated by the feature conversion step in a continuous manner; x2 is the natural logarithm conversion value of the minimum balance of the deposit account in the past three months generated by the feature conversion step in a continuous manner; x3 is the WOE conversion value of the maximum balance of the investment and wealth management account in the past three months generated by the feature conversion step; x4 is the WOE conversion value of the unit percentile of the current payroll payment generated by the feature conversion step; x5 is the WOE conversion value of the number of months from the observation point to the maximum balance of the deposit account in the past 12 months generated by the feature conversion step.

[0357] The model performance of some features is shown in Table 1-3 below:

[0358] Table 1-3

[0359]

[0360] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0361] Examples 1-8

[0362] The P value calculated by the above formula can be further used to calculate the score of any customer.

[0363] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0364] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0365]

[0366]

[0367] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0368] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0369] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0370] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0371] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0372] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Sigma_a model is 75.25. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0373] The following table shows the KS values ​​for the development and validation samples when the model was applied to the development and validation sets. Tables 1-4 show that the model achieved good discrimination in both the training and validation sets (where the training and validation sets had a 6:4 sample ratio), as well as across the entire sample set. This indicates that all sub-models achieved excellent discrimination.

[0374] Table 1-4

[0375] Sample Set Development samples (training set) Validation sample (validation set) Development sample + verification sample Distinguishing indicators KS KS KS Sigma_a 38.2 37.9 38.0 Sigma10 29.17 26.86 28.11 Sigma11 35.46 36.71 35.81 Sigma12 36.65 35.09 36.02

[0376] 2) Credit card business credit score (Credit score of Alpha_5 model)

[0377] Example 2-1 Collection of modeling samples

[0378] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample and selected 630 million data points from 2019 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 130 million and 500 million, respectively.

[0379] The specific Alpha model is a scenario scoring for credit card business. When designing the model, 130 million customers who have applied for credit business are analyzed and designed. More than 300,000 customers who have applied for credit business and have an account age of less than 3 are screened out from the 130 million customers to build a model to predict the probability of credit overdue for more than 60 days for people who have not applied for credit business. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0380] The model design includes: 1) Exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.5 million customers (this part of the population is the customer group that has applied for credit card business); 2) Time window setting: Determine the sample to use 2019 data as the modeling sample, and use the next 15 months as the performance period of Y; 3) Sample sampling: Use a good / bad sample number of 5:1 for sampling modeling.

[0381] Based on the above-mentioned sample of more than 300,000 customers who have applied for credit business and whose account age is less than 3 years, historical data information of borrowers is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0382] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods, including the following: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of their account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months, the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods in the last X months and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including "current repayment amount," "average number of overdue payments in the past six months," and "number of credit card installments in the past 12 months." In this application, "observation point" refers to the point in time at which samples were collected up to the time of modeling. "Current" also refers to the sampling cutoff time. These two terms have the same meaning.

[0383] Table 2-1 Summary of basic variables and derived variables used in this example

[0384]

[0385] Example 2-2 Feature Screening

[0386] A preliminary screening was conducted on the 157 features (variables) collected in Example 2-1 that have potential predictive power for customer overdue payments.

[0387] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 2-1 were first eliminated, resulting in a total of 37 variables being deleted, leaving 120 variables.

[0388] In the second round of preliminary screening, for the 120 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 11 variables were eliminated, leaving 109 variables.

[0389] In the third round of preliminary screening, the 109 features after the second round of preliminary screening were sorted according to the feature values ​​(the specific sorting method is: if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 18 features were eliminated, and 91 feature variables remain.

[0390] In the fourth round of preliminary screening, the 91 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 74 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0391] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0392] The fifth round of preliminary screening targets the 74 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0393] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0394] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0395] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0396] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0397] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0398] The criteria for approximate consistency are as follows:

[0399] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0400] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0401] After the fifth round of screening, 24 features were eliminated, leaving 50 features.

[0402] Example 2-3 Conversion of features after initial screening

[0403] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the feature initial screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numerical variables. For character variables, the dummy feature (dummy variable) conversion method is generally used for variable conversion; for numerical variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. The features using different conversion methods are grouped and divided into different data sets. Specifically,

[0404] First, the remaining 50 features in Example 2-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0405] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0406] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0407] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0408] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0409] Through this embodiment, 50 features are converted, of which 23 features are converted into WOE, 17 features are converted into dummy features (expanded into 51 variables), and 13 features are converted into continuous features.

[0410] Example 2-4 Feature Depth Screening (Feature Fine Screening Step)

[0411] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0412] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the total number of features from 87 to 47.

[0413] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 47 features to 23.

[0414] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0415] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0416] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0417] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0418] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0419] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0420] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 23 features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0421] Example 2-5 Construction of Alpha_5 model

[0422] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 300,000, and the customer base is mainly over 300,000 customers who have applied for credit services and have an account age of less than 3 years. The model is used to predict the probability of a credit delinquency of more than 60 days (the target variable) for people who have not applied for credit services.

[0423] In the present embodiment 2-5, it is finally confirmed that the final modeling result is described by taking the four features of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months, the minimum balance of the deposit account in the past 3 months, the number of months from the minimum balance of the financial management account in the past 3 months, and the number of salary payments in the past 12 months as an example. In the feature conversion step conversion mode selection, the four model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the four features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0424] The average monthly assets over the past six months divided by the average monthly assets over the past twelve months reflects changes in the borrower's asset level. For recently registered borrowers, an increase in their asset level indicates a continued increase in their asset size. A larger value (greater than 1) indicates a stronger and more stable borrower's financial situation, and a lower likelihood of default. A smaller value (less than 1) indicates a weaker and more unstable borrower's financial situation, and a higher likelihood of default. The conversion method is raw value conversion.

[0425] The minimum deposit balance over the past three months reflects the borrower's recent deposit status. For newly registered borrowers, the minimum deposit is a better indicator of asset quality. The higher the minimum deposit balance over the past three months, the more abundant the borrower's funds are and the lower the likelihood of a major default. Conversely, the lower the minimum deposit balance over the past three months, the less cash the borrower has and the higher the likelihood of a major default. Natural logarithm transformation is used.

[0426] The number of months from the minimum balance in the wealth management account over the past three months reflects recent changes in the borrower's wealth management assets. For newly registered clients, the higher the number of months from the minimum balance in the wealth management account over the past three months and the higher the recent wealth management assets, the lower the likelihood of a serious default; conversely, the higher the likelihood of a serious default. This conversion method uses dummy variable conversion.

[0427] The number of payroll payments made by agents in the past 12 months reflects the borrower's job stability. For recent new applicants, the lower the number of payroll payments made by agents in the past 12 months, the lower the borrower's stability and the higher the likelihood of a serious default. Conversely, the higher the number of payroll payments made by agents in the past 12 months, the greater the borrower's stability and the lower the likelihood of a serious default. The transformation method is natural logarithm.

[0428] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0429]

[0430] Where k is the number of features entering the model, and in Formula 2, k is 4.

[0431] α is the intercept term, with a range of values ​​from (-2.44505 to -2.74473), and an optimal value of -2.59489. β1 is the coefficient corresponding to the average monthly asset size in the past six months / the average monthly asset size in the past 12 months, with a range of values ​​from (-0.25311 to -0.41086), and an optimal value of -0.33199. β2 is the coefficient corresponding to the minimum balance of the deposit account in the past three months, with a range of values ​​from (-0.07435 to -0.46009), and an optimal value of -0.26722. β3 is the coefficient corresponding to the number of months from the minimum balance of the wealth management account in the past three months, with a range of values ​​from (-0.14379 to -0.14841), and an optimal value of -0.1461. β4 is the coefficient corresponding to the number of payroll payments in the past 12 months, with a range of values ​​from (-0.8194 to -1.10539), and an optimal value of -0.9624.

[0432] x1 is the original value of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months generated by the feature conversion step; x2 is the natural logarithm conversion value of the minimum balance of the deposit account in the past 3 months generated by the feature conversion step; x3 is the dummy variable conversion value of the number of months from the minimum balance of the wealth management account in the past 3 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the number of payroll payments in the past 12 months generated by the feature conversion step.

[0433] The model performance of some features is shown in Table 2-2 below.

[0434] Table 2-2

[0435]

[0436] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0437] Examples 2-6

[0438] The P value calculated by the above formula can be further used to calculate the score of any customer.

[0439] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0440] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0441]

[0442]

[0443] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0444] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0445] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0446] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0447] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0448] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Alpha model is 74. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0449] The following table shows the KS values ​​for the development and validation samples when the model was applied to the development and validation sets. Table 2-3 shows that the model achieved good discrimination in both the training and validation sets (where the training and validation sets had a sample size ratio of 6:4), as well as in the overall sample set. This indicates that all computational models achieved excellent discrimination.

[0450] Table 2-3

[0451]

[0452] 3) Credit card business credit score (Credit score of Alpha2_5 model)

[0453] Example 3-1 Collection of modeling samples

[0454] During the construction of this example, personal financial asset transaction data and personal customer information data for all retail customers of a large bank between 2017 and 2021 were collected, totaling 730 million individuals. A professional model design solution was used to confirm the modeling sample, and data from 620 million individuals in 2018 and 650 million individuals in 2019 were selected as the analysis samples. Because modeling requires performance variables and the performance period for credit cards and special installment loans is relatively short, 2019 was selected as the observation period. The 650 million analysis sample was then divided into two parts: those with registered credit applications and those without registered credit applications, with sample sizes of 130 million and 520 million, respectively.

[0455] The specific Alpha2 model is a scenario scoring for credit cards and special installment services. When designing the model, 110 million customers who have applied for credit cards and special installment services among the 130 million customers who have applied were analyzed and designed. From the 130 million customers, 430,000 customers who have applied for credit services and have an account age of less than 3 months were screened out to build a model to predict the probability of credit overdue for more than 60 days for people who have not applied for registered credit services. After the model is developed and put online, the results will be applied to the total number of 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0456] The model design includes 1) exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.59 million customers; 2) time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15-month time range as the performance period of Y; 3) sample sampling: use a good / bad sample number of 10:1 for sampling modeling.

[0457] Based on the above-mentioned sample of 430,000 customers who have applied for credit business and whose account age is less than 3 months, the borrowers' historical data information is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0458] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount', 'Average Number of Overdue Payments in the Last 6 Months', and 'Number of Credit Card Installments in the Last 12 Months'. In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0459] Table 3-1 Summary of basic variables and derived variables used in this example

[0460]

[0461] Example 3-2 Feature Screening

[0462] A preliminary screening was conducted on the 157 features (variables) collected in Example 3-1 that have potential predictive power for customer overdue payments.

[0463] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 3-1 were first eliminated, resulting in a total of 27 variables being deleted, leaving 130 variables.

[0464] In the second round of preliminary screening, for the 130 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 35 variables were eliminated, leaving 95 variables.

[0465] In the third round of preliminary screening, the 95 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 23 features were eliminated, and 72 feature variables remained.

[0466] In the fourth round of preliminary screening, the 72 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 43 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0467] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0468] The fifth round of preliminary screening targets the 43 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0469] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0470] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0471] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0472] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0473] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0474] The criteria for approximate consistency are as follows:

[0475] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0476] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0477] After the fifth round of screening, 43 features were eliminated and 24 features remained.

[0478] Example 3-3 Conversion of features after initial screening

[0479] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0480] First, the remaining 24 features in Example 3-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0481] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0482] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0483] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0484] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0485] Through this embodiment, 24 features are converted, of which 14 features are converted into WOE, 5 features are converted into dummy features (expanded into 20 variables), and 5 features are converted into continuous features, for a total of 39 features.

[0486] Example 3-4 Feature Depth Screening (Feature Fine Screening Step)

[0487] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0488] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. After this step, 39 features were reduced to 73.

[0489] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 39 features to 21.

[0490] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0491] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0492] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0493] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0494] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0495] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0496] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 21 features are eliminated as the final feature variables to be entered into the model, for example, 5 features, 6 features, 7 features, or 8 features.

[0497] Example 3-5 Construction of Alpha2_5 model

[0498] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 430,000, and the customer base is primarily new customers who have registered for credit services but whose account age is less than three months. The model is constructed to predict the probability of a credit delinquency of more than 60 days in the future for the sample group that has not registered for credit services.

[0499] In this embodiment 3-5, it is finally confirmed that the final modeling result is described by taking six features as an example: the number of consecutive months with the largest month-on-month decrease in the AUM value over the past three months, the current point-in-time deposit account balance, the average monthly AUM (personal financial assets) over the past six months, the number of months from the maximum point-in-time balance of the deposit account over the past 12 months, the minimum point-in-time balance of the deposit account over the past six months, and the maximum balance of the investment and financial management account over the past 12 months. In the feature conversion step, the six modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the six features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0500] The number of consecutive months with the largest month-on-month decrease in AUM over the past three months reflects the borrower's asset level and repayment ability. For new users who haven't registered for credit services, the AUM value reflects the borrower's recent changes in financial resources. Insufficient financial resources increase the likelihood of future defaults. This variable correlates with customer risk: the lower the number of consecutive months with the largest month-on-month decrease in AUM over the past three months, the stronger the borrower's asset level and repayment ability, and the lower the likelihood of a major default. Conversely, the higher the number of consecutive months with the largest month-on-month decrease in AUM over the past three months, the weaker the borrower's asset level and repayment ability, and the higher the likelihood of a major default. The conversion method is Word of Equity (WOE) conversion.

[0501] The current deposit account balance reflects the borrower's savings level and repayment ability. For new users who haven't registered for credit services, the greater the pressure on the borrower to repay their debt, the higher the likelihood of defaults. This variable correlates with customer risk: the larger the current deposit account balance, the better the borrower's savings level and repayment ability, and the lower the likelihood of a serious default. Conversely, the smaller the current deposit account balance, the lower the borrower's savings level and repayment ability, and the higher the likelihood of a serious default. Logarithmic transformation is used.

[0502] The average monthly AUM (personal financial assets) over the past six months reflects the borrower's repayment ability. For new users who haven't registered for credit services, the short-term average monthly AUM reflects the borrower's immediate financial capacity. A higher average monthly AUM indicates greater financial capacity and a lower likelihood of future defaults. This variable correlates with customer risk: the higher the average monthly AUM (personal financial assets) over the past six months, the greater the borrower's repayment capacity and the lower the likelihood of a serious default. Conversely, the lower the average monthly AUM (personal financial assets) over the past six months, the weaker the borrower's repayment capacity and the higher the likelihood of a serious default. Logarithmic transformation is used.

[0503] The number of months from the current month to the maximum deposit account balance over the past 12 months reflects changes in the borrower's financial resources. For new users who haven't registered for credit services, the smaller the maximum deposit balance over the past period, the more adequate the borrower's current financial resources are and the lower the likelihood of future loan defaults. This variable correlates with customer risk: the further the maximum deposit account balance over the past 12 months has been from the current month, the lower the borrower's recent income and repayment ability, and the higher the likelihood of a serious default. Conversely, the closer the maximum deposit account balance over the past 12 months has been from the current month to the current month, the higher the borrower's recent income and repayment ability, and the lower the likelihood of a serious default. The conversion method is raw value conversion.

[0504] The minimum balance in a deposit account over the past six months reflects a borrower's financial resources and repayment ability. For new users who haven't registered for credit services, the larger the minimum balance over the past period, the more secure the borrower's financial resources and the lower the likelihood of future loan defaults. This variable correlates with customer risk: the larger the minimum balance in a deposit account over the past six months, the more secure the borrower's financial resources and the lower the likelihood of a serious default. Conversely, the larger the minimum balance in a deposit account over the past six months, the weaker the borrower's financial resources and repayment ability, and the higher the likelihood of a serious default. The conversion method uses the square transformation within the continuous transformation.

[0505] The maximum balance in the investment and wealth management account over the past 12 months reflects the borrower's investment level and repayment ability. For new users who haven't registered for credit services, this maximum balance in the investment and wealth management account over the past 12 months reflects the adequacy of the borrower's secondary repayment source. A larger balance indicates stronger repayment ability and a lower likelihood of future defaults. This variable correlates with customer risk: the smaller the maximum balance in the investment and wealth management account over the past 12 months, the lower the borrower's investment level and repayment ability, and the higher the likelihood of a serious default. Conversely, the larger the maximum balance in the investment and wealth management account over the past 12 months, the higher the borrower's investment level and repayment ability, and the lower the likelihood of a serious default. The conversion method is the cube root conversion within the continuous conversion method.

[0506] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0507]

[0508] Where k is the number of features entering the model, and in Formula 3, k is 6.

[0509] α is the intercept, with a range of values ​​from (-2.2994, -2.0097), and an optimal value of -2.155; β1 is the maximum number of consecutive months of decrease in AUM over the past three months, with a range of values ​​from (-2.8509, -2.0665), and an optimal value of -2.459; β2 is the current point in time deposit account balance, with a range of values ​​from (-0.0925, -0.0588), and an optimal value of -0.076; β3 is the average monthly AUM (personal financial assets) over the past six months, with a range of values ​​from (-0.147, - β4 is the number of months from the current day to the maximum balance in the deposit account over the past 12 months, with a range of (-1.1717, -0.6868) and an optimal value of -0.929; β5 is the minimum balance in the deposit account over the past six months, with a range of (-0.119, -0.0689) and an optimal value of -0.094; β6 is the maximum balance in the investment and wealth management account over the past 12 months, with a range of (-0.634, -0.3404) and an optimal value of -0.487. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0510] x1 is the logarithmic transformation value of the maximum consecutive number of months of month-on-month decrease in AUM values ​​over the past three months generated by the feature conversion step; x2 is the logarithmic transformation value of the current deposit account balance generated by the feature conversion step; x3 is the WOE transformation value of the average monthly AUM (personal financial assets) over the past six months generated by the feature conversion step; x4 is the original value of the number of months from the current maximum balance of the deposit account over the past 12 months generated by the feature conversion step; x5 is the square transformation value of the minimum balance of the deposit account over the past six months generated by the feature conversion step; x6 is the cube root transformation value of the maximum balance of the investment and wealth management account over the past 12 months generated by the feature conversion step.

[0511] The model performance of some features is shown in Table 3-2 below:

[0512] Table 3-2

[0513]

[0514] The P-values ​​for all model characteristics are less than 0.05, indicating that these characteristics are significantly correlated with default performance. The Alpha2_5 model predicts borrowers with credit needs and a low risk of serious default based on information such as their asset level and repayment ability, deposit level, and changes in funding levels.

[0515] Examples 3-6

[0516] The P value calculated by the above formula can be further used to calculate the score of any customer.

[0517] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0518] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0519]

[0520]

[0521] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0522] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0523] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0524] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0525] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0526] According to the prediction model, a KS curve is drawn. The KS of this embodiment is 76.33. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0527] The following table shows the KS values ​​for the development and validation samples when the model was applied to the development and validation sets. Table 3-3 shows that the model performs well in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample set. This means that all sub-models have excellent discrimination.

[0528] Table 3-3

[0529]

[0530] 4) Credit card business credit score (Alphad_5 model score)

[0531] Example 4-1

[0532] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample and selected 630 million data points from 2019 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 130 million and 500 million, respectively.

[0533] When designing the specific model, the 130 million customers who have applied for it are analyzed and designed. From the 130 million customers, 20,000 customers who have applied for credit business and whose account age is less than 3 years are screened out to build a model to predict the probability of credit overdue for more than 90 days for people who have not applied for registered credit business. After the model is developed and put online, the results will be applied to the total number of 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0534] The model design includes 1) exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.5 million customers (credit card customers); 2) time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15-month time range as the performance period of Y; 3) sample sampling: use a good / bad sample number of 5:1 for sampling modeling.

[0535] Based on the above-mentioned sample of 20,000 customers who have applied for credit business and whose account age is less than 3 years, the borrowers' historical data information is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0536] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including "current repayment amount," "average number of overdue payments in the past six months," and "number of credit card installments in the past 12 months." In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0537] Table 4-1 Summary of basic variables and derived variables used in Example 4-1

[0538]

[0539] Example 4-2 Feature Screening

[0540] A preliminary screening was conducted on the 157 features (variables) collected in Example 4-1 that have potential predictive power for customer overdue payments.

[0541] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 4-1 were first eliminated, resulting in a total of 23 variables being deleted, leaving 134 variables.

[0542] In the second round of preliminary screening, for the 134 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 37 variables were eliminated, leaving 97 variables.

[0543] In the third round of preliminary screening, the 97 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box, and if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 43 features were eliminated, and 54 feature variables remained.

[0544] In the fourth round of preliminary screening, the 54 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 34 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, this embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0545] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0546] The fifth round of preliminary screening targets the 34 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0547] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0548] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0549] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0550] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0551] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0552] The criteria for approximate consistency are as follows:

[0553] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0554] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0555] After the fifth round of screening, 49 features were eliminated and 37 features remained.

[0556] Example 4-3 Conversion of features after initial screening

[0557] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0558] First, the remaining 34 features in Example 4-2 are judged on the conversion method of these features. The following three methods are selected based on the concentration of features, data type, etc.

[0559] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0560] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0561] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0562] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0563] Through this embodiment, 34 features are converted, of which 24 features are converted into WOE, 5 features are converted into dummy features (expanded into 20 variables), and 5 features are converted into continuous features.

[0564] Example 4-4 Feature Depth Screening (Feature Fine Screening Step)

[0565] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0566] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the number of features from 37 to 28.

[0567] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 28 features to 23.

[0568] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0569] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0570] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0571] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0572] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0573] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0574] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 23 features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0575] Example 4-5 Construction of Alphad_5 model

[0576] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 20,000, and the customer base is mainly customers who have applied for credit services and have an account age of less than 3 years. The model is used to predict the probability of credit delinquency of more than 90 days (the target variable) for people who have not applied for credit services.

[0577] In the present embodiment 4-5, it is finally confirmed that the final modeling result is described by taking the four features of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months, the maximum salary in the past 12 months, the maximum balance of the investment and financial management account in the past 6 months, and the minimum balance of the deposit account at the time point in the past 6 months as an example. In the feature conversion step conversion mode selection, the four model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the four features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0578] The average monthly assets over the past six months divided by the average monthly assets over the past twelve months reflects changes in the borrower's asset level. For recently registered borrowers, an increase in their asset level indicates a continued increase in their asset size. A larger ratio indicates stronger and more stable recent asset levels and a lower likelihood of default. A smaller ratio indicates weaker and more unstable recent asset levels and a higher likelihood of default. Conversion is to raw values.

[0579] The maximum salary over the past 12 months reflects the client's income level over the past period. For recent new applicants, higher average income levels and more assets available for credit repayments reduce the likelihood of serious defaults. Conversely, lower income levels increase the likelihood of serious defaults. The natural logarithm transformation is used.

[0580] The maximum balance in the investment and wealth management account over the past six months reflects the client's recent investment and wealth management skills. For newly registered clients, larger balances in investment and wealth management accounts over the past period indicate stronger financial management skills, better asset quality, and stronger debt repayment capacity, and a lower likelihood of serious overdue payments. Conversely, a lower balance indicates an increased likelihood of serious overdue payments. This transformation is performed using the natural logarithm.

[0581] The minimum deposit balance over the past six months reflects the borrower's recent deposit status. For newly registered borrowers, the minimum deposit amount can reflect their asset quality. The higher the minimum deposit balance over the past six months, the more abundant the borrower's funds are and the lower the likelihood of a major default. Conversely, the lower the minimum deposit balance over the past six months, the less cash the borrower has and the higher the likelihood of a major default. Natural logarithm transformation is used.

[0582] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0583]

[0584] Where k is the number of features entering the model, and in Formula 4, k is 4.

[0585] α is the intercept term, with a range of values ​​from (-2.04845, -2.31097), and the optimal value is -2.1797076; β1 is the coefficient corresponding to the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months, with a range of values ​​from (-0.20978, -0.34796), and the optimal value is -0.2788716; β2 is the coefficient corresponding to the maximum salary in the past 12 months, with a range of values ​​from (-0.055 51,-0.39342), and the optimal value is -0.2244648; β3 is the coefficient corresponding to the maximum balance of the investment and financial management account in the past 6 months, with a value range of (-0.1207,-0.12475), and the optimal value is -0.122724; β4 is the coefficient corresponding to the minimum balance of the deposit account in the past 6 months, with a value range of (-0.68315,-0.93368), and the optimal value is -0.808416.

[0586] x1 is the original value of the average monthly asset size in the past 6 months / the average monthly asset size in the past 12 months generated by the feature conversion step; x2 is the natural logarithm conversion value of the maximum salary in the past 12 months generated by the feature conversion step; x3 is the dummy variable conversion value of the maximum balance of the investment and financial management account in the past 6 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the minimum balance of the deposit account at a certain point in time in the past 6 months generated by the feature conversion step.

[0587] The model performance of some features is shown in Table 4-2 below.

[0588] Table 4-2

[0589]

[0590] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0591] Examples 4-6

[0592] The P value calculated by the above formula can be further used to calculate the score of any customer.

[0593] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0594] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0595]

[0596]

[0597] Where P is the borrower's calculation model, A is 54.2458, B is 115.4156, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0598] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0599] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0600] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0601] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0602] According to the prediction model, a KS curve is drawn. The KS of this embodiment is 80.77. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0603] The following table shows the KS values ​​of the model applied to the development and validation samples. Table 4-3 shows that the model has good discrimination performance in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample. This means that all sub-models have excellent discrimination performance.

[0604] Table 4-3

[0605]

[0606] 5) Credit card business credit score (Alpha2d_5 model score)

[0607] Example 5-1 Collection of modeling samples

[0608] During the construction of this example, personal financial asset transaction data and personal customer information data for all retail customers of a large bank between 2017 and 2021 were collected, totaling 730 million individuals. A professional model design solution was used to confirm the modeling sample, and data from 620 million individuals in 2018 and 650 million individuals in 2019 were selected as the analysis samples. Because modeling requires performance variables and the performance period for credit cards and special installment loans is relatively short, 2019 was selected as the observation period. The 650 million analysis sample was then divided into two parts: those with registered credit applications and those without registered credit applications, with sample sizes of 130 million and 520 million, respectively.

[0609] The specific Alpha2d model is a scenario scoring model for credit cards and special installment services. When designing the model, 110 million customers who have applied for credit cards and special installment services among the 130 million customers who have applied were analyzed and designed. More than 430,000 new customers who have applied for credit services and have an account age of less than 3 months were screened out from the 130 million customers to build a model to predict the probability of credit overdue for more than 90 days for people who have not applied for registered credit services. After the model is developed and put online, the results will be applied to the total number of 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0610] The model design includes 1) exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 58.59 million customers; 2) time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15-month time range as the performance period of Y; 3) sample sampling: use a good / bad sample number of 10:1 for sampling modeling.

[0611] Based on the above-mentioned sample of more than 430,000 customers who have applied for credit business and whose account age is less than 3 months, historical data information of borrowers is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0612] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount', 'Average Number of Overdue Payments in the Last 6 Months', and 'Number of Credit Card Installments in the Last 12 Months'. In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0613] Table 5-1 Summary of basic variables and derived variables used in this example

[0614]

[0615] Example 5-2 Feature Screening

[0616] A preliminary screening was conducted on the 157 features (variables) collected in Example 5-1 that have potential predictive power for customer overdue payments.

[0617] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 5-1 were first eliminated, resulting in a total of 23 variables being deleted, leaving 134 variables.

[0618] In the second round of preliminary screening, for the 134 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 35 variables were eliminated, leaving 99 variables.

[0619] In the third round of preliminary screening, the 99 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 23 features were eliminated, and 76 feature variables remain.

[0620] In the fourth round of preliminary screening, the 76 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 54 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0621] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0622] The fifth round of preliminary screening targets the 54 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0623] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0624] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0625] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0626] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0627] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0628] The criteria for approximate consistency are as follows:

[0629] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0630] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0631] After the fifth round of screening, 12 features were eliminated, leaving 42 features.

[0632] Example 5-3 Conversion of features after initial screening

[0633] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0634] First, the remaining 42 features in Example 5-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0635] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0636] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0637] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0638] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0639] Through this embodiment, 42 features are converted, of which 10 features are converted into WOE, 10 features are converted into dummy features (expanded into 40 variables), and 22 features are converted into continuous features, for a total of 72 features.

[0640] Example 5-4 Feature Depth Screening (Feature Fine Screening Step)

[0641] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0642] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the number of features from 72 to 56.

[0643] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 56 features to 46.

[0644] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0645] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0646] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0647] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0648] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0649] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0650] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 46 features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0651] Example 5-5 Construction of Alpha2d model

[0652] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 430,000, and the customer base is primarily new customers who have registered for credit services and have an account age of less than three months. The model is used to predict the probability of a credit delinquency of more than 90 days (the target variable) for people who have not registered for credit services.

[0653] In the present embodiment 5-5, it is finally confirmed that the six features of the maximum balance of the deposit account at the time point in the past 12 months is from the present month, the minimum balance of the deposit account at the time point in the past 6 months, the deposit account balance at the current time point, the monthly average AUM (personal financial assets) of the past 6 months, the average balance of the investment and financial management account in the past 6 months, and the maximum value of the salary paid on behalf of the fund in the past 12 months are used as examples to describe the final modeling result. In the feature conversion step conversion mode selection, the 6 model variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the 6 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0654] The number of months from the current month to the maximum deposit account balance over the past 12 months reflects changes in the borrower's financial resources. For new users who haven't registered for credit services, the smaller the maximum deposit balance over the past period, the more adequate the borrower's current financial resources are and the lower the likelihood of future loan defaults. This variable correlates with customer risk in the following way: the further the maximum deposit account balance over the past 12 months is from the current month, the lower the borrower's recent income and repayment ability, and the higher the likelihood of a serious default. Conversely, the closer the maximum deposit account balance over the past 12 months is from the current month to the current month, the higher the borrower's recent income and repayment ability, and the lower the likelihood of a serious default. This conversion method uses Word of Equity (WOE) conversion.

[0655] The minimum balance in a deposit account over the past six months reflects a borrower's financial resources and repayment ability. For new users who haven't registered for credit services, the larger the minimum balance over the past period, the more secure the borrower's financial resources and the lower the likelihood of future loan defaults. This variable correlates with customer risk: the larger the minimum balance in a deposit account over the past six months, the more secure the borrower's financial resources and the lower the likelihood of a serious default. Conversely, the larger the minimum balance in a deposit account over the past six months, the weaker the borrower's financial resources and repayment ability, and the higher the likelihood of a serious default. The transformation method is a logarithmic transformation within the continuous transformation model.

[0656] The current deposit account balance reflects the borrower's savings level and repayment ability. For new users who haven't registered for credit services, the greater the pressure on the borrower to repay their debt, the higher the likelihood of defaults. This variable correlates with customer risk: the larger the current deposit account balance, the better the borrower's savings level and repayment ability, and the lower the likelihood of a serious default. Conversely, the smaller the current deposit account balance, the lower the borrower's savings level and repayment ability, and the higher the likelihood of a serious default. Logarithmic transformation is used.

[0657] The average monthly AUM (personal financial assets) over the past six months reflects the borrower's repayment ability. For new users who haven't registered for credit services, the short-term average monthly AUM reflects the borrower's immediate financial capacity. A higher monthly AUM indicates greater financial capacity and a lower likelihood of future defaults. This variable correlates with customer risk: the higher the average monthly AUM (personal financial assets) over the past six months, the greater the borrower's repayment capacity and the lower the likelihood of a major default. Conversely, the lower the average monthly AUM (personal financial assets) over the past six months, the weaker the borrower's repayment capacity and the higher the likelihood of a major default. The conversion method is raw value conversion.

[0658] The average balance in investment and wealth management accounts over the past six months reflects the borrower's investment level and repayment ability. For new users who haven't registered for credit services, this average balance over the past six months reflects the adequacy of the borrower's secondary repayment source. Larger balances indicate stronger repayment ability and a lower likelihood of future defaults. This variable correlates with customer risk: the smaller the average balance in investment and wealth management accounts over the past six months, the lower the borrower's investment level and repayment ability, and the higher the likelihood of a serious default. Conversely, the larger the average balance in investment and wealth management accounts over the past six months, the higher the borrower's investment level and repayment ability, and the lower the likelihood of a serious default. The conversion method is square transformation.

[0659] The maximum amount of payroll payments over the past 12 months reflects the borrower's income level and repayment ability. For new users who haven't registered for credit services, the amount of payroll payments reflects the borrower's income and repayment ability. Financial constraints increase the likelihood of future defaults. This variable correlates with customer risk: the higher the maximum amount of payroll payments over the past 12 months, the stronger the borrower's income and repayment ability, and the lower the likelihood of a serious default. Conversely, the lower the maximum amount of payroll payments over the past 12 months, the weaker the borrower's income and repayment ability, and the higher the likelihood of a serious default. This conversion method uses cube root conversion.

[0660] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0661]

[0662] Where k is the number of features entering the model, and in Formula 5, k is 6.

[0663] α is the intercept, with a range of values ​​from (-2.6066, -2.2583), and the optimal value is -2.432; β1 is the number of months from the maximum balance of the deposit account in the past 12 months to the present, with a range of values ​​from (-3.0504, -2.2763), and the optimal value is -2.663; β2 is the minimum balance of the deposit account in the past 6 months, with a range of values ​​from (-0.1819, -0.1163), and the optimal value is -0.149; β3 is the balance of the deposit account at the current time, with a range of values ​​from (-0.1197, -0.0721), with an optimal value of -0.096; β4 is the average monthly AUM (personal financial assets) over the past six months, with a range of (-0.119, -0.0513), and an optimal value of -0.085; β5 is the average balance of investment and wealth management accounts over the past six months, with a range of (-0.5767, -0.2145), and an optimal value of -0.396; β6 is the maximum amount of payroll payments over the past 12 months, with a range of (-0.6459, -0.1799), and an optimal value of -0.413. (Note: The numerical range is derived from the 95% confidence interval, i.e., the 95% CI in the table below)

[0664] x1 is the logarithmic transformation value of the maximum balance of the deposit account in the past 12 months, generated by the feature conversion step; x2 is the logarithmic transformation value of the minimum balance of the deposit account in the past 6 months, generated by the feature conversion step; x3 is the WOE transformation value of the deposit account balance at the current time, generated by the feature conversion step; x4 is the original value of the average monthly AUM (personal financial assets) in the past 6 months, generated by the feature conversion step; x5 is the square transformation value of the average balance of the investment and wealth management account in the past 6 months, generated by the feature conversion step; x6 is the cube root transformation value of the maximum value of the salary payment in the past 12 months, generated by the feature conversion step.

[0665] The model performance of some features is shown in Table 5-2 below:

[0666] Table 5-2

[0667]

[0668] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0669] Examples 5-6

[0670] The P value calculated by the above formula can be further used to calculate the score of any customer.

[0671] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0672] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0673]

[0674]

[0675] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0676] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0677] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0678] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0679] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0680] According to the prediction model, a KS curve is drawn. The KS of this embodiment is 80.64. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0681] The following table shows the KS values ​​of the model applied to the development and validation samples. Table 5-3 shows that the model performs well in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample size. This means that all sub-models have excellent discrimination.

[0682] Table 5-3

[0683]

[0684] VI) Personal Loan Business Credit Scoring (Psi_6 Model Scoring)

[0685] Example 6-1

[0686] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample and selected 620 million data points from 2018 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 120 million and 500 million, respectively.

[0687] The specific PSI model is a scenario scoring for personal credit business. When designing the model, 33.14 million customers with personal credit business among the 120 million customers who have applied for it were analyzed and designed. Among the 33.14 million customers, more than 380,000 customers who have applied for credit business and whose account age is less than 3 months were used to build a model to predict the probability of credit overdue for more than 90 days in the future for the sample group that has not applied for registered credit business. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0688] The model design includes 1) exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 16.81 million customers; 2) time window setting: the sample is determined to use 2018 data as the modeling sample, and the time range of the next 24 months is used as the performance period of Y; 3) sample sampling: the sample number of good / bad is 5:1 for sampling modeling.

[0689] Based on the above-mentioned sample of more than 380,000 customers who have applied for credit business and whose account age is less than 3 months, historical data information of borrowers is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0690] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods. Basic methods include: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of a customer's account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months and the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including 'Current Repayment Amount', 'Average Number of Overdue Payments in the Last 6 Months', and 'Number of Credit Card Installments in the Last 12 Months'. In this article, "observation point" refers to the point in time from when samples were collected until modeling. "Current" also refers to the sampling cutoff time. "Observation point" and "Current" have the same meaning.

[0691] Table 6-1 Summary of basic variables and derived variables used in this example

[0692]

[0693] Example 6-2 Feature Screening

[0694] A preliminary screening was conducted on the 157 features (variables) collected in Example 6-1 that have potential predictive power for customer overdue payments.

[0695] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 6-1 were first eliminated, resulting in a total of 42 variables being deleted, leaving 105 variables.

[0696] In the second round of preliminary screening, for the 105 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 23 variables were eliminated, leaving 82 variables.

[0697] In the third round of preliminary screening, the 82 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is specifically that if the feature is a character variable, each value is a separate box, and if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 23 features were eliminated, and 59 feature variables remained.

[0698] In the fourth round of preliminary screening, the 59 features after the third round of screening were further screened based on the stepwise discrimination algorithm. After this round of screening, 47 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, this embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0699] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0700] The fifth round of preliminary screening targets the 47 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0701] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0702] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0703] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0704] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0705] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0706] The criteria for approximate consistency are as follows:

[0707] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0708] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0709] After the fifth round of screening, 7 features were eliminated, leaving 40 features.

[0710] Example 6-3 Conversion of features after initial screening

[0711] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the initial feature screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,

[0712] First, the remaining 40 features in Example 6-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0713] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0714] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0715] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0716] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0717] Through this embodiment, 40 features are converted, of which 20 features are converted into WOE, 10 features are converted into dummy features (expanded into 40 variables), and 10 features are converted into continuous features, for a total of 70 features.

[0718] Example 6-4 Feature Depth Screening (Feature Fine Screening Step)

[0719] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0720] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the total number of features from 70 to 53.

[0721] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 53 features to 45.

[0722] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0723] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0724] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0725] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0726] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0727] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0728] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 45 features are eliminated as the final feature variables to be entered into the model. For example, the final number of features can be 5, 6, 7, or 8.

[0729] Example 6-5 Construction of Psi_6 Model

[0730] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 380,000, and the customer base is primarily new customers who have applied for credit services with an age of less than three months. The model is constructed to predict the probability of a credit delinquency of more than 90 days (the target variable) for the sample population that has not applied for credit services.

[0731] In this embodiment 6-5, it is finally confirmed that the final modeling results are described using five features as examples: the current deposit account balance, the proportion of the maximum consecutive months of decrease in AUM value over the past 12 months, the minimum balance of the deposit account at the time point in the past 12 months, the number of months from the maximum balance of the deposit account at the time point in the past 12 months, and the number of investment and financial management accounts / total personal financial accounts in the past 12 months. In the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the features and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0732] The current deposit account balance reflects the borrower's savings level and repayment ability. For new users who haven't registered for credit services, when their income level exceeds their debt level, they will have a certain deposit balance, which reflects their repayment ability and reduces the likelihood of future loan defaults. This variable is associated with customer risk: the lower the current deposit account balance, the lower the borrower's savings level and repayment ability, and the higher the likelihood of a serious default. Conversely, the higher the current deposit account balance, the higher the borrower's savings level and repayment ability, and the lower the likelihood of a serious default. The natural logarithm transformation is used.

[0733] The percentage of months with the largest consecutive month-over-month decrease in AUM over the past 12 months reflects the borrower's asset level and repayment ability. For new users who haven't registered for credit services, the AUM value reflects the borrower's recent changes in financial capacity. Insufficient financial capacity increases the likelihood of future defaults. This variable correlates with customer risk: the lower the percentage of months with the largest consecutive month-over-month decrease in AUM over the past 12 months, the stronger the borrower's asset level and repayment ability, and the lower the likelihood of a major default. Conversely, the higher the percentage of months with the largest consecutive month-over-month decrease in AUM over the past 12 months, the weaker the borrower's asset level and repayment ability, and the higher the likelihood of a major default. The conversion method is Word of Equity (WOE).

[0734] The minimum balance in deposit accounts over the past 12 months reflects changes in a borrower's financial resources. For new users who haven't registered for credit services, the larger the minimum balance over the past period, the more adequate the borrower's historical average financial resources and the lower the likelihood of future loan defaults. This variable correlates with customer risk: the lower the minimum balance in deposit accounts over the past 12 months, the lower the borrower's repayment ability and the higher the likelihood of a serious default. Conversely, the higher the minimum balance in deposit accounts over the past 12 months, the stronger the borrower's repayment ability and the lower the likelihood of a serious default. This conversion method uses a cube root conversion.

[0735] The number of months from the current month to the maximum deposit account balance over the past 12 months reflects changes in the borrower's financial resources. For new users who haven't registered for credit services, the smaller the maximum deposit balance over the past period, the more adequate the borrower's current financial resources are and the lower the likelihood of future loan defaults. This variable correlates with customer risk in the following way: the further the maximum deposit account balance over the past 12 months is from the current month, the lower the borrower's recent income and repayment ability, and the higher the likelihood of a serious default. Conversely, the closer the maximum deposit account balance over the past 12 months is from the current month to the current month, the higher the borrower's recent income and repayment ability, and the lower the likelihood of a serious default. This conversion method uses Word of Equity (WOE) conversion.

[0736] The number of investment and wealth management accounts / total personal financial accounts over the past 12 months reflects the borrower's investment and repayment capabilities. For new users who haven't registered for credit services, the greater the proportion of investment and wealth management accounts over a period of time, the more secure the borrower's funds are and the lower the likelihood of future defaults. This variable correlates with customer risk: the lower the ratio of investment and wealth management accounts to total personal financial accounts over the past 12 months, the lower the borrower's investment and repayment capacity, and the higher the likelihood of a serious default. Conversely, the higher the ratio of investment and wealth management accounts to total personal financial accounts over the past 12 months, the higher the borrower's investment and repayment capacity, and the lower the likelihood of a serious default. The conversion method is the square transformation within the continuous transformation.

[0737] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0738]

[0739] Where k is the number of features entering the model, and in Formula 6, k is 5.

[0740] α is the intercept term, with a range of values ​​from (-2.7307, -2.3309) and an optimal value of -2.531; β1 is the current deposit account balance, with a range of values ​​from (-0.7567, -0.6587) and an optimal value of -0.708; β2 is the percentage of consecutive months with the largest month-on-month decrease in AUM over the past 12 months, with a range of values ​​from (-1.1549, -0.9041) and an optimal value of -1.03; β3 is the percentage of consecutive months with the largest month-on-month decrease in AUM over the past 12 months The minimum balance in the deposit account at that point in time ranges from (-0.5444, -0.415), with an optimal value of -0.48. β4 is the number of months from the current day to the maximum balance in the deposit account at that point in time over the past 12 months, with a range from (-0.6093, -0.4485), with an optimal value of -0.529. β5 is the total balance in the investment and wealth management account / individual financial account over the past 12 months, with a range from (-0.8953, -0.6327), with an optimal value of -0.764. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0741] x1 is the logarithmic conversion value of the deposit account balance at the current time point generated by the feature conversion step; x2 is the WOE conversion value of the proportion of the maximum consecutive months of decrease in AUM value over the past 12 months generated by the feature conversion step; x3 is the cube root conversion value of the minimum balance of the deposit account at a certain point in time in the past 12 months generated by the feature conversion step; x4 is the WOE conversion value of the number of months from the current maximum balance of the deposit account at a certain point in time in the past 12 months generated by the feature conversion step; x5 is the square conversion value of the number of investment and wealth management accounts / total personal finance accounts in the past 12 months generated by the feature conversion step.

[0742] The model performance of some features is shown in Table 6-2 below:

[0743] Table 6-2

[0744]

[0745] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0746] Example 6-6

[0747] The P value calculated by the above formula can be further used to calculate the sub-score and total score of any customer.

[0748] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0749] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0750]

[0751]

[0752] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0753] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0754] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0755] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0756] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0757] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Psi model is 75.54. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0758] The following table shows the KS values ​​of the model applied to the development and validation samples. Table 6-3 shows that the model has good discrimination performance in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample. This means that all sub-models have excellent discrimination performance.

[0759] Table 6-3

[0760]

[0761] 7) Personal Consumer Loan Business Credit Scores (Betai_7 Model Credit Scores)

[0762] Example 7-1 Collection of modeling samples

[0763] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample, selecting 620 million data points from 2018 and 650 million data points from 2019 as the analysis samples. Because modeling requires performance variables and the performance period for consumer credit is relatively short, 2019 was chosen as the observation period. The 650 million analysis sample was then divided into two groups: those with registered credit applications and those without. The sample sizes were 130 million and 520 million, respectively.

[0764] The Betai model is a scenario scoring model for consumer loan business. When designing the model, 40,000 customers with an account age of less than 3 months who have applied for credit business were screened out from the 130 million customers who have applied for credit business. The model is used to predict the probability of credit delinquency of more than 30 days in the future for the sample group that has not applied for credit business. After the model is developed and put online, the results will be applied to the entire 730 million people who have applied and those who have not applied, or all other potential customer groups.

[0765] Model design includes: 1) Exclusion rules: For example, excluding data from customers whose accounts have been settled, closed, have no performance, or have special circumstances. After these exclusions, the modeling population is 2.5 million customers (this group is the customer base that has registered for consumer loan services); 2) Time window setting: Determining the sample to use 2019 data as the modeling sample, and the next 15 months as the performance period for Y; 3) Sample sampling: Using a good / bad sample ratio of 10:1 for sampling modeling.

[0766] Based on the above-mentioned sample of 40,000 customers with an account age of less than 3 months who have applied for credit business, historical data information of borrowers is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0767] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods, including the following: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of their account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months, the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods in the last X months and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including "current repayment amount," "average number of overdue payments in the past six months," and "number of credit card installments in the past 12 months." In this application, "observation point" refers to the point in time at which samples were collected up to the time of modeling. "Current" also refers to the sampling cutoff time. These two terms have the same meaning.

[0768] Table 7-1 Summary of basic variables and derived variables used in this example

[0769]

[0770] Example 7-2 Feature Screening

[0771] A preliminary screening was conducted on the 157 features (variables) collected in Example 7-1 that have potential predictive power for customer overdue payments.

[0772] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 7-1 were first eliminated, resulting in a total of 25 variables being deleted, leaving 132 variables.

[0773] In the second round of preliminary screening, for the 132 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 25 variables were eliminated, leaving 107 variables.

[0774] In the third round of preliminary screening, the 107 features after the second round of preliminary screening were sorted according to the feature values ​​(the specific sorting method is: if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 34 features were eliminated, and 73 feature variables remained.

[0775] In the fourth round of preliminary screening, the 73 features after the third round of screening are further screened based on the stepwise discrimination algorithm. After this round of screening, 23 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean values ​​of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, this embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0776] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0777] The fifth round of preliminary screening is to further screen the 50 important features after the fourth round of screening based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0778] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0779] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0780] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0781] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0782] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0783] The criteria for approximate consistency are as follows:

[0784] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0785] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0786] After the fifth round of screening, 8 features were eliminated, leaving 42 features.

[0787] Example 7-3 Conversion of features after initial screening

[0788] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the feature initial screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numerical variables. For character variables, the dummy feature (dummy variable) conversion method is generally used for variable conversion; for numerical variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. The features using different conversion methods are grouped and divided into different data sets. Specifically,

[0789] First, the remaining 42 features in Example 7-2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data type, etc.

[0790] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0791] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0792] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion mode, and perform continuous conversion.

[0793] The feature merging module is used to horizontally splice the data of feature conversion method 1, feature conversion method 2 and feature conversion method 3.

[0794] Through this embodiment, 42 features are converted, of which 20 features are converted into WOE, 8 features are converted into dummy features (expanded into 24 variables), and 14 features are converted into continuous features, for a total of 58 features.

[0795] Example 7-4 Feature Depth Screening (Feature Fine Screening Step)

[0796] The feature fine screening step mainly performs the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability index (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.

[0797] Feature refinement step 1, based on a stepwise regression algorithm and employing F-tests and T-tests, introduces features in descending order of significance. Each time a feature is introduced, the selected features are individually tested. If a previously introduced feature becomes less significant due to the introduction of a subsequent feature, it is removed. This process is repeated until no features above the significance threshold are included in the equation, and no features below the significance threshold are removed from the regression equation. This step reduces the total number of features from 58 to 42.

[0798] Feature refinement step 2 further reduces multicollinearity in the model by removing features with high variance inflation factors. This step reduces the 42 features to 35.

[0799] Feature fine-screening step 3 compares the risk characteristics of each risk point (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend. Features whose model coefficients do not conform to the business trend are eliminated and re-iterate.

[0800] The specific implementation scheme of the feature fine screening module 3 is as follows:

[0801] 1. For features whose feature conversion method is WOE, the corresponding model training coefficient should be negative, and WOE conversion features with positive training coefficients should be eliminated.

[0802] 2. For continuous conversion methods, if the business logic indicates that the bad debt rate should increase as the value of this feature increases (for example, the credit limit utilization rate), the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the business logic indicates that the bad debt rate should decrease as the value of this feature increases (for example, the deposit amount), the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.

[0803] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the bad debt rate when the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.

[0804] Feature fine-screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated. For features with 0.25>PSI>0.1, they are carefully eliminated based on the impact of eliminating this feature on the model's discriminative ability.

[0805] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added or removed from the model. The final feature list and its conversion values ​​are obtained. After the above steps, 35 features are eliminated as the final feature variables to be entered into the model. For example, this could be 5 features, 6 features, 7 features, or 8 features.

[0806] Example 7-5 Construction of Betai_7 model

[0807] Generally speaking, the stronger the correlation between the feature variable and the target variable, the more accurate the final model. In this example, the sample size is approximately 40,000, and the customer base is primarily customers who have applied for credit services and have an account age of less than 3 months. The model is used to predict the probability of a credit delinquency of 30 days or more in the future (the target variable) for the sample group that has not applied for credit services.

[0808] In this embodiment 7-5, it is finally confirmed that the final modeling results are described using five features: the total number of wage payments in the past 12 months, the average monthly wage payment amount in the past 12 months, the minimum balance of the deposit account at a certain point in time in the past 6 months, the total number of financial management accounts in the current month, and the average monthly wage payment amount / average monthly AUM in the past 3 months. In the feature conversion step, the five modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the calculation model modeling step, the five features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) for logistic regression to obtain the calculation model.

[0809] The total number of payroll payments over the past 12 months reflects the borrower's income level and repayment ability. For new users who haven't registered for credit services, this number reflects the borrower's medium- to long-term income level and repayment ability. This variable correlates with customer risk: the lower the number of payroll payments over the past 12 months, the lower the borrower's income level and repayment ability, and the higher the likelihood of a serious default. Conversely, the greater the number of payroll payments over the past 12 months, the higher the borrower's income level and repayment ability, and the lower the likelihood of a serious default. The conversion method is the cube root conversion within the continuous conversion method.

[0810] The average monthly payroll payment amount over the past 12 months reflects the borrower's income level and repayment ability. For new users who have not registered for credit services, this amount reflects the borrower's medium- to long-term income level and repayment ability. This variable correlates with customer risk: the greater the average monthly payroll payment amount over the past 12 months, the stronger the borrower's assets and repayment ability, and the lower the likelihood of a serious default. Conversely, the smaller the average monthly payroll payment amount over the past 12 months, the weaker the borrower's assets and repayment ability, and the higher the likelihood of a serious default. The transformation method is the natural logarithm transformation within the continuous transformation.

[0811] The minimum balance in the deposit account over the past six months reflects the borrower's deposit level and repayment ability. For new users who have not registered for credit services, the minimum balance in the deposit account over the past six months reflects the borrower's short-term deposit level and repayment ability. If the borrower's income and expenditure are balanced, the larger the borrower's deposit, the lower the probability of future loan defaults. This variable is related to customer risk: the lower the minimum balance in the deposit account over the past six months, the lower the borrower's deposit level and repayment ability, and the higher the probability of a serious default. Conversely, the higher the minimum balance in the deposit account over the past six months, the higher the borrower's deposit level and repayment ability, and the lower the probability of a serious default. The transformation method is the natural logarithm transformation of the continuous transformation.

[0812] The total number of wealth management accounts in the current month reflects the borrower's investment level and repayment ability. For new users who have not registered for credit services, the total number of wealth management accounts reflects the borrower's short-term investment level and repayment ability. This variable is related to customer risk: the fewer the total number of wealth management accounts in the current month, the lower the borrower's investment level and repayment ability, and the higher the likelihood of a serious default. Conversely, the more the total number of wealth management accounts in the current month, the higher the borrower's investment level and repayment ability, and the lower the likelihood of a serious default. The conversion method is WOE conversion.

[0813] The average monthly payroll payment amount / average monthly AUM over the past three months reflects changes in the borrower's salary and asset levels. For new users who haven't registered for credit services, the greater the proportion of payroll payments to AUM over a period of time, the fewer secondary repayment sources the borrower has and the higher the likelihood of future defaults. This variable correlates with client risk: the lower the ratio of the average monthly payroll payment amount / average monthly AUM over the past three months, the higher the borrower's non-fixed income assets and the lower the likelihood of a major default. Conversely, the higher the ratio of the average monthly payroll payment amount / average monthly AUM over the past three months, the lower the borrower's non-fixed income assets and the higher the likelihood of a major default. The conversion method is the square transformation within the continuous transformation.

[0814] In the step of obtaining a calculation model, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the calculation model that represents the borrower:

[0815]

[0816] Where k is the number of features entering the model, and in Formula 7, k is 5.

[0817] α is the intercept term, with a range of values ​​from (0.2807 to 2.1858) and an optimal value of 1.233; β1 is the total number of wage payments in the past 12 months, with a range of values ​​from (-1.6949 to -0.9893) and an optimal value of -1.342; β2 is the average monthly wage payment amount in the past 12 months, with a range of values ​​from (-1.058 to -0.5641) and an optimal value of -0.811; β3 is the average monthly wage payment amount in the past 6 months. The minimum balance of a deposit account at a certain point in time ranges from (-0.1453, -0.0708), with an optimal value of -0.108. β4 is the total number of wealth management accounts in the current month, with a range from (-0.8792, -0.3186), with an optimal value of -0.599. β5 is the average monthly payroll amount / average monthly AUM over the past three months, with a range from (-0.596, -0.1754), with an optimal value of -0.386. (Note: The numerical ranges are derived from the 95% confidence interval, i.e., the 95% CI in the table below.)

[0818] x1 is the cube root conversion value of the total number of payroll payments in the past 12 months generated by the feature conversion step; x2 is the natural logarithm conversion value of the average monthly payroll payment amount in the past 12 months generated by the feature conversion step; x3 is the natural logarithm conversion value of the minimum balance of the deposit account at a certain point in time in the past 6 months generated by the feature conversion step; x4 is the WOE conversion value of the total number of wealth management accounts in the current month generated by the feature conversion step; x5 is the square conversion value of the average monthly payroll payment amount / average monthly AUM in the past 3 months generated by the feature conversion step.

[0819] The model performance of some features is shown in Table 7-2 below:

[0820] Table 7-2

[0821]

[0822] The P values ​​of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with default performance.

[0823] Example 7-6

[0824] The P value calculated by the above formula can be further used to calculate the sub-score and total score of any customer.

[0825] The scoring calculation step is used to convert the calculated calculation model into a score of 0-1000 using the pre-stored default score conversion code.

[0826] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:

[0827]

[0828]

[0829] Where P is the borrower's calculation model, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

[0830] In this embodiment, the P value calculated in the above embodiment can be substituted to calculate the score.

[0831] The Kolmogorov-Smirnov statistic was proposed by two Soviet mathematicians, A.N. Kolmogorov and N.V. Smirnov. In risk management, the KS statistic is often used to assess model discrimination. A greater degree of discrimination indicates a stronger risk ranking ability.

[0832] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:

[0833] KS=max(|cum(bad_rate)-cum(good_rate)|)

[0834] According to the prediction model, a KS curve is drawn. In this embodiment, the KS of the overall Betai model is 69.83. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk.

[0835] The following table shows the KS values ​​of the model applied to the development and validation samples. Table 7-3 shows that the model performs well in both the training and validation sets (where the training and validation sets have a sample size ratio of 6:4), as well as in the overall sample. This means that all sub-models have excellent discrimination.

[0836] Table 7-3

[0837]

[0838] 8) Personal Consumer Loan Business Credit Scores (Beta_6 Model Credit Scores)

[0839] Example 8-1 Collection of modeling samples

[0840] During the development of this example, we collected personal financial asset transaction data and personal customer information from all retail customers of a large bank between 2017 and 2021, totaling 730 million individuals. Using a professional model design solution, we confirmed the modeling sample and selected 630 million data points from 2019 as the analysis sample. Because modeling requires performance variables, the analysis sample was further divided into two groups: those with registered credit applications and those without. The sample sizes were 130 million and 500 million, respectively.

[0841] The Beta model is a scenario-based scoring system for consumer lending. During model design, the company analyzed and designed the model based on 130 million customers who had registered for credit services. From this 130 million customers, 30,000 customers with a credit age of less than three months were selected to construct a model that predicts the probability of a credit delinquency of 60 days or more for a sample group that has not registered for credit services. Once the model is developed and launched, the results will be applied to the entire 730 million customer base, both those with and without credit applications, as well as to all other potential customer groups.

[0842] The model design includes: 1) Exclusion rules: such as excluding customer data of settled accounts, closed accounts, no performance, and special circumstances. After exclusion, the modeling population is 2.48 million customers (this part of the population is the customer group that has applied for registration for consumer loan business); 2) Time window setting: Determine the sample to use 2019 data as the modeling sample, and use the next 15-month time range as the performance period of Y; 3) Sample sampling: Use a good / bad sample number of 5:1 for sampling modeling.

[0843] Based on the above-mentioned sample of 30,000 customers with an account age of less than 3 months who have applied for credit business, historical data information of borrowers is obtained: 1) Personal financial assets, including AUM, deposits, wealth management and payroll information.

[0844] Based on various risk points in the credit business, macro credit risk is segmented using information dimensions, time slices, and other methods. Variables are derived using professional characteristic variable construction methods, including the following: 1) Customer relationship length variables: such as the length of time a customer has had an account and the maximum age of their account; 2) Time interval variables: such as the number of months since the customer's most recent payment and the number of months since the customer's most recent overdue payment; 3) Behavior frequency variables: such as the number of times a customer's repayments exceed N in the last X months, the number of times a customer's credit limit utilization rate exceeds N in the last X months; 4) Current point variables: such as the customer's current monthly credit limit and current monthly balance; 5) Statistical variables: such as the maximum number of overdue periods in the last X months and the average credit limit utilization rate in the last X months; and 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments exceeding N in the last X months and the number of consecutive times a customer's repayment rate exceeds N in the last X months. Ultimately, 157 features with potential predictive power for customer delinquency were derived, including "current repayment amount," "average number of overdue payments in the past six months," and "number of credit card installments in the past 12 months." In this application, "observation point" refers to the point in time at which samples were collected up to the time of modeling. "Current" also refers to the sampling cutoff time. These two terms have the same meaning.

[0845] Table 8-1 Summary of basic variables and derived variables used in this example

[0846]

[0847] Example 8-2 Feature Screening

[0848] A preliminary screening was conducted on the 157 features (variables) collected in Example 8-1 that have potential predictive power for customer overdue payments.

[0849] In the first round of preliminary screening, features with a missing rate of more than 95% in the data collected in Example 8-1 were first eliminated, resulting in a total of 12 variables being deleted, leaving 145 variables.

[0850] In the second round of preliminary screening, for the 145 features that passed the first round of preliminary screening, the features with single values ​​exceeding 99% among the remaining features were eliminated, and a total of 23 variables were eliminated, leaving 122 variables.

[0851] In the third round of preliminary screening, the 122 features after the second round of preliminary screening were sorted according to the feature values ​​(the sorting method is as follows: if the feature is a character variable, each value is a separate box; if the feature is a numeric variable, it is sorted from small to large according to the value), and then divided into 10-20 boxes according to the quantile point. The feature IV value was calculated. If the feature IV value was lower than 0.02, it was eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values ​​lower than 0.05 are eliminated. A total of 16 features were eliminated, and 106 feature variables remained.

[0852] In the fourth round of preliminary screening, the 106 features after the third round of screening were further screened based on the stepwise discrimination algorithm. After this round of screening, 58 important feature variables can be quickly screened out. The following introduces the stepwise discrimination method. In actual data, there may be little difference in the mean of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of a stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.

[0853] The stepwise discriminant method uses the Wilks' Lambda criterion to measure the strength of feature capabilities, eliminating those that fail to meet the set threshold after three rounds of screening. During the stepwise discriminant process, the variable with the strongest discriminative ability is added first. As the number of variables in the model increases, the discriminative ability of the earlier introduced variables may also change. If the discriminative ability of a variable in the model falls below the threshold, the variable is removed. This process is repeated until all variables in the model meet the Wilks' Lambda similarity ratio criterion and no other variables meet the criteria for inclusion in the model.

[0854] The fifth round of preliminary screening targets the 58 important features after the fourth round of screening, and further screens the features based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.

[0855] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution does not conform to the business trend. Specifically,

[0856] (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median of each box and the corresponding bad debt rate.

[0857] (2) Calculate the rate of change (slope) using the median of the values ​​in each box compared to the previous box and the corresponding bad debt rate.

[0858] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.

[0859] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.

[0860] The criteria for approximate consistency are as follows:

[0861] First, if the business logic trend of this feature is that as the feature value increases, the bad debt rate increases (for example, the credit limit utilization rate), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0862] Second, if the business logic trend of this feature is that as the feature value increases, the bad debt rate decreases (for example, the deposit amount), this module will remove features where the number of boxes with a slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with a slope other than 0. This means that features that do not conform to the business trend are removed.

[0863] After the fifth round of screening, 20 features were eliminated, leaving 38 features.

[0864] Example 8-3 Conversion of features after initial screening

[0865] In the feature judgment step, the optimal conversion method is selected based on the concentration of features, data types, etc. after the feature initial screening step. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numerical variables. For character variables, the dummy feature (dummy variable) conversion method is generally used for variable conversion; for numerical variables, if the variable has less than 5 values, the WOE conversion method will be used. If the variable has more values, the WOE or continuous conversion method will be used (the optimal conversion form is selected based on the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has more values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. The features using different conversion methods are grouped and divided into different data sets. Specifically,

[0866] First, the remaining 38 features in Example 8-2 are judged for their conversion methods, and the following three methods are selected based on the concentration of the features, data type, etc.

[0867] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged by the feature judgment module as the optimal conversion method as WOE conversion.

[0868] Feature conversion mode 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and for which the optimal conversion mode is determined to be dummy feature conversion in the feature judgment module.

[0869] Feature conversion mode 3 is used to obtain the features in the data acquisition module and the features determined to be optimal for continu...

Claims

1. A method for calculating retail credit risk, comprising: Data collection step, which obtains retail credit prediction data of the sample to be predicted; a data processing step, which processes the acquired retail credit forecast data to obtain second-generation derived retail credit forecast data; The credit default probability calculation step is to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

2. The method according to claim 1, further comprising: After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.

3. The method according to claim 1 or 2, wherein: The retail credit prediction data includes original retail credit prediction data of the sample to be predicted and derived retail credit prediction data processed based on the original retail credit prediction data; Preferably, the original retail credit prediction data includes: Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions. Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

4. The method according to any one of claims 1 to 3, wherein The derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension; Preferably, the derived retail credit forecast data includes but is not limited to: Derivative retail credit forecast data obtained by processing based on sample relationship length, Derivative retail credit forecast data obtained by processing time interval variables, Derivative retail credit forecast data obtained based on the frequency of sample behavior. Derivative retail credit forecast data obtained by processing the sample at the current time point, Derivative retail credit forecast data obtained based on sample continuous behavior processing, Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

5. The method according to any one of claims 1 to 4, wherein In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data. The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

6. The method according to claim 5, wherein: The comprehensive business credit score includes a comprehensive business credit total score and three sub-scores; wherein the comprehensive business credit total score and the three sub-scores are calculated based on the retail credit prediction data using the three sub-models of the calculation model; wherein, based on the three sub-models, the three sub-scores of a certain sample to be tested are first calculated, and the sample to be tested is classified based on a decision tree method to determine a sub-model, and the score calculated by the sub-model is used as the comprehensive business credit total score; The credit card business credit score includes four credit card business credit scores; wherein the four credit card business credit scores are calculated based on the retail credit prediction data using the calculation model; The personal loan business credit score is calculated based on the retail credit prediction data using the calculation model; The personal consumer loan business credit score includes three types of personal consumer loan business credit scores; wherein the three types of personal consumer loan business credit scores are calculated based on the retail credit prediction data using the calculation model; The personal mortgage business credit score is calculated based on the retail credit prediction data using the calculation model; The personal quick loan business credit score is calculated based on the retail credit prediction data using the calculation model.

7. The method according to claim 6, wherein: The three sub-models of the integrated service class are shown in formulas (1-1) to (1-3). The four calculation models of the credit card business class are shown in Formula 2, Formula 3, Formula 4 and Formula 5 respectively; The calculation models of the personal loan business are shown in Formula 6 respectively; The three calculation models for the personal consumption loan business are shown in Formula 7, Formula 8, and Formula 9 respectively; The calculation model of the personal mortgage business is shown in Formula 10; The calculation models for the personal quick loan business are shown in Formula 11.

8. The method according to claim 6, wherein: The second-generation derivative retail credit prediction data is subjected to feature conversion and then substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes: Based on the feature type of the second-generation derivative retail credit prediction data that needs to be substituted into the credit default probability model, the WOE method, dummy feature method or continuous method is selected for feature conversion.

9. The method according to claim 8, wherein The continuous method for feature conversion includes the following methods: performing continuous feature conversion by calculating the cube root of the second-generation derivative retail credit forecast data or calculating the natural logarithm of the second-generation derivative retail credit forecast data.

10. The method according to claim 9, wherein: The credit default probability model is a model constructed based on the second-generation derivative retail credit forecast data and credit default probability using logistic regression based on the existing user population.

11. The method according to claim 10, wherein: The second-generation derivative retail credit prediction data is selected from one or more of the comprehensive business credit score, the personal mortgage business credit score, and the personal loan business credit score.

12. The method according to claim 11, wherein The credit default probability calculation step includes: Perform feature conversion on the comprehensive business credit score, personal mortgage business credit score and personal loan business credit score. It is preferred to use the continuous conversion method for the comprehensive business credit score; the continuous conversion method for the credit card business credit score; the continuous conversion method for the personal mortgage business credit score; and the WOE conversion method for the personal loan business credit score; Further preferred is to use a continuous conversion method to convert the total comprehensive business credit score into a calculation method of taking the cube root of the total comprehensive business credit score; and to use a continuous conversion method to convert the personal mortgage business credit score into a calculation method of taking the natural logarithm of the personal mortgage business credit score.

13. The method according to claim 12, wherein: The converted values ​​of the three features, namely the comprehensive business credit score, the personal mortgage business credit score and the personal loan business credit score, are substituted into the second-generation derivative retail credit prediction data and the credit default probability, and the credit default probability model constructed using logistic regression is used to calculate the credit default probability of the sample to be predicted.

14. The method according to claim 13, wherein: The credit default probability model is shown in the following formula 12-1: Where k is the number of features entering the model, and in Formula 12-1 k is 3; α is the intercept term, the value range is (5.5636, 5.4068), and the optimal value is 5.458; The value range of β1 is (0.10676, -0.12844), and the optimal value is -0.011; The β2 value range is (0.10474, -0.24806), and the optimal value is -0.072; The value range of β3 is (0.00984, -0.02936), and the optimal value is -0.010; x1 is the cube root transformation of the comprehensive business credit score generated in the feature transformation step; x2 is the WOE conversion value of the total credit score of personal loan business generated in the feature conversion step; x3 is the logarithmic transformation value of the total credit score of the personal mortgage business generated in the feature conversion step.

15. The method according to claim 14, wherein After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower: Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 443.9036, B is -72.1348, and the round function rounds the calculated score to the nearest integer. Finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.

16. A device for calculating retail credit risk, comprising: A data collection module, which is used to obtain retail credit prediction data of the sample to be predicted; A module for processing samples to be predicted, which is used to process the acquired retail credit prediction data to obtain second-generation derived retail credit prediction data; A credit default probability calculation module is used to substitute the second-generation derivative retail credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.

17. The device according to claim 16, wherein The device executes the steps of the method for calculating retail credit risk according to any one of claims 1 to 15.

18. A system for calculating retail credit risk, characterized in that: The system for calculating retail credit risk includes: a memory, a processor, and a program for the method for calculating retail credit risk stored in the memory and executable on the processor. When the program for calculating retail credit risk is executed by the processor, the steps of the method for calculating retail credit risk according to any one of claims 1 to 15 are implemented.

19. A computer storage medium, characterized in that The computer storage medium stores a program for calculating the retail credit risk method. When the program for calculating the retail credit risk method is executed by a processor, the steps of the method for calculating the retail credit risk according to any one of claims 1 to 15 are implemented.

20. A method for constructing a retail credit risk prediction model, comprising: A data collection step, which obtains the original retail credit prediction data of the sample used to build the model; A data derivation step, which processes the original retail credit forecast data into derived retail credit forecast data; a data processing step, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data; a feature initial screening step, which performs a preliminary screening of all categories, that is, all features, comprising the second-generation derivative retail credit forecast data to obtain features after the preliminary screening; The initial screening data conversion step determines the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method, and continuous conversion method for feature conversion, and uses the optimal method for feature conversion for each feature after the initial screening; Feature fine screening step, in which the features initially screened after feature conversion are deeply screened to obtain finely screened features; The credit default probability modeling step selects the logistic regression method to build the model based on the carefully screened features and the probability relationship between them and credit default, and confirms the method used to calculate the credit default probability.

21. The method according to claim 20, wherein In the data collection step, the original retail credit prediction data of the samples obtained for building the model include: Basic customer information data, which is based on the attributes of the sample users themselves but is not directly related to their behavior in financial institutions. Basic data on personal financial assets, which includes all other financial assets and financial transaction data of sample users in financial institutions that are not related to credit cards and loans.

22. The method according to claim 20, wherein In the data derivation step, the derived retail credit forecast data processed based on the original retail credit forecast data refers to data obtained by processing the collected original retail credit forecast data based on the time dimension, space dimension, frequency dimension, and statistical information dimension; Preferably, the derived retail credit forecast data includes but is not limited to: Derivative retail credit forecast data obtained by processing based on sample relationship length, Derivative retail credit forecast data obtained by processing time interval variables, Derivative retail credit forecast data obtained based on the frequency of sample behavior. Derivative retail credit forecast data obtained by processing the sample at the current time point, Derivative retail credit forecast data obtained based on sample continuous behavior processing, Derivative retail credit forecast data obtained by processing sample data based on statistical information dimensions.

23. The method according to claim 20, wherein In the data processing step, processing the acquired retail credit forecast data to obtain the second-generation derivative retail credit forecast data refers to processing the second-generation derivative retail credit forecast data using the original retail credit forecast data and the derivative retail credit forecast data using the calculation model of the second-generation derivative retail credit forecast data. The second-generation derivative retail credit prediction data is selected from comprehensive business credit scores, credit card business credit scores, personal loan business credit scores, personal quick loan business credit scores, personal consumer loan business credit scores and personal mortgage business credit scores.

24. The method according to any one of claims 20 to 23, wherein The feature initial screening step includes the following steps: The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model. The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high. The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features; The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order, The fourth preliminary screening step uses a stepwise discriminant algorithm to perform preliminary screening of features after the first to third preliminary screenings; The fifth preliminary screening step is to perform preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.

25. The method according to any one of claims 20 to 23, wherein In the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.

26. The method according to claim 25, wherein The initial screening data conversion steps include the following steps based on the judgment of concentration and data type: Classify the data type of each feature into character variables and numeric variables. For character variables, dummy feature conversion is used to convert the initial screening data. The process of further classifying numerical variables includes the following sub-steps: If the value of the numerical variable is less than n, the WOE conversion method is used to convert the initial screening data. If the value of the numerical variable is more than n, further judge if the conversion to a continuous variable has more values ​​and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used. Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.

27. The method according to claim 26, wherein Also includes: For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.

28. The method according to any one of claims 20 to 27, wherein The feature fine screening steps include: The first fine screening step is based on the stepwise regression algorithm, and the significance of the features is screened based on the F test and T test. The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate features with higher variance inflation factors to screen features. The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default in order to further perform feature screening.

29. The method according to any one of claims 20 to 28, wherein The credit default probability modeling step substitutes the features selected in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.

30. A device for constructing a retail credit risk prediction model, characterized in that: The device comprises: A data acquisition module, which is used to obtain raw retail credit prediction data of samples used to build a model; A data derivation module, which is used to process derived retail credit forecast data based on the original retail credit forecast data; A data processing module, which processes the original retail credit forecast data and the derived retail credit forecast data to produce second-generation derived retail credit forecast data; A feature initial screening module is used to perform initial screening on all categories, i.e., all features, of the original retail credit forecast data and the derived retail credit forecast data to obtain features after initial screening; The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, dummy feature conversion method and continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening; A feature fine-screening module is used to perform in-depth screening on the preliminarily screened features after feature conversion to obtain fine-screened features; The credit default probability modeling module is used to select the logistic regression method to construct a model based on the probability relationship between the carefully screened features and the credit default, and confirm the model used to calculate the credit default probability.

31. The device according to claim 30, wherein The device executes the steps of the method for constructing a retail credit risk prediction model according to any one of claims 20 to 29.

32. A system for constructing a retail credit risk prediction model, characterized in that: The system includes: a memory, a processor, and a program for constructing a retail credit risk prediction model method stored in the memory and executable on the processor. When the program for constructing a retail credit risk prediction model method is executed by the processor, the steps of the method for constructing a retail credit risk prediction model method as described in any one of claims 20 to 29 are implemented.

Citation Information

Patent Citations

  • A credit risk assessment method and apparatus based on logistic regression technology

    CN112686749B