Method for constructing common credit debt paying capability prediction model and common credit ScorezetaA model
By constructing an inclusive credit risk prediction model based on highly stable data, and combining various data transformation methods and logistic regression methods, the problem of insufficient model accuracy and stability in the credit business of small and medium-sized financial institutions has been solved, and more accurate credit risk prediction and risk management have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2026-04-10
AI Technical Summary
Small and medium-sized financial institutions lack data mining and analysis capabilities and risk modeling techniques in their lending business, which makes it impossible to effectively improve the accuracy and stability of their models. Existing scoring systems rely on a single data source and have insufficient modeling samples, resulting in discrepancies between predicted results and actual credit delinquency.
Using highly stable and comprehensive data samples, and combining three methods—WOE transformation, dummy feature transformation, and continuous transformation—a comprehensive credit risk prediction model is constructed through a stepwise discriminative variable selection method. Logistic regression is used to calculate the probability of credit default, thereby enhancing the efficiency and accuracy of model development.
It has improved the accuracy of financial institutions' credit risk prediction and the stability of their models, enhanced their risk management capabilities in the lending business, and supported financial institutions in increasing their business volume and competitive advantage on the basis of controllable risk.
Smart Images

Figure BDA0004605644350000081 
Figure BDA0004605644350000101 
Figure BDA0004605644350000131
Abstract
Description
Technical Field
[0001] This application relates to a credit risk management system and method. Utilizing the method and system of this application can assist financial institutions in making more accurate risk decisions and accelerate their digital transformation process. Specifically, this application relates to a method for constructing an inclusive credit risk prediction model and an inclusive credit risk prediction model. Background Technology
[0002] In the current environment of booming consumer credit, the manual approval mechanisms of some financial institutions are no longer able to cope with the increasing credit demand, thus creating an urgent need to improve their intelligent risk control capabilities. Financial institutions hope to build a comprehensive risk mitigation mechanism covering the entire credit process, from customer pre-screening, pre-loan review, loan approval, post-loan management to early collection stages.
[0003] If a scoring system can be developed based on the principles of early identification, early warning, early detection, and early handling, and credit business can be monitored and managed quickly and conveniently, then financial institutions can improve their business volume, competitive advantage, and asset quality while keeping risks under control.
[0004] However, building a scoring system is highly dependent on data and technology. The diversity and coverage of data dimensions, modeling techniques, and methodologies directly affect the final stability and ranking of the scoring system. Some financial institutions have limited experience in intelligent risk control for their business operations and possess weak risk control capabilities. In practical applications, factors such as insufficient data mining and analysis capabilities and weak risk modeling techniques prevent financial institutions from fully leveraging the value of their internal data and effectively improving model accuracy and stability, posing significant technical challenges. This is also one of the main obstacles faced by small and medium-sized financial institutions in their digital transformation efforts. Summary of the Invention
[0005] To address the shortcomings of the existing technologies mentioned above, this application provides a credit risk management system and method that can provide financial institutions with effective risk management. The credit risk prediction method and system of this application are based on massive amounts of data from large banks and utilize statistical principles to extract risk patterns, thus possessing broad applicability.
[0006] Currently, other popular scoring models on the market are often hampered by factors such as small sample sizes, limited data sources, and high homogeneity of data dimensions. Furthermore, since most current risk assessment models use data with weak financial attributes—that is, they are based on non-credit transaction data such as data from smart terminal devices, social media platforms, and online shopping malls, as well as non-delinquency prediction targets—their prediction results often deviate significantly from actual credit delinquency situations.
[0007] This application first provides a method applicable to the construction of credit risk models, and then provides financial institutions with an accurate method for calculating the credit risk of samples to be predicted based on the model constructed by this method.
[0008] The method and system in this application are developed based on highly stable and comprehensive data samples, and have made systematic innovations to the already relatively mature credit risk control system. They use credit samples containing various business forms to predict the potential credit risk of financial institutions.
[0009] Compared with existing models on the market, the method and system of this application retain the relatively mature basic framework for model building, and innovatively use a stepwise discriminative variable screening method, which greatly improves the efficiency of model development without affecting the overall model performance. It also innovatively combines three methods: continuous transformation, WOE transformation, and dummy feature transformation, which makes up for the potential problem of poor model performance when a single transformation occurs, enabling financial institutions to more accurately predict the credit risk of samples.
[0010] This application relates to the following technical solutions:
[0011] 1. A method for constructing an inclusive credit risk prediction model, comprising:
[0012] The data acquisition step involves obtaining raw corporate credit prediction data for samples used to build the model;
[0013] The data derivation step involves processing the original corporate credit prediction data to generate derived corporate credit prediction data.
[0014] The feature screening step involves preliminary screening of all categories, including both original and derived corporate credit prediction data, to obtain the features after preliminary screening.
[0015] The initial screening data transformation step involves determining the transformation method for the features after initial screening to confirm whether to use one of the following methods: WOE transformation, dummy feature transformation, or continuous transformation. For each feature after initial screening, the optimal method is determined to be used for feature transformation.
[0016] The feature refinement step involves performing a deep screening on the features that have undergone initial feature transformation to obtain refined features.
[0017] The steps for modeling the probability of credit default are as follows: based on the probabilistic relationship between the refined features and credit default, logistic regression is selected to build the model, and the method used to calculate the probability of credit default is confirmed.
[0018] 2. The method according to item 1, wherein,
[0019] In the data acquisition step, the raw corporate credit prediction data obtained for the samples used to build the model includes:
[0020] The basic data on corporate deposits is based on all available data regarding RMB deposits made by sample (corporate) users at financial institutions.
[0021] Basic enterprise information data refers to data based on the attributes of the sample (enterprise) users themselves, but which is not directly related to their behavior in financial institutions.
[0022] The basic data on the financial assets of business owners consists of all other financial assets and transactions held by the sample (actual controller of the enterprise) in financial institutions that are not related to credit cards and loans.
[0023] Basic information about business owners is based on the attributes of the sample (actual controller of the enterprise) users themselves, but is not directly related to their behavior in financial institutions.
[0024] 3. The method according to item 1, wherein,
[0025] In the data derivation step, the process of processing the original enterprise credit prediction data into derived enterprise credit prediction data refers to the data obtained by processing the collected original enterprise credit prediction data based on time dimension, spatial dimension, frequency dimension, and statistical information dimension.
[0026] Preferably, the derived corporate credit forecast data includes, but is not limited to:
[0027] Derivative corporate credit prediction data obtained by processing based on sample relationship length.
[0028] Derivative corporate credit prediction data obtained by processing time interval variables.
[0029] Derivative corporate credit prediction data is obtained by processing sample behavior frequency.
[0030] Derivative corporate credit prediction data is obtained by processing the data based on the current situation of the sample at the current point in time.
[0031] Derivative corporate credit prediction data obtained by processing the continuous behavior of samples.
[0032] Derivative corporate credit prediction data is obtained by processing sample data based on statistical information dimensions.
[0033] 4. The method according to any one of items 1 to 3, wherein,
[0034] The initial feature screening process includes the following steps:
[0035] The first initial screening step involves filtering features based on the data missing information for each feature of the samples used to build the model.
[0036] The second initial screening step involves filtering features based on whether a single value of a particular feature is excessively high.
[0037] The third initial screening step involves calculating the information value (IV) of each feature to perform preliminary screening of the features.
[0038] The order of the first, second, and third preliminary screening steps can be arbitrary.
[0039] The fourth preliminary screening step involves using a stepwise discrimination algorithm to perform preliminary screening of the features after the first to third preliminary screenings.
[0040] The fifth initial screening step involves preliminary screening of the features after the fourth initial screening step based on the consistency between the risk characteristics of each feature and the actual real results of the samples used for model construction.
[0041] 5. The method according to any one of items 1 to 4, further comprising:
[0042] The sample selection step is used to filter all users to obtain samples for model building before the data acquisition step.
[0043] Preferably, the sample selection step includes classifying all users in the sample based on a decision tree, and the classification criteria include, but are not limited to:
[0044] Is a user a customer holding corporate deposits?
[0045] Has a particular user been involved in a financial institution risk event?
[0046] 6. The method according to any one of items 1 to 5, wherein, in the initial screening data conversion step, the determination of the conversion method for the features after initial screening is based on the concentration and data type of the features after initial screening.
[0047] 7. The method according to item 6, wherein,
[0048] The initial data transformation process, based on determinations of concentration and data type, includes the following steps:
[0049] Classify each feature based on its data type, categorizing each feature into character variables and numeric variables.
[0050] For character-type variables, a dummy feature transformation method is used for initial data transformation.
[0051] The process of further classifying numerical variables includes the following sub-steps:
[0052] If the numerical variable has fewer than n values, use the WOE (Word of the Environment) transformation method for initial data transformation.
[0053] If the numerical variable has more than n values, further determine if converting it to a continuous variable would result in a large number of values and a concentration of any single value greater than m%. In this case, use the WOE (Word of the Entity) conversion method. If the concentration of any single value is less than or equal to m%, use the continuous conversion method.
[0054] Preferably, n and m are both positive integers, where n = 5 to 10 and m = 90 to 99.
[0055] 8. The method according to item 7, further comprising:
[0056] For features that have been confirmed to require continuous transformation, the optimal transformation method is selected based on the correlation between the feature and credit default under different continuous transformation methods.
[0057] The preferred method for continuous feature transformation is to directly select the original value, calculate the square of the original data, calculate the square root of the original data, calculate the cube root of the original data, or calculate the natural logarithm of the original data.
[0058] 9. The method according to any one of items 1 to 8, wherein the feature screening step comprises:
[0059] The first screening step, based on the stepwise regression algorithm, uses F-tests and T-tests to screen features based on their significance.
[0060] The second screening step involves calculating the variance inflation factor for each feature and removing features with high variance inflation factors to filter the features.
[0061] The third fine screening step involves analyzing the feature coefficients of the features after the first and second fine screening steps based on logistic regression to determine whether they conform to the trend of the predicted results for credit defaults, in order to further screen the features.
[0062] 10. The method according to any one of items 1 to 9, wherein the credit default probability modeling step substitutes the features selected by the feature screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.
[0063] 11. A method for calculating inclusive credit risk, comprising:
[0064] The data acquisition step involves obtaining enterprise credit prediction data for the sample to be predicted.
[0065] The step of classifying the samples to be predicted uses a decision tree-based method to classify the samples to be predicted to determine the sub-models used to calculate the probability of credit default.
[0066] The credit default probability calculation steps involve substituting enterprise credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted.
[0067] 12. The method according to item 11, further comprising:
[0068] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000.
[0069] 13. The method according to item 11 or 12, wherein,
[0070] The corporate credit prediction data includes the original corporate credit prediction data of the sample to be predicted and the derived corporate credit prediction data processed based on the original corporate credit prediction data.
[0071] Preferably, the original corporate credit prediction data includes:
[0072] The basic data on corporate deposits is based on all available data regarding RMB deposits made by sample (corporate) users at financial institutions.
[0073] Basic enterprise information data refers to data based on the attributes of the sample (enterprise) users themselves, but which is not directly related to their behavior in financial institutions.
[0074] The basic data on the financial assets of business owners consists of all other financial assets and transactions held by the sample (actual controller of the enterprise) in financial institutions that are not related to credit cards and loans.
[0075] Basic information about business owners is based on the attributes of the sample (actual controller of the enterprise) users themselves, but is not directly related to their behavior in financial institutions.
[0076] 14. The method according to any one of items 11 to 13, wherein,
[0077] Derived corporate credit forecast data, derived from original corporate credit forecast data, refers to data obtained by processing the collected original corporate credit forecast data based on time, space, frequency, and statistical information dimensions.
[0078] Preferably, the derived corporate credit forecast data includes, but is not limited to:
[0079] Derivative corporate credit prediction data obtained by processing based on sample relationship length.
[0080] Derivative corporate credit prediction data obtained by processing time interval variables.
[0081] Derivative corporate credit prediction data is obtained by processing sample behavior frequency.
[0082] Derivative corporate credit prediction data is obtained by processing the data based on the current situation of the sample at the current point in time.
[0083] Derivative corporate credit prediction data obtained by processing the continuous behavior of samples.
[0084] Derivative corporate credit prediction data is obtained by processing sample data based on statistical information dimensions.
[0085] 15. The method according to any one of items 11 to 14, wherein,
[0086] The enterprise credit forecast data is selected from one, two, three, four, five, six, seven, or eight of the following:
[0087] The following data points are considered: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount of the business over the past 6 months; the ratio of the number of credit transactions to the total number of transactions over the past 6 months; the ratio of the amount of credit transactions to the total amount of transactions over the past 3 months; the business's current monthly deposit balance; the business owner's current monthly AUM value; the ratio of the amount of debit transactions to the total amount of transactions over the past 3 months; the business owner's current balance at a given point in time; the minimum balance of the business owner's deposit account at a given point in time over the past 6 months; the difference between the number of months with the maximum balance of the business owner's deposit account at a given point in time over the past 12 months and the current month; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 3 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and the percentile of the business's current monthly deposit balance by industry.
[0088] 16. The method according to any one of items 11 to 15, wherein,
[0089] The steps for classifying the samples to be predicted include the following sub-steps:
[0090] Is the sample to be tested a customer holding corporate deposits?
[0091] Has the sample to be tested already experienced a financial institution risk event?
[0092] Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model used to calculate the probability of credit default. Under the premise of ensuring the rationality of business logic, the order of the above sub-steps can be arbitrarily set.
[0093] 17. The method according to any one of items 11 to 16, wherein,
[0094] The feature transformation of corporate credit prediction data is performed before being substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature transformation step includes:
[0095] Based on the feature type of the enterprise credit prediction data that needs to be substituted into the credit default probability model, either the WOE method or the continuous method is selected for feature transformation.
[0096] 18. The method according to item 17, wherein,
[0097] Continuous feature transformation can be performed in the following ways: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
[0098] 19. The method according to any one of items 11 to 18, wherein,
[0099] The credit default probability model is a model constructed based on the credit prediction data of sample enterprises and the credit default probability using logistic regression based on an existing user group, preferably a model constructed based on the method of any one of claims 1 to 10.
[0100] 20. The method according to any one of items 11 to 19, wherein,
[0101] The corporate credit forecast data is selected from:
[0102] One, two, three, four, or five of the following: the minimum balance of the business owner's deposit account at a given point in time over the past three months; the minimum balance of the business's current deposits over the past three months; the maximum monthly average amount of credit transactions over the past six months; the ratio of the number of credit transactions to the total number of transactions over the past six months; and the ratio of the amount of credit transactions to the total amount of transactions over the past three months.
[0103] 21. The method according to item 20, wherein,
[0104] Feature transformation is performed on the following data: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average amount of the business's credit transactions over the past 6 months, the ratio of the number of credit transactions to the total number of transactions over the past 6 months, and the ratio of the amount of credit transactions to the total amount of transactions over the past 3 months.
[0105] The preferred approach is to use a continuous conversion method for the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average amount of credit transactions over the past 6 months, and the ratio of the business's credit transaction amount to the total transaction amount over the past 3 months. The preferred approach is to use a dummy feature conversion method for the ratio of the number of credit transactions to the total number of transactions over the past 6 months.
[0106] Further optimization involves using a continuous transformation method for the minimum balance of a business owner's deposit account at a given point in time over the past three months (using natural logarithm transformation), a continuous transformation method for the minimum balance of the business owner's current account over the past three months (using cube root transformation), a continuous transformation method for the maximum monthly average credit transaction amount over the past six months (using cube root transformation), a continuous transformation method for the ratio of the business owner's credit transaction amount to the total transaction amount over the past three months (using natural logarithm transformation), and a dummy feature transformation method for the ratio of the number of credit transactions to the total number of transactions over the past six months (using dummy variable transformation).
[0107] 22. The method according to item 21, wherein,
[0108] The transformed values of five features—the minimum balance of the business owner's deposit account at a given point in time over the past three months, the minimum balance of the business's current deposits over the past three months, the maximum monthly average loan transaction amount over the past six months, the ratio of the number of loan transactions to the total number of transactions over the past six months, and the ratio of the loan transaction amount to the total transaction amount over the past three months—are substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probabilities to calculate the default probability of the sample to be predicted.
[0109] 23. The method according to item 22, wherein,
[0110] The sub-model is shown in Formula 1:
[0111]
[0112] Where k is the number of features entering the model, preferably k is 5.
[0113] α is the intercept term, with a preferred numerical range of (0.71157, 0.68367) and an optimal value of 0.697618;
[0114] β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a point in time over the past 3 months. The preferred value range is (-0.03417, -0.04716), and the optimal value is -0.040669.
[0115] β2 is the coefficient corresponding to the minimum balance of the enterprise's demand deposits in the past 3 months. The preferred value range is (-0.01744, -0.07704), and the optimal value is -0.047235.
[0116] β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months. The preferred value range is (0.00201, -0.08893), and the optimal value is -0.043456.
[0117] β4 is the coefficient corresponding to the ratio of the number of credit transactions to the total number of transactions in the past 6 months. The preferred value range is (0.08153, -0.06307), and the optimal value is 0.009227.
[0118] β5 is the ratio of a company's loan transaction amount to its total transaction amount in the past three months. The preferred value range is (-0.52827, -1.04399), and the optimal value is -0.786134.
[0119] x1 is the natural logarithm transformed value of the minimum balance of the business owner's deposit account at the point in time over the past 3 months, generated by the feature transformation step;
[0120] x2 is the cube root transformed value of the minimum current deposit balance of the enterprise in the past 3 months generated by the feature transformation step;
[0121] x3 is the square root of the maximum average monthly credit transaction amount of the enterprise over the past 6 months, generated by the feature transformation step.
[0122] x4 is the dummy variable transformed value of the ratio of the number of credit transactions to the total number of transactions in the past 6 months of the enterprise, generated by the feature transformation step;
[0123] x5 is the natural logarithm of the ratio of the company's credit transaction amount in the past 3 months to its total transaction amount in the past 3 months, generated by the feature transformation step.
[0124] 24. The method according to any one of items 11 to 19, wherein,
[0125] The corporate credit forecast data is selected from one, two, three, four, five, or six of the following: the current monthly deposit balance of the enterprise, the current monthly AUM value of the enterprise owner, the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, the current balance of the enterprise owner, the difference between the number of months of the minimum balance of the enterprise owner's deposit account in the past 6 months and the number of months of the maximum balance of the enterprise owner's deposit account in the past 12 months and the current month.
[0126] 25. According to the method described in item 24, the following features are applied to the enterprise's current monthly deposit balance, the enterprise owner's current monthly AUM value, the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, the enterprise owner's current balance, the enterprise owner's minimum deposit account balance in the past 6 months, and the enterprise owner's maximum deposit account balance in the past 12 months, and the difference between the current month and the number of months in which these features are applied.
[0127] The preferred approach is to use a continuous conversion method for the enterprise's current monthly deposit balance, the enterprise owner's current monthly AUM value, the enterprise owner's current balance at the current time, and the minimum balance of the enterprise owner's deposit account at the time of the past 6 months. The preferred approach is to use a dummy feature conversion method for the ratio of the enterprise's debit transaction amount to the total transaction amount of the past 3 months and the difference between the number of months with the maximum balance of the enterprise owner's deposit account at the time of the past 12 months and the current month.
[0128] Further optimization involves using a continuous transformation method for the company's current monthly deposit balance (natural logarithmic transformation), a continuous transformation method for the company owner's current monthly AUM value (square root transformation), a dummy transformation method for the ratio of the company's debit transaction amount to the total transaction amount in the past 3 months (dummy variable transformation), a continuous transformation method for the company owner's current balance (natural logarithmic transformation), a continuous transformation method for the minimum balance of the company owner's deposit account in the past 6 months (natural logarithmic transformation), and a dummy transformation method for the difference between the number of months with the maximum balance in the company owner's deposit account in the past 12 months and the current month (dummy variable transformation).
[0129] 26. According to the method described in item 25, the converted values of the following six features are substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probability to calculate the default probability of the sample to be predicted. These features include: the current monthly deposit balance of the enterprise; the current monthly AUM value of the business owner; the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months; the current balance of the business owner's deposit account in the past 6 months; and the difference between the number of months of the maximum balance of the business owner's deposit account in the past 12 months and the current month.
[0130] 27. The method according to item 26, wherein,
[0131] The sub-model is shown in Formula 1:
[0132]
[0133] Where k is the number of features entering the model, and k is 6 in Formula 2;
[0134] α is the intercept term, with a preferred numerical range of (1.18368, 0.87527) and an optimal value of 1.029473;
[0135] β1 is the coefficient corresponding to the company's current monthly deposit balance. The preferred value range is (-0.07368, -0.11077), and the optimal value is -0.092225.
[0136] β2 is the coefficient corresponding to the current monthly AUM value of the business owner. The preferred value range is (-0.07489, -0.07929), and the optimal value is -0.077093.
[0137] β3 is the coefficient corresponding to the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months. The preferred value range is (0.44374, 0.34763), and the optimal value is 0.395682.
[0138] β4 is the coefficient corresponding to the current balance of the business owner. The preferred value range is (-0.04003, -0.06985), and the optimal value is -0.054940.
[0139] β5 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 6 months. The preferred value range is (-0.01585, -0.03037), with the optimal value being -0.023109.
[0140] β6 is the coefficient corresponding to the difference between the number of months with the largest balance in the business owner's deposit account at a given point in time over the past 12 months and the number of months in the current month. The preferred value range is (-0.01482, -0.08255), and the optimal value is -0.048686.
[0141] x1 is the natural logarithm transformed value of the company's current monthly deposit balance generated by the feature transformation step;
[0142] x2 is the square root transformed value of the current monthly AUM value of the business owner generated by the feature transformation step;
[0143] x3 is the dummy variable transformed value of the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, generated by the feature transformation step;
[0144] x4 is the cube root transformed value of the business owner's current point-in-time balance generated by the feature transformation step;
[0145] x5 is the cube root transformed value of the minimum balance of the business owner's deposit account at any point in time over the past 6 months, generated by the feature transformation step;
[0146] x6 is the dummy variable transformed value of the difference between the number of months with the largest balance in the business owner's bank account at any point in the past 12 months and the number of months in the present, generated by the feature transformation step.
[0147] 28. The method according to any one of items 11 to 19, wherein,
[0148] The corporate credit forecast data is selected from:
[0149] The following parameters are considered: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and one, two, three, four, five, six, or seven percentiles of the business's current monthly deposit balance by industry.
[0150] 29. According to the method described in item 28, the following features are transformed: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average credit transaction amount of the business over the past 6 months, the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months, the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months, the minimum monthly product of the business's current deposits over the past 6 months, and the percentile of the business's current monthly deposit balance by industry.
[0151] The preferred conversion method uses a continuous conversion approach for the following metrics: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and the percentile of the business's current monthly deposit balance by industry.
[0152] Further optimization involves using a continuous conversion method for the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, and the maximum monthly average credit transaction amount over the past 6 months. This involves performing a natural logarithmic transformation on the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, or the maximum monthly average credit transaction amount over the past 6 months. Additionally, the ratio of the business's average monthly debit transaction amount to its average monthly balance over the past 12 months and the minimum monthly product of the business's current deposits over the past 6 months are used. The continuous conversion method involves taking the cube root of the ratio of the company's average monthly debit transaction amount to its average monthly balance over the past 12 months, or the minimum monthly accumulation of the company's current deposits over the past 6 months. The continuous conversion method for the ratio of the company's average monthly credit transaction number to its average monthly total credit transaction number over the past 12 months, and the percentage of the company's current monthly deposit balance by industry, involves taking the square root of the ratio of the company's average monthly credit transaction number to its average monthly total credit transaction number over the past 12 months, or the percentage of the company's current monthly deposit balance by industry.
[0153] 30. According to the method described in item 29, the following seven characteristics are transformed and substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probability to calculate the default probability of the sample to be predicted. These characteristics are: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the business's average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions over the past 3 months to the total number of credit transactions over the past 12 months; the minimum monthly product of the business's current deposits over the past 6 months; and the quantile of the business's current monthly deposit balance by industry.
[0154] 31. The method according to item 30, wherein,
[0155] The sub-model is shown in Formula 1:
[0156]
[0157] Where k is the number of features entering the model, and k is 7 in Formula 1;
[0158] α is the intercept term, with a preferred numerical range of (4.90128, 4.79548) and an optimal value of 4.848383;
[0159] β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 3 months. The preferred value range is (-0.04771, -0.07097), with the optimal value being -0.059338.
[0160] β2 is the coefficient corresponding to the minimum balance of the enterprise's demand deposits in the past 3 months. The preferred value range is (-0.04223, -0.04769), and the optimal value is -0.044961.
[0161] β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months. The preferred value range is (0.00699, -0.01024), and the optimal value is -0.001625.
[0162] β4 is the coefficient corresponding to the ratio of the company's average monthly debit transaction amount to the average monthly balance over the past 12 months. The preferred value range is (0.01636, 0.01505), and the optimal value is 0.0157047.
[0163] β5 is the coefficient corresponding to the ratio of the average number of monthly credit transactions in the past 3 months to the total number of monthly credit transactions in the past 12 months. The preferred value range is (-0.29326, -0.34022), and the optimal value is -0.316739.
[0164] β6 is the coefficient corresponding to the quantile of the current month's deposit balance of enterprises by industry. The preferred value range is (-0.00684, -0.00892), and the optimal value is -0.007878.
[0165] β7 is the coefficient corresponding to the minimum monthly accumulation of the company's current deposits over the past 6 months. The preferred value range is (-0.12165, -0.19573), and the optimal value is -0.158691.
[0166] x1 is the natural logarithm transformed value of the minimum balance of the business owner's deposit account at the point in time over the past 3 months, generated by the feature transformation step;
[0167] x2 is the natural logarithm transformed value of the minimum demand deposit balance of the enterprise over the past 3 months, generated by the feature transformation step;
[0168] x3 is the natural logarithm transformed value of the maximum monthly average credit transaction amount of the enterprise over the past 6 months, generated by the feature transformation step;
[0169] x4 is the cube root transformed value of the ratio of the company's average monthly debit transaction amount to its average monthly balance over the past 12 months, generated by the feature transformation step.
[0170] x5 is the square root of the ratio of the company's average monthly number of loan transactions in the past 3 months to the average monthly total number of loan transactions in the past 12 months, generated by the feature transformation step.
[0171] x6 is the cube root of the minimum monthly product of the company's current deposits over the past 6 months.
[0172] x7 is the square root conversion value of the current deposit balance percentile of the enterprise by industry.
[0173] 32. The method according to item 23, 27 or 31, wherein,
[0174] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to generate a credit score to characterize the borrower using the following formula:
[0175]
[0176]
[0177] Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 54.2458; B is 115.4156, the round function rounds the calculated score to the nearest integer; finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.
[0178] 33. An apparatus for constructing an inclusive credit risk prediction model, wherein the apparatus comprises:
[0179] The data acquisition module is used to acquire raw corporate credit prediction data for samples used to build the model;
[0180] The data derivation module is used to process the original corporate credit prediction data to generate derived corporate credit prediction data.
[0181] The feature screening module is used to perform preliminary screening on all categories, i.e. all features, including original corporate credit prediction data and derived corporate credit prediction data, to obtain the features after preliminary screening.
[0182] The initial screening data transformation module is used to determine the transformation method for the features after initial screening to confirm whether to use WOE transformation method, dummy feature transformation method, or continuous transformation method for feature transformation, and to use the optimal method for each feature after initial screening.
[0183] The feature refinement module is used to perform in-depth filtering on the features that have undergone initial feature transformation to obtain refined features.
[0184] The credit default probability modeling module is used to select logistic regression as the model construction method based on the probabilistic relationship between the refined feature combination and credit default, and to confirm the method used to calculate the credit default probability.
[0185] 34. The apparatus according to claim 33, wherein the apparatus performs the steps of the method for constructing an inclusive credit risk prediction model as described in any one of claims 1 to 10.
[0186] 35. A system for constructing an inclusive credit risk prediction model, wherein the system comprises: a memory, a processor, and a program for constructing an enterprise credit risk prediction model stored in the memory and executable on the processor, wherein when the program for constructing an enterprise credit risk prediction model is executed by the processor, it implements the steps of the method for constructing an inclusive credit risk prediction model as described in any one of items 1 to 10.
[0187] 36. An apparatus for calculating inclusive credit risk, comprising:
[0188] The data acquisition module is used to acquire enterprise credit prediction data for the sample to be predicted.
[0189] The module classifies the samples to be predicted, which is used to classify the samples to be predicted based on the decision tree method to determine the model used to calculate the probability of credit default.
[0190] The credit default probability calculation module is used to input corporate credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.
[0191] 37. The apparatus according to claim 36, wherein the apparatus performs the steps of the method for calculating inclusive credit risk as described in any one of claims 11 to 32.
[0192] 38. A system for calculating inclusive credit risk, wherein the system for calculating inclusive credit risk comprises: a memory, a processor, and a program for the method of calculating inclusive credit risk stored in the memory and executable on the processor, wherein when the program for calculating inclusive credit risk is executed by the processor, it implements the steps of the method for calculating inclusive credit risk as described in any one of items 11 to 32.
[0193] 39. A computer storage medium, wherein the computer storage medium stores a program for calculating inclusive credit risk, the program for calculating inclusive credit risk, when executed by a processor, implementing the steps of the method for calculating inclusive credit risk as described in any one of items 11 to 32.
[0194] Invention Effects
[0195] The method and system for constructing an inclusive credit risk prediction model in this application incorporate a large sample from a major financial institution when constructing the model. Customer data that has applied for inclusive credit business is selected from the sample as the original data. The original data obtained from the sample is processed and derived in depth. Advanced statistical analysis methods and interpretable machine learning techniques are used to construct a general scoring model for inclusive credit based on the characteristics of the original data and derived data and the strong financial attribute information contained in the data that is difficult to obtain in the market.
[0196] Furthermore, this application first utilizes decision trees to classify samples in the most reasonable way at the beginning of model construction, and then constructs sub-classification models based on the classified samples. Combining decision trees to construct financial models can effectively classify samples according to actionable categories, thereby better covering the credit risk characteristics of different customer groups and avoiding the problem of the model lacking subgroup representativeness due to using all samples to construct the model.
[0197] Furthermore, the use of stepwise discriminant analysis in the risk prediction model construction of this application significantly improves the overall efficiency of model development by employing a preliminary screening method for model features. The stepwise discriminant method can better identify more important feature variables within the same dimension, greatly reducing the workload required for developers to judge and screen variables one by one based on trends in the next step. This significantly improves model development efficiency without affecting the overall model performance, making the initial variable screening more efficient and accurate.
[0198] This application innovatively combines three methods in model construction: continuous transformation, WOE transformation, and dummy feature transformation. It further processes some features that have undergone initial screening, combines the advantages and disadvantages of the three transformation methods, and innovatively designs a transformation judgment method. Based on parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic judgment, the optimal feature transformation method is selected. This model construction method can avoid the technical shortcomings of overfitting when building a model with only WOE variables and the inability to well adapt to categorical variables when building a model with only continuous variables.
[0199] Furthermore, the method and system for calculating credit risk probability or scoring credit risk constructed using this application, namely the Zeta A model (Zeta A series scoring cards) (including Zeta_a 1, Zeta_a 2, and Zeta_a 3 sub-models or sub-scoring cards), first splits the model based on decision trees during its construction. Therefore, the initial calculation of credit risk involves effectively classifying customer samples to select the most suitable sub-model or sub-scoring card. At the same time, since the model construction method used in the sub-models or sub-scoring cards is the method described in this application, it avoids the technical shortcomings of overfitting when constructing models with complete WOE variables and the inability to well adapt categorical variables when constructing models with completely continuous variables. Therefore, it has a significantly better effect than existing models in predicting credit risk. In addition, since this application targets inclusive credit business, the sample data is customer data of enterprises that have applied for credit business selected from massive data. It not only covers predictive variables of Internet business, but also includes overdue performance data of Internet business. Therefore, the risk patterns extracted by using statistical principles through specific business scenario data have very significant adaptability and distinguishability in terms of credit risk of inclusive credit business. Attached Figure Description
[0200] Figure 1 This diagram illustrates whether the data used for binning is consistent with business trends.
[0201] Figure 2 The diagram illustrates the process of classifying all user samples based on a decision tree.
[0202] Figure 3 This is a schematic diagram illustrating the model differentiation effect of Embodiment 1 of this application. Detailed Implementation
[0203] Credit risk is the risk arising from changes in a borrower's economic capacity, reduced willingness to repay, or inability to fulfill loan obligations, rather than the risk of default due to intentional fraud. Credit defaults occur in various inclusive lending scenarios and are particularly related to the borrower's personal economic situation. The reasons for borrowers' credit delinquency can be mainly divided into four categories: 1. Short credit history: These borrowers have limited experience in financial management; 2. Temporary forgetting to repay; 3. Over-borrowing: These borrowers have relatively low repayment capacity due to a large amount of debt; 4. Impact of significant negative factors: These borrowers experience long-term impacts on their repayment ability due to factors such as reduced income, unemployment, or divorce. These different reasons can all lead to borrowers defaulting to varying degrees, potentially resulting in more serious default situations. Credit risk scoring aims to uncover the inherent mathematical relationship between a customer's historical information and the probability of future default, and to quantify the probability of default through a scoring system.
[0204] Currently, credit risk scoring technologies mainly utilize historical credit data and data from business and judicial authorities to develop and reflect payment behavior, willingness to pay, and corporate business and judicial information.
[0205] The sample data used to build the model in this application not only includes transaction behavior information of credit business, but also adds corporate credit data that is usually difficult to obtain. Corporate credit risk management usually needs to cover both corporate owner risk management and corporate risk management. However, it is difficult for general institutions to obtain information on both aspects at the same time, making it difficult to make a comprehensive prediction of inclusive credit business.
[0206] Inclusive lending refers to low-interest loans provided by banks or other financial institutions to individuals or micro and small enterprises that meet certain conditions. These loans aim to support economically disadvantaged groups, promote social equity, and drive economic development.
[0207] Inclusive credit risk refers to the risk arising from borrowers' failure to repay debts as agreed in inclusive lending businesses.
[0208] Solvency refers to a company's ability to repay its debts when they fall due. A company's ability to pay cash and repay debts is crucial to its healthy survival and development. Solvency is an important indicator reflecting a company's financial condition and operational capabilities. It represents a company's capacity or guarantee to repay its debts when they fall due, including the ability to repay both short-term and long-term debts.
[0209] The point in time refers to the end of each month. For example, the minimum balance in a business owner's bank account over the past three months refers to the minimum balance at the end of the past three months.
[0210] The difference between the number of months in which a business owner's bank account had the largest balance in the past 12 months and the current month refers to the difference between the month in which the account had the largest balance in the past 12 months and the current month. For example, if a business owner had the largest balance in March of the past 12 months, and the current month is May, then the difference is 2 months.
[0211] The accumulated amount refers to the balance multiplied by the number of days.
[0212] Quantiles are continuous and can be any value from 0 to 100%.
[0213] <Overall Description of Model Construction Methods>
[0214] Specifically, this application relates to a method for constructing an inclusive credit risk prediction model, comprising: a data acquisition step, which acquires original corporate credit prediction data for the sample used to construct the model; a data derivation step, which processes the original corporate credit prediction data into derived retail credit prediction data; a feature initial screening step, which performs preliminary screening on all categories, i.e., all features, including the original corporate credit prediction data and the derived corporate credit prediction data, to obtain the features after preliminary screening; a preliminary screening data transformation step, which judges the transformation method of the features after preliminary screening to determine whether to use WOE transformation, dummy feature transformation, or continuous transformation method for feature transformation, and performs feature transformation using the optimal method judged for each feature after preliminary screening; a feature fine screening step, which performs deep screening on the features after feature transformation to obtain the features after fine screening; and a credit default probability modeling step, which selects logistic regression to construct the model based on the probability relationship between the features after fine screening and credit default, and confirms the method used to calculate the credit default probability.
[0215] In one specific embodiment, the method for constructing an inclusive credit risk prediction model according to this application may further include a sample selection step before the data collection step. This step is used to screen all users to obtain samples for model construction before the data collection step. Those skilled in the art will understand that, based on the total number of users and data available for model construction, they may choose whether to exclude user data that does not meet the requirements for model construction. Furthermore, those skilled in the art may also initially select a subset of users as samples and then continuously increase the sample size for modeling according to appropriate rules.
[0216] The samples used to build the model in this application can be individual users, who may also be the actual controllers of corporate users, or the companies themselves. The data used to build the model samples can be data from individual and corporate users of large financial institutions. Such data not only includes transaction behavior information of credit business, but also adds asset data that is usually difficult to obtain. While reflecting the borrower's payment behavior and willingness to pay, it provides a more comprehensive assessment of their debt repayment ability and personal qualifications, and can provide more accurate prediction results.
[0217] In one specific implementation, the selected sample is a sample of customers who have applied for inclusive credit business within a certain time period.
[0218] In one specific implementation, behavioral data of enterprises and their actual controllers, loan application and behavior data, financial asset transaction data, and information data of enterprises and their actual controllers can be collected over a certain period. Subsequently, based on the collected sample data, some samples can be eliminated. Specifically, for example, the method in this application first collected raw enterprise credit prediction data from 2018 to 2021 to build the model sample. Then, through a professional model design scheme, the modeling sample was confirmed, and data from 15 million enterprises in 2019 were selected as the analysis sample.
[0219] <Sample Selection Steps>
[0220] In one specific implementation, the sample selection step involves screening all users before the data collection step to obtain samples for model building. Specifically, the sample selection step includes classifying all users based on a decision tree, with classification criteria including, but not limited to: whether a user is a customer who has already registered for credit services with a financial institution; whether a user is a customer with a financial institution but has not registered for credit services; the geographical region where a user conducts business; whether a user has experienced any financial institution risk events (e.g., whether there has been a current default, or whether there is a current mortgage); whether a user holds a credit card issued by a financial institution and / or whether the credit card is continuously used and / or whether the credit card or personal loan is being used on a revolving basis; whether a user holds a corporate deposit account; the aging period of a user's corporate deposit account; and the current balance of a user's corporate deposit account.
[0221] If the sample is the actual controller of an enterprise, i.e., an individual user, the sample selection steps include classifying all users in the sample based on a decision tree. The classification criteria include, but are not limited to: whether a user is a customer who has applied for and registered for credit business with a financial institution; whether a user is a customer who has not applied for and registered for credit business with a financial institution; the geographical area to which a user conducts business; whether a user has experienced a risk event with a financial institution; whether a user holds a credit card issued by a financial institution and / or whether the credit card is continuously used and / or whether the credit card or personal loan is used for revolving withdrawal.
[0222] Furthermore, taking a sample of individuals from a financial institution as an example, all users in the sample are classified based on a decision tree. Therefore, the sample is further divided into whether or not they have applied for and registered for corporate credit business with the relevant bank. In other words, users can be divided into two parts: those who have applied for registration and those who have not.
[0223] Credit business is a fundamental and crucial asset operation for banks. It involves issuing loans to recover principal and interest, and then generating profit after deducting costs. Generally, bank credit business is a vital means of profitability. From a classification perspective, bank credit business can be divided into corporate credit business and personal credit business. Corporate credit business includes project loans, working capital loans, small business loans, and real estate enterprise loans; personal credit business includes personal housing loans, personal consumer loans, and personal business loans.
[0224] In one specific implementation, the method for building the model in this application first divides the user sample into a sample of 1.35 million households that have applied for corporate credit business and a sample of 13.65 million households that have not applied for credit business. The specific model design is based on the analysis and design of the 1.35 million households that have applied. After the model is developed and launched, the results will be applied to the entire customer group of 18 million households that have applied and those that have not applied, and even to all customer groups that will use the services of financial institutions in the future.
[0225] Decision trees are a decision analysis method that, based on the known probabilities of various scenarios, constructs a decision tree to calculate the probability that the expected net present value is greater than or equal to zero, evaluating project risk and determining its feasibility. It is a graphical method that intuitively applies probability analysis. In this application, based on the characteristics of data from large financial institutions, decision trees are used to classify the samples used for modeling. After separating the samples into those with applied corporate credit business and those without, the samples with applied corporate credit business can be further classified according to whether they hold corporate deposit accounts. The samples with applied credit business are classified into two categories: those holding corporate deposit accounts and those with overdue payments. The decision regarding whether a corporate deposit account is held can be determined based on the needs of those skilled in the art. For example, based on the aging of the funds, the samples can be classified into customers holding corporate deposit accounts with an aging of more than m months and those holding... Customers with corporate deposits with an aging period of less than or equal to m months are divided into two categories: those with an aging period of less than or equal to m months and those with an aging period of more than m months. Further classification can be made based on the average balance of the corporate deposit accounts over the past n months for the sample of customers with an aging period of more than m months. The determination of the average balance over the past n months can be based on the needs of those skilled in the art. For example, the sample of customers with an aging period of more than or equal to m months can be divided into customers with an average balance of less than p yuan over the past n months and customers with an average balance of more than or equal to p yuan over the past m months.
[0226] In one specific implementation, m can be set to any integer greater than 1, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc., for example, 6.
[0227] In one specific implementation, n can be set to any integer greater than 1, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc., and n must be less than or equal to m, for example, n is 6.
[0228] In one specific implementation, p can be set to any value greater than 0, such as 1, 5, 10, 15, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, etc. For example, p can be 5000.
[0229] Users can also be categorized based on their spending habits after joining the network. For example, in one specific implementation, the modeling samples can be categorized based on whether a user holds a corporate deposit account.
[0230] Those skilled in the art will fully understand that other splitting methods can be chosen when splitting samples. For example, the sample can be split based on whether a user has experienced a financial institution risk event. The customer sample splitting method during model building can be considered and carried out according to the needs of modeling.
[0231] Furthermore, based on the specific details required to build the model, this application can further segment the samples according to the region to which the customer belongs. For example, if the customer has already applied for credit business, the sample can be further divided into different sample groups from economically underdeveloped regions, moderately developed regions, and relatively developed regions based on decision trees.
[0232] For another group that does not apply for credit business, the sample can be further divided into different sample groups from economically underdeveloped regions, moderately developed regions, and relatively developed regions based on decision trees. In this application, the definitions of economically underdeveloped regions, moderately developed regions, and relatively developed regions can be based on standards commonly recognized by those skilled in the art. These standards can be derived from data published by authoritative statistical departments, classifications made by a rating agency, or classifications based on separately constructed financial models. Those skilled in the art can fully understand that after selecting the classification criteria, the user group can be divided into three different regions without overlap.
[0233] In one specific approach, further classification can be based on regional development and bad debt rates to divide users into different sample groups from economically underdeveloped, moderately developed, and relatively developed regions.
[0234] In this application, decision-making can also be categorized based on whether a customer has defaulted. For example, for a sample of users who have applied for inclusive credit business, the current status of default behavior can be classified, such as into a group that has defaulted and a group that has not defaulted.
[0235] Furthermore, it should be noted that the decision-making and classification process described above is merely an example. Those skilled in the art can make decisions and classifications based on the sample used for modeling. For instance, in a specific implementation, classification can be based on whether a customer has a mortgage or the geographical region of the customer. The modeling method of this application does not impose any limitations on this. For example, customers can be divided into different sample groups based on whether they have applied for inclusive credit, whether they are overdue, whether they have a mortgage, whether they have inclusive loans, or the region where the customer is located. As long as the logic between each layer of the decision tree is reasonable, the order of these decision-making methods can be arbitrarily changed and combined. The sample selection step can fully reflect the control variables closely related to risk characteristics in the business process and can comprehensively match various customer groups in the market.
[0236] In addition, as mentioned above, this application can also skip the above sample selection step and directly build the model of this application based on all samples.
[0237] For a financial institution's enterprise user samples, multi-head data samples, and operator samples, those skilled in the art can fully understand that similar methods can be used to classify the samples in order to identify the samples used to build the model or sub-model.
[0238] In the process of constructing the model for this application, since it is necessary to build a model for inclusive credit business, the selected sample is customers who have applied for inclusive credit business within a certain period of time as the modeling sample.
[0239] In one specific implementation, 1.35 million enterprises that had applied for corporate credit services were selected from the 15 million enterprise data points in 2019 as the modeling sample. Such a sample is more targeted and representative when predicting the credit risk of inclusive lending.
[0240] <Original Credit Forecast Data>
[0241] The original credit forecast data includes original inclusive credit forecast data and corporate credit forecast data. Among them, the original data used to forecast corporate credit risk includes original corporate credit forecast data and original corporate controlling shareholder credit forecast data.
[0242] In this application, the original corporate credit forecast data refers to data related to the enterprise. However, in order to more reasonably predict corporate credit risk, when constructing a forecasting model for corporate credit risk, in addition to the relevant data of the enterprise itself, it is also necessary to add variable information of the actual controller of the enterprise. This is because the risk of the actual controller of the enterprise largely reflects the credit risk of the enterprise. A corporate credit risk forecasting model that includes data on the actual controller will be more accurate in predicting corporate credit risk.
[0243] The data types involved in the original inclusive credit forecast data or the original corporate actual controller credit forecast data are consistent.
[0244] When constructing the model used in this application, approximately 120 types of raw credit prediction data (i.e., more than 120 basic features) can be obtained based on the basic data of financial institutions. Furthermore, the inclusive credit risk points are split or processed to the greatest extent possible under the premise of optimal effect according to different dimensions, generating more than 3,000 derivative features.
[0245] When building the model, firstly, based on all historical data of large financial institutions, the original corporate credit prediction data for building the model sample was initially collected between 2018 and 2021, including various data information of enterprises and their actual controllers, totaling 18 million corporate users. Each enterprise has corresponding data to process every month. It can be seen that the data system used to build the model of this application is comprehensive and the amount of data is very large. Based on such a data system, the modeling methodology needs to be considered when building the model. Otherwise, it will get stuck in the huge amount of data, causing the special sample groups that need to be focused on to be covered in the huge amount of data and unable to be effectively identified, resulting in the computer program running slowly or even unable to run, so as to be unable to accurately build the most suitable prediction model.
[0246] In one specific implementation, the original corporate credit prediction data used to build the model is based on data from 15 million enterprises in 2019, representing over 100 types of corporate data and over 120 types of personal data (i.e., basic variables or basic features). These include, but are not limited to: corporate self-credit prediction data and corporate controlling shareholder credit prediction data. Corporate self-credit prediction data includes: 1) Basic customer information such as industry, size, and administrative region of business; 2) Corporate bank deposits such as balance and debit / credit transaction amounts. Corporate controlling shareholder credit prediction data includes: 1) Basic customer information such as gender, age, and administrative region of business; 2) Personal financial assets such as AUM (asset under management), deposits, wealth management products, and payroll information.
[0247] In the data collection step, the original corporate credit prediction data acquired for the sample used to build the model includes: basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions.
[0248] Basic enterprise information data refers to data based on the attributes of the sample (enterprise) users themselves, but which is not directly related to their behavior in financial institutions.
[0249] The basic data on the financial assets of business owners consists of all other financial assets and transactions held by the sample (actual controller of the enterprise) in financial institutions that are not related to credit cards and loans.
[0250] Basic information about business owners is based on the attributes of the sample (actual controller of the enterprise) users themselves, but is not directly related to their behavior in financial institutions.
[0251] In one specific implementation, the basic data for corporate deposits includes information such as the balance of corporate deposits and the amount of debit and credit transactions. The basic data is not limited to the specific categories listed. As corporate deposit business changes, those skilled in the art can further cover newly emerging data types in implementation. That is, all types of data that are available based on the user's deposit situation and behavior can be used as the basic data for corporate deposits.
[0252] In one specific implementation, the basic information data of an enterprise includes information such as industry, size, and administrative region to which the business is located. The basic data is not limited to the specific categories listed. As customer situations and social relationships develop, those skilled in the art can further cover other or newly emerging data types in the implementation within the scope of the relevant business application scenarios. That is, all types of data based on the attributes of the user sample itself but not directly related to the behavior in financial institutions can be used as basic information data of an enterprise.
[0253] In one specific implementation, the basic data of business owners' financial assets includes AUM, deposits, wealth management and payroll information. The above basic data is not limited to the specific categories listed. As financial assets change, those skilled in the art can further cover newly emerging data types in the implementation. That is, based on the sample, all other financial assets and financial transaction data that are not related to credit cards and loans in financial institutions can be used as the basic data of business owners' financial assets.
[0254] In one specific implementation, the basic information of business owners includes information on gender, age, and the administrative region where the business is located. The basic data is not limited to the specific categories listed. As the circumstances and social relationships of business owners develop, those skilled in the art can further cover other or newly emerging data types in the scope of the relevant business application scenarios. That is, all types of data based on the attributes of the user sample itself but not directly related to the behavior in financial institutions can be used as the basic information of business owners.
[0255] In one specific implementation, the original corporate credit prediction data obtained for the sample used to build the model is based on more than 120 data types (i.e., basic variables or basic features) obtained from 18 million corporate users, including but not limited to: deposit balance at a certain time, deposit account type, average daily balance, transaction type, transaction amount, business location, industry and other basic information.
[0256] <Derivative Enterprise Credit Forecast Data>
[0257] In this application, in the data derivation step, processing the original enterprise credit prediction data into derived enterprise credit prediction data refers to the data obtained by processing the collected original enterprise credit prediction data based on time dimension, spatial dimension, frequency dimension, and statistical information dimension.
[0258] In one specific implementation, the derivative enterprise credit prediction data includes, but is not limited to: derivative enterprise sales credit prediction data processed based on sample relationship length; derivative enterprise credit prediction data processed based on time interval variables; derivative enterprise credit prediction data processed based on sample behavior frequency; derivative enterprise credit prediction data processed based on the current time point of the sample; derivative enterprise credit prediction data processed based on continuous sample behavior; or, derivative enterprise credit prediction data processed based on statistical information dimensions. For example, monthly customer data can be obtained and processed based on monthly data. In this application, processing based on statistical information dimensions includes obtaining maximum, minimum, and average values of the data to describe the data situation.
[0259] In one specific implementation, for example, from the time dimension, customer relationship length-related variables such as customer account opening time and maximum customer account age can be considered as types of derivative corporate credit prediction data, that is, as derivative features or derivative variables.
[0260] In one specific implementation, for example, from the time dimension, the time interval is considered, such as the number of months since the customer's most recent repayment and the number of months since the customer's most recent overdue payment, as a derived feature or derived variable.
[0261] In a specific implementation, for example, starting from the frequency level, consider the frequency level variable of behavior: such as the number of times a customer has made repayments greater than N in the last X months, the number of times a customer's credit limit utilization rate has been greater than N in the last X months, etc., as derived features or derived variables. There is no limit to X, as long as the business logic is reasonable, it can be any positive integer greater than 0.
[0262] In one specific implementation, for example, from a time dimension, current point-in-time variables such as the customer's current monthly credit limit and the customer's current monthly balance can be considered as derived features or derived variables.
[0263] In one specific approach, from the perspective of statistical information, statistical variables are considered, such as the customer's maximum number of overdue periods in the most recent X months and the customer's average credit limit utilization rate in the most recent X months, as derived features or derived variables. There is no limit to X; as long as the business logic is reasonable, it can be any positive integer above 0.
[0264] In a specific implementation, for example, from the perspective of frequency, consider continuous behavioral variables: such as the maximum number of consecutive overdue payments > N in the customer's most recent X months, the number of times the customer's repayment rate is > N in the customer's most recent X months, etc., as derived features or derived variables. There is no limit to X, as long as the business logic is reasonable, it can be any positive integer above 0.
[0265] Those skilled in the art will understand that the methods described above for processing derived variables are merely examples and can be used arbitrarily. For instance, 120 types of basic data can be processed into more than 3,000 derived data. In this application, data, data type, data category, variable, or feature are sometimes used interchangeably, and those skilled in the art can understand this based on common statistical knowledge.
[0266] In this application, derived data can be derived from original inclusive credit prediction data through simple processing, or it can be derived from derived data through complex processing. Simply processed derived data, such as current monthly credit limit or current monthly balance, can be used directly after data aggregation. Complexly processed derived data requires time slicing and logical processing based on variables of the current category, and can generate derived variables such as the customer's maximum number of overdue periods in the last X months or the customer's average credit limit utilization rate in the last X months.
[0267] In one specific implementation, when constructing the model used in this application, the original credit prediction data of more than 100 types of enterprises and more than 120 types of individuals can be obtained based on the basic data of financial institutions (i.e., more than 100 types of original enterprise data and more than 120 types of original individual basic characteristics). Based on different dimensions of predicting credit risk, the data is split or processed to the maximum extent possible under the premise of optimal effect, resulting in a total of 427 enterprise derivative features and 143 individual derivative features.
[0268] <Preliminary Feature Screening>
[0269] In the initial screening data transformation step of this application, the determination of the transformation method for the initially screened features is based on the concentration and data type of the initially screened features. The initial screening data transformation step based on the determination of concentration and data type includes the following steps: classifying the data type of each feature into character variables and numerical variables; using a dummy feature transformation method for character variables for initial screening data transformation; and further classifying numerical variables including the following sub-steps: if the numerical variable has fewer than n values, using the WOE transformation method for initial screening data transformation; if the numerical variable has more than n values, further determining whether to convert to a continuous variable with many values and a concentration of a single value greater than m%, then using the WOE transformation method; if the concentration of a single value is less than or equal to m%, then using the continuous transformation method. Preferably, n and m are both positive integers, where n = 5 to 10 and m = 90 to 99.
[0270] For example, n can be 5, 6, 7, 8, 9 or 10, and m can be 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.
[0271] Specifically, this application employs multiple different feature screening methods to filter numerous features, thereby effectively selecting from the most dimensional features. Existing credit risk scoring often uses missing rate, concentration, and information value (IV) for feature screening.
[0272] In one specific approach, features are filtered based on the data missingness of each feature in the samples used to build the model. When filtering by missing rate, variables with a missing rate greater than 90%, 91%, 92%, 93%, 94%, or 95% are generally considered to be removed. For example, features with a data missing rate of more than 95% or more can be removed.
[0273] In one specific approach, features are filtered based on the high percentage of a single value in a particular feature sample. When filtering by concentration, variables with a single value percentage of 99%, 98%, 97%, 96%, or 95% or higher are generally considered for deletion. For example, features with a single value exceeding 99% or 95% can be removed.
[0274] In one specific approach, the Informative Value (IV) value for each feature is calculated to perform initial feature screening. The IV value can be used to measure the predictive power of a feature; a higher IV value indicates stronger predictive power. The calculation method for the IV value of a single feature is as follows:
[0275]
[0276] Where k is the number of groups after discretizing this feature; yi is the number of non-defaulting customers in the i-th group; ys is the total number of non-defaulting customers; ni is the total number of defaulting customers in the i-th group; and ns is the total number of defaulting customers. The quantitative index of IV has the following meanings: when the calculated value of IV is less than 0.02, it indicates that the predictive power of this feature is extremely weak; when the calculated value of IV is above 0.02 but less than 0.1, it indicates that the predictive power of this feature is weak; when the calculated value of IV is above 0.1 but less than 0.3, it indicates that the predictive power of this feature is relatively good; and when the calculated value of IV is above 0.3, it indicates that the predictive power of this feature is strong.
[0277] Of course, the threshold for deleting the calculated IV value can also be set to 0.03, 0.04, or 0.05, etc.
[0278] In the model construction method of this application, the steps of using missing rate, concentration, and information value (IV) for initial feature screening can be performed in any order. For example, screening can be performed first based on missing rate, then based on concentration, and finally based on information value (IV). Alternatively, screening can be performed first based on concentration, then based on missing rate, and finally based on information value (IV). Those skilled in the art can choose the appropriate method based on the sample data. Therefore, these three methods can effectively remove features with obvious defects in a certain aspect, thus effectively reducing data dimensionality and improving the model construction effect.
[0279] In one specific implementation, features that have been screened using missing rate, concentration, and information value (IV) are initially screened using a stepwise discriminant algorithm. Then, the initial screening of features is performed based on the consistency between the risk characteristics of each feature and the actual real results of the samples used for model construction.
[0280] Building upon this, the technical solution of this application introduces a stepwise discriminant analysis method to improve the overall efficiency of model development, making the initial variable screening more efficient and accurate. In real-world data, good and bad samples may have similar distributions on a certain variable, resulting in weak distinguishing ability. Alternatively, there may be a class of variables, each capable of effectively distinguishing between good and bad samples, but including all of them in the model would lead to redundancy due to overlapping data dimensions. To address this, this application uses a stepwise discriminant analysis method, employing the Wilks's Lambda value as the inclusion and exclusion statistic to remove variables with weak discriminative power or redundancy from the data.
[0281] The stepwise discriminant method uses the Wilks's Lambda criterion to measure the strength of features, eliminating those that do not meet the set threshold from the remaining features after three rounds of screening. In the stepwise discriminant process, the variable with the strongest discriminant power is added first. As the number of variables in the model gradually increases, the discriminant power of earlier introduced variables may also change. If the discriminant power of a variable in the model falls below the threshold, that variable is removed. This process is repeated until all variables in the model satisfy the Wilks's Lambda similarity ratio criterion, and no other variables meet the criteria for inclusion in the model.
[0282] In this application, the stepwise discriminant method for screening is a crucial step. To capture as many credit risk points as possible, the model involved in this application uses a large amount of data with numerous dimensions; therefore, variable screening is necessary to further reduce the time cost of model development. Existing financial risk model evaluations generally use variable importance methods for variable reduction, such as calculating variable importance using algorithms like the Gini index and information entropy, selecting variables with higher importance. Stepwise discriminant methods for screening are rarely used. Compared to commonly used variable importance screening schemes in the industry, the methodology used in this application can retain a large number of variables with relatively weak importance but relatively independent information dimensions.
[0283] In one specific approach, over 3000 derived data points can be generated from, for example, 120 basic data categories. After three rounds of screening based on missing rate, central tendency, and information value (IV), approximately 20-30% of the poorer features can be removed. However, more than 50 features will still remain. If subsequent variable refinement is performed based on all of these features, it will severely impact development efficiency. Therefore, after comparing various variable reduction schemes, stepwise discriminant analysis was ultimately determined to be the optimal choice.
[0284] For the important features after the stepwise discrimination algorithm, further screening of features is carried out based on the risk characteristics of various risk points themselves (system preset) and the actual bad debt rate of the sample. It is judged whether the actual bad debt rate distribution of the remaining features conforms to the business trend, and features whose actual bad debt rate distribution does not conform to the business trend are removed. Specifically, (1) the remaining target features are divided into 10-20 boxes according to the quantile, and the median value of each box and the corresponding bad debt rate are calculated. (2) the change rate (slope) is calculated using the median value of each box and the previous box and the corresponding bad debt rate. (3) the number of boxes with a change rate greater than 0 and non-0 (non-greater than 0 boxes) between adjacent boxes is counted, and the percentage of boxes with a change rate greater than 0 in the change rate is calculated to the percentage of boxes with a change rate greater than 0 in the change rate. (4) the preset risk characteristics of this feature are obtained, and based on the percentage of boxes with a change rate greater than 0 in the change rate calculated above to the percentage of boxes with a change rate greater than 0 in the change rate, whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend. The criteria for approximate consistency are as follows: First, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate increases (e.g., credit line utilization rate), then in this module, features where the percentage of boxes with a value greater than 0 in the above-calculated slope is less than 70% of the total number of boxes with a slope that is not zero are removed; that is, features whose behavior does not conform to the business trend are removed. Second, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate decreases (e.g., deposit amount), then in this module, features where the percentage of boxes with a value greater than 0 in the above-calculated slope is greater than 30% of the total number of boxes with a slope that is not zero are removed; that is, features whose behavior does not conform to the business trend are removed.
[0285] For example, the table below (Table 1) provides an example of binning. In this example, the characteristic values are divided into 10 bins, and the value ranges for each bin are summarized in the table below. A graph showing the change in the bad debt rate of each characteristic bin as a function of the median of the characteristic bins is also plotted based on the data in the table below. Figure 1 As shown. In Figure 1 In the example, there are 7 slopes greater than 0 and 2 slopes less than 0. Using the above method, features can be further filtered based on the risk characteristics of various risk points (system preset) and the actual bad debt rate of the sample. Utilizing this binning method for further feature filtering can effectively identify the features that best align with business development trends, thereby obtaining features suitable for modeling.
[0286] Table 1. Packing Information
[0287]
[0288] <Eigenvalue Transformation>
[0289] Current credit risk scoring methods primarily employ Word of Entity (WOE) transformation (multi-class classification) and dummy feature transformation (binary classification) to discretize continuous features (such as age, debt aging, etc.). WOE transformation uses the optimal binning scheme based on the modeling samples, discretizing continuous features according to the optimal cut-off point and embedding good and bad sample data into the WOE values. Therefore, it performs better in model building. However, the binning results may overfit the modeling samples, leading to a significant decline in model performance when applied to the overall population (poor generalization ability). Furthermore, because the normalization operation used in binning converts the original features falling into different bins into single values corresponding to each bin, it loses the ability to distinguish the risk of individuals falling within the same interval. Dummy feature transformation is mainly used for grouped features. Its advantage is that it can eliminate the distinction between good and bad features, but it becomes exceptionally complex and redundant when processing continuous variables.
[0290] In the model construction method of this application, the transformation method of the initially screened features is determined to confirm that one of the WOE transformation method, dummy feature transformation method, and continuous transformation method is used for feature transformation, and the optimal method is determined for each initially screened feature.
[0291] The initial screening data transformation steps, based on concentration and data type judgment, include the following steps: Classifying the data type of each feature into character variables and numerical variables; using dummy feature transformation for character variables; and further classifying numerical variables, including the following sub-steps: If the numerical variable has fewer than n values, using the Word of Entity (WOE) transformation method for initial screening; if the numerical variable has more than n values, further determining whether to convert to a continuous variable with many values and a concentration of a single value greater than m%, then using the WOE transformation method; if the concentration of a single value is less than or equal to m%, then using the continuous transformation method. Preferably, n and m are both positive integers, where n = 5–10 and m = 90–99.
[0292] Specifically, taking a character-type variable like education level as an example, the value of this feature variable can be primary school, middle school, university, postgraduate, etc. For a numerical variable, if the numerical variable is the number of overdue months in the past 3 months, then the value would be 0, 1, 2, or 3.
[0293] In this application, n can be 5, 6, 7, 8, 9 or 10, and m can be 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.
[0294] In one specific implementation, m is chosen as 5 and n is chosen as 95.
[0295] Specifically, the WOE transformation works as follows: The optimal cut point for the features is found; the original feature value range is segmented into multiple bins; the WOE transformation value for each bin is calculated and output based on the good / bad performance of each bin; and the original features are then divided according to the binning results, and the WOE transformation values are output. The WOE value for each bin is calculated as follows:
[0296]
[0297] Where yi is the number of customers in the i-th group who have not defaulted; ys is the total number of customers who have not defaulted; ni is the total number of customers in the i-th group who have defaulted; and ns is the total number of customers who have defaulted.
[0298] Let's take age as an example to illustrate the WOE binning process. The original feature, age, contains values for those aged 18 to 50. After binning, five bins are obtained: 18 to 24, 25 to 30, 31 to 35, 36 to 42, and 43 to 50. Then, the WOE transformation value for each bin is calculated based on the number of good and bad customers in each bin. Finally, each piece of user data assigned to each bin is output according to the corresponding transformation value for each bin. For example, a 23-year-old is output with the WOE transformation value corresponding to the 18 to 24 bin, and a 46-year-old is output with the WOE transformation value corresponding to the 43 to 50 bin.
[0299] The conversion method for dummy features is as follows: a single categorical feature is converted into an equal number of dummy features according to the number of values it contains. If a customer belongs to the value corresponding to the generated dummy feature, the corresponding dummy feature value is 1, and the values of the other dummy features are 0.
[0300] Let's take gender as an example to illustrate the dummy feature transformation method. The original features include: male and female. After dummy feature transformation, two dummy features are generated: 'gender-male' and 'gender-female'. If the customer's gender is male, then 'gender-male' is recorded as 1, and 'gender-female' as 0. Taking education level as an example, the original features include: associate degree or below, bachelor's degree, and master's degree or above. After dummy feature transformation, three dummy features are generated: 'education level-associate degree or below', 'education level-bachelor's degree', and 'education level-master's degree or above'. If the customer's education level is bachelor's degree, then 'education level-associate degree or below' is recorded as 0, 'education level-bachelor's degree' as 1, and 'education level-master's degree or above' as 0.
[0301] The continuous transformation method involves performing various continuous transformations on the original features (including but not limited to: directly selecting the original values, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, and calculating the natural logarithm of the original data). The correlation coefficient r between the transformed feature values and the overdue labels is calculated. The transformation method with the largest absolute value of the correlation coefficient is selected, and the transformation method corresponding to the original features is output. The formula for calculating the correlation coefficient is as follows:
[0302]
[0303] Where Σ is the summation symbol in mathematics; n is the total number of observations; and xi is the transformed value of the original feature of the i-th observation after continuous transformation. This is the average of the transformed values; where yi is the binary classification feature representing whether the i-th observation is in default; This is the average value of the binary classification features. The closer the absolute value of the correlation coefficient is to 1, the more relevant the transformed value is to the default situation, and the better the effect of this transformation method. The quantitative meaning of the correlation coefficient r is as follows: when the absolute value of the correlation coefficient is above 0 and less than 0.3, it indicates low correlation; when the absolute value of the correlation coefficient is above 0.3 and less than 0.8, it indicates moderate correlation; and when the absolute value of the correlation coefficient is above 0.8 and less than 1, it indicates high correlation.
[0304] In the model construction of this application, for features that are confirmed to adopt a continuous transformation method, the optimal transformation method is selected based on the correlation between the feature and credit default under different continuous transformation methods. Preferably, the continuous feature transformation is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
[0305] Let's take age as an example to introduce the continuous transformation method. The original feature age contains values from 18 to 50 years old. The continuous transformation yields the original age value, square, square root, cube root, natural logarithm, etc. Then, the WOE value and the absolute value of the correlation coefficient between each transformed value and the good / bad label (prediction result) are calculated. The transformation method with the largest absolute value of the correlation coefficient is selected for the output. If the cube root of the age feature has the largest absolute value of the correlation coefficient with the good / bad label compared to other transformation methods, the cube root of the age is output as the transformed value.
[0306] Existing technical solutions primarily employ one of two feature transformation methods—WOE (Word of Excellence) or dummy features—for model building. For example, Chinese patent CN112686749B uses WOE to transform feature values. WOE is an optimal binning scheme based on the modeling samples, discretizing continuous features according to the optimal cut-off point. Therefore, it performs better in model building. However, the binning results may overfit the modeling samples, leading to a significant decrease in model performance (poor generalization ability) when applied to the overall population. Furthermore, because the normalization operation used in binning converts the original features falling into different bins into single values corresponding to each bin, it loses the ability to differentiate the risk of individuals falling within the same interval.
[0307] Dumb feature transformation is mainly applied to grouped features. Its advantage is that it can eliminate the distinction between good and bad values of different features. For example, in the customer's industry, since there is no obvious superiority or inferiority between the retail and wholesale industries, the dummy feature processing method is more suitable; there is a hierarchical difference between junior college and university, so although dummy features can be used, the WOE processing method is actually more appropriate.
[0308] On the other hand, continuous transformation can avoid overfitting of the modeled samples in WOE transformation, has a stronger generalization ability to the overall sample, and because it does not perform interval mapping, it is less likely that most customer groups will fall into a single value. However, it cannot be applied to some features with poor monotonicity or discreteness (such as occupation, job title, etc.).
[0309] As described above, the technical solution of this application takes a different approach, innovatively combining three methods: continuous transformation, WOE transformation, and dummy feature transformation. It reprocesses some carefully selected features, combines the advantages and disadvantages of the three transformation methods, and innovatively designs a transformation judgment method. Based on parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic judgment, the optimal feature transformation method is selected.
[0310] <Logistic Regression and Deep Feature Selection Based on Logistic Regression>
[0311] In existing credit risk scoring, logistic regression models are mainly used for model development due to the requirement of model interpretability. The software that can be used is generally SAS, R, Python, etc.
[0312] In one specific implementation, this application uses SAS software for model development.
[0313] Specifically, logistic regression uses the sigmoid function to fit the probability of predicting default. The sigmoid function is:
[0314]
[0315] Where Z is a linear combination of model coefficients and feature transformation values, and Z is defined as follows:
[0316] Z = α + β1x1 + β2x2 + ... + β k-1 x k-1 +β k x k
[0317] The predicted probability of default is:
[0318] P=P(Y=1|x1,x2,x3,...,x k-1 x k )
[0319] The probability of the fitted prediction being a default is:
[0320]
[0321] From the above formula, we can further derive:
[0322] Substituting the Z value into the above formula allows us to calculate the probability P of predicting default.
[0323] The core of logistic regression model construction is feature selection, which involves the following steps: First, batch-select features based on missing rate, concentration, and information value (IV). Second, select all remaining features one by one based on whether they align with business trends, retaining those that correctly reflect the business trend. For example, if it's found that the default rate decreases as the loan balance increases, this feature is considered inconsistent with the business trend. In credit business understanding, a higher loan balance corresponds to a higher customer default risk exposure (EAD), and thus greater risk. This feature would be removed from the feature list. Third, use the stepwise regression function of logistic regression to remove less important features that are highly correlated with other features. Fourth, select coefficients based on the sign of the training coefficients and the business trend of the feature transformation values, retaining those whose coefficient signs align with business logic. In the Y-label, where 0 represents good customers and 1 represents bad customers, for a feature whose bad debt rate monotonically increases with the feature value (e.g., loan balance), its training coefficient should be positive; conversely, for a feature whose bad debt rate monotonically decreases with the feature value (e.g., deposit balance), its training coefficient should be negative. If these criteria are not met, the feature should be removed. Fifth, further remove highly correlated features using variance inflation factor (VIF) and correlation coefficient: for variance inflation factor, remove features with the highest VIF and greater than 4 one by one; for highly correlated features, remove features with lower IV values from the feature group with the highest correlation coefficient and greater than 0.80 one by one. Sixth, use the population stability index (PSI) to remove features with large differences in distribution at different time points, leading to instability. Features with PSI > 0.25 are directly removed; for features with 0.25 > PSI > 0.1, remove them cautiously based on the impact of removing these features on the model's discriminative ability.
[0324] This application strictly adheres to the above rules in feature selection, ensuring the model's interpretability, stability, and ability to distinguish between good and bad customers.
[0325] The feature screening steps of this application include: a first screening step, which uses a stepwise regression algorithm to screen features based on the significance of F-tests and T-tests; a second screening step, which calculates the variance inflation factor for each feature and removes features with high variance inflation factors; and a third screening step, which uses logistic regression to analyze whether the feature coefficients of the features after the first and second screening steps conform to the trend of the prediction results for credit defaults in order to further screen features.
[0326] The first step of feature refinement is based on a stepwise regression algorithm, using F-tests and T-tests. Features are introduced sequentially from highest to lowest significance. Each introduced feature is tested individually. A feature that becomes insignificant due to the introduction of subsequent features is removed. This process is repeated until no feature with a significance higher than the threshold is added to the equation, and no feature with a significance lower than the threshold is removed from the regression equation.
[0327] The second feature screening step further reduces multicollinearity in the model by eliminating features with high variance inflation factors.
[0328] The third step of feature refinement compares the risk characteristics (system preset) of various risk points with the positive or negative sign of the model training coefficients to determine whether the feature coefficients of the remaining features in the model conform to business trends. Features whose model coefficients do not conform to business trends are removed, and the process is iterated again. The specific implementation plan for the third step of feature refinement is as follows: 1. For features with a WOE (Word of the Exchange) transformation method, the corresponding model training coefficient should be negative, and WOE transformation features with positive training coefficients should be removed. 2. For continuous transformation methods, if an increase in the value of this feature in business logic should lead to an increase in the bad debt rate (e.g., credit line utilization rate), the corresponding model training coefficient should be positive, and continuous transformation features with negative training coefficients should be removed; if an increase in the value of this feature in business logic should lead to a decrease in the bad debt rate (e.g., deposit amount), the corresponding model training coefficient should be negative, and continuous transformation features with positive training coefficients should be removed. 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive; otherwise, it should be negative.
[0329] The fourth step in the feature screening process is the data stability monitoring step, which is used to evaluate whether there is a significant shift in the distribution of individual features and overall scores at different time points. In this embodiment, features with PSI > 0.25 will be directly removed, and features with 0.25 > PSI > 0.1 will be removed prudently based on the impact of removing these features on the model's discriminative ability.
[0330] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added to or removed from the model, at which point the iteration stops, yielding the final feature list and its transformed values. Through these steps, the input variables used to construct the model of this application can finally be obtained.
[0331] The credit default probability modeling step in this application involves substituting the features selected through the feature screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.
[0332] In existing technologies, the core of all scoring models lies in the representativeness of the data they use. Due to information security and cost considerations, most existing scoring models have small sample sizes and few bad labels, making it difficult to guarantee model stability and enabling them to independently model specific customer groups. Furthermore, because the information available for modeling in the market is mostly multi-borrowing data and data with weak financial attributes (such as data from smart terminal devices, social media platforms, online shopping malls, and other non-credit transaction data), it cannot accurately reflect a customer's asset status and repayment ability. This application uses full business data from a large bank, resulting in a massive amount of modeling samples and bad labels. This application designs modeling samples based on different sample groups, allowing for more refined differentiation of risk differences among these customers. Simultaneously, cross-time validation and PSI validation ensure the stability of the data and model. This application uses historical personal credit data and asset data for model development, thus providing a better reflection of borrowers' repayment willingness and ability.
[0333] For some small and medium-sized financial institutions, their internally built scoring models, due to the small scale of their retail lending business or the late start of the business, lack sufficient historical data to develop a stable and highly discriminative credit risk model. Therefore, they rely heavily on manual approval when conducting credit checks. The efficiency limitations of manual approval restrict the development of their retail lending business, and the subjectivity of manual approval increases operational risks in the credit check process. The model constructed in this application can assist such financial institutions in making digital decisions, enhancing the accuracy and speed of their approval processes, and mitigating the aforementioned adverse effects.
[0334] Regarding feature transformation, compared to the traditional WOE transformation method, which requires coarse binning and discretization of continuous features based on data performance and the modeler's experience, the results of coarse binning are greatly affected by the modeler's subjective factors. Furthermore, the discretization process of continuous features may result in too many customers falling into the same interval, leading to a large number of single-valued scores. This application combines continuous feature transformation, WOE transformation, and dummy feature transformation to encode the original features, reducing the impact of human factors and single-valued scores while enhancing the scoring's discriminative ability.
[0335] <Methods for calculating inclusive credit risk or methods for predicting the debt repayment capacity of inclusive loans>
[0336] This application further relates to a method for calculating the inclusive credit risk (i.e., the ability to repay inclusive credit) of a sample to be tested using the inclusive credit risk model (i.e., the model used to predict the probability of default of inclusive credit risk) constructed in this application.
[0337] In this application, inclusive credit risk refers to the risk that the tested sample will be unable to repay its debts as agreed. In a specific context, the probability of credit default refers to the probability of a loan becoming overdue, specifically, for example, the probability of a loan becoming overdue for more than 30 days.
[0338] This application relates to a method for calculating inclusive credit risk or for predicting the debt repayment capacity of inclusive credit, comprising: a data acquisition step, which acquires corporate credit prediction data of a sample to be predicted; a classification step, which classifies the sample to be predicted to determine a sub-model for calculating the probability of credit default; and a credit default probability calculation step, which substitutes the corporate credit prediction data into the credit default probability sub-model to calculate the probability of credit default of the sample to be predicted.
[0339] In this application, the credit default probability covers the credit default probability of various types of credit business.
[0340] In this application, the sample to be predicted can be a sample used to build the model, meaning that at the time of model construction, the sample was already a customer of a financial institution holding credit business (including but not limited to credit card or personal loan business). The method for calculating inclusive credit risk in this application can be used to calculate its future potential inclusive credit risk and the risk of its future potential debt repayment ability. Alternatively, the sample to be predicted can be a sample not used to build the model, meaning that at the time of model construction, it was not a customer of a financial institution holding credit business, but is now a customer. The method for calculating inclusive credit risk in this application can be used to calculate its future potential inclusive credit risk. Furthermore, the sample can also be a customer of a financial institution that was not, and is not currently, a customer of the financial institution holding credit business, but is a customer of other financial asset classes. The method for calculating inclusive credit risk in this application can be used to calculate its future potential inclusive credit risk.
[0341] This application relates to a method for calculating inclusive credit risk or for predicting the debt repayment capacity of inclusive loans, comprising: a data acquisition step, which acquires corporate credit prediction data of a sample to be predicted; a classification step, which classifies the sample to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; a credit default probability calculation step, which substitutes the corporate credit prediction data into the credit default probability sub-model to calculate the probability of credit default of the sample to be predicted; and a step of calculating the credit score of the sample to be predicted after calculating the probability of credit default.
[0342] The step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000.
[0343] In the method of this application, the corporate credit prediction data includes original corporate credit prediction data of the sample to be predicted and derived corporate credit prediction data processed based on the original corporate credit prediction data. The descriptions of the original corporate credit prediction data and the derived corporate credit prediction data are consistent with the descriptions in the model construction section.
[0344] In this application, the step of classifying the sample to be predicted includes the following sub-steps:
[0345] Does the sample to be tested hold a corporate deposit account?
[0346] Is the sample to be tested an overdue loan transaction?
[0347] Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model used to calculate the probability of credit default, and the order of the above sub-steps can be arbitrarily set.
[0348] In this application, Figure 2 The flowchart illustrates the classification of samples to be predicted. Based on the classification in this step, a sub-model can be selected for calculating the probability of credit default. Figure 2 The sample to be predicted is classified in the following order: First, determine whether the sample is a customer whose corporate deposit account has been outstanding for more than *a* months; then, determine whether the sample is a customer whose average corporate deposit account balance over the past *b* months is less than *c* yuan. Using this classification process, the most suitable sub-model for predicting the sample can be identified.
[0349] In this application, corporate deposits (also known as "corporate deposits") refer to RMB deposits made by enterprises, institutions, government agencies, military units, and social organizations in financial institutions, including time deposits, demand deposits, call deposits, negotiated deposits, and other deposits approved by the People's Bank of China.
[0350] Corporate deposit accounts include, but are not limited to, basic deposit accounts, general deposit accounts, temporary deposit accounts, special deposit accounts, time deposit accounts, notice deposit accounts, negotiated deposit accounts, etc.
[0351] In this application, the corporate credit prediction data is feature-transformed and then substituted into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted. The feature transformation step includes selecting WOE method or continuous method for feature transformation based on the feature type of the corporate credit prediction data to be substituted into the credit default probability sub-model.
[0352] The methods for feature transformation using WOE or continuous methods are well known to those skilled in the art, and specific transformation methods can be found in the description of the model construction section of this application. However, existing models in this field only use WOE for parameter transformation.
[0353] WOE (Word of the Edge) transformation is an optimal binning scheme based on the modeling samples. It discretizes continuous features according to the optimal cut-off point, thus performing better in model building. However, the binning results may overfit the modeling samples, leading to a significant decrease in model performance when applied to the overall population (poor generalization ability). Furthermore, because the normalization operation used in binning converts the original features falling into different bins into single values for each bin, it loses the ability to distinguish the risk of people falling into the same interval. Dumb feature transformation is mainly used for grouped features. Its advantage is that it can eliminate the distinction between good and bad values of different features. For example, in the customer's industry, since there is no obvious superiority or inferiority between retail and wholesale, dumb feature processing is more suitable; however, there is a hierarchical difference between vocational and undergraduate degrees, so although dumb features can be used, WOE processing is actually more appropriate. Continuous transformation avoids the overfitting of modeling samples that occurs with WOE transformation, has strong generalization ability for the overall sample, and because it does not perform interval mapping, most customer groups are less likely to fall into a single value. However, this approach is not applicable to features with poor monotonicity or discrete characteristics (such as occupation, job title, etc.). This application uses a combination of continuous transformation and WOE transformation to process the initial features, combining the advantages and disadvantages of both methods. A transformation judgment module is designed to select the optimal feature transformation method based on parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic. This application's technical solution takes a novel approach, innovatively combining continuous transformation, WOE transformation, and dummy feature transformation to reprocess carefully selected features. It combines the advantages and disadvantages of the three transformation methods and uniquely designs a transformation judgment method to select the optimal feature transformation method based on parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic. Substituting the feature transformation method determined in this way into the model for calculating inclusive credit risk can more accurately predict the inclusive credit risk of the sample.
[0354] The sub-model used in this application to calculate the probability of credit default is a model constructed based on the credit prediction data of sample enterprises and the probability of credit default using logistic regression based on an existing user group, that is, a model constructed using the method described in this application.
[0355] <Device, system, and computer storage medium for calculating inclusive credit risk>
[0356] This application relates to an apparatus for calculating inclusive credit risk, comprising: a data acquisition module for acquiring corporate credit prediction data of a sample to be predicted; a module for classifying the sample to be predicted using a decision tree method to classify the sample to be predicted to determine a model for calculating the probability of credit default; and a credit default probability calculation module for substituting the corporate credit prediction data into the credit default probability model to calculate the probability of credit default of the sample to be predicted.
[0357] The apparatus for calculating inclusive credit risk of this application can perform the steps of the method for calculating inclusive credit risk of this application.
[0358] This application also relates to a system for calculating inclusive credit risk, wherein the system for calculating inclusive credit risk includes: a memory, a processor, and a program for calculating inclusive credit risk stored in the memory and executable on the processor, wherein when the program for calculating inclusive credit risk is executed by the processor, it implements the steps of the method for calculating inclusive credit risk as described in this application.
[0359] This application relates to a computer storage medium storing a program for calculating inclusive credit risk, wherein the program for calculating inclusive credit risk, when executed by a processor, implements the steps of the method for calculating inclusive credit risk as described in this application.
[0360] All the contents described above in the method for calculating inclusive credit risk are fully applicable to the apparatus, system and computer storage medium for calculating inclusive credit risk.
[0361] The method for calculating inclusive credit risk in this application can avoid the technical shortcomings of overfitting when building a model with all WOE variables and the inability to well adapt to categorical variables when building a model with all continuous variables. Therefore, when this model is used to calculate inclusive credit risk, it can better cover the inclusive credit risk prediction needs of various types of customers and can meet the need to obtain more accurate prediction results when used as a general scoring method to predict credit risk.
[0362] Example
[0363] Example 1: Collection of Modeling Samples
[0364] In the construction of this embodiment, raw corporate credit prediction data from 2018 to 2021 was collected to build the model sample, including various data information of enterprises and their actual controllers, totaling 18 million corporate users. Further, the modeling sample was confirmed through a professional model design scheme, selecting 15 million corporate data from 2019 as the analysis sample. The user sample was then divided into 1.35 million users who had applied for corporate credit services and 13.65 million users who had not applied for credit services. The specific model design focused on the 1.35 million users who had applied. After the model is developed and launched, the results will be applied to the entire 18 million customer groups, both those who have applied and those who haven't, and even to all customers who will use financial institution services in the future.
[0365] The model design includes: 1) Exclusion rules: excluding customers with excessively large contract amounts, those who have settled their accounts, closed their accounts, written off their accounts, or have no performance data during the performance period; the modeled population after exclusion is 750,000 customers; 2) Time window setting: using 2019 data as the modeling sample, and a 24-month timeframe as the performance period for Y; 3) Sample sampling: using a good / bad sample ratio of 4:1 for modeling. Different segmentation schemes are designed using decision trees, and a comparison method between parent and child models is used to confirm the final segmentation scheme, followed by separate modeling for each segment. The model segmentation scheme in this application is designed based on whether customers hold corporate deposit accounts and whether corporate credit business is overdue, fully reflecting the control variables closely related to risk characteristics in the process of corporate business operations, and can better cover the corporate credit customer group in the market.
[0366] In this embodiment, the modeling samples are first divided into sub-models based on the decision tree to facilitate the construction of subsequent sub-models. The first layer of the decision tree, when determining the modeling samples, considers whether the corporate deposit aging is greater than 6 years, used to distinguish whether a customer holds corporate deposits, the length of their deposit history, and the richness of their deposit information. In this embodiment, for corporate customers with corporate deposit accounts aging greater than 6 years, they are further divided into Zeta_a 2 and Zeta_a 3 models based on their average monthly corporate deposit balance over the past 6 months (the division method is whether the average monthly corporate deposit balance over the past 6 months is less than 5000). Customers with corporate deposit accounts aging less than or equal to 6 years are assigned to the Zeta_a 1 model.
[0367] For customers with corporate deposit accounts aged over 6 months, if the average balance of corporate deposits over the past 6 months is less than 5,000 yuan, they will be classified into the Zeta_a 2 segmentation model; if the average balance of corporate deposits over the past 6 months is greater than 5,000 yuan, they will be classified into the Zeta_a 3 segmentation model.
[0368] Specifically, a sub-model used to build the Zeta_a 1 segmentation model has a sample size of about 90,000, and its customer base is mainly customers with corporate deposit accounts with an aging of less than or equal to 6 years. The model is built to predict the probability of them having a credit delinquency of more than 30 days.
[0369] Specifically, a sub-model used to build the Zeta_a 2 segmentation model has a sample size of about 70,000. Its customer base mainly consists of corporate deposit accounts with an age of more than 6 years and an average monthly balance of less than 5,000 in the past 6 months. The model is used to predict the probability of them being overdue for more than 30 days.
[0370] Specifically, a sub-model used to build the Zeta_a 3 segmentation model has a sample size of about 470,000. Its customer base mainly consists of corporate deposit accounts with an age of more than 6 years and an average monthly balance of corporate deposits of more than or equal to 5,000 over the past 6 months. The model is used to predict the probability of them being overdue for more than 30 days.
[0371] Based on the different customer group samples of each sub-model confirmed above, the original corporate credit prediction data used to construct the models includes: corporate credit prediction data and corporate controlling shareholder credit prediction data. Corporate credit prediction data includes: 1) basic corporate customer information, with basic fields including industry, size, and administrative region of business; 2) corporate bank deposits, with basic fields including corporate bank deposit balance and debit / credit transaction amounts. Corporate controlling shareholder credit prediction data includes: 1) basic customer information, including gender, age, and administrative region of business; 2) personal financial assets, including AUM, deposits, wealth management products, and payroll information.
[0372] Based on various risk points in credit business, macro credit risk is broken down according to information dimensions, time slices, and other methods. Variables are derived through professional characteristic variable construction methods, including: 1) Customer relationship length variables: such as customer account opening duration, maximum account age, etc.; 2) Time interval variables: such as the number of months since the customer's last repayment, the number of months since the customer's last overdue payment, etc.; 3) Behavioral frequency variables: such as the number of times the customer has made repayments greater than N in the last X months, the number of times the customer's credit limit utilization rate has exceeded N in the last X months, etc.; 4) Current point-in-time variables: the customer's current monthly credit limit, the customer's current monthly balance, etc.; 5) Statistical value variables: such as the maximum number of overdue periods in the last X months, the average credit limit utilization rate in the last X months, etc.; 6) Continuous behavior variables: such as the maximum number of consecutive overdue payments greater than N in the last X months, the number of consecutive repayment rates greater than N in the last X months, etc. This ultimately led to 3067 individual characteristic variables with potential predictive power for customer delinquency, such as 'current repayment amount', 'average number of overdue periods in the past 6 months', and 'number of credit card installment payments in the past 12 months'. It also led to 3557 corporate characteristic variables with potential predictive power for customer delinquency, such as 'average of overdue amount / total loan amount for the past 12 months' and 'average monthly loan transaction amount - average monthly debit transaction amount for corporate deposits'. In this paper, the observation point refers to the time point from the start of modeling when samples were collected, and 'current' refers to the sampling cutoff time; observation point and current have the same meaning.
[0373] Table 2 summarizes the enterprise-related basic and derived variables used in the embodiments.
[0374]
[0375] Example 2: Initial Screening of Features
[0376] A preliminary screening was conducted on the 427 enterprise-derived characteristics and 143 individual-derived characteristics (variables) collected in Example 1 that have the potential to predict customer delinquency.
[0377] The initial screening of features for the three sub-models Zeta_a1 to Zeta_a3 as defined in Example 1 was performed as follows:
[0378] In the first round of preliminary screening, features with a missing rate of over 95% in the data collected in Example 1 were removed, resulting in the deletion of 64 variables and the remaining 506 variables.
[0379] In the second round of preliminary screening, features with a single value exceeding 99% were removed from the 506 features that had passed the first round of preliminary screening. A total of 36 variables were removed, leaving 470 variables.
[0380] In the third round of preliminary screening, the 470 features selected after the second round of preliminary screening were sorted according to their numerical values (specifically, if the feature is a character variable, each value is placed in a separate bin; if the feature is a numerical variable, it is sorted by numerical value from smallest to largest). Then, the features were divided into 10-20 bins based on their quantiles. The feature IV value was calculated, and features with IV values below 0.02 were removed. In this embodiment, because the variables are strongly financial and asset-related, the predictive effect is strong; therefore, variables with IV values below 0.05 were removed, resulting in 86 features being removed, leaving 384 feature variables.
[0381] The fourth round of preliminary screening further filters the 384 features identified in the third round based on a stepwise discrimination algorithm. This round of screening quickly identifies 342 important feature variables.
[0382] The following describes the stepwise discriminant method. In real-world data, different categories may have similar mean values for a particular variable. In such cases, using this variable for classification will not be very effective. Another type of variable, when considered independently, can effectively distinguish different categories in the data, but including all of these variables in the model may become redundant. Therefore, this implementation innovatively employs a stepwise discriminant method to better select more important feature variables within the same dimension. This significantly reduces the workload of the fifth round of screening, which requires judging and selecting variables one by one based on their trends, thus greatly improving model development efficiency without affecting the overall model performance.
[0383] The stepwise discriminant method uses the Wilks's Lambda criterion to measure the strength of features, eliminating those that do not meet the set threshold from the remaining features after three rounds of screening. In the stepwise discriminant process, the variable with the strongest discriminant power is added first. As the number of variables in the model gradually increases, the discriminant power of earlier introduced variables may also change. If the discriminant power of a variable in the model falls below the threshold, that variable is removed. This process is repeated until all variables in the model satisfy the Wilks's Lambda similarity ratio criterion, and no other variables meet the criteria for inclusion in the model.
[0384] The fifth round of preliminary screening targets the 342 key features identified in the fourth round of screening. This screening is based on the risk characteristics of each risk point (system preset) and the actual bad debt rate of the sample.
[0385] Determine whether the actual bad debt rate distribution of the remaining features aligns with business trends, and remove features whose actual bad debt rate distribution does not conform to business trends. Specifically,
[0386] (1) Divide the remaining target features into 10-20 bins according to quantiles, and calculate the median value of each bin and the corresponding bad debt rate.
[0387] (2) Calculate the rate of change (slope) using the median of the values of each box and the previous box and the corresponding bad debt rate.
[0388] (3) Count the number of boxes with a change rate greater than 0 and boxes with a change rate of non-zero between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to boxes with a change rate of non-zero.
[0389] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of boxes with a value greater than 0 in the above-calculated change rate to the percentage of boxes with a value greater than 0 in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.
[0390] The criteria for approximate consistency are as follows:
[0391] First, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate increases (e.g., credit line utilization rate, etc.), then in this module, features where the percentage of boxes with a slope greater than 0 in the above calculation is less than 70% will be removed; that is, features whose performance does not conform to the business trend will be removed.
[0392] Second, if the business logic trend for this feature should be: as the value of this feature increases, the bad debt rate decreases (e.g., deposit amount, etc.), then in this module, features where the percentage of boxes with a slope greater than 0 in the above calculation is greater than 30% will be removed; that is, features whose behavior does not conform to the business trend will be removed.
[0393] After the fifth round of screening, 125 features were eliminated, leaving 217 features.
[0394] Example 3: Feature Conversion After Initial Screening
[0395] The feature evaluation step selects the optimal transformation method based on the feature concentration and data type observed in the initial feature screening step. When determining the transformation method, data types are generally divided into two categories: character variables and numeric variables. For character variables, dummy variable transformation is typically used. For numeric variables, if the variable has fewer than five values, Word of Entity (WOE) transformation is used; if the variable has many values, WOE or continuous transformation is used (the optimal transformation form is selected based on the correlation with the target variable). During this process, the concentration of variables needs to be considered. For example, if a continuous variable has many values, but the concentration of a single value exceeds 95%, continuous transformation is not performed, and WOE transformation is used directly. Features using different transformation methods are grouped and divided into different datasets. Specifically…
[0396] First, for the remaining 217 features in Example 2, the conversion method of these features is determined, and the following three methods are selected based on the concentration of features, data type, etc.
[0397] Feature conversion method 1 is used to convert features obtained from the data acquisition module and determined by the feature judgment module to be the optimal conversion method of WOE conversion into WOE conversion.
[0398] Feature transformation method 2 is used to perform dummy feature transformation on features that are determined by the feature judgment module to be the optimal transformation method in the data acquisition module.
[0399] Feature transformation method 3 is used to select the optimal continuous transformation method from the features determined by the feature judgment module in the data acquisition module and perform continuous transformation.
[0400] The feature merging module is used to horizontally concatenate the data from feature transformation method 1, feature transformation method 2, and feature transformation method 3.
[0401] In this embodiment, 217 features were transformed, including 54 features that underwent WOE transformation, 23 features that underwent dummy feature transformation (expanded to 43 variables), and 140 features that underwent continuous transformation.
[0402] Example 4: Feature Depth Screening (Feature Refinement Screening Steps)
[0403] The feature screening process in this embodiment mainly involves the following four steps: Steps 1 and 2 use stepwise regression and variance inflation factor calculation to remove features with high multicollinearity from the feature merging module, enhancing model robustness. Step 3 removes features whose training coefficients do not conform to business trends. Step 4 uses the Group Stability Index (PSI) to remove unstable features. This embodiment uses the LOGISTIC procedure in SAS for deep screening.
[0404] Feature refinement step 1, based on a stepwise regression algorithm and using F-tests and T-tests, introduces features in descending order of significance. Each introduced feature is tested sequentially. If a previously introduced feature becomes insignificant due to the introduction of subsequent features, it is removed. This process is repeated until no feature with a significance higher than the threshold is added to the equation, and no feature with a significance lower than the threshold is removed from the regression equation. After this step, 237 features are reduced to 126 features.
[0405] The second step of feature refinement further reduces multicollinearity in the model by eliminating features with high variance inflation factors. After this step, 126 features are reduced to 82 features.
[0406] The feature screening step 3 compares the risk characteristics of various risk points (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model conform to the business trend. Features whose model coefficients do not conform to the business trend are removed and the process is iterated again.
[0407] The specific implementation plan for feature-based fine screening module 3 is as follows:
[0408] 1. For features whose feature transformation method is WOE, the corresponding model training coefficient should be negative, and WOE transformation features with positive training coefficients should be removed.
[0409] 2. For continuous conversion methods, if an increase in the value of this feature in business logic should lead to a higher bad debt rate (e.g., credit line utilization rate), then the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be removed. If an increase in the value of this feature in business logic should lead to a lower bad debt rate (e.g., deposit amount), then the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be removed.
[0410] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive; otherwise, it should be negative.
[0411] The fourth step in the feature screening process is the data stability monitoring step, which is used to evaluate whether there is a significant shift in the distribution of individual features and overall scores at different time points. In this embodiment, features with PSI > 0.25 will be directly removed, and features with 0.25 > PSI > 0.1 will be removed prudently based on the impact of removing these features on the model's discriminative ability.
[0412] Based on steps 3 and 4 above, the feature list is iterated multiple times until no new features are added to or removed from the model, at which point the iteration stops, yielding the final feature list and its transformation values. After these steps, 64 features are removed as the final feature variables to be included in the model; for example, this could be 5, 6, 7, or 8 features.
[0413] Example 5: Construction of the Zeta_a 1 model
[0414] Generally, the stronger the correlation between the feature variable and the target variable, the higher the accuracy of the final model. In this embodiment, Zeta_a1, as a sub-model of the segmentation model, adopts the decision tree method described above. The sample size obtained from the classification is about 90,000. Its customer base mainly consists of corporate customers whose corporate deposit accounts have an aging of less than or equal to 6 years. The model is constructed to predict the probability of them having a credit overdue period of more than 30 days (target variable).
[0415] In Example 5, the final modeling result is described using five features as examples: the minimum balance of the business owner's deposit account at any given time in the past three months, the minimum balance of the business's current deposits in the past three months, the maximum monthly average loan transaction amount in the past six months, the ratio of the number of loan transactions to the total number of transactions in the past six months, and the ratio of the loan transaction amount to the total transaction amount in the past three months. In the feature transformation step, the six input variables have been transformed according to their correlation with the target variable. In the credit default probability modeling step, the six features selected in the feature screening step are substituted into the Sigmoid function (using the LOGISTIC procedure in SAS software) to perform logistic regression to calculate the credit default probability model.
[0416] The minimum balance in a business owner's bank account over the past three months reflects the borrowing company's asset level. For companies with corporate bank accounts aged 6 years or less, the smaller the balance in the business owner's bank account over the past three months, the weaker their debt repayment ability, and the higher the likelihood of serious defaults by their companies in the future. Conversely, the higher the balance, the lower the likelihood of future defaults. The conversion method is a natural logarithm transformation.
[0417] The minimum balance of a company's current accounts over the past three months reflects its asset level. For companies with corporate deposit accounts aged 6 years or less, the smaller the minimum balance of their current accounts over the past three months, the lower their asset level, the lower their solvency, and the higher the probability of future serious defaults. Conversely, the higher the minimum balance, the lower the probability of future serious defaults. The conversion method is a cube root conversion within continuous conversions.
[0418] The maximum average monthly credit transaction amount for a company over the past 6 months reflects the inflow of funds into its corporate settlement account. For companies with corporate deposit accounts aged 6 years or less, the higher the inflow amount in the settlement account over the past 6 months, the better their operating level and asset condition, and the lower the probability of serious default in the future; conversely, the lower the inflow amount, the higher the probability of serious default in the future. The conversion method is square root conversion.
[0419] The ratio of a company's number of credit transactions to the total number of transactions in the past six months reflects the frequency of transactions in the company's settlement account over the past six months. For companies with corporate deposit accounts aged 6 months or less, the higher the proportion of inflow transactions in the settlement account to the total number of transactions in the past six months, the more assets the company has received in the past period, suggesting a potentially better operating condition, stronger debt repayment ability, and a lower probability of serious default in the future. Conversely, a lower proportion of inflow transactions increases the probability of serious default in the future. The conversion method is dummy variable conversion.
[0420] The ratio of a company's credit transaction amount to its total transaction amount over the past three months reflects the company's settlement account transaction activity over the past three months. For companies with corporate deposit accounts aged 6 months or less, the higher the proportion of inflow transactions in the settlement account to the total transaction amount over the past three months, the more assets the company has received in the past period, potentially indicating better operating conditions, stronger debt repayment ability, and a lower likelihood of serious default in the future. Conversely, a lower proportion indicates a higher likelihood of serious default in the future. The conversion method is a natural logarithm transformation.
[0421] In the step of calculating the probability of credit default, based on the data and coefficients obtained in the feature transformation step, the following model formula is used to predict the probability (P) that the borrower will default:
[0422]
[0423] Where k is the number of features entering the model, and k is 6 in Formula 1.
[0424] α is the intercept term, with a value range of (0.71157, 0.68367), and the optimal value is 0.697618;
[0425] β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, with a value range of (-0.03417, -0.04716) and an optimal value of -0.040669;
[0426] β2 is the coefficient corresponding to the minimum balance of a company's demand deposits over the past three months, with a value range of (-0.01744, -0.07704) and an optimal value of -0.047235.
[0427] β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months, with a value range of (0.00201, -0.08893) and an optimal value of -0.043456;
[0428] β4 is the coefficient corresponding to the ratio of the number of credit transactions of an enterprise in the past 6 months to the total number of transactions in the past 6 months, with a value range of (0.08153, -0.06307) and an optimal value of 0.009227;
[0429] β5 is the coefficient corresponding to the ratio of a company's credit transaction amount to its total transaction amount in the past three months, with a value range of (-0.52827, -1.04399), and an optimal value of -0.786134. (Note: The value range is from the 95% confidence interval, i.e., the 95% CI in the table below)
[0430] x1 is the natural logarithm of the minimum balance of the business owner's deposit account at any point in time over the past 3 months, generated by the feature transformation step; x2 is the cube root of the minimum current deposit balance of the business over the past 3 months, generated by the feature transformation step; x3 is the square root of the maximum monthly average credit transaction amount of the business over the past 6 months, generated by the feature transformation step; x4 is the dummy variable transformation of the ratio of the number of credit transactions to the total number of transactions over the past 6 months, generated by the feature transformation step; x5 is the natural logarithm of the ratio of the credit transaction amount to the total transaction amount over the past 3 months, generated by the feature transformation step. The model representation of the features is shown in Table 3 below.
[0431] Table 3
[0432]
[0433]
[0434] The p-values for all model features were less than 0.05, indicating that the above features were significantly correlated with default performance.
[0435] Example 6: Construction of the Zeta_a 2 model
[0436] For Zeta_a 2 as a sub-model of the segmentation model, the above decision tree method was adopted. The sample size obtained by classification is about 70,000. Its customer base is mainly corporate customers whose corporate deposits are greater than 6 and whose average monthly balance of corporate deposits in the past 6 months is less than 5,000. A model was built to predict the probability of their credit overdue for more than 30 days (target variable).
[0437] In this embodiment 6, the final modeling result is described using six features as examples: the company's current monthly deposit balance, the business owner's current monthly AUM value, the ratio of the company's debit transaction amount to the total transaction amount of the past three months, the business owner's current balance, the business owner's minimum deposit account balance in the past six months, and the difference between the number of months with the maximum deposit account balance in the past twelve months and the current month. In the feature transformation step, the six input variables have been transformed according to their correlation with the target variable. In the credit default probability modeling step, the six features selected in the feature screening step are substituted into the Sigmoid function (using the LOGISTIC procedure in SAS software) to perform logistic regression to calculate the credit default probability model.
[0438] The current monthly deposit balance of a company reflects its current asset level. For clients whose corporate deposit accounts have been outstanding for more than 6 months and whose average monthly balance over the past 6 months is less than 5,000, the higher the company's current deposits, the higher its asset level, the more assets it can use to fulfill its credit obligations, and the stronger its repayment ability. This reduces the likelihood of future credit defaults, while conversely, it increases the likelihood of serious future defaults. The conversion method is a natural logarithmic transformation.
[0439] The current monthly AUM value of a business owner reflects the overall asset level of the company's controlling shareholder. For clients whose corporate deposit accounts have been outstanding for more than 6 months and whose average monthly balance over the past 6 months is less than 5,000, their own corporate asset level is relatively low. The higher the overall asset level of the controlling shareholder, the more assets they have available for their companies to fulfill credit repayment obligations, the stronger their repayment ability, and the lower the likelihood of future credit defaults. Conversely, the likelihood of serious defaults increases. The conversion method is square root conversion.
[0440] The ratio of a company's debit transaction amount to its total transaction amount over the past three months reflects the borrower's recent cash flow situation. For clients with corporate deposit accounts older than six months and an average monthly balance of less than 5,000 over the past six months, a higher ratio of debit transaction amount to total transaction amount over the past three months indicates more cash outflows in recent transactions, while a lower ratio of cash inflows suggests greater cash needs, fewer assets available for credit repayment, and a higher likelihood of serious future default. Conversely, a lower ratio indicates a lower likelihood of serious future default. The conversion method is a dummy variable conversion within a continuous conversion model.
[0441] The current deposit balance of a business owner reflects the level of deposit assets held by the actual controller of the enterprise. For clients whose corporate deposit accounts have been outstanding for more than 6 years and whose average monthly balance over the past 6 months is less than 5,000, their own asset level is relatively low. The higher the overall asset level of the actual controller, the more deposits they can use to fulfill credit repayment obligations for their companies, the stronger their repayment ability, and the lower the likelihood of future credit defaults. Conversely, the likelihood of serious defaults increases. The conversion method is a natural logarithmic transformation.
[0442] The minimum balance in a business owner's bank account over the past six months reflects the level of short- to medium-term deposit assets held by the company's controlling shareholder. For clients whose corporate bank accounts are older than six months and whose average monthly balance over the past six months is less than 5,000, their corporate asset level is relatively low. The higher the controlling shareholder's short- to medium-term asset level, the more deposits they have available for their businesses to fulfill credit repayment obligations, resulting in stronger repayment ability and a lower likelihood of future credit defaults. Conversely, a lower level of short- to medium-term assets increases the likelihood of serious future defaults. The conversion method is a natural logarithmic transformation.
[0443] The difference between the number of months with the highest balance in a business owner's bank account over the past 12 months and the current month reflects the changes in the business owner's bank account balance over a period of time. For clients with corporate bank accounts older than 6 years and an average monthly balance of less than 5,000 over the past 6 months, the number of months since the highest balance in the past 12 months reflects the changes in the business owner's assets. The larger the number of months since the highest balance, the lower the business owner's recent asset level, the greater the possibility of a recent decline in their asset situation, the lower their repayment ability, and the higher the probability of serious default. Conversely, the lower the number of months since the highest balance, the lower the probability of serious default. The conversion method is dummy variable conversion.
[0444] In the step of calculating the probability of credit default, based on the data and coefficients obtained in the feature transformation step, the following model formula is used to predict the probability (P) that the borrower will default:
[0445]
[0446] Where k is the number of features entering the model, and k is 6 in Formula 1.
[0447] α is the intercept term, with a value range of (1.18368, 0.87527) and an optimal value of 1.029473; β1 is the coefficient corresponding to the company's current monthly deposit balance, with a value range of (-0.07368, -0.11077) and an optimal value of -0.092225; β2 is the coefficient corresponding to the company owner's current monthly AUM value, with a value range of (-0.07489, -0.07929) and an optimal value of -0.077093; β3 is the coefficient corresponding to the ratio of the company's debit transaction amount to its total transaction amount in the past 3 months, with a value range of (0.44374, 0.34763). The optimal value is 0.395682; β4 is the coefficient corresponding to the current deposit balance of the business owner, with a value range of (-0.04003, -0.06985), and an optimal value of -0.054940; β5 is the coefficient corresponding to the minimum deposit account balance of the business owner in the past 6 months, with a value range of (-0.01585, -0.03037), and an optimal value of -0.023109; β6 is the coefficient corresponding to the difference between the number of months with the maximum deposit account balance in the past 12 months and the current month, with a value range of (-0.01482, -0.08255), and an optimal value of -0.048686. (Note: The value ranges are from the 95% confidence interval, i.e., the 95% CI in the table below)
[0448] x1 is the natural logarithm of the company's current monthly deposit balance generated by the feature transformation step; x2 is the square root of the current monthly AUM value generated by the feature transformation step; x3 is the dummy variable of the ratio of the company's debit transaction amount to the total transaction amount in the past 3 months generated by the feature transformation step; x4 is the cube root of the company's current point-in-time deposit balance generated by the feature transformation step; x5 is the cube root of the minimum balance of the company's deposit account in the past 6 months generated by the feature transformation step; x6 is the dummy variable of the difference between the number of months with the maximum balance in the company's deposit account in the past 12 months and the current month generated by the feature transformation step. The model representation of some features is shown in Table 4 below.
[0449] Table 4
[0450]
[0451] The p-values for all model features were less than 0.05, indicating that the above features were significantly correlated with default performance.
[0452] Example 7: Construction of the Zeta_a3 model
[0453] For Zeta_a 3 as a sub-model of the segmentation model, the above decision tree method was adopted, and the sample size obtained by classification was about 470,000. Its customer base is mainly corporate customers whose corporate deposit accounts have an age of more than 6 months and whose average monthly balance of corporate deposits in the past 6 months is greater than or equal to 5,000. A model was built to predict the probability of them having a credit overdue period of more than 30 days (target variable).
[0454] In this embodiment 7, the seven input variables were ultimately transformed accordingly: the minimum balance of the business owner's deposit account at a given point in time over the past three months; the minimum balance of the business's current deposits over the past three months; the maximum monthly average credit transaction amount over the past six months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past three months; the minimum monthly product of the business's current deposits over the past six months; and the quantile of the business's current month's deposit balance by industry. In the credit default probability modeling step, the seven features selected through the feature screening step were substituted into the Sigmoid function (using the LOGISTIC procedure in SAS software) to perform logistic regression to calculate the credit default probability model.
[0455] The minimum balance in a business owner's bank account over the past three months reflects the level of short- to medium-term deposit assets held by the company's controlling shareholder. For clients whose corporate bank accounts have been outstanding for more than six months and whose average monthly balance over the past six months is greater than or equal to 5,000, their corporate asset level is relatively low. The higher the short- to medium-term asset level of the controlling shareholder, the more deposits they have available for their companies to fulfill credit repayment obligations, resulting in stronger repayment ability and a lower likelihood of future credit defaults. Conversely, a lower level of short- to medium-term assets increases the likelihood of serious future defaults. The conversion method is a natural logarithmic transformation.
[0456] The minimum balance of a company's current accounts over the past three months reflects its recent liquidity situation. For clients with corporate deposit accounts older than six months and an average monthly balance of 5,000 or more over the past six months, a higher recent current account balance indicates more liquid assets, more funds available to repay credit debts, and a lower likelihood of serious default. Conversely, a lower balance increases the likelihood of serious default. The conversion method is a natural logarithmic transformation.
[0457] The maximum average monthly loan transaction amount for a company over the past 6 months reflects the borrower's transaction activity. For clients with corporate deposit accounts older than 6 years and an average monthly balance of 5,000 or more over the past 6 months, a larger loan transaction amount and greater cash inflow over a period of time indicate more assets available for repaying credit transactions, stronger repayment ability, and a lower likelihood of serious default. Conversely, a smaller loan transaction amount and a smaller average monthly balance increase the likelihood of serious default. The conversion method is a natural logarithm transformation.
[0458] The ratio of a company's average monthly debit transaction amount to its average monthly balance over the past 12 months reflects the borrower's financial transactions over the past year. For clients with corporate deposit accounts older than 6 years and an average monthly balance of corporate deposits greater than or equal to 5,000 over the past 6 months, a higher ratio of debit transaction amount to total transaction amount over the past 12 months indicates more recent cash outflows, while a lower ratio of cash inflows suggests greater potential funding needs, fewer assets available for credit repayment, and a higher likelihood of future serious defaults. Conversely, a lower ratio indicates a lower likelihood of future serious defaults. The conversion method is a cube root conversion within continuous conversion.
[0459] The ratio of the average monthly number of credit transactions over the past 3 months to the average monthly total number of credit transactions over the past 12 months reflects the changes in the borrowing company's transactions over a period of time. For clients whose corporate deposit accounts have an aging of more than 6 months and whose average monthly balance over the past 6 months is greater than or equal to 5,000, the higher the ratio of the average monthly number of credit transactions over the past 3 months to the average monthly total number of credit transactions over the past 12 months, the more frequent the recent cash inflow transactions, the better the company's asset condition, the better its repayment ability, and the lower the probability of serious default. Conversely, the lower the ratio, the higher the probability of serious default. The conversion method is a square conversion within the continuous conversion.
[0460] The minimum monthly accumulated balance of a company's current accounts over the past 6 months reflects its short- to medium-term deposit situation. For clients with corporate deposit accounts older than 6 months and an average monthly balance of corporate deposits greater than or equal to 5,000 over the past 6 months, a smaller monthly accumulated balance of current accounts over the past 6 months indicates a lower overall asset level, weaker repayment ability, and a higher probability of serious default. Conversely, a larger accumulated balance indicates a lower probability of serious default. The conversion method is cube root conversion.
[0461] The percentile of a company's current month's deposit balance by industry reflects its deposit level within the same industry. For clients with corporate deposit accounts older than 6 months and an average monthly balance of 5,000 or more over the past 6 months, a higher percentile of the current remaining months' deposit balance by industry indicates a higher deposit ranking within the industry, better overall asset condition, stronger repayment ability, and a lower likelihood of serious default. Conversely, a lower percentile increases the likelihood of serious default. The conversion method is the square root method.
[0462] In the step of calculating the probability of credit default, based on the data and coefficients obtained in the feature transformation step, the following model formula is used to predict the probability (P) that the borrower will default:
[0463]
[0464] Where k is the number of features entering the model, and k is 5 in Formula 1.
[0465] α is the intercept term, with a value range of (4.90128, 4.79548), and the optimal value is 4.848383;
[0466] β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, with a value range of (-0.04771, -0.07097) and an optimal value of -0.059338;
[0467] β2 is the coefficient corresponding to the minimum balance of a company's demand deposits over the past three months, with a value range of (-0.04223, -0.04769) and an optimal value of -0.044961.
[0468] β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months, with a value range of (0.00699, -0.01024) and an optimal value of -0.001625;
[0469] β4 is the coefficient corresponding to the ratio of the company's average monthly debit transaction amount to the average monthly balance over the past 12 months, with a value range of (0.01636, 0.01505) and an optimal value of 0.0157047.
[0470] β5 is the coefficient corresponding to the ratio of the average number of monthly credit transactions of an enterprise in the past 3 months to the total number of monthly credit transactions in the past 12 months. The value range is (-0.29326, -0.34022), with the optimal value being -0.316739.
[0471] β6 is the coefficient corresponding to the quantile of the current month's deposit balance of enterprises by industry, with a value range of (-0.00684, -0.00892), and the optimal value is -0.007878;
[0472] β7 is the coefficient corresponding to the minimum monthly accumulation of a company's demand deposits over the past 6 months, with a value range of (-0.12165, -0.19573) and an optimal value of -0.158691. (Note: The value range is from the 95% confidence interval, i.e., the 95% CI in the table below)
[0473] x1 is the natural logarithm of the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, generated by the feature transformation step; x2 is the natural logarithm of the minimum balance of the business's current deposits over the past 3 months, generated by the feature transformation step; x3 is the natural logarithm of the maximum monthly average credit transaction amount over the past 6 months, generated by the feature transformation step; x4 is the cube root of the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months, generated by the feature transformation step; x5 is the square of the ratio of the average monthly number of credit transactions to the average total number of credit transactions over the past 12 months, generated by the feature transformation step; x6 is the cube root of the minimum monthly product of the business's current deposits over the past 6 months; x7 is the quantile of the business's current deposit balance by industry. The model representation of the features is shown in Table 5 below:
[0474] Table 5
[0475]
[0476] The p-values for all model features were less than 0.05, indicating that the above features were significantly correlated with default performance.
[0477] Example 8
[0478] The P-values calculated from the above formulas can be used to further calculate the rating of any customer.
[0479] The scoring step is used to convert the calculated credit default probability into a score of 0-1000 using a pre-stored default score conversion code.
[0480] In the scoring module, a credit score representing the borrower is generated using the following formula:
[0481]
[0482]
[0483] Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 54.2458; B is 115.4156, the round function rounds the calculated score to the nearest integer; finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.
[0484] In this embodiment, the P-values calculated in the above embodiments can be substituted to calculate the score.
[0485] The KS (Kolmogorov-Smirnov) statistic was proposed by two Soviet mathematicians, AN Kolmogorov and NVSmirnov. In risk control, KS is often used to evaluate the model's discrimination ability. The higher the discrimination ability, the stronger the model's risk ranking ability.
[0486] The KS statistic is based on the Empirical Cumulative Distribution Function (ECDF) and is generally defined as follows:
[0487] KS=max(|cum(bad_rate)-cum(good_rate)|)
[0488] Based on the prediction model, the KS curve was plotted. The KS of the zeta_a model was 52.31, indicating that the model constructed in this embodiment performs excellently in assessing customer credit risk. Figure 3 An example with a KS value of 52 was given in the KS curve graph.
[0489] Table 6 shows the KS values of the model applied to the development and reserved validation samples. As can be seen from Table 6, the model performs well in both the training and validation sets (where the ratio of training to validation samples is 6:4) and in the overall sample set; that is, it demonstrates excellent discrimination performance across all sub-models.
[0490] Table 6
[0491] Classification Development Sample Validation Sample Development + Validation Samples zeta_a 52.31 52.53 52.39 zeta_a1 40.38 39.58 39.91 zeta_a2 44.83 44.88 44.82 zeta_a3 43.33 43.59 43.49
Claims
1. A method for constructing an inclusive credit risk prediction model, comprising: The data acquisition step involves obtaining raw corporate credit prediction data for samples used to build the model; The data derivation step involves processing the original corporate credit prediction data to generate derived corporate credit prediction data. The feature screening step involves preliminary screening of all categories, including both original and derived corporate credit prediction data, to obtain the features after preliminary screening. The initial screening data transformation step involves determining the transformation method for the features after initial screening to confirm whether to use one of the following methods: WOE transformation, dummy feature transformation, or continuous transformation. For each feature after initial screening, the optimal method is determined to be used for feature transformation. The feature refinement step involves performing a deep screening on the features that have undergone initial feature transformation to obtain refined features. The steps for modeling the probability of credit default are as follows: based on the probabilistic relationship between the refined features and credit default, logistic regression is selected to build the model, and the method used to calculate the probability of credit default is confirmed.
2. The method according to claim 1, wherein, In the data acquisition step, the raw corporate credit prediction data obtained for the samples used to build the model includes: The basic data on corporate deposits is based on all available data regarding RMB deposits made by sample (corporate) users at financial institutions. Basic enterprise information data refers to data based on the attributes of the sample (enterprise) users themselves, but which is not directly related to their behavior in financial institutions. The basic data on the financial assets of business owners consists of all other financial assets and transactions held by the sample (actual controller of the enterprise) in financial institutions that are not related to credit cards and loans. Basic information about business owners is based on the attributes of the sample (actual controller of the enterprise) users themselves, but is not directly related to their behavior in financial institutions.
3. The method according to claim 1, wherein, In the data derivation step, the process of processing the original enterprise credit prediction data into derived enterprise credit prediction data refers to the data obtained by processing the collected original enterprise credit prediction data based on time dimension, spatial dimension, frequency dimension, and statistical information dimension. Preferably, the derived corporate credit forecast data includes, but is not limited to: Derivative corporate credit prediction data obtained by processing based on sample relationship length. Derivative corporate credit prediction data obtained by processing time interval variables. Derivative corporate credit prediction data is obtained by processing sample behavior frequency. Derivative corporate credit prediction data is obtained by processing the data based on the current situation of the sample at the current point in time. Derivative corporate credit prediction data obtained by processing the continuous behavior of samples. Derivative corporate credit prediction data is obtained by processing sample data based on statistical information dimensions.
4. The method according to any one of claims 1 to 3, wherein, The initial feature screening process includes the following steps: The first initial screening step involves filtering features based on the data missing information for each feature of the samples used to build the model. The second initial screening step involves filtering features based on whether a single value of a particular feature is excessively high. The third initial screening step involves calculating the information value (IV) of each feature to perform preliminary screening of the features. The order of the first, second, and third preliminary screening steps can be arbitrary. The fourth preliminary screening step involves using a stepwise discrimination algorithm to perform preliminary screening of the features after the first to third preliminary screenings. The fifth initial screening step involves preliminary screening of the features after the fourth initial screening step based on the consistency between the risk characteristics of each feature and the actual real results of the samples used for model construction.
5. The method according to any one of claims 1 to 4, wherein, Also includes: The sample selection step is used to filter all users to obtain samples for model building before the data acquisition step. Preferably, the sample selection step includes classifying all users in the sample based on a decision tree, and the classification criteria include, but are not limited to: Is a user a customer holding corporate deposits? Has a particular user been involved in a financial institution risk event? 6. The method according to any one of claims 1 to 5, wherein, In the initial screening data transformation step, the determination of the transformation method for the features after preliminary screening is based on the concentration and data type of the features after preliminary screening.
7. The method according to claim 6, wherein, The initial data transformation process, based on determinations of concentration and data type, includes the following steps: Classify each feature based on its data type, categorizing each feature into character variables and numeric variables. For character-type variables, a dummy feature transformation method is used for initial data transformation. The process of further classifying numerical variables includes the following sub-steps: If the numerical variable has fewer than n values, use the WOE (Word of the Environment) transformation method for initial data transformation. If the numerical variable has more than n values, further determine if converting it to a continuous variable would result in a large number of values and a concentration of any single value greater than m%. In this case, use the WOE (Word of the Entity) conversion method. If the concentration of any single value is less than or equal to m%, use the continuous conversion method. Preferably, n and m are both positive integers, where n = 5 to 10 and m = 90 to 99.
8. The method of claim 7, wherein, Also includes: For features that have been confirmed to require continuous transformation, the optimal transformation method is selected based on the correlation between the feature and credit default under different continuous transformation methods. The preferred method for continuous feature transformation is to directly select the original value, calculate the square of the original data, calculate the square root of the original data, calculate the cube root of the original data, or calculate the natural logarithm of the original data.
9. The method according to any one of claims 1 to 8, wherein, The feature screening steps include: The first screening step, based on the stepwise regression algorithm, uses F-tests and T-tests to screen features based on their significance. The second screening step involves calculating the variance inflation factor for each feature and removing features with high variance inflation factors to filter the features. The third fine screening step involves analyzing the feature coefficients of the features after the first and second fine screening steps based on logistic regression to determine whether they conform to the trend of the predicted results for credit defaults, in order to further screen the features.
10. The method according to any one of claims 1 to 9, wherein, The credit default probability modeling process involves substituting the features selected through the feature screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability.
11. A method for calculating inclusive credit risk, comprising: The data acquisition step involves obtaining enterprise credit prediction data for the sample to be predicted. The step of classifying the samples to be predicted uses a decision tree-based method to classify the samples to be predicted to determine the sub-models used to calculate the probability of credit default. The credit default probability calculation steps involve substituting enterprise credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted.
12. The method of claim 11, further comprising: After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000.
13. The method according to claim 11 or 12, wherein, The corporate credit prediction data includes the original corporate credit prediction data of the sample to be predicted and the derived corporate credit prediction data processed based on the original corporate credit prediction data. Preferably, the original corporate credit prediction data includes: The basic data on corporate deposits is based on all available data regarding RMB deposits made by sample (corporate) users at financial institutions. Basic enterprise information data refers to data based on the attributes of the sample (enterprise) users themselves, but which is not directly related to their behavior in financial institutions. The basic data on the financial assets of business owners consists of all other financial assets and transactions held by the sample (actual controller of the enterprise) in financial institutions that are not related to credit cards and loans. Basic information about business owners is based on the attributes of the sample (actual controller of the enterprise) users themselves, but is not directly related to their behavior in financial institutions.
14. The method according to any one of claims 11 to 13, wherein, Derived corporate credit forecast data, derived from original corporate credit forecast data, refers to data obtained by processing the collected original corporate credit forecast data based on time, space, frequency, and statistical information dimensions. Preferably, the derived corporate credit forecast data includes, but is not limited to: Derivative corporate credit prediction data obtained by processing based on sample relationship length. Derivative corporate credit prediction data obtained by processing time interval variables. Derivative corporate credit prediction data is obtained by processing sample behavior frequency. Derivative corporate credit prediction data is obtained by processing the data based on the current situation of the sample at the current point in time. Derivative corporate credit prediction data obtained by processing the continuous behavior of samples. Derivative corporate credit prediction data is obtained by processing sample data based on statistical information dimensions.
15. The method according to any one of claims 11 to 14, wherein, The enterprise credit forecast data is selected from one, two, three, four, five, six, seven, or eight of the following: The following data points are considered: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount of the business over the past 6 months; the ratio of the number of credit transactions to the total number of transactions over the past 6 months; the ratio of the amount of credit transactions to the total amount of transactions over the past 3 months; the business's current monthly deposit balance; the business owner's current monthly AUM value; the ratio of the amount of debit transactions to the total amount of transactions over the past 3 months; the business owner's current balance at a given point in time; the minimum balance of the business owner's deposit account at a given point in time over the past 6 months; the difference between the number of months with the maximum balance of the business owner's deposit account at a given point in time over the past 12 months and the current month; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 3 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and the percentile of the business's current monthly deposit balance by industry.
16. The method according to any one of claims 11 to 15, wherein, The steps for classifying the samples to be predicted include the following sub-steps: Is the sample to be tested a customer holding corporate deposits? Has the sample to be tested already experienced a financial institution risk event? Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model used to calculate the probability of credit default. Under the premise of ensuring the rationality of business logic, the order of the above sub-steps can be arbitrarily set.
17. The method according to any one of claims 11 to 16, wherein, The feature transformation of corporate credit prediction data is performed before being substituted into the credit default probability model to calculate the credit default probability of the sample to be predicted. The feature transformation step includes: Based on the feature type of the enterprise credit prediction data that needs to be substituted into the credit default probability model, either the WOE method or the continuous method is selected for feature transformation.
18. The method according to claim 17, wherein, Continuous feature transformation can be performed in the following ways: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
19. The method according to any one of claims 11 to 18, wherein, The credit default probability model is a model constructed based on the credit prediction data of sample enterprises and the credit default probability using logistic regression based on an existing user group, preferably a model constructed based on the method of any one of claims 1 to 10.
20. The method according to any one of claims 11 to 19, wherein, The corporate credit forecast data is selected from: One, two, three, four, or five of the following: the minimum balance of the business owner's deposit account at a given point in time over the past three months; the minimum balance of the business's current deposits over the past three months; the maximum monthly average amount of credit transactions over the past six months; the ratio of the number of credit transactions to the total number of transactions over the past six months; and the ratio of the amount of credit transactions to the total amount of transactions over the past three months.
21. The method according to claim 20, wherein, Feature transformation is performed on the following data: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average amount of the business's credit transactions over the past 6 months, the ratio of the number of credit transactions to the total number of transactions over the past 6 months, and the ratio of the amount of credit transactions to the total amount of transactions over the past 3 months. The preferred approach is to use a continuous conversion method for the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average amount of credit transactions over the past 6 months, and the ratio of the business's credit transaction amount to the total transaction amount over the past 3 months. The preferred approach is to use a dummy feature conversion method for the ratio of the number of credit transactions to the total number of transactions over the past 6 months. Further optimization involves using a continuous transformation method for the minimum balance of a business owner's deposit account at a given point in time over the past three months (using natural logarithm transformation), a continuous transformation method for the minimum balance of the business owner's current account over the past three months (using cube root transformation), a continuous transformation method for the maximum monthly average credit transaction amount over the past six months (using cube root transformation), a continuous transformation method for the ratio of the business owner's credit transaction amount to the total transaction amount over the past three months (using natural logarithm transformation), and a dummy feature transformation method for the ratio of the number of credit transactions to the total number of transactions over the past six months (using dummy variable transformation).
22. The method according to claim 21, wherein, The transformed values of five features—the minimum balance of the business owner's deposit account at a given point in time over the past three months, the minimum balance of the business's current deposits over the past three months, the maximum monthly average loan transaction amount over the past six months, the ratio of the number of loan transactions to the total number of transactions over the past six months, and the ratio of the loan transaction amount to the total transaction amount over the past three months—are substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probabilities to calculate the default probability of the sample to be predicted.
23. The method according to claim 22, wherein, The sub-model is shown in Formula 1: Where k is the number of features entering the model, preferably k is 5. α is the intercept term, with a preferred numerical range of (0.71157, 0.68367) and an optimal value of 0.697618; β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a point in time over the past 3 months. The preferred value range is (-0.03417, -0.04716), and the optimal value is -0.040669. β2 is the coefficient corresponding to the minimum balance of the enterprise's demand deposits in the past 3 months. The preferred value range is (-0.01744, -0.07704), and the optimal value is -0.047235. β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months. The preferred value range is (0.00201, -0.08893), and the optimal value is -0.043456. β4 is the coefficient corresponding to the ratio of the number of credit transactions to the total number of transactions in the past 6 months. The preferred value range is (0.08153, -0.06307), and the optimal value is 0.009227. β5 is the ratio of a company's loan transaction amount to its total transaction amount in the past three months. The preferred value range is (-0.52827, -1.04399), and the optimal value is -0.786134. x1 is the natural logarithm transformed value of the minimum balance of the business owner's deposit account at the point in time over the past 3 months, generated by the feature transformation step; x2 is the cube root transformed value of the minimum current deposit balance of the enterprise in the past 3 months generated by the feature transformation step; x3 is the square root of the maximum average monthly credit transaction amount of the enterprise over the past 6 months, generated by the feature transformation step. x4 is the dummy variable transformed value of the ratio of the number of credit transactions to the total number of transactions in the past 6 months of the enterprise, generated by the feature transformation step; x5 is the natural logarithm of the ratio of the company's credit transaction amount in the past 3 months to its total transaction amount in the past 3 months, generated by the feature transformation step.
24. The method according to any one of claims 11 to 19, wherein, The corporate credit forecast data is selected from one, two, three, four, five, or six of the following: the current monthly deposit balance of the enterprise, the current monthly AUM value of the enterprise owner, the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, the current balance of the enterprise owner, the difference between the number of months of the minimum balance of the enterprise owner's deposit account in the past 6 months and the number of months of the maximum balance of the enterprise owner's deposit account in the past 12 months and the current month.
25. The method according to claim 24, wherein, The following features are transformed: the current monthly deposit balance of the enterprise, the current monthly AUM value of the business owner, the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, the current balance of the business owner, the minimum balance of the business owner's deposit account in the past 6 months, and the difference between the number of months and the current month of the maximum balance of the business owner's deposit account in the past 12 months. The preferred approach is to use a continuous conversion method for the enterprise's current monthly deposit balance, the enterprise owner's current monthly AUM value, the enterprise owner's current balance at the current time, and the minimum balance of the enterprise owner's deposit account at the time of the past 6 months. The preferred approach is to use a dummy feature conversion method for the ratio of the enterprise's debit transaction amount to the total transaction amount of the past 3 months and the difference between the number of months with the maximum balance of the enterprise owner's deposit account at the time of the past 12 months and the current month. Further optimization involves using a continuous transformation method for the company's current monthly deposit balance (natural logarithmic transformation), a continuous transformation method for the company owner's current monthly AUM value (square root transformation), a dummy transformation method for the ratio of the company's debit transaction amount to the total transaction amount in the past 3 months (dummy variable transformation), a continuous transformation method for the company owner's current balance (natural logarithmic transformation), a continuous transformation method for the minimum balance of the company owner's deposit account in the past 6 months (natural logarithmic transformation), and a dummy transformation method for the difference between the number of months with the maximum balance in the company owner's deposit account in the past 12 months and the current month (dummy variable transformation).
26. The method of claim 25, wherein, The converted values of six features—the company's current monthly deposit balance, the business owner's current monthly AUM value, the ratio of the company's debit transaction amount to the total transaction amount in the past three months, the business owner's current balance, the difference between the number of months with the minimum balance in the business owner's deposit account in the past six months and the number of months with the maximum balance in the business owner's deposit account in the past twelve months—are substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probability to calculate the default probability of the sample to be predicted.
27. The method according to claim 26, wherein, The sub-model is shown in Formula 1: Where k is the number of features entering the model, and k is 6 in Formula 2; α is the intercept term, with a preferred numerical range of (1.18368, 0.87527) and an optimal value of 1.029473; β1 is the coefficient corresponding to the company's current monthly deposit balance. The preferred value range is (-0.07368, -0.11077), and the optimal value is -0.092225. β2 is the coefficient corresponding to the current monthly AUM value of the business owner. The preferred value range is (-0.07489, -0.07929), and the optimal value is -0.077093. β3 is the coefficient corresponding to the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months. The preferred value range is (0.44374, 0.34763), and the optimal value is 0.395682. β4 is the coefficient corresponding to the current balance of the business owner. The preferred value range is (-0.04003, -0.06985), and the optimal value is -0.054940. β5 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 6 months. The preferred value range is (-0.01585, -0.03037), with the optimal value being -0.023109. β6 is the coefficient corresponding to the difference between the number of months with the largest balance in the business owner's deposit account at a given point in time over the past 12 months and the number of months in the current month. The preferred value range is (-0.01482, -0.08255), and the optimal value is -0.048686. x1 is the natural logarithm transformed value of the company's current monthly deposit balance generated by the feature transformation step; x2 is the square root transformed value of the current monthly AUM value of the business owner generated by the feature transformation step; x3 is the dummy variable transformed value of the ratio of the enterprise's debit transaction amount in the past 3 months to the total transaction amount in the past 3 months, generated by the feature transformation step; x4 is the cube root transformed value of the business owner's current point-in-time balance generated by the feature transformation step; x5 is the cube root transformed value of the minimum balance of the business owner's deposit account at any point in time over the past 6 months, generated by the feature transformation step; x6 is the dummy variable transformed value of the difference between the number of months with the largest balance in the business owner's bank account at any point in the past 12 months and the number of months in the present, generated by the feature transformation step.
28. The method according to any one of claims 11 to 19, wherein, The corporate credit forecast data is selected from: The following parameters are considered: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and one, two, three, four, five, six, or seven percentiles of the business's current monthly deposit balance by industry.
29. The method according to claim 28, wherein, The following features are transformed: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposit over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months; the minimum monthly product of the business's current deposits over the past 6 months; and the percentile of the business's current monthly deposit balance by industry. The preferred conversion method uses a continuous conversion approach for the following metrics: the minimum balance of the business owner's deposit account at a given point in time over the past 3 months; the minimum balance of the business's current deposits over the past 3 months; the maximum monthly average credit transaction amount over the past 6 months; the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months; the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months; the minimum monthly accumulation of the business's current deposits over the past 6 months; and the percentile of the business's current monthly deposit balance by industry. Further optimization involves using a continuous conversion method for the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, and the maximum monthly average credit transaction amount over the past 6 months. This involves performing a natural logarithmic transformation on the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, or the maximum monthly average credit transaction amount over the past 6 months. Additionally, the ratio of the business's average monthly debit transaction amount to its average monthly balance over the past 12 months and the minimum monthly product of the business's current deposits over the past 6 months are used. The continuous conversion method involves taking the cube root of the ratio of the company's average monthly debit transaction amount to its average monthly balance over the past 12 months, or the minimum monthly accumulation of the company's current deposits over the past 6 months. The continuous conversion method for the ratio of the company's average monthly credit transaction number to its average monthly total credit transaction number over the past 12 months, and the percentage of the company's current monthly deposit balance by industry, involves taking the square root of the ratio of the company's average monthly credit transaction number to its average monthly total credit transaction number over the past 12 months, or the percentage of the company's current monthly deposit balance by industry.
30. The method according to claim 29, wherein, The converted values of seven features—the minimum balance of the business owner's deposit account at a given point in time over the past 3 months, the minimum balance of the business's current deposits over the past 3 months, the maximum monthly average credit transaction amount over the past 6 months, the ratio of the average monthly debit transaction amount to the average monthly balance over the past 12 months, the ratio of the average monthly number of credit transactions to the average monthly total number of credit transactions over the past 12 months, the minimum monthly product of the business's current deposits over the past 6 months, and the quantile of the business's current monthly deposit balance by industry—are substituted into a sub-model constructed using logistic regression based on sample retail credit prediction data and credit default probabilities to calculate the default probability of the sample to be predicted.
31. The method according to claim 30, wherein, The sub-model is shown in Formula 1: Where k is the number of features entering the model, and k is 7 in Formula 1; α is the intercept term, with a preferred numerical range of (4.90128, 4.79548) and an optimal value of 4.848383; β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account at a given point in time over the past 3 months. The preferred value range is (-0.04771, -0.07097), with the optimal value being -0.059338. β2 is the coefficient corresponding to the minimum balance of the enterprise's demand deposits in the past 3 months. The preferred value range is (-0.04223, -0.04769), and the optimal value is -0.044961. β3 is the coefficient corresponding to the maximum monthly average loan transaction amount of the enterprise in the past 6 months. The preferred value range is (0.00699, -0.01024), and the optimal value is -0.001625. β4 is the coefficient corresponding to the ratio of the company's average monthly debit transaction amount to the average monthly balance over the past 12 months. The preferred value range is (0.01636, 0.01505), and the optimal value is 0.0157047. β5 is the coefficient corresponding to the ratio of the average number of monthly credit transactions in the past 3 months to the total number of monthly credit transactions in the past 12 months. The preferred value range is (-0.29326, -0.34022), and the optimal value is -0.316739. β6 is the coefficient corresponding to the quantile of the current month's deposit balance of enterprises by industry. The preferred value range is (-0.00684, -0.00892), and the optimal value is -0.007878. β7 is the coefficient corresponding to the minimum monthly accumulation of the company's current deposits over the past 6 months. The preferred value range is (-0.12165, -0.19573), and the optimal value is -0.158691. x1 is the natural logarithm transformed value of the minimum balance of the business owner's deposit account at the point in time over the past 3 months, generated by the feature transformation step; x2 is the natural logarithm transformed value of the minimum demand deposit balance of the enterprise over the past 3 months, generated by the feature transformation step; x3 is the natural logarithm transformed value of the maximum monthly average credit transaction amount of the enterprise over the past 6 months, generated by the feature transformation step; x4 is the cube root transformed value of the ratio of the company's average monthly debit transaction amount to its average monthly balance over the past 12 months, generated by the feature transformation step. x5 is the square root of the ratio of the company's average monthly number of loan transactions in the past 3 months to the average monthly total number of loan transactions in the past 12 months, generated by the feature transformation step. x6 is the cube root of the minimum monthly product of the company's current deposits over the past 6 months. x7 is the square root conversion value of the current deposit balance percentile of the enterprise by industry.
32. The method according to claim 23, 27 or 31, wherein, After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to generate a credit score to characterize the borrower using the following formula: Where P is the borrower's default probability (P) generated in the credit default probability calculation module, A is 54.2458; B is 115.4156, the round function rounds the calculated score to the nearest integer; finally, scores greater than 1000 are set to 1000, and scores less than 0 are set to 0.
33. An apparatus for constructing an inclusive credit risk prediction model, wherein, The device includes: The data acquisition module is used to acquire raw corporate credit prediction data for samples used to build the model; The data derivation module is used to process the original corporate credit prediction data to generate derived corporate credit prediction data. The feature screening module is used to perform preliminary screening on all categories, i.e. all features, including original corporate credit prediction data and derived corporate credit prediction data, to obtain the features after preliminary screening. The initial screening data transformation module is used to determine the transformation method for the features after initial screening to confirm whether to use WOE transformation method, dummy feature transformation method, or continuous transformation method for feature transformation, and to use the optimal method for each feature after initial screening. The feature refinement module is used to perform in-depth filtering on the features that have undergone initial feature transformation to obtain refined features. The credit default probability modeling module is used to select logistic regression as the model construction method based on the probabilistic relationship between the refined feature combination and credit default, and to confirm the method used to calculate the credit default probability.
34. The apparatus according to claim 33, wherein, The apparatus performs the steps of the method for constructing an inclusive credit risk prediction model as described in any one of claims 1 to 10.
35. A system for constructing an inclusive credit risk prediction model, wherein, The system includes: a memory, a processor, and a program for constructing a corporate credit risk prediction model stored in the memory and executable on the processor. When the program for constructing a corporate credit risk prediction model is executed by the processor, it implements the steps of constructing an inclusive credit risk prediction model as described in any one of claims 1 to 10.
36. An apparatus for calculating inclusive credit risk, comprising: The data acquisition module is used to acquire enterprise credit prediction data for the sample to be predicted. The module classifies the samples to be predicted, which is used to classify the samples to be predicted based on the decision tree method to determine the model used to calculate the probability of credit default. The credit default probability calculation module is used to input corporate credit prediction data into the credit default probability model to calculate the credit default probability of the sample to be predicted.
37. The apparatus according to claim 36, wherein, The apparatus performs the steps of the method for calculating inclusive credit risk according to any one of claims 11 to 32.
38. A system for calculating inclusive credit risk, wherein, The system for calculating inclusive credit risk includes: a memory, a processor, and a program for calculating inclusive credit risk stored in the memory and executable on the processor, wherein when the program for calculating inclusive credit risk is executed by the processor, it implements the steps of the method for calculating inclusive credit risk as described in any one of claims 11 to 32.
39. A computer storage medium, wherein, The computer storage medium stores a program for calculating inclusive credit risk, which, when executed by a processor, implements the steps of the method for calculating inclusive credit risk as described in any one of claims 11 to 32.
Citation Information
Patent Citations
A credit risk assessment method and apparatus based on logistic regression technology
CN112686749B