Method for constructing common credit risk prediction model and common credit Scorezeta model
By building a high-stability and high coverage credit risk prediction model in financial institutions, using statistical principles and multiple feature conversion methods, the technical difficulties of financial institutions in credit risk prediction in inclusive credit business have been solved, and more accurate risk management and stronger competitive advantages have been achieved.
Patent Information
- Application Number
- CN202311629584.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-06
AI Technical Summary
In the inclusive credit business, existing financial institutions face the problems of insufficient data mining and analysis capabilities and weak risk modeling technology, which leads to the inability to effectively improve the accuracy and stability of the credit risk prediction model, hindering the digital transformation process.
Provide a credit risk management system and method, and build a high-stability and high-coverage credit risk prediction model through massive data from large banks and using the risk laws extracted from statistical principles. This method uses a combination of decision tree classification, step-by-step discriminant variable screening, continuous conversion, WOE conversion and dummy feature conversion to improve the efficiency and effect of model development.
It has achieved more accurate prediction of credit risks, improved the risk management capabilities of financial institutions, and enhanced their competitive advantages and asset quality in inclusive credit business.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and specifically to a credit risk management system and method. The method and system of the present application can assist financial institutions in making more accurate risk decisions and accelerate their digital transformation process. In detail, the present application relates to a method for constructing an inclusive credit risk prediction model for inclusive credit business and an inclusive credit risk prediction model for inclusive credit business. Background Art
[0002] In the current environment of booming inclusive credit, the manual approval mechanism of some financial institutions can no longer cope with the increasing demand for credit, so there is an urgent need to improve the intelligent risk control capabilities of financial institutions. Financial institutions hope to build a risk mitigation mechanism for the entire credit process, from customer pre-screening, pre-loan review, loan approval, post-loan management to early collection stages.
[0003] If a scoring system can be developed based on the principles of early identification, early warning, early detection and early disposal, and credit business can be monitored and managed quickly and conveniently, the financial institutions' own business volume, competitive advantage and asset quality can be improved while risks are controllable.
[0004] However, the construction of a scoring system is highly dependent on data and technology. The diversity and coverage of data dimensions, modeling techniques and methodology directly affect the final stability and ranking of the scoring system. Some financial institutions have little experience in intelligent risk control of inclusive business and have weak risk control capabilities. In the actual application process, factors such as lack of data mining and analysis capabilities and weak risk modeling technology have led to financial institutions being unable to fully utilize the value of internal data and effectively improve the accuracy and stability of models, and other control technology problems. This is also one of the main obstacles faced by small and medium-sized financial institutions in the direction of digital transformation. Summary of the invention
[0005] In view of the above shortcomings in the prior art, this application is intended to provide a credit risk management system and method that can provide effective risk management for financial institutions. The credit risk prediction method and system of this application is based on the massive data of large banks and the risk rules extracted by statistical principles, which is essentially of promotional significance.
[0006] The construction process of other popular scoring models in the market is often limited by unfavorable factors such as small number of modeling samples, relatively single data sources, and high homogeneity of data dimensions. At the same time, since most of the risk assessment models on the market currently use data with weak financial attributes, that is, they are based on non-credit transaction data such as smart terminal device data, social platform data, online shopping mall data, and non-overdue prediction targets for modeling, their prediction results often have a large deviation from the actual credit overdue situation.
[0007] This application first provides a method that can be applied to the construction of a credit risk model, and based on the model constructed by this method, provides financial institutions with a set of accurate methods for calculating the credit risk of samples to be predicted.
[0008] The method and system of the present application are developed based on highly stable and high-coverage data samples, and systematically innovate the relatively mature credit risk control system. They use credit samples containing various business forms to predict the potential credit risks of financial institutions.
[0009] Compared with existing models on the market, the method and system of the present application retain a relatively mature basic framework for model construction, innovatively use a stepwise discriminant variable screening method, fully improve the model development efficiency without affecting the overall model effect, and innovatively combine continuous conversion, WOE conversion, and dummy feature conversion to make up for the potential poor model effect of a single conversion, enabling financial institutions to more accurately predict the credit risk of samples.
[0010] Specifically, this application adopts the following technical solutions:
[0011] 1. A method for calculating inclusive credit risk, comprising:
[0012] A data collection step, which obtains inclusive credit prediction data of the sample to be predicted;
[0013] A step of classifying samples to be predicted, which classifies the samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default;
[0014] The credit default probability calculation step is to substitute the inclusive credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted.
[0015] 2. The method according to claim 1, further comprising:
[0016] After calculating the credit default probability, a step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.
[0017] 3. The method according to item 1 or 2, wherein:
[0018] The inclusive credit prediction data includes the original inclusive credit prediction data of the sample to be predicted and the derived inclusive credit prediction data processed based on the original inclusive credit prediction data;
[0019] Preferably, the original inclusive credit prediction data includes:
[0020] Basic data on corporate credit, which is all available data based on the loan application and usage behavior of sample (corporate) users.
[0021] Basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions.
[0022] Basic data of enterprise basic information, which is based on the attributes of the sample (enterprise) users themselves but is not directly related to their behavior in financial institutions.
[0023] Basic data on financial assets of business owners, which includes all other financial assets and financial transaction data of sample (actual controller of the business) users in financial institutions that are not related to credit cards and loans.
[0024] Basic data on business owner credit, which is all available data based on the loan application and usage behavior of sample users (actual controllers of enterprises).
[0025] Basic data on the basic information of business owners is based on the attributes of the sample users (actual controllers of the enterprises) themselves, but is not directly related to their behavior in financial institutions.
[0026] 4. The method according to any one of items 1 to 3, wherein
[0027] The derived inclusive credit prediction data processed based on the original inclusive credit prediction data refers to the data obtained by processing the collected original inclusive credit prediction data based on the time dimension, space dimension, frequency dimension and statistical information dimension;
[0028] Preferably, the derived inclusive credit forecast data includes but is not limited to:
[0029] Derived inclusive credit forecast data obtained by processing based on sample relationship length,
[0030] Derived inclusive credit forecast data obtained by processing time interval variables,
[0031] Derived inclusive credit forecast data obtained based on the frequency of sample behavior.
[0032] Derivative inclusive credit forecast data obtained by processing the current time point of the sample,
[0033] Derived inclusive credit forecast data obtained based on the continuous behavior of samples,
[0034] Derived inclusive credit forecast data is obtained by processing sample data based on statistical information dimensions.
[0035] 5. The method according to any one of items 1 to 4, wherein the inclusive credit prediction data is selected from one, two, three, four, five, six, seven or eight of the following:
[0036] The minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, the (credit transaction amount - debit transaction amount) of the business owner in the past six months, the current month's RMB account deposit balance of the business owner, the minimum monthly average credit transaction amount of the business owner in the past six months, the minimum monthly accumulation of the business owner's current deposits in the past six months, the number of consecutive months of increase in the business owner's bill balance in the past 12 months, and the business owner's credit card balance in the past 12 months Whether the recycling rate is greater than 0, the percentile of the average deposit balance of the enterprise in the past 6 months by industry, the average monthly accumulation of the enterprise's current deposit accounts in the past 3 months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance of the past 12 months, the average monthly number of credit transactions of the enterprise in the past 3 months / the average monthly number of credit transactions in the past 12 months, the percentile of the credit transaction amount of the enterprise in the past 3 months by region, the percentile of the current deposit balance of the enterprise by industry, the current remaining balance of the business owner's credit card, the number of months from the present when the maximum balance of the business owner's deposit account in the past 12 months was the maximum, and whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0.
[0037] 6. The method according to any one of items 1 to 5, wherein
[0038] The step of classifying the sample to be predicted includes the following sub-steps:
[0039] Whether the sample to be tested is a customer with a corporate deposit account age greater than a months;
[0040] Whether the sample to be tested is a customer whose average balance of corporate deposit account is less than c yuan in the past b months;
[0041] Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model for calculating the probability of credit default. Under the premise of ensuring the rationality of business logic, the order of the above sub-steps can be set arbitrarily;
[0042] It is preferred to classify the predicted samples in the following order:
[0043] First, determine whether the sample to be tested is a customer with a corporate deposit age greater than a month;
[0044] Then determine whether the sample to be tested is a customer whose average balance of corporate deposit accounts in the past b months is less than c yuan.
[0045] 7. The method according to any one of items 1 to 6, wherein
[0046] After feature conversion is performed on the inclusive credit prediction data, the data is substituted into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes:
[0047] Based on the feature type of the inclusive credit prediction data that needs to be substituted into the credit default probability sub-model, the WOE method or the continuous method is selected for feature conversion.
[0048] 8. The method according to claim 7, wherein:
[0049] The continuous method for feature conversion includes the following methods: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data to perform continuous feature conversion.
[0050] 9. The method according to any one of items 1 to 8, wherein
[0051] The credit default probability sub-model is a model constructed based on sample inclusive credit prediction data and credit default probability using logistic regression based on the existing user population.
[0052] 10. The method according to any one of items 1 to 9, wherein
[0053] The inclusive credit prediction data is selected from: the minimum balance of the business owner's deposit account in the past three months, the maximum average monthly credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the average monthly debit transaction amount of the business owner in the past six months / the average monthly balance of the past six months, and one, two, three, four, five or six of the average of (credit transaction amount-debit transaction amount) of the business owner in the past six months.
[0054] 11. The method according to item 10, wherein the step of calculating the probability of credit default comprises:
[0055] The minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, and the average value of (credit transaction amount-debit transaction amount) of the business owner in the past six months are transformed into features.
[0056] Preferably, the minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, and the average value of (credit transaction amount - debit transaction amount) of the business owner in the past six months are converted in a continuous manner;
[0057] Further preferred calculation methods include taking the natural logarithm of the minimum balance in the business owner's deposit account in the past three months, taking the square root of the maximum average monthly credit transaction amount in the past six months, taking the cube root of the average utilization rate of the business owner's revolving loan in the past 12 months, taking the natural logarithm of the minimum value of the business's demand deposit balance in the past three months, taking the square root of the business owner's monthly average debit transaction amount in the past six months / the average monthly balance in the past six months, and taking the cube root of the average value of (credit transaction amount - debit transaction amount) in the past six months.
[0058] 12. The method according to claim 11, wherein:
[0059] The converted values of the six features, namely, the minimum balance of the business owner's deposit account in the past three months, the maximum average monthly credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the average monthly debit transaction amount of the business owner in the past six months / the average monthly balance of the past six months, and the average of (credit transaction amount - debit transaction amount) of the business owner in the past six months, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data and credit default probability to calculate the default probability of the sample to be predicted.
[0060] 13. The method according to claim 12, wherein:
[0061] The sub-model is shown in the following formula 1:
[0062]
[0063] Where k is the number of features entering the model, preferably k is 6,
[0064] α is the intercept term, the preferred value range is (2.4046, 2.3654), and the optimal value is 2.385;
[0065] β1 is the coefficient corresponding to the minimum balance of the business owner’s deposit account in the past three months. The preferred value range is (-0.085, -0.281), and the optimal value is -0.183;
[0066] β2 is the coefficient corresponding to the maximum value of the average monthly credit transaction amount of the enterprise in the past 6 months. The preferred value range is (-0.4232, -0.5408), and the optimal value is -0.482;
[0067] β3 is the coefficient corresponding to the average utilization rate of the enterprise's main revolving loan in the past 12 months, the preferred value range is -0.5442, -0.6618), and the optimal value is -0.603;
[0068] β4 is the coefficient corresponding to the minimum value of the current deposit balance of the enterprise in the past three months. The preferred value range is (-0.0262, -0.3398), and the optimal value is -0.183;
[0069] β5 is the corresponding coefficient of the average monthly debit transaction amount of the enterprise in the past 6 months / the average monthly balance in the past 6 months. The preferred value range is (0.3516, 0.1164), and the optimal value is 0.234;
[0070] β6 is the corresponding coefficient of the mean of (credit transaction amount - debit transaction amount) of the enterprise in the past 6 months. The preferred value range is (-0.8732, -0.9908), and the optimal value is -0.932;
[0071] x1 is the natural logarithm transformation value of the minimum balance of the business owner’s deposit account in the past three months generated by the feature conversion step;
[0072] x2 is the square root conversion value of the maximum average monthly credit transaction amount of the enterprise in the past 6 months generated by the feature conversion step;
[0073] x3 is the cube root conversion value of the average utilization rate of the enterprise's main revolving loan in the past 12 months generated by the feature conversion step;
[0074] x4 is the natural logarithm transformation value of the minimum demand deposit balance of the enterprise in the past three months generated in the feature transformation step;
[0075] x5 is the square root conversion value of the enterprise's average monthly debit transaction amount in the past six months / average monthly balance in the past six months generated in the feature conversion step;
[0076] x6 is the cube root conversion value of the mean of (credit transaction amount - debit transaction amount) of the enterprise in the past 6 months generated by the feature conversion step.
[0077] 14. The method according to claim 13, wherein:
[0078] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower:
[0079]
[0080]
[0081] Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
[0082] 15. The method according to any one of items 1 to 9, wherein
[0083] The inclusive credit prediction data is selected from: the current month's RMB account deposit balance of the enterprise, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months, the number of consecutive months of increased bill balances of the business owner in the past 12 months, whether the revolving utilization rate of the business owner's credit card in the past 12 months is greater than 0, and one, two, three, four, five or six percentiles of the average deposit balance of the enterprise by industry in the past 6 months.
[0084] 16. The method according to item 15, wherein the step of calculating the probability of credit default comprises:
[0085] The current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's current deposits in the past 6 months, the number of consecutive months of increase in the bill balance of the business owner in the past 12 months, whether the recycling rate of the business owner's credit card in the past 12 months is greater than 0, and the average deposit balance quantile of the enterprise by industry in the past 6 months are converted into features.
[0086] It is preferred to use a continuous conversion method for the current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's current deposits in the past 6 months, the number of consecutive months of increase in the bill balance of the enterprise owner in the past 12 months, and the average deposit balance quantile of the enterprise by industry in the past 6 months; and use a dummy feature method to convert whether the recycling rate of the enterprise owner's credit card in the past 12 months is greater than 0;
[0087] Further optimization includes taking the cube root of the enterprise's current month's RMB account deposit balance; taking the square root of the enterprise's minimum monthly average credit transaction amount in the past six months; taking the natural logarithm of the minimum monthly product of the enterprise's demand deposits in the past six months; taking the original value of the number of consecutive months of increase in the business owner's bill balance in the past 12 months; and taking the original value of the percentile of the enterprise's average deposit balance by industry in the past six months.
[0088] 17. The method according to claim 15, wherein:
[0089] The converted values of the six features, namely, the current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months, the number of consecutive months of increase in the owner's bill balance in the past 12 months, whether the owner's credit card revolving rate in the past 12 months is greater than 0, and the average deposit balance quantile of the enterprise by industry in the past 6 months, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data to calculate the default probability of the sample to be predicted.
[0090] 18. The method according to claim 17, wherein:
[0091] The sub-model is shown in the following formula 2:
[0092]
[0093] Where k is the number of features entering the model, and k is preferably 6;
[0094] α is the intercept term, the preferred value range is (1.48351, 1.4688), and the optimal value is 1.476;
[0095] β1 is the coefficient corresponding to the balance of the company's RMB account in the current month. The preferred value range is (-0.0351, -0.09687), and the optimal value is -0.066;
[0096] β2 is the coefficient corresponding to the minimum value of the average monthly credit transaction amount of the enterprise in the past 6 months. The preferred value range is (-0.22636, -0.2822), and the optimal value is -0.254;
[0097] β3 is the coefficient corresponding to the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months. The preferred value range is (-0.33275, -0.39923), and the optimal value is -0.366;
[0098] β4 is the coefficient corresponding to the number of consecutive months of increase in the bill balance of the business owner in the past 12 months. The preferred value range is (0.01777, -0.05296), and the optimal value is -0.018;
[0099] β5 is the coefficient corresponding to whether the revolving utilization rate of the business owner's credit card in the past 12 months is greater than 0. The preferred value range is (-0.17879, -0.1862), and the optimal value is -0.182;
[0100] β6 is the corresponding coefficient of the average deposit balance quantile of the enterprise by industry in the past 6 months. The preferred value range is (-0.34864, -0.35845), and the optimal value is -0.354;
[0101] x1 is the natural logarithm transformation value of the enterprise’s RMB account deposit balance in the current month generated by the feature conversion step;
[0102] x2 is the square root conversion value of the minimum average monthly credit transaction amount of the enterprise in the past 6 months generated by the feature conversion step;
[0103] x3 is the natural logarithm transformation value of the minimum monthly product of the enterprise's demand deposits in the past 6 months generated by the feature conversion step;
[0104] x4 is the original value of the number of consecutive months of increase in the bill balance of the business owner in the past 12 months generated in the feature conversion step;
[0105] x5 is the dummy variable conversion value of the recurring usage rate of the business owner’s credit card in the past 12 months generated in the feature conversion step;
[0106] x6 is the original value of the average deposit balance quantile of enterprises by industry in the past six months generated by the feature transformation step.
[0107] 19. The method according to claim 18, wherein:
[0108] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower:
[0109]
[0110]
[0111] Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
[0112] 20. The method according to any one of items 1 to 9, wherein
[0113] The inclusive credit forecast data is selected from:
[0114] The average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months, the average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months, the percentile of the enterprise's credit transaction amount in the past three months by region, the percentile of the enterprise's current deposit balance by industry, the current remaining balance of the enterprise owner's credit card, the number of months from the present when the enterprise owner's deposit account had the maximum balance in the past 12 months, and whether the maximum (interest / limit) of the enterprise owner's credit card in the past six months is greater than 0, one, two, three, four, five, six, seven or eight of the following.
[0115] 21. The method according to claim 20, wherein the step of calculating the probability of credit default comprises:
[0116] The average monthly accumulation of the company's current deposit accounts in the past three months, the average monthly debit transaction amount of the company in the past 12 months / the average monthly balance in the past 12 months, the average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months, the quantile of the company's credit transaction amount in the past three months by region, the quantile of the company's current deposit balance by industry, the current remaining balance of the business owner's credit card, the number of months from the maximum balance of the business owner's deposit account in the past 12 months, and whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0 are converted into features.
[0117] It is preferred to use a continuous conversion method for the average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount / average monthly balance of the past 12 months, the average monthly number of credit transactions / average monthly number of credit transactions in the past three months, the quantile of the credit transaction amount of the enterprise by region in the past three months, the quantile of the current deposit balance of the enterprise by industry, the current remaining balance of the enterprise owner's credit card, and the number of months from the maximum balance of the enterprise owner's deposit account in the past 12 months; whether the maximum (interest / limit) of the enterprise owner's credit card in the past six months is greater than 0 is converted using a dummy feature method;
[0118] Further optimization includes taking the square root of the average monthly product of the enterprise's current deposit accounts in the past three months; taking the original value of the enterprise's average monthly debit transaction amount in the past 12 months / the average monthly balance in the past 12 months; taking the square root of the enterprise's average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months; taking the cube root of the enterprise's credit transaction amount quantiles by region in the past three months; taking the natural logarithm of the enterprise's credit transaction amount quantiles by region in the past three months; taking the natural logarithm of the business owner's current credit card balance; and taking the natural logarithm of the number of months from the current maximum balance of the business owner's deposit account in the past 12 months.
[0119] 22. The method according to claim 20, wherein:
[0120] The average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance of the past 12 months, the average monthly number of credit transactions of the enterprise in the past three months / the average monthly number of credit transactions in the past 12 months, the quantiles of the enterprise's credit transaction amount in the past three months by region, the quantiles of the enterprise's current deposit balance by industry, the current remaining balance of the enterprise owner's credit card, and the number of months from the maximum balance of the enterprise owner's deposit account in the past 12 months are converted continuously; the converted values of the eight features, such as whether the maximum (interest / limit) of the enterprise owner's credit card in the past six months is greater than 0, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data and credit default probability to calculate the default probability of the sample to be predicted.
[0121] 23. The method according to claim 22, wherein:
[0122] The sub-model is shown in Formula 3 below:
[0123]
[0124] Where k is the number of features entering the model, preferably k is 8;
[0125] α is the intercept term, the preferred value range is (1.61086, 1.53246), and the optimal value is 1.572;
[0126] β1 is the corresponding coefficient of the average monthly accumulation of the enterprise's current deposit accounts in the past three months. The preferred value range is (-0.06973, -0.07035), and the optimal value is -0.070;
[0127] β2 is the corresponding coefficient of the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months. The preferred value range is (-0.25804, -0.27992), and the optimal value is -0.269;
[0128] β3 is the coefficient of the average number of credit transactions in the past three months / the average number of credit transactions in the past 12 months. The preferred value range is (-0.12274, -0.13067), and the optimal value is -0.127;
[0129] β4 is the corresponding coefficient of the credit transaction amount quantile of the enterprise by region in the past three months. The preferred value range is (0.02505, -0.00954), and the optimal value is 0.008;
[0130] β5 is the corresponding coefficient of the current deposit balance quantile of the enterprise by industry, the preferred value range is (-0.02632, -0.05308), and the optimal value is -0.040;
[0131] β6 is the coefficient corresponding to the current remaining balance of the business owner’s credit card. The preferred value range is (-0.02632, -0.05308), and the optimal value is -0.346;
[0132] β7 is the coefficient corresponding to the number of months from the current point in time of the maximum balance of the business owner's deposit account in the past 12 months. The preferred value range is (0.34439, 0.34161), and the optimal value is 0.343;
[0133] β8 is the coefficient of whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0. The preferred value range is (0.066, -0.13), and the optimal value is -0.032;
[0134] x1 is the square root conversion value of the average monthly product of the company's current deposit accounts in the past three months generated by the feature conversion step;
[0135] x2 is the original value of the average monthly debit transaction amount / average monthly balance of the enterprise in the past 12 months generated by the feature conversion step;
[0136] x3 is the square root conversion value of the average number of monthly credit transactions of the enterprise in the past three months / the average number of monthly credit transactions in the past 12 months generated by the feature conversion step;
[0137] x4 is the natural logarithm transformation value of the credit transaction amount quantile of the enterprise by region in the past three months generated in the feature conversion step;
[0138] x5 is the cube root conversion value of the current deposit balance quantile of the enterprise by industry generated in the feature conversion step;
[0139] x6 is the natural logarithm transformation value of the business owner’s current credit card balance generated in the feature transformation step;
[0140] x7 is the natural logarithm transformation value of the maximum balance of the business owner’s deposit account in the past 12 months, which is generated in the feature conversion step;
[0141] x8 is the dummy variable conversion value generated in the feature conversion step, indicating whether the maximum (interest / limit) of the business owner’s credit card in the past 6 months is greater than 0.
[0142] 24. The method according to claim 23, wherein:
[0143] After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower:
[0144]
[0145]
[0146] Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
[0147] 25. A device for calculating inclusive credit risk, comprising:
[0148] A data collection module, which is used to obtain inclusive credit prediction data of samples to be predicted;
[0149] A module for classifying samples to be predicted, which is used to classify the samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default;
[0150] A credit default probability calculation module is used to substitute the inclusive credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted, and preferably the credit default probability is the credit default probability for the inclusive credit business.
[0151] 26. The apparatus according to item 25, wherein the apparatus executes the steps of the method for calculating inclusive credit risk described in any one of items 1 to 24.
[0152] 27. A system for calculating inclusive credit risk, characterized in that the system for calculating inclusive credit risk comprises: a memory, a processor, and a program for the method for calculating inclusive credit risk stored in the memory and executable on the processor, wherein the program for calculating inclusive credit risk implements the steps of the method for calculating inclusive credit risk as described in any one of items 1 to 24 when executed by the processor.
[0153] 28. A computer storage medium, characterized in that a program for calculating inclusive credit risk is stored on the computer storage medium, and when the program for calculating inclusive credit risk is executed by a processor, the steps of the method for calculating inclusive credit risk as described in any one of items 1 to 24 are implemented.
[0154] 29. A method for constructing an inclusive credit risk prediction model for inclusive credit business, comprising:
[0155] A data collection step, which obtains the original inclusive credit forecast data of the sample used to build the model;
[0156] A data derivation step, which processes derived inclusive credit forecast data based on the original inclusive credit forecast data;
[0157] The feature initial screening step is to perform initial screening on all categories, that is, all features, including the original inclusive credit forecast data and the derived inclusive credit forecast data, to obtain features after initial screening;
[0158] The initial screening data conversion step is to judge the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and use the best method to convert each feature after the initial screening;
[0159] The feature fine screening step is to perform in-depth screening on the initially screened features after feature conversion to obtain finely screened features;
[0160] The credit default probability modeling step is to select the logistic regression method to build the model based on the relationship between the carefully screened features and the probability of credit default, and confirm the method used to calculate the credit default probability;
[0161] The data samples obtained in the data collection step are samples of customers who have used inclusive credit services.
[0162] 30. The method according to claim 29, wherein:
[0163] In the data collection step, the original inclusive credit prediction data of the samples obtained for building the model include:
[0164] Basic data on corporate credit, which is all available data based on the loan application and usage behavior of sample (corporate) users.
[0165] Basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions.
[0166] Basic data of enterprise basic information, which is based on the attributes of the sample (enterprise) users themselves but is not directly related to their behavior in financial institutions.
[0167] Basic data on financial assets of business owners, which includes all other financial assets and financial transaction data of sample (actual controller of the business) users in financial institutions that are not related to credit cards and loans.
[0168] Basic data on business owner credit, which is all available data based on the loan application and usage behavior of sample users (actual controllers of enterprises).
[0169] Basic data on the basic information of business owners is based on the attributes of the sample users (actual controllers of the enterprises) themselves, but is not directly related to their behavior in financial institutions.
[0170] 31. The method according to claim 29, wherein:
[0171] In the data derivation step, the derived inclusive credit prediction data processed based on the original inclusive credit prediction data refers to the data obtained by processing the collected original inclusive credit prediction data based on the time dimension, space dimension, frequency dimension, and statistical information dimension;
[0172] Preferably, the derived inclusive credit forecast data includes but is not limited to:
[0173] Derived inclusive credit forecast data obtained by processing based on sample relationship length,
[0174] Derived inclusive credit forecast data obtained by processing time interval variables,
[0175] Derived inclusive credit forecast data obtained based on the frequency of sample behavior.
[0176] Derivative inclusive credit forecast data obtained by processing the current time point of the sample,
[0177] Derived inclusive credit forecast data obtained based on the continuous behavior of samples,
[0178] Derived inclusive credit forecast data is obtained by processing sample data based on statistical information dimensions.
[0179] 32. The method according to any one of items 29 to 31, wherein
[0180] The feature initial screening step includes the following steps:
[0181] The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model.
[0182] The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high.
[0183] The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features;
[0184] The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order.
[0185] The fourth preliminary screening step uses a stepwise discrimination algorithm to perform preliminary screening of features after the first to third preliminary screenings;
[0186] The fifth preliminary screening step is to conduct preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.
[0187] 33. The method according to any one of items 29 to 32, further comprising:
[0188] The sample selection step is used to screen all users and obtain samples of inclusive credit business for model construction before the data collection step.
[0189] Preferably, the sample selection step includes classifying all sample users based on a decision tree, and the classification basis includes but is not limited to:
[0190] Whether a user is a customer with a corporate deposit account age greater than a months;
[0191] Whether a user is a customer whose average balance of corporate deposit accounts in the past b months is less than c yuan.
[0192] 34. A method according to any one of items 29 to 33, wherein, in the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.
[0193] 35. The method according to claim 34, wherein:
[0194] The initial screening data conversion step includes the following steps based on the judgment of concentration and data type:
[0195] Classify the data type of each feature into character variables and numeric variables.
[0196] For character variables, dummy feature conversion is used to perform initial screening data conversion.
[0197] The process of further classifying numerical variables includes the following sub-steps:
[0198] If the value of the numeric variable is less than n, the WOE conversion method is used to convert the initial screening data.
[0199] If the value of the numerical variable is more than n, further judge if the value of the continuous variable is large and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used.
[0200] Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.
[0201] 36. The method according to item 35, further comprising:
[0202] For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature.
[0203] Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
[0204] 37. The method according to any one of items 29 to 36, wherein the feature fine screening step comprises:
[0205] The first fine screening step is to screen the features based on the stepwise regression algorithm, the F test and the T test to determine the significance of the features.
[0206] The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate the features with higher variance inflation factors to screen the features.
[0207] The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default so as to further perform feature screening.
[0208] 38. A method according to any one of items 29 to 37, wherein the credit default probability modeling step substitutes the features screened in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.
[0209] 39. A device for constructing an inclusive credit risk prediction model for inclusive credit business, characterized in that the device comprises:
[0210] A data collection module, which is used to obtain the original inclusive credit prediction data of the samples used to build the model;
[0211] A data derivation module, which is used to process derived inclusive credit prediction data based on the original inclusive credit prediction data;
[0212] A feature initial screening module, which is used to perform initial screening on all categories, that is, all features, including the original inclusive credit prediction data and the derived inclusive credit prediction data, to obtain features after initial screening;
[0213] The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening;
[0214] A feature fine screening module is used to perform in-depth screening on the initially screened features after feature conversion to obtain finely screened features;
[0215] The credit default probability modeling module is used to select the logistic regression method to build a model based on the probability relationship between the finely screened features and credit default, and confirm the method used to calculate the credit default probability.
[0216] The data samples obtained by the data collection module are samples of customers who have used inclusive credit services.
[0217] 40. An apparatus according to item 39, wherein the apparatus executes the steps of the method for constructing an inclusive credit risk prediction model as described in any one of items 29 to 38.
[0218] 41. A system for constructing an inclusive credit risk prediction model for inclusive credit business, characterized in that the system comprises: a memory, a processor, and a program for constructing a method for constructing an inclusive credit risk prediction model stored in the memory and executable on the processor, wherein the program for constructing a method for constructing an inclusive credit risk prediction model, when executed by the processor, implements the steps of the method for constructing an inclusive credit risk prediction model as described in any one of items 29 to 38.
[0219] Effects of the Invention
[0220] The method and system for constructing an inclusive credit risk prediction model in this application combines a large number of samples from a large financial institution when constructing the model, and selects customer data that have applied for inclusive credit business as original data, and deeply processes and derives the original data obtained from the samples. It uses advanced statistical analysis methods and explainable machine learning technology to construct a general scoring model for inclusive credit based on the characteristics of the original data and derived data and the strong financial attribute information contained in the data that is difficult to obtain on the market.
[0221] In addition, this application uses the decision tree method to classify samples in the most reasonable way at the beginning of model construction and builds a sub-classification model based on the classified samples. Combining decision trees to build financial models can effectively classify samples according to executable categories, thereby better covering the credit risk characteristics of different customer groups and avoiding the problem of using all samples to build models, resulting in the model lacking sub-group representativeness.
[0222] Furthermore, when constructing the risk prediction model in this application, the overall efficiency of model development is fully improved by using the stepwise discriminant analysis method for the preliminary screening of model features. The stepwise discriminant method can better select more important feature variables under the same dimension, greatly reducing the workload of developers in the next step who need to judge and screen one by one according to the trend of the variables, and fully improving the efficiency of model development without affecting the overall model effect, making the initial screening of variables more efficient and accurate.
[0223] This application innovatively combines continuous conversion, WOE conversion, and dummy feature conversion when building a model, reprocesses some features that have passed the initial screening, combines the advantages and disadvantages of the three conversion methods, and creatively designs a conversion judgment method. According to parameters such as feature data attributes, missing rate, concentration, etc., supplemented by business logic judgment, the optimal feature conversion method is selected. Through such a model building method, it is possible to avoid the technical shortcomings of overfitting when building a model with complete WOE variables and the inability to adapt to categorical variables when building a model with complete continuous variables.
[0224] In addition, the method and system for calculating credit risk probability or scoring credit risk constructed by the present application, namely the Zeta model (Zeta series scoring cards) (including Zeta 1, Zeta 2 and Zeta 3 sub-models or sub-scoring cards), because its model is first split based on the decision tree when it is constructed, so the credit risk is initially calculated by effectively classifying the customer samples and selecting the most suitable sub-model or sub-scoring card for processing. At the same time, because the model construction method used in the sub-model or sub-scoring card is the method described in the present application, the construction process also avoids the technical shortcomings of overfitting when constructing a model with complete WOE variables and the inability to adapt to categorical variables when constructing a model with complete continuous variables. Therefore, it has a significantly better effect than existing models in predicting credit risk. In addition, since this application is aimed at inclusive credit business, the sample data is customer data that has applied for inclusive credit business screened out from massive data. It not only covers the predictive variables of Internet business, but also includes the overdue performance data of Internet business. Therefore, the risk rules extracted from specific business scenario data using statistical principles have very significant adaptability and distinctiveness in terms of credit risk for inclusive credit business. BRIEF DESCRIPTION OF THE DRAWINGS
[0225] The accompanying drawings are used to better understand the present application and do not constitute an improper limitation on the present application.
[0226] Figure 1 Instances of binning to confirm whether data conforms to business trends;
[0227] Figure 2This is a typical flow chart for classifying all sample users based on a decision tree;
[0228] Figure 3 It is a graphical representation of the model differentiation effect of Example 1 of the present application. DETAILED DESCRIPTION
[0229] The following is a description of the exemplary embodiments of the present application, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0230] Credit risk is the risk caused by the borrower's reduced willingness to repay or inability to perform the loan contract due to changes in economic ability, rather than the risk of default caused by the borrower's deliberate fraud. Credit default occurs in all kinds of inclusive credit business scenarios, and is particularly related to the borrower's personal economic situation. The reasons for borrowers' credit defaults can be divided into four categories: 1. Short credit history, such borrowers have little experience in financial management; 2. Borrowers temporarily forget to repay; 3. Over-borrowing, such borrowers have relatively low repayment ability due to a large amount of debt; 4. Affected by major negative factors, such borrowers have long-term impacts on their repayment ability due to major factors such as reduced income, unemployment or divorce. The above different reasons will more or less cause borrowers to default, which may lead to more serious defaults. The purpose of credit risk scoring is to dig out the inherent mathematical relationship between various historical information of customers and the probability of future default, and convert it into a scoring method to quantify the probability of default.
[0231] Currently, the credit risk scoring in existing technologies is mainly developed using credit history data and industrial and commercial judicial data, which reflects their payment behavior, payment willingness and corporate industrial and commercial judicial information.
[0232] The sample data used to construct the model in this application not only includes transaction behavior information of credit business, but also adds inclusive credit data that is usually difficult to obtain. The management of inclusive credit risk usually needs to cover both business owner risk management and enterprise risk management. It is difficult for general institutions to obtain information on both aspects at the same time, making it difficult to make a more comprehensive prediction of inclusive credit business.
[0233] Inclusive credit refers to low-interest loans provided by banks or other financial institutions to individuals or small and micro enterprises that meet certain conditions, all of which are aimed at supporting economically disadvantaged groups, promoting social equity and economic development. Inclusive loans can be divided into small loans for individuals and loan funds for multiple purposes such as production and operation activities issued to micro-enterprises, low-income residents, the disabled, farmers, and the poor.
[0234] Inclusive credit risk refers to the risk caused by the borrower's failure to repay the debt as agreed in inclusive credit business.
[0235] <Overall description of model building method>
[0236] Specifically, the present application relates to a method for constructing an inclusive credit risk prediction model for inclusive credit business, which includes: a data collection step, which obtains original inclusive credit prediction data of samples for building a model; a data derivation step, which processes derived inclusive credit prediction data based on the original inclusive credit prediction data; a feature initial screening step, which performs preliminary screening on all categories including the original inclusive credit prediction data and the derived inclusive credit prediction data, that is, all features, to obtain features after preliminary screening; a preliminary screening data conversion step, which determines the conversion method of the features after preliminary screening to confirm that one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method is used for feature conversion, and the best method is used for feature conversion for each feature after preliminary screening; a feature fine screening step, which performs in-depth screening on the features after the feature conversion to obtain the finely screened features; a credit default probability modeling step, which selects the logistic regression method for model construction based on the finely screened features combined with the probability relationship between the features and the credit default, and confirms the method for calculating the credit default probability.
[0237] In a specific manner, the method for constructing an inclusive credit risk prediction model for inclusive credit business involved in the present application may also include a sample selection step before the data collection step, which is used to screen all users before the data collection step to obtain samples for model construction. Those skilled in the art understand that based on the total number and data of users who can build the model, those skilled in the art can choose whether to eliminate user data that does not meet the requirements for building the model. In addition, those skilled in the art can also first select some users as samples, and then continuously increase the samples for modeling according to the corresponding rules.
[0238] The samples used to build the model in this application can be individual users, who are the actual controllers of the following corporate users, or the companies themselves. The data used to build the model samples can be data from individual users and corporate users of large financial institutions. The data using such samples not only contains transaction behavior information of credit business, but also adds asset data that is usually difficult to obtain. While reflecting the borrower's payment behavior and willingness to pay, it also provides a more comprehensive measurement of their debt repayment ability, personal qualifications, etc., which can provide more accurate prediction results.
[0239] In a specific implementation, the selected sample is a sample of customers who have applied for inclusive credit services among customer samples within a certain period of time.
[0240] <Sample selection steps>
[0241] In a specific implementation, the sample selection step is to screen all users before the data collection step to obtain samples for model construction. Specifically, the sample selection step includes classifying all sample users based on a decision tree, and the classification basis includes but is not limited to: whether a certain user is a customer who has applied for credit business registration at a financial institution; whether a certain user is a customer who has not applied for credit registration at a financial institution; the geographical area where a certain user handles business; whether a certain user has experienced a financial institution risk event (for example, whether a default has occurred, whether there is a mortgage); whether a certain user holds a credit card issued by a financial institution and / or whether the credit card is continuously used and / or whether the credit card or personal loan is revolving; whether a certain user holds a public deposit account; the age of a certain user's public deposit account; and the current balance of a certain user's public deposit account.
[0242] Among them, if the sample is the actual controller of the enterprise, that is, an individual user, the sample selection step includes classifying all sample users based on a decision tree, and the classification basis includes but is not limited to: whether a certain user is a customer who has applied for registration of credit business at a financial institution; whether a certain user is a customer who has not applied for registration of credit business at a financial institution; the geographical area where a certain user handles business; whether a certain user has had a financial institution risk event; whether a certain user holds a credit card issued by a financial institution and / or whether the credit card is used continuously and / or whether the credit card or personal loan is used in a revolving manner.
[0243] Furthermore, taking the personal sample of a financial institution as an example, all sample users are classified based on the decision tree. Since performance variables are required for modeling, the analysis sample is further divided into whether or not they have applied for registration for credit business in the relevant bank, that is, the users can be divided into two parts: those who have applied for registration and those who have not applied for registration.
[0244] Credit business is a basic and important asset business of banks. It recovers the principal and interest by issuing bank loans and obtains profits after deducting costs. Generally speaking, bank credit business is an important means for banks to make profits. From the classification of bank credit business, it can be divided into corporate credit business and personal credit business. Among them, corporate credit business includes project loans, working capital loans, small business loans, real estate enterprise loans, etc.; personal credit business includes personal housing loans, personal consumption loans, personal business loans, etc.
[0245] In a specific implementation, in the method of building a model in this application, the user samples are divided into 1.35 million households that have applied for inclusive credit services and 13.65 million households that have not applied for inclusive credit services. When designing the specific model, the 1.35 million households that have applied are analyzed and designed. After the model is developed and launched, the results will be applied to the full 18 million corporate customer groups and 600 million retail customer groups that have applied for inclusive credit services and have not applied for inclusive credit services.
[0246] Decision Tree is a decision analysis method that uses a decision tree to construct a decision tree based on the known probabilities of various situations to obtain the probability that the expected value of the net present value is greater than or equal to zero, evaluate project risks, and determine its feasibility. It is a graphical method that intuitively uses probability analysis. Taking the corporate samples of a certain financial institution as an example, when constructing the model, this application uses a decision tree to classify the samples used for modeling based on the characteristics of the data of large financial institutions. After distinguishing between samples that have applied for credit business and samples that have not applied for credit business, the public deposit age of the samples that have applied for credit business can be further classified, wherein the decision on the public deposit age can be determined based on the needs of technical personnel in this field. For example, the samples that have applied for credit business are classified into two categories: customers with a public deposit age of less than or equal to a months and customers with a public deposit age of more than a months. After distinguishing between customers with a public deposit age of less than or equal to a month and customers with a public deposit age of more than a month, the average balance of public deposit accounts in the past b months for the samples of customers with a public deposit age of more than a month can be further classified, wherein the decision on the average balance of public deposit accounts in the past b months can be determined based on the needs of technical personnel in this field. For example, the samples of customers with a public deposit age of less than or equal to a month are divided into customers with an average balance of public deposit accounts of less than c yuan in the past b months and customers with an average balance of public deposit accounts of more than c yuan in the past b months.
[0247] In a specific embodiment, a can be set to any integer greater than 1, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc.
[0248] In a specific implementation, b can be set to any integer greater than 1, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc., and b needs to be less than or equal to a.
[0249] In a specific embodiment, c can be set to any value greater than 0, for example, 1, 5, 10, 15, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, etc.
[0250] The users may also be classified based on their consumption after joining the network. For example, in a specific implementation, the modeling samples may be classified based on whether a user holds a corporate deposit account.
[0251] Those skilled in the art can fully understand that other splitting methods can also be selected when performing sample splitting. For example, the sample can be split first based on whether a certain user has experienced a financial institution risk event. The customer sample splitting method when building the model can be considered and carried out entirely according to the needs of modeling.
[0252] Furthermore, based on the segmentation situation for which the model needs to be built, this application can also further divide the customers according to their regions. For example, if the customers have applied for credit business, the samples can be further divided into different sample groups from economically underdeveloped areas, moderately developed areas and relatively developed areas based on the decision tree.
[0253] For another group that has not applied for credit business, the samples can also be further divided into different sample groups from economically underdeveloped areas, economically moderately developed areas, and economically relatively developed areas based on the decision tree. In this application, the definitions of economically underdeveloped areas, economically moderately developed areas, and economically relatively developed areas can be divided based on the commonly recognized standards known to those skilled in the art. The standards can be derived from data released by an authoritative statistical department, or can be divided based on a rating agency, or can be divided based on another constructed financial model. Those skilled in the art can fully understand that after selecting the division standard, the user population can be divided into three different areas without duplication.
[0254] In a specific way, when further classification is carried out, it can be divided comprehensively based on the regional development situation and the bad debt rate, and the user population can be divided into different sample groups from economically underdeveloped areas, moderately developed areas and relatively developed areas.
[0255] In the present application, decision classification can also be made based on whether the customer has overdue payments. For example, for a sample group of users who have applied for inclusive credit services, classification can be carried out based on whether there is any overdue behavior, for example, dividing them into a group that has already incurred overdue behavior and a group that has not incurred any overdue behavior.
[0256] The group that has already been overdue can be further divided into the sample group that has been seriously defaulted and the sample group that has been overdue but not seriously defaulted. For example, when the number of overdue months exceeds n months, it can be considered as seriously overdue, and if it is within n months, it can be considered as overdue but not seriously defaulted. In a specific method, n can be selected as 2, 3, 4, 5, 6 or a larger integer.
[0257] More specifically, a serious default may be, for example, a customer's credit is overdue for 4, 5 or 6 consecutive periods, etc. A mild to moderate default may be, for example, a customer's credit is overdue for 1, 2 or 3 consecutive periods, etc.
[0258] For the group that has not yet defaulted, decisions can be further made based on whether they have a mortgage. For example, customers can be further divided into a sample group that has not defaulted and has a mortgage and a sample group that has not defaulted and has no mortgage.
[0259] The sample group with no overdue payments and no mortgage loans can be further differentiated based on whether they only hold credit card businesses. For example, the customer samples can be divided into samples with no overdue payments and no mortgage loans but with inclusive loans, that is, samples with businesses other than credit cards, and samples with no overdue payments and no mortgage loans but no inclusive loans, that is, samples with no businesses other than credit cards.
[0260] For samples that have not been overdue and have no mortgages but no business other than credit cards, decisions can be made based on whether the customer is a revolving credit card user, and further divided into a sample group that revolves around credit cards and a sample group that uses credit cards normally. In this application, users are divided into a sample group that revolves around credit cards and a sample group that uses credit cards normally based on whether the monthly bill is fully paid when the customer uses the credit card. Normal use is considered when the monthly bill is fully paid, otherwise it is revolving use.
[0261] The purpose of classifying and dividing samples is to find the best group separation so that a set of scoring models built on this basis can maximize the predictive ability of the entire scoring system. There are many ways to segment. The samples selected based on the decision tree algorithm of this application can cover the customer segmentation factors that are usually of great concern in the development of credit business, and the credit default rates in different customer segments can effectively distinguish samples (the division method of the decision tree itself is already an effective method for preliminary sorting of customers), and lay the foundation for building a reasonable and effective model for a certain type of sample group in the future, which is suitable for the application needs of sorting customer credit risks at various business products and credit stages.
[0262] In addition, it should be pointed out that the decision-making and classification process described above is only an example. Those skilled in the art can make decisions and classifications based on the samples used for modeling. For example, in a specific implementation, classification can be performed based on whether there is a mortgage loan or based on the geographical area where the customer is located. The modeling method of this application does not impose any restrictions on this. For example, customers can be divided into different sample groups by dividing them into different groups based on whether the customer has applied, whether the customer has applied for inclusive credit business, whether the payment is overdue, whether the customer has a mortgage loan, whether the customer has an inclusive loan, and the customer's region. As long as the logic between each layer of the decision tree is reasonable, the order of these decision-making methods can be changed and combined at will. The sample selection step can fully reflect the control variables that are closely related to risk characteristics in the process of business development, and can more comprehensively match various customer groups in the market.
[0263] In addition, as mentioned above, the present application may not perform the above-mentioned sample selection step, and the model of the present application may be directly established based on all samples.
[0264] For the corporate user samples, multi-head data samples and operator samples of a certain financial institution, those skilled in the art can fully understand that similar methods can be used to classify the samples to confirm the samples used to build a model or sub-model.
[0265] In the process of building this application model, since it is necessary to build a model for inclusive credit business, the selected samples are customers who have applied for inclusive credit business within a certain period of time as modeling samples.
[0266] In a specific implementation, 1.35 million households that have applied for inclusive credit business are selected from the data of 15 million enterprises in 2019 as the sample for modeling. Such samples are more targeted and representative when predicting the credit risk of inclusive credit business.
[0267] <Original credit prediction data>
[0268] The original credit prediction data includes original inclusive credit prediction data and enterprise credit prediction data, among which the original data used to predict enterprise credit risk includes original enterprise credit prediction data and original enterprise actual controller credit prediction data.
[0269] In this application, the original enterprise credit prediction data refers to enterprise-related data, but in order to more reasonably predict the credit risk of the enterprise, when building a prediction model for the credit risk of the enterprise, in addition to the relevant data of the enterprise itself, it is necessary to further add the variable information of the actual controller of the enterprise. This is because the risk of the actual controller of the enterprise largely reflects the credit risk of the enterprise. The enterprise credit risk prediction model with the data of the actual controller of the enterprise will be more accurate in predicting the credit risk of the enterprise.
[0270] The data types involved in the original inclusive credit prediction data or the original corporate actual controller credit prediction data are consistent.
[0271] When constructing the model for this application, firstly, based on the basic data of financial institutions, about 120 types of original inclusive credit prediction data (i.e. more than 120 basic features) can be obtained, and the inclusive credit risk points are split or processed as much as possible according to different dimensions under the premise of optimal effect, generating more than 3,000 derivative features in total.
[0272] When building the model, first of all, based on all historical data of large financial institutions, the original corporate credit forecast data for building model samples between 2017 and 2021 were initially collected when building the model of this application, including various data information of enterprises and actual controllers of enterprises, totaling 18 million corporate users, of which each enterprise has corresponding data to be processed every month. It can be seen that the data system used to build the model of this application is comprehensive and the data volume is very large. When building a model based on such a data system, it is necessary to consider the modeling methodology, otherwise it will be trapped in the huge data, resulting in the special sample groups that need to be paid attention to being covered in the huge data volume and unable to be effectively identified, causing the computer program to run slowly or even unable to run, so that the most suitable prediction model cannot be accurately constructed.
[0273] In another specific implementation, the original inclusive credit prediction data of the sample used to build the model is based on more than 100 types of enterprise data and more than 120 types of personal data (i.e., basic variables or basic features) obtained from the corporate data of 15 million households in 2019 as the analysis sample, including but not limited to: Enterprises: enterprise loan accounts, guarantee status of enterprise loans, overdue status of enterprise loans, balance of enterprise loans, total amount of enterprise loans, repayment and actual repayment of enterprise loans, etc.; basic information of corporate customers such as industry, scale, administrative region to which the business belongs; corporate deposit transaction information such as corporate deposit balances, debit and credit transaction amounts. Actual controller of the enterprise: credit card bills, credit card cash withdrawals, credit card installments, interest generated by credit cards, and other information on accounts, overdue payments, balances, amounts due, actual payments, etc. under different dimensions (such as time dimension, space dimension, and frequency dimension); personal loan accounts, overdue personal loan status, personal loan balances, total amount of personal loans, personal loan payments due, actual payments, etc.; basic customer information including gender, age, and administrative region to which the business belongs; basic information such as AUM (i.e., assets under management), deposits, wealth management, and payroll payment.
[0274] In the data collection step, the original inclusive credit prediction data of the samples used to build the model include: basic data on corporate credit, which is all available data based on the loan application and usage behavior of sample (enterprise) users,
[0275] Basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions.
[0276] Basic data of enterprise basic information, which is based on the attributes of the sample (enterprise) users themselves but is not directly related to their behavior in financial institutions.
[0277] Basic data on financial assets of business owners, which includes all other financial assets and financial transaction data of sample (actual controller of the business) users in financial institutions that are not related to credit cards and loans.
[0278] Basic data on business owner credit is all available data based on the loan application status and usage behavior of sample users (actual controllers of the enterprise).
[0279] Basic data on the basic information of business owners is based on the attributes of the sample users (actual controllers of the enterprises) themselves, but is not directly related to their behavior in financial institutions.
[0280] In a specific implementation, the basic data of corporate credit includes, but is not limited to, basic fields including repayment, overdue, credit, credit use, balance, aging, collateral, application rejection, etc. The above basic data is not limited to the specific categories listed. With the changes in corporate credit business, those skilled in the art can further cover new data types in implementation, that is, all types of data that can be obtained based on the user's loan application and behavior can be used as basic data of corporate credit.
[0281] In a specific implementation, the basic data of corporate public deposits include, but are not limited to, deposit balance, inflow amount, outflow amount, deposit fluctuation, industry deposit level, industry transaction level, etc. The above basic data are not limited to the specific categories listed. With the changes in corporate public deposit business, those skilled in the art can further cover new data types in implementation, that is, all types of data that can be obtained based on the user's deposit situation and behavior can be used as basic data of corporate public deposits.
[0282] In a specific implementation, the basic data of enterprise basic information includes but is not limited to: industry, region, registration and other information. The above basic data is not limited to the specific categories listed. With the development of customer conditions and social relationships, those skilled in the art can further cover other or emerging data types in the implementation within the scope of the business application scenario, that is, all types of data based on the attributes of the user sample itself but not directly related to the behavior in the financial institution can be used as basic data of enterprise basic information.
[0283] In a specific implementation, the basic data of the financial assets of the business owner include, but are not limited to: asset size, deposits, financial management and payroll, deposit account transaction flow, etc. The above basic data are not limited to the specific categories listed. With the changes in financial assets, those skilled in the art can also further cover new data types in implementation, that is, all other financial assets and financial transaction data that are not related to credit cards and loans in financial institutions based on samples can be used as the basic data of the financial assets of the business owner.
[0284] In a specific implementation, the basic data of business owners' credit includes, but is not limited to: overdue, repayment, balance, account age, credit utilization rate, credit, etc. The above basic data is not limited to the specific categories listed. With the changes in corporate credit business, those skilled in the art can also further cover new data types in implementation, that is, all types of data that can be obtained based on the user's loan application and behavior can be used as basic data of corporate credit.
[0285] In a specific implementation, the basic data of the basic information of the business owner includes but is not limited to: gender, age, and the administrative region where the business belongs. The above basic data is not limited to the specific categories listed. With the development of the business owner's situation and social relations, those skilled in the art can also further cover other or emerging data types in the implementation within the scope of the business application scenario, that is, all types of data based on the attributes of the user sample itself but not directly related to the behavior in the financial institution can be used as the basic data of the basic information of the business owner.
[0286] In a specific implementation, the original inclusive credit prediction data of the samples obtained for building the model is based on more than 120 data types (i.e., basic variables or basic features) obtained from 18 million corporate users, including but not limited to: deposit point balance, customer point asset size value, wealth management product point balance, summary description of wage payment, transaction amount, age, loan balance, consumption interest, cash withdrawal interest, installment amount, installment balance, overdue interest, current repayment amount, product code, account status code, guarantee method code, product code, account status code, guarantee method code, loan amount, product code, account status code, guarantee method code, total amount of overdue payment and other basic information.
[0287] <Derivative Credit Forecast Data>
[0288] In the present application, in the data derivation step, processing derived credit prediction data based on original credit prediction data refers to data obtained by processing the collected original credit prediction data based on the time dimension, space dimension, frequency dimension, and statistical information dimension.
[0289] In a specific embodiment, the derived credit prediction data includes, but is not limited to: derived credit prediction data processed based on the length of sample relationships, derived credit prediction data processed based on time interval variables, derived credit prediction data processed based on the frequency of sample behavior, derived credit prediction data processed based on the current time point of the sample, derived credit prediction data processed based on the continuous behavior of the sample, or derived credit prediction data processed based on the statistical information dimension. For example, the monthly data of the customer can be obtained and processed based on the monthly data. In the present application, processing based on the statistical information dimension includes obtaining the maximum value, minimum value, average value, etc. of the data to describe the data situation.
[0290] In a specific implementation, for example, starting from the time dimension, customer relationship length variables are considered: such as customer account opening time, customer maximum account age, etc. as types of derived inclusive credit prediction data, that is, as derived features or derived variables.
[0291] In a specific implementation, for example, starting from the time dimension, time intervals are considered, such as the number of months from the customer's most recent repayment to the observation point, the number of months from the customer's most recent overdue payment, etc. as derived features or derived variables.
[0292] In a specific implementation, for example, starting from the time dimension, the current time point variables are considered: the customer's current monthly credit limit, the customer's current monthly balance, etc. as derived features or derived variables.
[0293] In a specific method, starting from the dimension of statistical information, statistical variables are considered, such as the maximum number of overdue payments of a customer in the last X months and the average credit limit utilization rate of a customer in the last X months as derived features or derived variables. There is no limit on X, and it can be any positive integer greater than 0 as long as the business logic is reasonable.
[0294] In a specific implementation, for example, starting from the dimension of statistical information, ratio variables are considered, such as call duration / number of calls in the last X months, consumption amount in the last month / consumption amount in the last three months, etc. as derived features or derived variables. There is no limit on X, and it can be any positive integer greater than 0 as long as the business logic is reasonable.
[0295] In a specific implementation, for example, starting from the frequency level, behavioral frequency level variables are considered: such as the maximum number of consecutive overdue payments of a customer in the last X months > N, the number of consecutive repayment rates of a customer in the last X months > N, etc. as derived features or derived variables. There is no limit on X, and it can be any positive integer greater than 0 as long as the business logic is reasonable.
[0296] It is clear to those skilled in the art that the above methods for processing derived variables are merely enumerated and can be selected arbitrarily. For example, 120 types of basic data can be processed into more than 3,000 derived data. In this application, data, data types, data types, variables or features are sometimes mixed, and those skilled in the art can understand them based on common sense in statistics.
[0297] In this application, the derivative data can be derived data obtained by simple processing of the original inclusive credit prediction data, or it can be derived data obtained by complex processing. The simply processed derivative data, such as the customer's current monthly credit limit and the customer's current monthly balance, can be used directly after the data is aggregated. The complex processed derivative data needs to be based on the current class variables based on time slicing and logical processing, and can generate derivative variables such as the customer's maximum number of overdue periods in the last X months and the customer's average credit limit utilization rate in the last X months.
[0298] In a specific implementation, when constructing a model for this application, firstly, based on the basic data of financial institutions, original credit prediction data of approximately more than 100 types of enterprises and more than 120 types of individuals (i.e., more than 100 types of original enterprise data and more than 120 types of original personal basic features) can be obtained, and according to the different dimensions of predicted credit risk, the greatest possible splitting or processing is performed under the premise of optimal effect, generating a total of more than 3,500 enterprise derivative features and more than 3,000 personal derivative features.
[0299] In a specific implementation, when constructing a model for multiple data, the multiple application volume data is first obtained based on the bank's internal authorization, with about 40 basic variables. Subsequently, based on the different dimensions of predicted credit risk, the model is split and processed as much as possible while achieving the best effect, generating a total of about 130 multiple data variables (including original variables + derived variables).
[0300] In a specific implementation, when modeling operator data, based on the above raw data from the operator, the maximum possible splitting and processing is performed under the best effect, and about 3,600 features are derived, plus about 800 existing features of the operator, a total of more than 4,400. It fully covers the eight dimensions of information contained in the raw data, including basic information, communication behavior, SMS behavior, Internet behavior, consumption information, terminal information, social information and location information.
[0301] <Preliminary feature screening>
[0302] In the initial screening data conversion step of the present application, the judgment of the conversion method of the features after the initial screening is based on the concentration and data type of the features after the initial screening. The initial screening data conversion step includes the following steps based on the judgment of the concentration and data type: classify each feature into character variables and numerical variables according to the data type of each feature, use the dummy feature conversion method to perform initial screening data conversion on the character variables, and further classify the numerical variables. The process includes the following sub-steps: if the value of the numerical variable is less than n, use the WOE conversion method to perform initial screening data conversion, if the value of the numerical variable is more than n, further judge if the value of the continuous variable is more and the concentration of the single value is greater than m%, then use the WOE conversion method, if the concentration of the single value is less than or equal to m%, then use the continuous conversion method, preferably, n and m are both positive integers, where n = 5 to 10, m = 90 to 99.
[0303] For example, n is 5, 6, 7, 8, 9 or 10, and m is 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.
[0304] Specifically, in this application, a variety of different feature initial screening methods are used to screen a large number of features, so that the screening can be effectively performed from the largest dimension of features. Existing credit risk scoring often uses missing rate, concentration and information value IV for feature initial screening.
[0305] In a specific method, features are screened based on the data missingness of each feature of the sample used to build the model. When screening the missing rate, it is generally considered to delete variables with a missing rate greater than 90%, 91%, 92%, 93%, 94%, or 95%. For example, features with a data missing rate exceeding 95% can be eliminated, and features with a data missing rate exceeding 90% can also be eliminated.
[0306] In a specific method, features are screened based on the fact that a single value of a feature sample is too high. When screening by concentration, it is generally considered to delete variables whose single values account for more than 99%, 98%, 97%, 96%, or 95%. For example, features whose single values exceed 99% can be eliminated, and features whose single values exceed 95% can also be eliminated.
[0307] In a specific method, the information IV value of each feature is calculated to perform preliminary screening of the features. During IV screening, the IV value can be used to measure the predictive power of the feature. The larger the IV value, the stronger the predictive power of the feature. The calculation method of the IV value of a single feature is as follows:
[0308]
[0309] Where k is the number of groups after this feature is discretized; y i is the number of non-defaulting customers in group i; s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s The quantitative index of IV has the following meanings: when the calculated value of IV is less than 0.02, it means that the predictive power of this feature is very weak; when the calculated value of IV is above 0.02 and below 0.1, it means that the predictive power of this feature is weak; when the calculated value of IV is above 0.1 and below 0.3, it means that the predictive power of this feature is good; when the calculated value of IV is above 0.3, it means that the predictive power of this feature is strong.
[0310] Of course, the IV calculated value deletion threshold may also be set to 0.03, 0.04, 0.05, etc.
[0311] In the model building method of the present application, the step of using missing rate, concentration and information value IV for feature initial screening can be performed in any order, for example, the screening can be performed based on the missing rate first, then based on the concentration, and finally based on the information value IV for feature initial screening. It is also possible to first screen based on the concentration, then screen based on the missing rate, and finally based on the information value IV for feature initial screening. It is also possible to first screen based on the missing rate, then screen based on the information value IV, and finally based on the concentration for feature initial screening. It is also possible to first screen based on the information value IV, then screen based on the missing rate, and finally based on the concentration for feature initial screening. It is also possible to first screen based on the information value IV, then screen based on the missing rate, and finally based on the concentration for feature initial screening. It is also possible to first screen based on the information value IV, then screen based on the concentration, and finally based on the missing rate for feature initial screening. It is also possible to first screen based on the concentration, then screen based on the information value IV, and finally based on the missing rate for feature initial screening. Those skilled in the art can make a choice based on the data situation of the sample, so based on these three methods, features with obvious defects in one aspect can be effectively removed, which can effectively reduce the data dimension and improve the effect of model building.
[0312] In one specific method, after the features have been screened using missing rate, concentration and information value IV respectively, a stepwise discriminant algorithm is used to perform preliminary screening of the features, and then preliminary screening of the features is performed based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used to build the model.
[0313] On this basis, the technical solution of the present application introduces a stepwise discriminant method to make the initial screening of variables more efficient and accurate in order to improve the overall efficiency of model development. In actual data, there may be a close distribution of good and bad samples on a certain variable, and the ability to distinguish between good and bad is weak. There may also be a class of variables, each of which can distinguish good and bad samples well, but if all are included in the model, they will be redundant due to the duplication of data dimensions covered by the variables. In this regard, a stepwise discriminant analysis method is used in this application, and WILKS'S LAMBDA value is used as the statistical criterion for entry and removal, and variables with weak or redundant discriminant effects in the data are deleted.
[0314] The stepwise discrimination method uses the WILKS'S LAMBDA criterion to measure the strength of feature capabilities and eliminates features that do not meet the set threshold among the remaining features after three rounds of screening. In the stepwise discrimination process, the variable with the strongest discriminant ability is added first. As the variables in the model gradually increase, the discriminant ability of the variables introduced earlier may also change. If the discriminant ability of a variable in the model is less than the threshold, the variable is removed, and then the process is repeated until the variables contained in the model meet the WILKS'S LAMBDA similarity ratio criterion and other variables do not meet the criteria for entering the model.
[0315] In this application, the step of screening using the stepwise discriminant method is a more critical step. For the purpose of capturing as many credit risk points as possible, the model involved in this application uses a huge amount of data and many data dimensions, so variable screening must be performed to further reduce the time cost of model development. In the existing financial risk model evaluation, the variable importance method is generally used to reduce variables, such as calculating the importance of variables through algorithms such as the Gini index and information entropy, and selecting variables with high importance. The modeling method of screening using the stepwise discriminant method is rarely used. Compared with the variable importance screening scheme commonly used in the industry, the methodology used in this application can retain a large number of variables with relatively weak importance but relatively independent information dimensions.
[0316] In a specific way, more than 3,000 derivative data can be processed based on, for example, 120 types of basic data. After three rounds of screening of missing rate, concentration and information value IV, about 20-30% of the features with poor data can be deleted, but more than 2,000 features will still be retained. If subsequent variable fine screening is performed based on all the above features, it will seriously affect the development efficiency. Therefore, after comparing and testing a variety of different variable deletion schemes, it was finally determined that stepwise discrimination was the best choice.
[0317] For the important features selected by the stepwise discrimination algorithm, further feature selection is performed based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample. Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and eliminate the features whose actual bad debt rate distribution does not conform to the business trend. Specifically, (1) Divide the remaining target features into 10-20 boxes according to the quantile points, and calculate the median value of each box and the corresponding bad debt rate. (2) Use the median value of each box and the previous box and the corresponding bad debt rate to calculate the rate of change (slope). (3) Count the number of boxes greater than 0 and non-0 boxes (non-greater than 0 boxes) in the rate of change between two adjacent boxes, and calculate the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change. (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the rate of change to the number of non-0 boxes in the rate of change, check whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend. The criteria for whether they are approximately consistent are as follows: First, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate increases (such as the credit line utilization rate, etc.). In this module, the features whose number of boxes greater than 0 in the above calculated slope accounts for less than 70% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated. Second, if the business logic trend of this feature should be: as the value of this feature increases, the bad debt rate decreases (such as the deposit amount, etc.). In this module, the features whose number of boxes greater than 0 in the above calculated slope accounts for more than 30% of the number of boxes with a slope other than 0 are eliminated, that is, the features whose feature performance does not conform to the business trend are eliminated.
[0318] For example, Table 1 below gives an example of binning. In this example, the feature values are divided into 10 bins. The value ranges of the bins are also summarized in the table below. At the same time, a graph of the bad debt rate of the feature bins versus the median of the feature bins is drawn based on the data in the table below, as shown in Figure 1 As shown. Figure 1 In the example, there are 7 with slopes greater than 0 and 2 with slopes less than 0. Through the above method, the features can be further screened based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample. By further screening features using such a binning method, the features that best match the business development trend can be effectively screened out, thereby further obtaining features suitable for modeling.
[0319] Table 1
[0320]
[0321] <Conversion of eigenvalues>
[0322] The initial feature processing of existing credit risk scoring methods mainly uses WOE conversion (multi-classification) and dummy feature conversion (binary classification) to discretize continuous features (such as age, account age, etc.). The WOE conversion method is an optimal binning scheme based on modeling samples. It discretizes continuous features according to the optimal cutting point and embeds the data of good and bad samples into the WOE value. Therefore, it performs better in model construction, but the binning results may overfit the modeling samples, resulting in a serious decline in model effect when applied to the whole (poor generalization ability). At the same time, due to the normalization operation used in binning, the original features falling into different bins are converted into a single value corresponding to each bin, thus losing the ability to distinguish the risks of people falling into the same interval. The dummy feature conversion method is mainly applied to grouping features. Its advantage is that it can eliminate the difference between good and bad values of different feature values, but it will become extremely complex and redundant when processing continuous variables.
[0323] In the model building method of the present application, the conversion method of the features after preliminary screening is judged to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and the optimal method is used to perform feature conversion for each feature after preliminary screening.
[0324] The initial screening data conversion step includes the following steps based on the judgment of concentration and data type: classifying each feature into character variables and numeric variables according to the data type of each feature, using dummy feature conversion method to perform initial screening data conversion on character variables, and further classifying the numeric variables includes the following sub-steps: if the value of the numeric variable is less than n, the WOE conversion method is used for initial screening data conversion; if the value of the numeric variable is more than n, further judging that if the value converted to a continuous variable is large and the concentration of a single value is greater than m%, the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, a continuous conversion method is used; preferably, n and m are both positive integers, wherein n=5-10, m=90-99.
[0325] Specifically, taking the character variable as education level, the value of this feature variable can be primary school, middle school, university, graduate school, etc. For numeric variables, if the numeric variable is the number of overdue months in the past three months, the value is 0, 1, 2, 3.
[0326] In the present application, the above n can be 5, 6, 7, 8, 9 or 10, and m can be 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99.
[0327] In a specific implementation, m is selected as 5 and n is selected as 95.
[0328] Specifically, the WOE conversion method is: find the optimal cutting point of the feature, divide the value range of the original feature into multiple bins, calculate the WOE conversion value corresponding to each bin after segmentation according to the good or bad situation in each bin and output it, divide the original feature according to the binning result and output the WOE conversion value. For each bin, the WOE value is calculated as follows:
[0329]
[0330] where y i is the number of non-defaulting customers in group i; s is the total number of customers who have not defaulted; n i is the total number of defaulting customers in group i; n s is the total number of defaulting customers.
[0331] Let's take age as an example to introduce the WOE binning process. The original feature age includes values from 18 to 50 years old. After binning, 5 bins are obtained, namely 18 to 24 years old, 25 to 30 years old, 31 to 35 years old, 36 to 42 years old, and 43 to 50 years old. Then, the WOE conversion value of each bin is calculated according to the number of good and bad customers in each bin. Finally, each user data in each bin is output according to the corresponding conversion value of each bin. For example, 23 years old outputs the WOE conversion value corresponding to the 18 to 24 bins, and 46 years old outputs the WOE conversion value corresponding to the 43 to 50 bins.
[0332] The conversion method of dummy features is: convert a single classification feature into an equal number of dummy features according to the number of values it contains. If a customer belongs to the corresponding value of the generated dummy feature, the corresponding dummy feature value is 1, and the remaining dummy feature values are 0.
[0333] Taking gender as an example, we will introduce the method of dummy feature conversion. The original features include: male, female. After dummy feature conversion, two dummy features, 'gender-male' and 'gender-female', are generated. If the customer's gender is male, 'gender-male' is recorded as 1, and 'gender-female' is recorded as 0. If we take education as an example, the original features include: junior college and below, undergraduate, master and above. After dummy feature conversion, three dummy features, 'education-junior college and below', 'education-bachelor', and 'education-master and above', are generated. If the customer's education is undergraduate, 'education-junior college and below' is recorded as 0, 'education-bachelor' is recorded as 1, and 'education-master and above' is recorded as 0.
[0334] The continuous conversion method is: perform a variety of continuous conversions on the original features (continuous conversion methods include but are not limited to: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, and calculating the natural logarithm of the original data). Calculate the correlation coefficient r (Correlation Coefficient) between the feature value and the overdue label after continuous conversion, select the conversion method with the largest absolute value of the correlation coefficient, and output the conversion method corresponding to the original feature. The calculation formula of the correlation coefficient is as follows:
[0335]
[0336] Where Σ is the summation symbol in mathematics; n is the total number of observations; x i is the conversion value of the original feature of the i-th observation after continuous conversion; This is the average of the transformed values; where y i is a binary feature indicating whether the i-th observation is in default; It is the average value of this binary feature. The closer the absolute value of the correlation coefficient is to 1, the more relevant the converted value is to the default situation, and the better the effect of this conversion method. The quantitative index of the correlation coefficient r has the following meanings: when the absolute value of the correlation coefficient is above 0 and less than 0.3, it indicates a low correlation; when the absolute value of the correlation coefficient is above 0.3 and less than 0.8, it indicates a moderate correlation; when the absolute value of the correlation coefficient is above 0.8 and less than 1, it indicates a high correlation.
[0337] In the model construction of the present application, for the features that are confirmed to adopt a continuous conversion method, the optimal conversion method is selected to perform continuous feature conversion of the feature based on the correlation of the feature with credit default under different continuous conversion methods. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
[0338] Take age as an example to introduce the continuous conversion method. The original feature age contains values from 18 to 50 years old. The continuous conversion obtains the original value, square, square root, cube root, natural logarithm, etc. of age, and then calculates the WOE value and the absolute value of the correlation coefficient between each conversion value and the good or bad label (prediction result). The conversion method with the largest absolute value of the correlation coefficient is selected for conversion and output. If the cube root of the age feature has the largest absolute value of the correlation coefficient with the good or bad label compared with other conversion methods, the cube root of age is output as the conversion value.
[0339] In the prior art solutions, one of the feature conversion methods, WOE and dummy features, is mainly selected for model construction. For example, in the Chinese patent CN112686749B, the WOE method is used to convert the feature values. The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes the continuous features according to the optimal cutting point, so it performs better in model construction, but the binning results may overfit the modeling samples, resulting in a serious decrease in the model effect when applied to the whole (poor generalization ability). At the same time, due to the normalization operation used in the binning, the original features falling into different bins are converted into a single value corresponding to each bin, so the risk differentiation ability of people falling into the same interval is lost.
[0340] The conversion method of dummy features is mainly used for grouping features. Its advantage is that it can eliminate the difference between good and bad values of different feature values. For example, for the industry in which the customer is located, because there is no obvious advantage or disadvantage between the retail industry and the wholesale industry, it is more suitable to use the dummy feature processing method; there is a hierarchical difference between junior colleges and undergraduates, so although dummy features can be used, it is actually more appropriate to use the WOE processing method.
[0341] On the other hand, the continuous conversion method can avoid the overfitting of the modeling samples in the WOE conversion method, has a strong generalization ability for the overall sample, and because no interval mapping is performed, most customer groups rarely fall into a single value. However, it cannot be applied to some features with poor monotonicity or discrete characteristics (such as occupation, position, etc.).
[0342] As described above, the technical solution of the present application takes a different approach and creatively combines three methods: continuous conversion, WOE conversion, and dumb feature conversion. It reprocesses some carefully screened features, combines the advantages and disadvantages of the three conversion methods, and creatively designs a conversion judgment method. It selects the optimal feature conversion method based on parameters such as feature data attributes, missing rate, concentration, and business logic judgment.
[0343] <Logistic regression and deep feature screening based on logistic regression>
[0344] In existing credit risk scoring, due to the requirement of model interpretability, logistic regression model is mainly used for model development, and the software that can be used are generally: SAS, R, Python, etc.
[0345] In a specific embodiment, the present application develops models based on SAS software.
[0346] Specifically, the Sigmoid function is used in logistic regression to fit the probability of predicted default. The Sigmoid function is:
[0347]
[0348] Where Z is a linear combination of the model coefficients and the feature transformation values, and Z is defined as follows:
[0349] Z=α+β 1 x 1 +β 2 x 2 +...+β k-1 x k-1 +β k x k
[0350] The predicted probability of default is:
[0351] P=P(Y=1|x 1 ,x 2 ,x 3 ,...,x k-1 ,x k )
[0352] The fitted prediction for the probability of default is:
[0353]
[0354] From the above formula we can further deduce:
[0355] Substituting the Z value into the above formula can calculate the probability P of predicted default.
[0356] The core of logistic regression model construction is feature screening. The steps of feature screening are as follows: First, batch screen the features according to missing rate, concentration and information value IV. Second, screen all remaining features one by one according to whether the features are in line with business trends, and retain the features with correct business trends. For example: if it is found that the bad debt rate of the customer group decreases with the increase of the characteristic loan balance, it is considered that this feature does not conform to the business trend. In the understanding of credit business, the higher the loan balance, the higher the customer's default risk exposure (EAD) level, and the greater the risk. At this time, this feature will be removed from the feature list. Third, use the stepwise regression function of logistic regression to eliminate the less important features that are highly correlated with other features. Fourth, filter the coefficients according to the positive and negative signs of the training coefficients and the business trend of the feature conversion value, and retain the features whose feature coefficient signs conform to the business logic. In the case where 0 is defined as a good customer and 1 is defined as a bad customer in the Y label, for a feature whose bad debt rate increases monotonically with the increase of the feature value (such as loan balance), the sign of its training coefficient should be positive; conversely, for a feature whose bad debt rate decreases monotonically with the increase of the feature value (such as deposit balance), the sign of its training coefficient should be negative. If it does not meet the above standards, it should be eliminated. Fifth, use variance inflation factor (VIF), correlation coefficient, etc. to further eliminate features with high correlation: for variance inflation factor, eliminate features with the highest VIF and greater than 4 one by one; for features with high correlation, eliminate features with low IV values in the feature group with the highest correlation coefficient and greater than 0.80 one by one. Sixth, use the population stability index (PSI) to eliminate features with large distribution differences at different time points, which lead to instability. For features with PSI>0.25, directly eliminate them, and for features with 0.25>PSI>0.1, carefully eliminate them according to the impact of eliminating this feature on the model's distinguishing ability.
[0357] This application strictly adheres to the above rules to screen features, ensuring the model's interpretability, stability, and ability to distinguish between good and bad customers.
[0358] The feature fine-screening steps of the present application include: a first fine-screening step, based on a stepwise regression algorithm, to screen features based on the F test and the T test for the significance of features; a second fine-screening step, to screen features based on calculating the variance inflation factor for each feature and eliminating features with higher variance inflation factors; a third fine-screening step, based on logistic regression, to analyze whether the feature coefficients of the features after the first fine-screening step and the second fine-screening step are in line with the trend of the prediction results for credit default so as to further perform feature screening.
[0359] Feature fine screening step 1, based on the stepwise regression algorithm, based on the F test and T test method, introduces features in order from high to low significance. Each time a feature is introduced, the selected features are tested one by one. When the originally introduced feature becomes no longer significant due to the introduction of the subsequent features, it is eliminated. This process is repeated until no features above the significance threshold are selected into the equation, and no features below the significance threshold are eliminated from the regression equation.
[0360] Feature fine-screening step 2 is based on the method of eliminating features with high variance inflation factors to further reduce multicollinearity in the model.
[0361] Feature fine screening step 3, based on the risk characteristics of each risk point itself (system preset) and the positive and negative signs of the model training coefficient, determines whether the feature coefficients of the remaining features in the model in the feature fine screening step 3 are consistent with the business trend, removes the features whose model coefficients do not meet the business trend, and iterates again. The specific implementation plan of feature fine screening module 3 is as follows: 1. For features whose feature conversion method is WOE type, the corresponding model training coefficient should be a negative value, and the WOE conversion type features with positive training coefficients should be eliminated. 2. For continuous conversion methods, if the value of this feature increases in business logic, the bad debt rate should increase (such as the credit line utilization rate, etc.), then the corresponding model training coefficient should be a positive value, and such continuous conversion type features with negative training coefficients should be eliminated; if the value of this feature increases in business logic, the bad debt rate should decrease (such as the deposit amount, etc.), then the corresponding model training coefficient should be a negative value, and such continuous conversion type features with positive training coefficients should be eliminated. 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.
[0362] Feature fine screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated, and features with 0.25>PSI>0.1 are carefully eliminated based on the impact of eliminating this feature on the model's distinguishing ability.
[0363] Based on the above steps 3 and 4, the feature list is iterated multiple times until no new features are added or removed from the model, and the iteration is stopped to obtain the final feature list and its conversion value. After the above steps, the input variables for constructing the model of this application can be finally obtained.
[0364] The credit default probability modeling step of the present application substitutes the features screened in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.
[0365] In the prior art, the core of all scoring models built is whether the data they use is representative. For the existing scoring models in the prior art, due to information security and cost reasons, the sample size and bad labels of most models are small, and the stability of the model cannot be guaranteed, let alone independent modeling of a certain type of customer. At the same time, because the scoring models on the market can obtain information dimensions in the modeling process, mostly multi-loan data and weak financial attribute data (such as smart terminal device data, social platform data, online shopping mall data and other non-credit transaction data), it is impossible to accurately reflect the customer's asset status and repayment ability. The data source used in this application is the full business data of large banks, and the data volume of modeling samples and bad labels is very large. This application designs modeling samples based on different sample groups, which can more finely distinguish the risk differences between such customers, and at the same time ensure the stability of data and models based on cross-time verification and PSI verification. This application uses personal credit history data and asset data for model development, so it has a good reflection of the borrower's repayment willingness and repayment ability.
[0366] For some small and medium-sized financial institutions, the internally built scoring models are relatively small or started late in their inclusive credit business, and the accumulated historical data is insufficient to develop a credit risk model with stable data and strong differentiation ability. Therefore, they are highly dependent on manual approval when used for credit review business. The efficiency limitations of manual approval have restricted the development of their inclusive credit business. At the same time, the subjectivity of manual approval has increased the operational risks in the credit review process. The model constructed in this application can assist such financial institutions in making digital decisions, enhance their approval accuracy and speed, and reduce the above-mentioned adverse effects.
[0367] In terms of feature conversion, compared with the traditional WOE conversion method, it is necessary to roughly bin and discretize continuous features according to data performance and the experience of modelers. The results of rough binning are greatly affected by the subjective factors of modelers; and the process of discretizing continuous features may result in too many customers falling into the same interval, which leads to more single values of the calculated scores. This application combines the continuous conversion method, the WOE conversion method and the dummy feature conversion method to encode the original features, which reduces the impact of human factors and single values while enhancing the ability to distinguish scores.
[0368] <Credit risk model building device, system and computer storage medium>
[0369] The present application relates to a device for predicting the construction of a credit risk default probability model, the device comprising: a data acquisition module, which is used to obtain original credit prediction data of samples used to build the model; a data derivation module, which is used to process derived credit prediction data based on the original credit prediction data; a feature initial screening module, which is used to perform preliminary screening of all categories including original credit prediction data and derived credit prediction data, that is, all features, to obtain features after preliminary screening; a preliminary screening data conversion module, which is used to judge the conversion method of the features after preliminary screening to confirm that one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method is used for feature conversion, and for each feature after preliminary screening, the best method is used for feature conversion; a feature fine screening module, which is used to perform in-depth screening of the features after preliminary screening after feature conversion to obtain finely screened features; a credit default probability modeling module, which is used to select a logistic regression method for model construction based on the probability relationship between the finely screened features and credit default, and confirm the method for calculating the credit default probability.
[0370] Furthermore, the model building device involved in the present application includes a sample selection module, which is used to screen all users before the data collection step to obtain samples for model building.
[0371] The present application also relates to a system for constructing a model for predicting credit risk default probability, the system comprising: a memory, a processor, and a program for constructing a model for predicting credit risk default probability stored in the memory and executable on the processor, wherein the program for constructing a model for predicting credit risk default probability, when executed by the processor, implements the steps of the method for constructing a model for predicting credit risk default probability described in the present application.
[0372] The present application relates to a computer storage medium, on which is stored a program for constructing a model for predicting the probability of default of credit risk. When the program for constructing a model for predicting the probability of default of credit risk is executed by a processor, the steps of the method for constructing a model for predicting the probability of default of credit risk described in the present application are implemented.
[0373] All the contents described above for constructing a method for predicting a credit risk default probability model can be fully applicable to a device, system and computer storage medium for constructing a credit risk default probability model.
[0374] <Methods for calculating credit risk or methods for predicting the probability of default under credit risk>
[0375] The present application further relates to a method for calculating the credit risk of a sample to be tested (i.e., for predicting the credit risk default probability) by using the credit risk model constructed in the present application (i.e., a model for predicting the credit risk default probability).
[0376] In this application, credit risk refers to the risk that the sample to be tested cannot repay the debt as agreed. In a specific manner, the probability of credit default refers to the probability of credit default, specifically, for example, whether a credit default of more than 90 days will occur.
[0377] The present application relates to a method for calculating credit risk or predicting the probability of default of credit risk, which comprises: a data collection step, which obtains credit prediction data of samples to be predicted; a step of classifying the samples to be predicted, which classifies the samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; a step of calculating the probability of credit default, which substitutes the credit prediction data into the credit default probability sub-model to calculate the probability of credit default of the samples to be predicted.
[0378] In the present application, the sample to be predicted may be a sample used to build a model, that is, when the model is built, the sample is already a customer of a financial institution holding a credit business (including but not limited to credit cards or personal loan business), and the method for calculating credit risk of the present application can be used to calculate its potential credit risk in the future. In the present application, the sample to be predicted may be a sample that is not used to build a model, that is, when the model is built, it is not a customer of a financial institution holding a credit business, but now it is a customer of a financial institution holding a credit business, and the method for calculating credit risk of the present application can be used to calculate its potential credit risk in the future. In the present application, the sample may also be a customer who was not a customer of the financial institution holding a credit business when the model was built in the past, but is not a customer of the financial institution holding other financial asset businesses, and the method for calculating credit risk of the present application can be used to calculate its potential credit risk in the future.
[0379] The present application relates to a method for calculating credit risk or predicting the probability of default of credit risk, which comprises: a data collection step, which obtains credit prediction data of samples to be predicted; a step of classifying samples to be predicted, which classifies samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; a step of calculating the probability of credit default, which substitutes the credit prediction data into the credit default probability sub-model to calculate the probability of credit default of the sample to be predicted; and a step of calculating the credit score of the sample to be predicted after calculating the probability of credit default.
[0380] The step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.
[0381] In the method of the present application, the credit prediction data includes the original credit prediction data of the sample to be predicted and the derived credit prediction data processed based on the original credit prediction data. The description of the original credit prediction data and the derived credit prediction data is consistent with the description of the model building part.
[0382] In this application, the step of classifying the sample to be predicted includes the following sub-steps:
[0383] Whether the sample to be tested is a customer with a corporate deposit account age greater than a months;
[0384] Whether the sample to be tested is a customer whose average balance of corporate deposit account is less than c yuan in the past b months;
[0385] Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model for calculating the probability of credit default, and the order of performing the above sub-steps can be set arbitrarily.
[0386] In this application, Figure 2 The flowchart for classifying the samples to be predicted is shown in Figure 1. Based on the classification in this step, the sub-model for calculating the probability of credit default can be selected. Figure 2 The sample to be predicted is classified in the following order: first, determine whether the sample to be tested is a customer whose corporate deposit age is greater than a months; then determine whether the sample to be tested is a customer whose average corporate deposit balance in the past b months is less than c yuan. Using this step classification, the most suitable sub-model for predicting the sample to be tested can be determined.
[0387] Public deposits (also known as "unit deposits") refer to RMB deposits made by enterprises, institutions, government agencies, military units and social groups in financial institutions, including time deposits, demand deposits, notice deposits, agreement deposits and other deposits approved by the People's Bank of China.
[0388] Corporate deposit accounts include but are not limited to basic deposit accounts, general deposit accounts, temporary deposit accounts, special deposit accounts, time deposit accounts, notice deposit accounts, agreement deposit accounts, etc.
[0389] In the present application, the credit prediction data is subjected to feature conversion and then substituted into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes selecting the WOE method or the continuous method for feature conversion based on the feature type of the credit prediction data to be substituted into the credit default probability sub-model.
[0390] The method of converting features in WOE mode or continuous mode is well known to those skilled in the art, and the specific conversion method can be referred to the description in the model construction section of this application. However, the existing models in the art only use WOE mode to convert parameters.
[0391] The WOE conversion method is an optimal binning scheme based on the modeling sample. It discretizes the continuous features according to the optimal cutting point. Therefore, it performs better in model construction, but the binning results may overfit the modeling sample, resulting in a serious decline in the model effect when applied to the overall population (poor generalization ability). At the same time, due to the normalization operation used in binning, the original features falling into different bins are converted into a single value corresponding to each bin, thus losing the ability to distinguish the risks of people falling into the same interval. The conversion method of dummy features is mainly applied to grouping features. Its advantage is that it can eliminate the difference between good and bad values of features. For example: for the industry in which the customer is located, because there is no obvious advantage or disadvantage between the retail industry and the wholesale industry, it is more suitable for the dummy feature processing method; there is a hierarchical difference between junior colleges and undergraduates, so although dummy features can be used, it is actually more appropriate to use the WOE processing method. The continuous conversion method can avoid the overfitting of the modeling sample in the WOE conversion method, and has a strong generalization ability for the overall sample. Since no interval mapping is performed, most customer groups rarely fall into a single value. However, it cannot be applied to some features with poor monotonicity or discreteness (such as occupations, positions, etc.). The present application uses a combination of continuous conversion and WOE conversion to process the initial features, combines the advantages and disadvantages of the two conversion methods, designs a conversion judgment module, and selects the optimal feature conversion method based on the parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic judgment. The technical solution of the present application takes a different approach and creatively combines the three methods of continuous conversion, WOE conversion, and dummy feature conversion to reprocess some carefully screened features, combines the advantages and disadvantages of the three conversion methods, and creatively designs a conversion judgment method, and selects the optimal feature conversion method based on the parameters such as feature data attributes, missing rate, and concentration, supplemented by business logic judgment. The feature conversion method determined in this way is then substituted into the model for calculating credit risk, which can more accurately predict the credit risk of the sample.
[0392] The sub-model used in the present application to calculate the probability of credit default is a model constructed based on sample credit prediction data and the probability of credit default using logistic regression based on an existing user population, that is, a model constructed using the method described in the present application.
[0393] <Device, system and computer storage medium for predicting credit risk default probability>
[0394] The present application relates to a device for predicting the probability of default of credit risk, which comprises: a data acquisition module, which is used to obtain credit prediction data of samples to be predicted; a module for classifying samples to be predicted, which is used to classify samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; a credit default probability calculation module, which is used to substitute the credit prediction data into the credit default probability sub-model to calculate the credit default probability of the samples to be predicted.
[0395] The device for predicting the credit risk default probability of the present application can execute the steps of the method for calculating the credit risk or predicting the credit risk default probability of the present application.
[0396] The present application also relates to a system for calculating credit risk or for predicting the probability of default of credit risk, characterized in that the system for calculating credit risk or for predicting the probability of default of credit risk comprises: a memory, a processor, and a program for calculating credit risk stored in the memory and executable on the processor, wherein the program for calculating credit risk implements the steps of the method for calculating credit risk as described in the present application when executed by the processor.
[0397] The present application relates to a computer storage medium, characterized in that a program for predicting a credit risk default probability is stored on the computer storage medium, and when the program for predicting a credit risk default probability is executed by a processor, the steps of the method for predicting a credit risk default probability as described in the present application are implemented.
[0398] All the contents described above for the method of calculating credit risk or predicting the probability of default of credit risk can be fully applied to the device, system and computer storage medium for calculating credit risk or predicting the probability of default of credit risk.
[0399] The method for calculating credit risk or predicting the probability of default of credit risk in the present application can avoid the technical shortcomings of overfitting when building a model with complete WOE variables and the inability to adapt well to categorical variables when building a model with complete continuous variables. Therefore, when the model is used for credit risk calculation, it can better cover the credit risk prediction needs of various types of customer groups and meet the demand for obtaining more accurate prediction results when predicting credit risk as a general score.
[0400] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0401] Example
[0402] Example 1 Collection of modeling samples
[0403] In the process of building this embodiment, the original corporate credit prediction data for building model samples between 2017 and 2021 were collected, including various data information of enterprises and actual controllers of enterprises, totaling 18 million corporate users. The modeling samples were further confirmed through a professional model design scheme, and the enterprise data of 15 million households in 2019 were selected as analysis samples. The user samples were then divided into 1.35 million households that have applied for inclusive credit business and 13.65 million households that have not applied for inclusive credit business. When designing the specific model, the 1.35 million households that have applied were analyzed and designed. After the model is developed, the online results will be applied to the full 18 million corporate customer groups and 600 million retail customer groups that have applied for inclusive credit business and have not applied for inclusive credit business.
[0404] The model design includes 1) exclusion rules: such as excluding customer data that has been settled, closed, has no performance, or has special circumstances. After exclusion, the modeling population is 740,000 customers (inclusive credit customers); 2) time window setting: determine the sample to use 2019 data as the modeling sample, and use the next 15 months as the performance period of Y; 3) sample sampling: use a good / bad sample number of 4:1 for sampling modeling. Different segmentation schemes are designed through the decision tree method, and the parent-child model comparison method is used to confirm the final model segmentation plan, and then each segmentation model is modeled separately. The model segmentation scheme of this application is designed by dividing whether the corporate customer's corporate deposit age is greater than 6 months, whether the average monthly balance of corporate deposits in the past 6 months is less than 1,000, etc. It fully reflects the control variables that are closely related to risk characteristics in the process of business development, and can more comprehensively match various customer groups in the market.
[0405] In this embodiment, the modeling samples are first divided into the following sub-model modeling samples based on the decision tree to facilitate the construction of subsequent sub-models. The first layer of the decision tree when determining the modeling samples is overdue, which is used to distinguish the customer's historical repayment behavior. In this embodiment, for customers with a corporate deposit age of more than 6 months, the customer group is divided into two models, Zeta2 and Zeta3, based on their average monthly balance of corporate deposits in the past 6 months; the customer group with a corporate deposit age of less than or equal to 6 is divided into Zeta1, where the bad debt rate refers to the ratio of bad customers in a certain type of sample to the total number of samples in this type.
[0406] For customers with corporate deposit age greater than 6 years, if the average balance of corporate deposits in the past 6 months is less than 1,000 yuan, they will enter the Zeta2 segmentation model; if the average balance of corporate deposits in the past 6 months is greater than or equal to 1,000 yuan, they will enter the Zeta3 segmentation model.
[0407] Specifically, the sample size of a sub-model used to construct the Zeta 1 segmentation model is about 90,000. Its customer base is mainly corporate deposit accounts with an age of less than or equal to 6 years. The model is used to predict the probability of credit delinquency of more than 30 days in this group of people.
[0408] Specifically, the sample size of a sub-model used to build the Zeta 2 segmentation model is about 70,000. Its customer base is mainly corporate deposit accounts with an age of less than or equal to 6 years and an average balance of corporate deposit accounts of less than 1,000 yuan in the past 6 months. The model is used to predict the probability of credit delinquency of more than 30 days in this group of people.
[0409] Specifically, the sample size of a sub-model used to build the Zeta 3 segmentation model is about 470,000. Its customer base is mainly corporate deposit accounts with an age of less than or equal to 6 years and an average balance of corporate deposit accounts greater than or equal to 1,000 yuan in the past 6 months. The model is built to predict the probability of credit delinquency of more than 30 days in this group of people.
[0410] Based on the different customer samples of each sub-model confirmed above, historical data information of borrowing enterprises is obtained: 1) Enterprise credit data, basic fields include repayment, overdue, credit, credit use, balance, account age, guarantee mortgage, application rejection and other information; 2) Enterprise public deposit category, including deposit balance, inflow amount, outflow amount, deposit fluctuation, industry deposit level, industry transaction level and other information; 3) Enterprise basic information, including industry, region, registration and other information; 4) Business owner financial assets, including AUM, deposits, wealth management and payroll, deposit account transaction flow and other information; 5) Business owner credit, including overdue, repayment, balance, account age, credit utilization rate, credit and other information; 6) Including information such as gender, age and administrative region to which the business belongs.
[0411] According to various risk points in the credit business, macro credit risk is split according to information dimensions, time slices, etc. Variables are derived through professional characteristic variable construction methods. The basic methods are: 1) Customer relationship length variables: such as customer account opening time, customer maximum account age, etc. 2) Time interval variables: such as the number of months from the customer's last repayment to the observation point, the number of months from the customer's last overdue payment, etc.; 3) Behavior frequency degree variables: such as the number of times the customer's repayments are > N in the last X months, the number of times the customer's credit utilization rate is > N in the last X months, etc.; 4) Current time point variables: customer's current monthly credit limit, customer's current monthly balance, etc.; 5) Statistical value variables: such as the customer's maximum number of overdue periods in the last X months, the customer's average credit utilization rate in the last X months; 6) Continuous behavior variables: such as the maximum number of consecutive overdue periods > N in the last X months, the number of consecutive repayment rates > N in the last X months, and other characteristic variables. Finally, 3067 features with potential predictive power for customer overdue payment were derived, such as 'current repayment amount', 'average number of overdue periods in the past 6 months', and 'number of credit card installments in the past 12 months'. In this article, observation point refers to the time point when samples are collected until modeling, and current also refers to the sampling cutoff time. Observation point and current have the same meaning.
[0412] Table 1 Summary of basic variables and derived variables used in the examples
[0413]
[0414] Example 2 Feature Screening
[0415] A preliminary screening was performed on the 6753 features (variables) collected in Example 1 that have potential predictive power for customer overdue payments.
[0416] The preliminary screening of the three sub-model features from Zeta 1 to Zeta 3 divided in Example 1 is carried out according to the following method:
[0417] In the first round of preliminary screening, the features with a missing rate of more than 95% in the data collected in Example 1 were first eliminated, and a total of 1,392 variables were deleted, leaving 5,361 variables.
[0418] In the second round of preliminary screening, for the 5361 features that passed the first round of preliminary screening, the features with single values exceeding 99% among the remaining features were eliminated, a total of 1023 variables were eliminated, and 4338 variables remained.
[0419] In the third round of preliminary screening, the 4338 features after the second round of preliminary screening are sorted according to the feature values (the sorting method is specifically that if the feature is a character variable, each value is a separate box, if the feature is a numeric variable, it is sorted from small to large according to the value size), and then divided into 10-20 boxes according to the quantile point, and the feature IV value is calculated. If the feature IV value is lower than 0.02, it is eliminated. In this embodiment, since the variable attributes are strong financial and asset variables, the prediction effect is strong. Here, variables with IV values lower than 0.05 are eliminated. A total of 1895 features are eliminated, and 2443 feature variables remain.
[0420] In the fourth round of preliminary screening, the 2443 features after the third round of screening were further screened based on the stepwise discrimination algorithm. After this round of screening, 583 important feature variables can be quickly screened out. The following is an introduction to the stepwise discrimination method. In actual data, there may be little difference in the mean values of different categories on a certain variable. In this case, the effect of using this variable for classification will not be very good; there is also a type of variable that can indeed distinguish different categories in the data well when considered independently, but after including these variables in the model variables, they may appear redundant. Therefore, the embodiment pioneered the use of the stepwise discrimination method to better select more important feature variables under the same dimension, greatly reducing the workload of judging and screening one by one according to the trend of the variables in the fifth round of screening, and fully improving the efficiency of model development without affecting the overall model effect.
[0421] The stepwise discrimination method uses the WILKS'S LAMBDA criterion to measure the strength of feature capabilities and eliminates features that do not meet the set threshold among the remaining features after three rounds of screening. In the stepwise discrimination process, the variable with the strongest discriminant ability is added first. As the variables in the model gradually increase, the discriminant ability of the variables introduced earlier may also change. If the discriminant ability of a variable in the model is less than the threshold, the variable is removed, and then the process is repeated until the variables contained in the model meet the WILKS'S LAMBDA similarity ratio criterion and other variables do not meet the criteria for entering the model.
[0422] The fifth round of preliminary screening is to further screen the 583 important features after the fourth round of screening based on the risk characteristics of each risk point itself (system preset) and the actual bad debt rate of the sample.
[0423] Determine whether the actual bad debt rate distribution of the remaining features is consistent with the business trend, and remove the features whose actual bad debt rate distribution of the samples does not conform to the business trend. Specifically,
[0424] (1) Divide the remaining target features into 10-20 boxes according to the quantiles, and calculate the median of each box and the corresponding bad debt rate.
[0425] (2) Calculate the rate of change (slope) using the median of the values of each box compared to the previous box and the corresponding bad debt rate.
[0426] (3) Count the number of boxes with a change rate greater than 0 and the number of boxes with a change rate greater than 0 between two adjacent boxes, and calculate the percentage of boxes with a change rate greater than 0 to the number of boxes with a change rate greater than 0.
[0427] (4) Obtain the risk characteristics preset for this feature, and based on the percentage of the number of boxes greater than 0 in the above-calculated change rate to the number of non-zero boxes in the change rate, determine whether the actual bad debt rate change trend is approximately consistent with the business logic bad debt rate change trend.
[0428] The criteria for approximate consistency are as follows:
[0429] First, if the business logic trend of this feature is: as the value of this feature increases, the bad debt rate increases (such as the credit line utilization rate, etc.), in this module, the features whose number of boxes with slope greater than 0 in the above calculation accounts for less than 70% of the number of boxes with slope non-0 are eliminated, that is, the features that do not conform to the business trend are eliminated.
[0430] Second, if the business logic trend of this feature is: as the value of this feature increases, the bad debt rate decreases (for example, the amount of deposits, etc.), in this module, the features whose number of boxes with slope greater than 0 in the above calculation accounts for more than 30% of the number of boxes with slope not 0 are eliminated, that is, the features that do not conform to the business trend are eliminated.
[0431] After the fifth round of screening, 187 features were eliminated, leaving 396 features.
[0432] Example 3 Conversion of features after initial screening
[0433] In the feature judgment step, the optimal conversion method is selected according to the concentration of features after the initial feature screening step, data type, etc. When judging the conversion method, data types are generally divided into two categories, one is character variables, and the other is numeric variables. For character variables, the dummy variable conversion method is generally used for variable conversion; for numeric variables, if the variable value is less than 5 or so, the WOE conversion method will be used. If the variable value is more, the WOE or continuous conversion method will be used (the optimal conversion form is selected by the correlation with the target variable). In this process, it is necessary to comprehensively consider the situation of variable concentration. For example, if the continuous variable has many values, but the concentration of a single value exceeds 95%, the continuous processing will not be performed, and the WOE conversion method can be used directly. Features using different conversion methods are grouped and divided into different data sets. Specifically,
[0434] First, the remaining 396 features in Example 2 are used to determine the conversion methods of these features. The following three methods are selected based on the concentration of the features, data types, etc.
[0435] Feature conversion method 1 is used to perform WOE conversion on the features obtained in the data acquisition module and judged in the feature judgment module that the optimal conversion method is WOE conversion.
[0436] Feature conversion method 2 is used to perform dummy feature conversion on the features obtained in the data acquisition module and judged in the feature judgment module that the optimal conversion method is dummy feature conversion.
[0437] Feature conversion method 3 is used to obtain the features in the data acquisition module that are determined to be optimal for continuous conversion in the feature judgment module, select the optimal continuous conversion method, and perform continuous conversion.
[0438] The feature merging module is used to horizontally splice the data of feature conversion mode 1, feature conversion mode 2 and feature conversion mode 3.
[0439] Through this embodiment, 396 features are converted, of which 56 features are converted into WOE, 73 features are converted into dummy features (expanded into 187 variables), and 267 features are converted into continuous types.
[0440] Example 4 Feature Depth Screening (Feature Fine Screening Step)
[0441] The feature fine screening step mainly includes the following four steps in this embodiment. Steps 1 and 2 are based on stepwise regression and calculation of variance inflation factor to eliminate features with large multicollinearity in the feature merging module to enhance the robustness of the model. Step 3 eliminates features whose positive and negative signs of the training coefficients in the model do not conform to the business trend. Step 4 uses the population stability coefficient (PSI) to eliminate unstable features. This embodiment uses the LOGISTIC process in SAS for deep screening.
[0442] Feature fine screening step 1, based on the stepwise regression algorithm, based on the F test and T test method, introduces features in order from high to low significance. Each time a feature is introduced, the selected features are tested one by one. When the originally introduced feature becomes no longer significant due to the introduction of the subsequent features, it is eliminated. This process is repeated until no features above the significance threshold are selected into the equation, and no features below the significance threshold are eliminated from the regression equation. After this step, 396 features are eliminated to 278 features.
[0443] Feature refinement step 2, which is based on the method of removing features with high variance inflation factors, further reduces multicollinearity in the model. After this step, 278 features were eliminated to 145 features.
[0444] Feature fine screening step 3 compares the risk characteristics of each risk point itself (system preset) with the positive and negative signs of the model training coefficients to determine whether the feature coefficients of the remaining features in the model in step 3 are consistent with the business trend, eliminates the features whose model coefficients do not conform to the business trend, and iterates again.
[0445] The specific implementation scheme of the feature fine screening module 3 is as follows:
[0446] 1. For features whose feature conversion method is WOE type, the corresponding model training coefficient should be a negative value, and the WOE conversion type features with positive training coefficients should be eliminated.
[0447] 2. For continuous conversion methods, if the value of this feature increases in business logic, the bad debt rate should increase (such as the credit limit utilization rate, etc.), then the corresponding model training coefficient should be positive, and such continuous conversion features with negative training coefficients should be eliminated; if the value of this feature increases in business logic, the bad debt rate should decrease (such as the deposit amount, etc.), then the corresponding model training coefficient should be negative, and such continuous conversion features with positive training coefficients should be eliminated.
[0448] 3. For the dummy feature conversion method, it is necessary to determine the bad debt rate when the value is 1 and the value is 0. If the bad debt rate when the value is 1 is greater than the bad debt rate when the value is 0, the coefficient should be positive, otherwise it should be negative.
[0449] Feature fine screening step 4, data stability monitoring step, is used to evaluate whether the distribution of individual features and overall scores at different time points has obvious deviations. In this embodiment, features with PSI>0.25 are directly eliminated, and features with 0.25>PSI>0.1 are carefully eliminated based on the impact of eliminating this feature on the model's distinguishing ability.
[0450] Based on the above steps 3 and 4, the feature list is iterated multiple times until no new features are added or removed from the model, and the iteration is stopped to obtain the final feature list and its conversion value. After the above steps, 65 features are eliminated as the feature variables finally entered into the model, for example, 5 features, 6 features, 7 features, and 8 features.
[0451] Example 5 Construction of Zeta 1 Model
[0452] Generally speaking, the stronger the correlation between the characteristic variable and the target variable, the more accurate the final model will be. In this embodiment, Zeta 1 is used as a sub-model of the segmentation model. The above decision tree method is adopted to obtain a sample size of about 90,000. Its customer base is mainly customers with a corporate deposit age of less than or equal to 6 years. The model is constructed to predict the probability of credit overdue for more than 30 days (target variable) in this group.
[0453] In this embodiment 5, the minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the enterprise in the past six months, the average utilization rate of the enterprise owner's revolving loan in the past 12 months, the minimum current deposit balance of the enterprise in the past three months, the average monthly debit transaction amount of the enterprise in the past six months / the average monthly balance of the past six months, and the average value of (credit transaction amount-debit transaction amount) of the enterprise in the past six months are used as examples to describe the final modeling results. In the feature conversion step, the conversion method selection has been carried out according to the correlation between the feature and the target variable. The 6 modeling variables have been converted in a corresponding form. In the credit overdue probability modeling step, the 6 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) to perform logistic regression to calculate the credit overdue probability model.
[0454]
[0455] The maximum monthly average credit transaction amount of the enterprise in the past 6 months reflects the capital inflow of the enterprise. For inclusive corporate credit customers with public deposits of less than 6 months, the higher the capital inflow amount of the enterprise in the past 6 months, the higher the liquidity and income of the enterprise in the past six months, and the higher the debt repayment ability, the lower the possibility of overdue payment; conversely, the higher the possibility of overdue payment. The conversion method is the square root conversion value.
[0456] The average utilization rate of the revolving loan of business owners in the past 12 months reflects the credit demand of business owners in the past 12 months. For inclusive corporate credit customers with public deposits of less than 6 months, the higher the credit demand of their business owners in the past year, the higher their capital demand, and the greater the possibility of overdue payments of their small and micro enterprises in the future, and vice versa. The conversion method is cube root conversion.
[0457] The minimum current deposit balance of an enterprise in the past three months reflects the level of assets that the enterprise can flexibly spend in the near future. For inclusive enterprise credit customers with corporate deposit age less than 6 months, the higher the level of assets that can flexibly spend in the near future, the stronger their debt repayment ability, the lower the possibility of overdue in the future, and the higher the possibility of overdue in the future. The conversion method is natural logarithmic conversion.
[0458] The average monthly debit transaction amount of the enterprise in the past 6 months / the average monthly balance in the past 6 months reflects the proportion of the average monthly outflow amount and the average monthly balance of the enterprise in the past 6 months. For inclusive corporate credit customers with public deposits of less than 6 months, the higher the proportion of outflow amount in the past 6 months, the higher the recent capital demand, the greater the possibility of overdue payment in the future, and vice versa. The possibility of overdue payment in the future is smaller. The conversion method is square root conversion.
[0459] The average of (credit transaction amount - debit transaction amount) in the past six months reflects the net capital inflow of the enterprise in the past six months. For inclusive corporate credit customers with public deposits of less than six months, the more net capital inflow in the past six months, the better the asset status, the lower the possibility of future overdue payments, and vice versa. The conversion method is cube root conversion.
[0460] In the step of calculating the credit overdue probability, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the probability (P) that the borrower is overdue:
[0461]
[0462] Where k is the number of features entering the model, and in Formula 1, k is 6. The model performance of some features is shown in Table 1 below.
[0463] Table 1
[0464] Label Estimated value 95% CI Intercept 2.385 (2.4046,2.3654) The minimum balance of the business owner's deposit account at a certain point in time in the past three months -0.183 (-0.085,-0.281) The maximum average monthly credit transaction amount of the enterprise in the past 6 months -0.482 (-0.4232,-0.5408) Average utilization rate of corporate revolving loans in the past 12 months -0.603 (-0.5442,-0.6618) The minimum balance of the company's current deposits in the past three months -0.183 (-0.0262,-0.3398) Average monthly debit transaction amount / average monthly balance of the enterprise in the past 6 months 0.234 (0.3516,0.1164) Average of (credit transaction amount - debit transaction amount) of the enterprise in the past 6 months -0.932 (-0.8732,-0.9908)
[0465] α is the intercept term, with a value range of (2.4046, 2.3654), and the optimal value is 2.385; β1 is the coefficient corresponding to the minimum balance of the business owner's deposit account in the past three months, with a value range of (-0.085, -0.281), and the optimal value is -0.183; β2 is the coefficient corresponding to the maximum value of the average monthly credit transaction amount of the enterprise in the past six months, with a value range of (-0.4232, -0.5408), and the optimal value is -0.482; β3 is the coefficient corresponding to the average value of the utilization rate of the enterprise owner's revolving loan in the past 12 months, with a value range of (-0.5442, -0.6618), and the optimal value is -0.603; β 4 is the coefficient corresponding to the minimum value of the current deposit balance of the enterprise in the past three months, with a range of values of (-0.0262, -0.3398), and the optimal value is -0.183; β5 is the coefficient corresponding to the average monthly debit transaction amount of the enterprise in the past six months / the average monthly balance of the past six months, with a range of values of (0.3516, 0.1164), and the optimal value is 0.234; β6 is the coefficient corresponding to the mean value of (credit transaction amount - debit transaction amount) of the enterprise in the past six months, with a range of values of (-0.8732, -0.9908), and the optimal value is -0.932 (Note: The numerical range comes from the 95% confidence interval, that is, the 95% CI in the table below)
[0466] x1 is the natural logarithm conversion value of the minimum balance of the deposit account of the business owner in the past three months generated by the feature conversion step; x2 is the square root conversion value of the maximum monthly average credit transaction amount of the enterprise in the past six months generated by the feature conversion step; x3 is the cube root conversion value of the average utilization rate of the revolving loan of the enterprise owner in the past 12 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the minimum current deposit balance of the enterprise in the past three months generated by the feature conversion step; x5 is the square root conversion value of the average monthly debit transaction amount of the enterprise in the past six months / the average monthly balance of the past six months generated by the feature conversion step; x6 is the cube root conversion value of the mean value of (credit transaction amount-debit transaction amount) of the enterprise in the past six months generated by the feature conversion step. The model performance of some features is shown in Table 2 below.
[0467] Table 2
[0468] The P values of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with overdue performance.
[0469]
[0470] Example 6 Construction of Zeta 2 Model
[0471] For Zeta 2 as a sub-model of the segmentation model, the above-mentioned decision tree method is adopted, and the sample size obtained for classification is about 70,000. Its customer base is mainly customers with corporate deposit account age greater than 6 years and an average corporate deposit account balance of less than 1,000 yuan in the past 6 months. The model is built to predict the probability of credit delinquency of more than 30 days in this group (target variable).
[0472] In this embodiment 6, the final confirmation is to describe the final modeling results using six features as examples: the current month's RMB account deposit balance of the enterprise, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's current deposits in the past 6 months, the number of consecutive months of increase in the bill balance of the business owner in the past 12 months, whether the recycling rate of the business owner's credit card in the past 12 months is greater than 0, and the average deposit balance quantile of the enterprise by industry in the past 6 months. In the feature conversion step, the 6 modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the credit overdue probability modeling step, the 6 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) to perform logistic regression to calculate the credit overdue probability model.
[0473] The current monthly RMB account balance of an enterprise reflects the current asset level of the borrowing enterprise. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit account balance of less than 1,000 yuan in the past 6 months, the higher the current RMB account balance, the higher the asset level, the stronger the debt repayment ability, and the lower the possibility of future overdue payments. Conversely, the greater the possibility of future overdue payments. The conversion method is natural logarithm conversion.
[0474] The minimum average monthly credit transaction amount of the enterprise in the past 6 months reflects the minimum average monthly capital inflow amount of the borrowing enterprise in the past 6 months. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit account balance of less than 1,000 yuan in the past 6 months, the smaller the inflow of funds in the past 6 months, the worse the asset condition is, and the greater the possibility of overdue payments in the future, and vice versa. The conversion method is square root conversion.
[0475] The minimum monthly accumulation of the company's demand deposits in the past six months reflects the minimum average daily deposit of the borrowing company in the past six months. For companies with a corporate deposit age of more than 6 years and an average balance of corporate deposits in the past six months of less than 1,000 yuan, the smaller the average daily deposit in the past six months, the worse its asset level and debt repayment ability, and the greater the possibility of future overdue payments, and vice versa. The conversion method is natural logarithm conversion.
[0476] The number of consecutive months in which the balance of the business owner's bills increased in the past 12 months reflects the change in the loan balance of the actual controller of the borrowing enterprise in the past year. For enterprises with a corporate deposit age of more than 6 years and an average balance of corporate deposit accounts of less than 1,000 yuan in the past 6 months, the more consecutive months the actual controller's bill balance increased in the past 12 months, the stronger the increase in its credit demand, the higher the debt level, and the greater the possibility of future overdue payments. Conversely, the smaller the possibility of future overdue payments. The conversion method is the original value.
[0477] Whether the recycle rate of the business owner's credit card in the past 12 months is greater than 0 reflects the repayment situation of the actual controller of the borrowing enterprise in the past 12 months. For enterprises with a corporate deposit age of more than 6 years and an average balance of corporate deposit accounts of less than 1,000 yuan in the past 6 months, if the recycle rate of the actual controller in the past 12 months is greater than 0, it means that the actual controller cannot repay the full amount and has poor repayment behavior, the greater the possibility of overdue payment in the future, and vice versa, the smaller the possibility of overdue payment in the future. The conversion method is dummy variable conversion.
[0478] The percentile of the average deposit balance of enterprises in the past six months by industry reflects the overall asset ranking of the borrowing enterprises in their respective industries in the past six months. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit account balance of less than 1,000 yuan in the past six months, the higher the deposit percentile in the industry, the better the asset status, the stronger the debt repayment ability, and the smaller the possibility of future overdue payments. Conversely, the possibility of future overdue payments is relatively high.
[0479] In the step of calculating the credit overdue probability, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the probability (P) that the borrower is overdue:
[0480]
[0481] Among them, k is the number of features entering the model, and k is 6 in Formula 2. The model performance of some features is shown in Table 3 below.
[0482] Table 3
[0483] Label Estimated value 95% CI Intercept 1.476 (1.48351,1.4688) The company's RMB account deposit balance for the current month -0.066 (-0.0351,-0.09687) Minimum amount of credit transactions of the enterprise in the past 6 months -0.254 (-0.22636,-0.2822) The minimum monthly accumulation of the company's demand deposits in the past six months -0.366 (-0.33275,-0.39923) The number of consecutive months in which the business owner's bill balance increased in the past 12 months -0.018 (0.01777,-0.05296) Is the revolving usage rate of the business owner's credit card greater than 0 in the past 12 months? -0.182 (-0.17879,-0.1862) Average deposit balance quantiles of enterprises by industry in the past six months -0.354 (-0.34864,-0.35845)
[0484] α is the intercept term, with a range of values of (1.48351, 1.4688), and the optimal value is 1.476158; β1 is the corresponding coefficient of the current month's RMB account deposit balance of the enterprise, with a range of values of (-0.0351, -0.09687), and the optimal value is -0.065986; β2 is the corresponding coefficient of the minimum credit transaction amount of the enterprise in the past 6 months, with a range of values of (-0.22636, -0.2822), and the optimal value is -0.254282; β3 is the corresponding coefficient of the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months, with a range of values of (-0.33275, -0.39923) , the optimal value is -0.365991; β4 is the coefficient corresponding to the number of consecutive months of increase in the balance of the business owner's bill in the past 12 months, the value range is (0.01777, -0.05296), and the optimal value is -0.017597; β5 is the coefficient corresponding to whether the revolving rate of the business owner's credit card in the past 12 months is greater than 0, the value range is (-0.17879, -0.1862), and the optimal value is -0.182496; β6 is the coefficient corresponding to the quantile of the average deposit balance of the enterprise by industry in the past 6 months, the value range is (-0.34864, -0.35845), and the optimal value is -0.353545. (Note: The value range comes from the 95% confidence interval, that is, the 95% CI in the table below)
[0485] x1 is the cube root conversion value of the current month's RMB account balance generated by the feature conversion step; x2 is the square root conversion value of the minimum credit transaction amount of the enterprise in the past 6 months generated by the feature conversion step; x3 is the natural logarithm conversion value of the minimum monthly product of the enterprise's current deposits in the past 6 months generated by the feature conversion step; x4 is the original value of the number of consecutive months of increase in the bill balance of the business owner in the past 12 months generated by the feature conversion step; x5 is the dummy variable conversion value of whether the recycling rate of the business owner's credit card in the past 12 months is greater than 0 generated by the feature conversion step; x6 is the original value of the average deposit balance quantile of the enterprise by industry in the past 6 months generated by the feature conversion step. The model performance of some features is shown in Table 4 below. The P values of all model features are less than 0.05, indicating that the above features are significantly correlated with overdue performance.
[0486] Table 4
[0487]
[0488] Example 7 Construction of Zeta3 Model
[0489] For Zeta 3 as a sub-model of the segmentation model, the above-mentioned decision tree method is adopted, and the sample size obtained for classification is about 420,000. Its customer base is mainly customers with corporate deposit age greater than 6 years and corporate deposit account balance greater than or equal to 1,000 yuan in the past 6 months. The model is built to predict the probability of credit overdue for more than 30 days in this group (target variable).
[0490] In this embodiment 7, the final confirmation is to describe the final modeling result by taking eight features as examples: the average monthly accumulation of the enterprise's current deposit accounts in the past 3 months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months, the average monthly number of credit transactions of the enterprise in the past 3 months / the average monthly number of credit transactions in the past 12 months, the quantile of the credit transaction amount of the enterprise in the past 3 months by region, the quantile of the current deposit balance of the enterprise by industry, the current credit card balance of the business owner, the maximum balance of the deposit account of the business owner in the past 12 months, and whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0. In the conversion mode selection of the feature conversion step, the 8 modeling variables have been converted in a corresponding form according to the correlation between the feature and the target variable. In the credit overdue probability modeling step, the 8 features screened by the feature fine screening step are substituted into the Sigmoid function (using the LOGISTIC process of SAS software) to perform logistic regression to calculate the credit overdue probability model.
[0491] The average monthly accumulation of the company's current deposit accounts in the past three months reflects the average daily current deposits of the borrowing company in the past three months. For companies with a corporate deposit age of more than 6 years and a corporate deposit balance of more than or equal to RMB 1,000 in the past six months, the higher the average daily current deposit balance in the past three months, the better its debt repayment ability and the lower the possibility of future overdue payments. Conversely, the greater the possibility of future overdue payments. Conversion method Square root conversion value.
[0492] The average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months reflects the proportion of outflow funds of the borrowing enterprise in the past 12 months. For enterprises with a corporate deposit age of more than 6 years and a corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the higher the proportion of outflow funds in the past year, the higher the capital demand, the worse the debt repayment ability, and the greater the possibility of overdue payment, and vice versa. The possibility of overdue payment is smaller. The conversion method is the original value.
[0493] The average number of credit transactions per month in the past three months / the average number of credit transactions per month in the past 12 months reflects the changes in the capital inflow of the borrowing enterprise. For enterprises with a corporate deposit age of more than 6 years and a corporate deposit balance of more than or equal to RMB 1,000 in the past six months, compared with the average of the past year, the more recent capital inflows, the higher the recent asset level, the stronger the debt repayment ability, and the lower the possibility of overdue payment, and vice versa. The conversion method is square root conversion.
[0494] The current deposit balance quantile of enterprises by industry reflects the current asset level of the borrowing enterprises in the industry. For enterprises with a corporate deposit age of more than 6 years and a corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the higher the current deposit level in the industry, the better the asset status, the stronger the debt repayment ability, and the lower the possibility of future overdue payments. Conversely, the greater the possibility of future overdue payments. The conversion method is natural logarithm conversion.
[0495] The percentile of the credit transaction amount of the enterprise in the past three months by region reflects the level of capital inflow of the borrowing enterprise in the region in the past three months. For enterprises with a corporate deposit age of more than 6 years and a corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the more capital inflow it currently has in its region, the better its operating conditions, and the lower the possibility of overdue payments. Conversely, the greater the possibility of overdue payments. The conversion method is cube root conversion.
[0496] The current remaining balance of the business owner's credit card reflects the current available credit limit of the actual controller of the borrowing enterprise. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the more available credit the actual controller has, the less debt pressure there is, and the less likely it is to be overdue. Conversely, the more likely it is to be overdue. The conversion method is natural logarithm conversion.
[0497] The number of months from the maximum balance of the deposit account of the business owner in the past 12 months reflects the change in the asset level of the actual controller of the borrowing enterprise in the past year. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the smaller the number of months from the maximum balance of the deposit account of the actual controller in the past year, the better the recent asset status, the better the debt repayment ability, and the lower the possibility of overdue payment, and vice versa. The conversion method is natural logarithm conversion.
[0498] Whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0 reflects the credit card interest of the actual controller of the borrowing enterprise in the past 6 months. For enterprises with a corporate deposit age of more than 6 years and an average corporate deposit balance of more than or equal to 1,000 yuan in the past 6 months, the interest generated by the actual controller's credit card in the past six months indicates that he has not paid in full and has poor credit performance. The greater the possibility of overdue payment, the smaller the possibility of overdue payment. The conversion method is dummy variable conversion. The model performance of some characteristics is shown in Table 5 below.
[0499] Table 5
[0500]
[0501]
[0502] In the step of calculating the credit overdue probability, based on the data and coefficients obtained in the feature conversion step, the following model formula is used to predict the probability (P) that the borrower is overdue:
[0503]
[0504] Where k is the number of features entering the model, and in Formula 3, k is 8.
[0505] α is the intercept term, with a range of values of (1.61086, 1.53246), and the optimal value is 1.571658; β1 is the corresponding coefficient of the average monthly accumulation of the enterprise's current deposit accounts in the past three months, with a range of values of (-0.06973, -0.07035), and the optimal value is -0.070037; β2 is the corresponding coefficient of the enterprise's average monthly debit transaction amount in the past 12 months / average monthly balance in the past 12 months, with a range of values of (-0.25804, -0.27992), and the optimal value is -0.268979; β3 is the corresponding coefficient, with a range of values of (-0.12274, -0.13067), and the optimal value of J26 is -0.126706; β4 is the corresponding coefficient of the current deposit balance quantile of the enterprise by industry, with a range of values of (0.02505, -0. 00954), the optimal value is 0.007753; β5 is the corresponding coefficient of the credit transaction amount of the balance of the enterprise in the past three months by region, the value range is (-0.02632,-0.05308), the optimal value is -0.039701; β6 is the corresponding coefficient of the current credit card balance of the business owner, the value range is (-0.34253,-0.34889), the optimal value is -0.34571; β7 is the corresponding coefficient of the maximum balance of the deposit account of the business owner in the past 12 months, the value range is (0.34439,0.34161), the optimal value is 0.343; β8 is the corresponding coefficient of whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0, the value range is (0.066,-0.13), the optimal value is -0.032. (Note: The value range comes from the 95% confidence interval, that is, the 95% CI in the table below)
[0506] x1 is the square root conversion value of the average monthly product of the company's current deposit accounts in the past three months generated by the feature conversion step; x2 is the original value of the company's average monthly debit transaction amount in the past 12 months / average monthly balance in the past 12 months generated by the feature conversion step; x3 is the square root conversion value of the company's average monthly credit transaction number in the past three months / average monthly credit transaction number in the past 12 months generated by the feature conversion step; x4 is the natural logarithm conversion value of the company's current deposit balance quantile by industry generated by the feature conversion step; x5 is the cube root conversion value of the company's balance credit transaction amount quantile by region in the past three months generated by the feature conversion step; x6 is the natural logarithm conversion value of the business owner's current credit card balance generated by the feature conversion step; x7 is the natural logarithm conversion value of the number of months from the current maximum balance of the business owner's deposit account at a point in time generated by the feature conversion step; x8 is the dummy variable conversion value of whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0 generated by the feature conversion step
[0507] The model performance of some features is shown in Table 6 below:
[0508] Table 6
[0509]
[0510] The P values of all model characteristics are less than 0.05, indicating that the above characteristics are significantly correlated with overdue performance.
[0511] Example 8
[0512] The P values calculated by the above formulas can be further used to calculate the score of any customer.
[0513] The scoring calculation step is used to convert the calculated credit overdue probability into a score of 0-1000 points using a pre-stored default score conversion code.
[0514] In the scoring calculation module, the following formula is used to calculate and generate the credit score used to characterize the borrower:
[0515]
[0516]
[0517] Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
[0518] In this embodiment, the P value calculated in the above embodiments can be substituted to calculate the score.
[0519] The KS (Kolmogorov-Smirnov) statistic was proposed by two Soviet mathematicians, ANKolmogorov and NVSmirnov. In risk control, KS is often used to evaluate the discrimination of models. The greater the discrimination, the stronger the risk ranking ability of the model.
[0520] The KS statistic is based on the empirical cumulative distribution function (ECDF), which is generally defined as:
[0521] KS=max(|cum(bad_rate)-cum(good_rate)|)
[0522] According to the prediction model, a KS curve is drawn. The KS of this embodiment is 56. It can be seen that the model constructed in this embodiment performs well in assessing customer credit risk. Figure 3 shown.
[0523] The following shows the results of the KS values of the model applied to the development sample and the reserved validation sample in the development sample and the validation sample. From Table 7, it can be seen that the model has a good discrimination effect in the training set and the validation set (where the sample size ratio of the training set and the validation set is 6:4) and the overall sample, that is, the discrimination effect is very good in all sub-models.
[0524] Table 7
[0525] Classification Development samples Verification Sample Development + Validation Samples Zeta 56.01 56.48 56.17 Zeta1 42.43 42.75 42.50 Zeta2 50.86 51.07 50.88 Zeta3 48.94 48.24 48.63
[0526] Although the embodiments of the present application are described above, the present application is not limited to the above specific embodiments and application fields, and the above specific embodiments are merely illustrative and instructive, rather than restrictive. A person of ordinary skill in the art can make many forms under the guidance of this specification and without departing from the scope of protection of the claims of the present application, all of which belong to the protection of the present application.
Claims
1. A method for calculating inclusive credit risk, wherein include: A data collection step, which obtains inclusive credit prediction data of the sample to be predicted; A step of classifying samples to be predicted, which classifies the samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; The credit default probability calculation step is to substitute the inclusive credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted.
2. The method according to claim 1, further comprising: include: After calculating the credit default probability, a step of calculating the credit score of the sample to be predicted is used to calibrate the calculated credit default probability to a standardized score of 0-1000 points.
3. The method according to claim 1 or 2, in, The inclusive credit prediction data includes the original inclusive credit prediction data of the sample to be predicted and the derived inclusive credit prediction data processed based on the original inclusive credit prediction data; Preferably, the original inclusive credit prediction data includes: Basic data on corporate credit, which is all available data based on the loan application and usage behavior of sample (corporate) users. Basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions. Basic data of enterprise basic information, which is based on the attributes of the sample (enterprise) users themselves but is not directly related to their behavior in financial institutions. Basic data on financial assets of business owners, which includes all other financial assets and financial transaction data of sample (actual controller of the business) users in financial institutions that are not related to credit cards and loans. Basic data on business owner credit, which is all available data based on the loan application and usage behavior of sample users (actual controllers of enterprises). Basic data on the basic information of business owners is based on the attributes of the sample users (actual controllers of the enterprises) themselves, but is not directly related to their behavior in financial institutions.
4. The method according to any one of claims 1 to 3, in, The derived inclusive credit prediction data processed based on the original inclusive credit prediction data refers to the data obtained by processing the collected original inclusive credit prediction data based on the time dimension, space dimension, frequency dimension and statistical information dimension; Preferably, the derived inclusive credit forecast data includes but is not limited to: Derived inclusive credit forecast data obtained by processing based on sample relationship length, Derivative inclusive credit forecast data obtained by processing time interval variables, Derived inclusive credit forecast data obtained based on the frequency of sample behavior. Derivative inclusive credit forecast data obtained by processing the current time point of the sample, Derived inclusive credit forecast data obtained based on the continuous behavior of samples, Derived inclusive credit forecast data is obtained by processing sample data based on statistical information dimensions.
5. The method according to any one of claims 1 to 4, in, The inclusive credit prediction data is selected from one or more of the following: The minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, the (credit transaction amount - debit transaction amount) of the business owner in the past six months, the current month's RMB account deposit balance of the business owner, the minimum monthly average credit transaction amount of the business owner in the past six months, the minimum monthly accumulation of the business owner's current deposits in the past six months, the number of consecutive months of increase in the business owner's bill balance in the past 12 months, and the business owner's credit card balance in the past 12 months Whether the recycling rate is greater than 0, the percentile of the average deposit balance of the enterprise in the past 6 months by industry, the average monthly accumulation of the enterprise's current deposit accounts in the past 3 months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance of the past 12 months, the average monthly number of credit transactions of the enterprise in the past 3 months / the average monthly number of credit transactions in the past 12 months, the percentile of the credit transaction amount of the enterprise in the past 3 months by region, the percentile of the current deposit balance of the enterprise by industry, the current remaining balance of the business owner's credit card, the number of months from the present when the maximum balance of the business owner's deposit account in the past 12 months was the maximum, and whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0.
6. The method according to any one of claims 1 to 5, in, The step of classifying the sample to be predicted includes the following sub-steps: Whether the sample to be tested is a customer with a corporate deposit account age greater than a months; Whether the sample to be tested is a customer whose average balance of corporate deposit account is less than c yuan in the past b months; Based on the above sub-steps, the samples to be predicted are classified to determine the sub-model for calculating the probability of credit default. Under the premise of ensuring the rationality of business logic, the order of the above sub-steps can be set arbitrarily; It is preferred to classify the predicted samples in the following order: First, determine whether the sample to be tested is a customer with a corporate deposit age greater than a month; Then determine whether the sample to be tested is a customer whose average balance of corporate deposit accounts in the past b months is less than c yuan.
7. The method according to any one of claims 1 to 6, in, After feature conversion is performed on the inclusive credit prediction data, the data is substituted into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted. The feature conversion step includes: Based on the feature type of the inclusive credit prediction data that needs to be substituted into the credit default probability sub-model, the WOE method or the continuous method is selected for feature conversion.
8. The method according to claim 7, in, The continuous method for feature conversion includes the following methods: directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data to perform continuous feature conversion.
9. The method according to any one of claims 1 to 8, in, The credit default probability sub-model is a model constructed based on sample inclusive credit prediction data and credit default probability using logistic regression based on the existing user population.
10. The method according to any one of claims 1 to 9, in, The inclusive credit prediction data is selected from: the minimum balance of the business owner's deposit account at a certain point in time in the past three months, the maximum average monthly credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the average monthly debit transaction amount of the business owner in the past six months / the average monthly balance of the past six months, and the average of (credit transaction amount-debit transaction amount) of the business owner in the past six months.
11. The method according to claim 10, in, The steps to calculate the probability of credit default include: The minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, and the average value of (credit transaction amount-debit transaction amount) of the business owner in the past six months are transformed into features. Preferably, the minimum balance of the business owner's deposit account in the past three months, the maximum monthly average credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the monthly average debit transaction amount of the business owner in the past six months / the monthly average balance of the past six months, and the average value of (credit transaction amount - debit transaction amount) of the business owner in the past six months are converted in a continuous manner; Further preferred calculation methods include taking the natural logarithm of the minimum balance in the business owner's deposit account in the past three months, taking the square root of the maximum average monthly credit transaction amount in the past six months, taking the cube root of the average utilization rate of the business owner's revolving loan in the past 12 months, taking the natural logarithm of the minimum value of the business's demand deposit balance in the past three months, taking the square root of the business owner's monthly average debit transaction amount in the past six months / the average monthly balance in the past six months, and taking the cube root of the average value of (credit transaction amount - debit transaction amount) in the past six months.
12. The method according to claim 11, in, The converted values of the six features, namely, the minimum balance of the business owner's deposit account in the past three months, the maximum average monthly credit transaction amount of the business owner in the past six months, the average utilization rate of the business owner's revolving loan in the past 12 months, the minimum current deposit balance of the business owner in the past three months, the average monthly debit transaction amount of the business owner in the past six months / the average monthly balance of the past six months, and the average of (credit transaction amount - debit transaction amount) of the business owner in the past six months, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data and credit default probability to calculate the default probability of the sample to be predicted.
13. The method according to claim 12, in, The sub-model is shown in the following formula 1: Where k is the number of features entering the model, preferably k is 6, α is the intercept term, the preferred value range is (2.4046, 2.3654), and the optimal value is 2.385; β1 is the coefficient corresponding to the minimum balance of the business owner’s deposit account in the past three months. The preferred value range is (-0.085, -0.281), and the optimal value is -0.183; β2 is the coefficient corresponding to the maximum value of the average monthly credit transaction amount of the enterprise in the past 6 months. The preferred value range is (-0.4232, -0.5408), and the optimal value is -0.482; β3 is the coefficient corresponding to the average utilization rate of the enterprise's main revolving loan in the past 12 months, the preferred value range is -0.5442, -0.6618), and the optimal value is -0.603; β4 is the coefficient corresponding to the minimum value of the current deposit balance of the enterprise in the past three months. The preferred value range is (-0.0262, -0.3398), and the optimal value is -0.183; β5 is the corresponding coefficient of the average monthly debit transaction amount of the enterprise in the past 6 months / the average monthly balance in the past 6 months. The preferred value range is (0.3516, 0.1164), and the optimal value is 0.234; β6 is the corresponding coefficient of the mean of (credit transaction amount - debit transaction amount) of the enterprise in the past 6 months. The preferred value range is (-0.8732, -0.9908), and the optimal value is -0.932; x1 is the natural logarithm transformation value of the minimum balance of the business owner’s deposit account in the past three months generated by the feature conversion step; x2 is the square root conversion value of the maximum monthly average credit transaction amount of the enterprise in the past 6 months generated by the feature conversion step; x3 is the cube root conversion value of the average utilization rate of the enterprise's main revolving loan in the past 12 months generated by the feature conversion step; x4 is the natural logarithm transformation value of the minimum demand deposit balance of the enterprise in the past three months generated in the feature transformation step; x5 is the square root conversion value of the enterprise's average monthly debit transaction amount in the past six months / average monthly balance in the past six months generated in the feature conversion step; x6 is the cube root conversion value of the mean of (credit transaction amount - debit transaction amount) of the enterprise in the past 6 months generated by the feature conversion step.
14. The method according to claim 13, in, After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower: Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
15. The method according to any one of claims 1 to 9, in, The inclusive credit prediction data is selected from: the current month's RMB account deposit balance of the enterprise, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months, the number of consecutive months of increased bill balances of the business owner in the past 12 months, whether the revolving utilization rate of the business owner's credit card in the past 12 months is greater than 0, and one or more of the average deposit balance percentiles of the enterprise by industry in the past 6 months.
16. The method according to claim 15, in, The steps to calculate the probability of credit default include: The current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's current deposits in the past 6 months, the number of consecutive months of increase in the bill balance of the business owner in the past 12 months, whether the recycling rate of the business owner's credit card in the past 12 months is greater than 0, and the average deposit balance quantile of the enterprise by industry in the past 6 months are converted into features. It is preferred to use a continuous conversion method for the current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's current deposits in the past 6 months, the number of consecutive months of increase in the bill balance of the enterprise owner in the past 12 months, and the average deposit balance quantile of the enterprise by industry in the past 6 months; and use a dummy feature method to convert whether the recycling rate of the enterprise owner's credit card in the past 12 months is greater than 0; Further optimization includes taking the cube root of the enterprise's current month's RMB account deposit balance; taking the square root of the enterprise's minimum monthly average credit transaction amount in the past six months; taking the natural logarithm of the minimum monthly product of the enterprise's demand deposits in the past six months; taking the original value of the number of consecutive months of increase in the business owner's bill balance in the past 12 months; and taking the original value of the percentile of the enterprise's average deposit balance by industry in the past six months.
17. The method according to claim 15, in, The converted values of the six features, namely, the current month's RMB account deposit balance, the minimum monthly average credit transaction amount of the enterprise in the past 6 months, the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months, the number of consecutive months of increase in the owner's bill balance in the past 12 months, whether the owner's credit card revolving rate in the past 12 months is greater than 0, and the average deposit balance quantile of the enterprise by industry in the past 6 months, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data to calculate the default probability of the sample to be predicted.
18. The method according to claim 17, in, The sub-model is shown in the following formula 2: Where k is the number of features entering the model, and k is preferably 6; α is the intercept term, the preferred value range is (1.48351, 1.4688), and the optimal value is 1.476; β1 is the coefficient corresponding to the balance of the company's RMB account in the current month. The preferred value range is (-0.0351, -0.09687), and the optimal value is -0.066; β2 is the coefficient corresponding to the minimum value of the average monthly credit transaction amount of the enterprise in the past 6 months. The preferred value range is (-0.22636, -0.2822), and the optimal value is -0.254; β3 is the coefficient corresponding to the minimum monthly accumulation of the enterprise's demand deposits in the past 6 months. The preferred value range is (-0.33275, -0.39923), and the optimal value is -0.366; β4 is the coefficient corresponding to the number of consecutive months of increase in the bill balance of the business owner in the past 12 months. The preferred value range is (0.01777, -0.05296), and the optimal value is -0.018; β5 is the coefficient corresponding to whether the revolving utilization rate of the business owner's credit card in the past 12 months is greater than 0. The preferred value range is (-0.17879, -0.1862), and the optimal value is -0.182; β6 is the corresponding coefficient of the average deposit balance quantile of the enterprise by industry in the past 6 months. The preferred value range is (-0.34864, -0.35845), and the optimal value is -0.354; x1 is the natural logarithm transformation value of the enterprise’s RMB account deposit balance in the current month generated by the feature conversion step; x2 is the square root conversion value of the minimum average monthly credit transaction amount of the enterprise in the past 6 months generated by the feature conversion step; x3 is the natural logarithm transformation value of the minimum monthly product of the enterprise's demand deposits in the past 6 months generated by the feature conversion step; x4 is the original value of the number of consecutive months of increase in the bill balance of the business owner in the past 12 months generated in the feature conversion step; x5 is the dummy variable conversion value of the recurring usage rate of the business owner’s credit card in the past 12 months generated in the feature conversion step; x6 is the original value of the average deposit balance quantile of enterprises by industry in the past six months generated by the feature transformation step.
19. The method according to claim 18, in, After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower: Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
20. The method according to any one of claims 1 to 9, in, The inclusive credit forecast data is selected from: The average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months, the average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months, the percentile of the enterprise's credit transaction amount by region in the past three months, the percentile of the enterprise's current deposit balance by industry, the current remaining balance of the enterprise owner's credit card, the number of months from the present when the enterprise owner's deposit account had the maximum balance in the past 12 months, and whether the maximum (interest / limit) of the enterprise owner's credit card in the past 6 months is greater than 0 or one or more of the following.
21. The method according to claim 20, in, The steps to calculate the probability of credit default include: The average monthly accumulation of the company's current deposit accounts in the past three months, the average monthly debit transaction amount of the company in the past 12 months / the average monthly balance in the past 12 months, the average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months, the quantile of the company's credit transaction amount in the past three months by region, the quantile of the company's current deposit balance by industry, the current remaining balance of the business owner's credit card, the number of months from the maximum balance of the business owner's deposit account in the past 12 months, and whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0 are converted into features. It is preferred to use a continuous conversion method for the average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount / average monthly balance of the past 12 months, the average monthly number of credit transactions / average monthly number of credit transactions in the past three months, the quantile of the credit transaction amount of the enterprise by region in the past three months, the quantile of the current deposit balance of the enterprise by industry, the current remaining balance of the enterprise owner's credit card, and the number of months from the maximum balance of the enterprise owner's deposit account in the past 12 months; whether the maximum (interest / limit) of the enterprise owner's credit card in the past six months is greater than 0 is converted using a dummy feature method; Further optimization includes taking the square root of the average monthly product of the enterprise's current deposit accounts in the past three months; taking the original value of the enterprise's average monthly debit transaction amount in the past 12 months / the average monthly balance in the past 12 months; taking the square root of the enterprise's average monthly number of credit transactions in the past three months / the average monthly number of credit transactions in the past 12 months; taking the cube root of the enterprise's credit transaction amount quantiles by region in the past three months; taking the natural logarithm of the enterprise's credit transaction amount quantiles by region in the past three months; taking the natural logarithm of the business owner's current credit card balance; and taking the natural logarithm of the number of months from the current maximum balance of the business owner's deposit account in the past 12 months.
22. The method according to claim 20, in, The average monthly accumulation of the enterprise's current deposit accounts in the past three months, the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance of the past 12 months, the average monthly number of credit transactions of the enterprise in the past three months / the average monthly number of credit transactions in the past 12 months, the quantiles of the enterprise's credit transaction amount in the past three months by region, the quantiles of the enterprise's current deposit balance by industry, the current remaining balance of the enterprise owner's credit card, and the number of months from the maximum balance of the enterprise owner's deposit account in the past 12 months are converted continuously; the converted values of the eight features, such as whether the maximum (interest / limit) of the enterprise owner's credit card in the past six months is greater than 0, are substituted into the sub-model constructed by logistic regression based on the sample inclusive credit prediction data and credit default probability to calculate the default probability of the sample to be predicted.
23. The method according to claim 22, in, The sub-model is shown in Formula 3 below: Where k is the number of features entering the model, preferably k is 8; α is the intercept term, the preferred value range is (1.61086, 1.53246), and the optimal value is 1.572; β1 is the corresponding coefficient of the average monthly accumulation of the enterprise's current deposit accounts in the past three months. The preferred value range is (-0.06973, -0.07035), and the optimal value is -0.070; β2 is the corresponding coefficient of the average monthly debit transaction amount of the enterprise in the past 12 months / the average monthly balance in the past 12 months. The preferred value range is (-0.25804, -0.27992), and the optimal value is -0.269; β3 is the coefficient of the average number of credit transactions in the past three months / the average number of credit transactions in the past 12 months. The preferred value range is (-0.12274, -0.13067), and the optimal value is -0.127; β4 is the corresponding coefficient of the credit transaction amount quantile of the enterprise by region in the past three months. The preferred value range is (0.02505, -0.00954), and the optimal value is 0.008; β5 is the corresponding coefficient of the current deposit balance quantile of the enterprise by industry, the preferred value range is (-0.02632, -0.05308), and the optimal value is -0.040; β6 is the coefficient corresponding to the current remaining balance of the business owner’s credit card. The preferred value range is (-0.02632, -0.05308), and the optimal value is -0.346; β7 is the coefficient corresponding to the number of months from the current point in time of the maximum balance of the business owner's deposit account in the past 12 months. The preferred value range is (0.34439, 0.34161), and the optimal value is 0.343; β8 is the coefficient of whether the maximum (interest / limit) of the business owner's credit card in the past 6 months is greater than 0. The preferred value range is (0.066, -0.13), and the optimal value is -0.032; x1 is the square root conversion value of the average monthly product of the company's current deposit accounts in the past three months generated by the feature conversion step; x2 is the original value of the average monthly debit transaction amount / average monthly balance of the enterprise in the past 12 months generated by the feature conversion step; x3 is the square root conversion value of the average number of monthly credit transactions of the enterprise in the past three months / the average number of monthly credit transactions in the past 12 months generated by the feature conversion step; x4 is the natural logarithm transformation value of the credit transaction amount quantile of the enterprise by region in the past three months generated in the feature conversion step; x5 is the cube root conversion value of the current deposit balance quantile of the enterprise by industry generated in the feature conversion step; x6 is the natural logarithm transformation value of the business owner’s current credit card balance generated in the feature transformation step; x7 is the natural logarithm transformation value of the maximum balance of the business owner’s deposit account in the past 12 months, which is generated in the feature conversion step; x8 is the dummy variable conversion value generated in the feature conversion step, indicating whether the maximum (interest / limit) of the business owner’s credit card in the past 6 months is greater than 0.
24. The method according to claim 23, in, After calculating the credit default probability, the step of calculating the credit score of the sample to be predicted is to use the following formula to calculate and generate a credit score for characterizing the borrower: Among them, P is the borrower's overdue probability (P) generated in the credit overdue probability calculation module, A is 54.2475; B is 115.4156, and the round function rounds the calculated score to the integer value; finally, the scores greater than or equal to 1000 are set to 1000 points, and the scores less than 0 points are set to 0 points.
25. A device for calculating inclusive credit risk, wherein include: A data collection module, which is used to obtain inclusive credit prediction data of samples to be predicted; A module for classifying samples to be predicted, which is used to classify the samples to be predicted based on a decision tree method to determine a sub-model for calculating the probability of credit default; A credit default probability calculation module is used to substitute the inclusive credit prediction data into the credit default probability sub-model to calculate the credit default probability of the sample to be predicted, and preferably the credit default probability is the credit default probability for the inclusive credit business.
26. The device according to claim 25, in, The device executes the steps of the method for calculating inclusive credit risk according to any one of claims 1 to 24.
27. A system for calculating inclusive credit risk, It is characterized in that The system for calculating inclusive credit risk includes: a memory, a processor, and a program for the method for calculating inclusive credit risk stored in the memory and executable on the processor. When the program for calculating inclusive credit risk is executed by the processor, the steps of the method for calculating inclusive credit risk as described in any one of claims 1 to 24 are implemented.
28. A computer storage medium, It is characterized in that The computer storage medium stores a program for calculating inclusive credit risk, and when the program for calculating inclusive credit risk is executed by a processor, the steps of the method for calculating inclusive credit risk according to any one of claims 1 to 24 are implemented.
29. A method for constructing an inclusive credit risk prediction model for inclusive credit business, wherein include: A data collection step, which obtains the original inclusive credit forecast data of the sample used to build the model; A data derivation step, which processes derived inclusive credit forecast data based on the original inclusive credit forecast data; The feature initial screening step is to perform initial screening on all categories, that is, all features, including the original inclusive credit forecast data and the derived inclusive credit forecast data, to obtain features after initial screening; The initial screening data conversion step is to judge the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and use the best method to convert each feature after the initial screening; The feature fine screening step is to perform in-depth screening on the initially screened features after feature conversion to obtain finely screened features; The credit default probability modeling step is to select the logistic regression method to build the model based on the relationship between the carefully screened features and the probability of credit default, and confirm the method used to calculate the credit default probability; The data samples obtained in the data collection step are samples of customers who have used inclusive credit services.
30. The method according to claim 29, in, In the data collection step, the original inclusive credit prediction data of the samples obtained for building the model include: Basic data on corporate credit, which is all available data based on the loan application and usage behavior of sample (corporate) users. Basic data on corporate deposits, which is based on all available data on RMB deposits made by sample (corporate) users in financial institutions. Basic data of enterprise basic information, which is based on the attributes of the sample (enterprise) users themselves but is not directly related to their behavior in financial institutions. Basic data on financial assets of business owners, which includes all other financial assets and financial transaction data of sample (actual controller of the business) users in financial institutions that are not related to credit cards and loans. Basic data on business owner credit, which is all available data based on the loan application and usage behavior of sample users (actual controllers of enterprises). Basic data on the basic information of business owners is based on the attributes of the sample users (actual controllers of the enterprises) themselves, but is not directly related to their behavior in financial institutions.
31. The method according to claim 29, in, In the data derivation step, the derived inclusive credit prediction data processed based on the original inclusive credit prediction data refers to the data obtained by processing the collected original inclusive credit prediction data based on the time dimension, space dimension, frequency dimension, and statistical information dimension; Preferably, the derived inclusive credit forecast data includes but is not limited to: Derived inclusive credit forecast data obtained by processing based on sample relationship length, Derivative inclusive credit forecast data obtained by processing time interval variables, Derived inclusive credit forecast data obtained based on the frequency of sample behavior. Derivative inclusive credit forecast data obtained by processing the current time point of the sample, Derived inclusive credit forecast data obtained based on the continuous behavior of samples, Derived inclusive credit forecast data is obtained by processing sample data based on statistical information dimensions.
32. The method according to any one of claims 29 to 31, in, The feature initial screening step includes the following steps: The first preliminary screening step is to screen the features based on the missing data of each feature of the sample used to build the model. The second initial screening step is to screen the features based on the fact that a single value of a feature sample is too high. The third preliminary screening step is to calculate the information IV value of each feature to perform preliminary screening of the features; The order of the first preliminary screening step, the second preliminary screening step and the third preliminary screening step can be any order. The fourth preliminary screening step uses a stepwise discrimination algorithm to perform preliminary screening of features after the first to third preliminary screenings; The fifth preliminary screening step is to conduct preliminary screening of the features after the fourth preliminary screening step based on the consistency between the risk characteristics of each feature itself and the actual real results of the samples used for model construction.
33. The method according to any one of claims 29 to 32, in, Also includes: The sample selection step is used to screen all users and obtain samples of inclusive credit business for model construction before the data collection step. Preferably, the sample selection step includes classifying all sample users based on a decision tree, and the classification basis includes but is not limited to: Whether a user is a customer with a corporate deposit account age greater than a months; Whether a user is a customer whose average balance of corporate deposit accounts in the past b months is less than c yuan.
34. The method according to any one of claims 29 to 33, in, In the initial screening data conversion step, the conversion method of the features after the initial screening is determined based on the concentration and data type of the features after the initial screening.
35. The method according to claim 34, in, The initial screening data conversion step includes the following steps based on the judgment of concentration and data type: Classify the data type of each feature into character variables and numeric variables. For character variables, dummy feature conversion is used to perform initial screening data conversion. The process of further classifying numerical variables includes the following sub-steps: If the value of the numeric variable is less than n, the WOE conversion method is used to convert the initial screening data. If the value of the numerical variable is more than n, further judge if the value of the continuous variable is large and the concentration of a single value is greater than m%, then the WOE conversion method is used; if the concentration of a single value is less than or equal to m%, then the continuous conversion method is used. Preferably, n and m are both positive integers, wherein n=5-10, and m=90-99.
36. The method according to claim 35, in, Also includes: For the features that are confirmed to adopt the continuous conversion method, the optimal conversion method is selected based on the correlation between the feature and credit default under different continuous conversion methods to perform the continuous feature conversion of the feature. Preferably, the continuous feature conversion is performed by directly selecting the original value, calculating the square of the original data, calculating the square root of the original data, calculating the cube root of the original data, or calculating the natural logarithm of the original data.
37. The method according to any one of claims 29 to 36, in, The feature fine screening steps include: The first fine screening step is to screen the features based on the stepwise regression algorithm, the F test and the T test to determine the significance of the features. The second fine screening step is to calculate the variance inflation factor based on each feature and eliminate the features with higher variance inflation factors to screen the features. The third fine screening step is to analyze the characteristics after the first fine screening step and the second fine screening step based on logistic regression to see whether the characteristic coefficients are consistent with the trend of the prediction results for credit default so as to further perform feature screening.
38. The method according to any one of claims 29 to 37, in, The credit default probability modeling step substitutes the features selected in the feature fine screening step into the Sigmoid function to perform logistic regression to calculate the credit default probability model.
39. A device for constructing an inclusive credit risk prediction model for inclusive credit business, It is characterized in that The device comprises: A data collection module, which is used to obtain the original inclusive credit prediction data of the samples used to build the model; A data derivation module, which is used to process derived inclusive credit prediction data based on the original inclusive credit prediction data; A feature initial screening module, which is used to perform initial screening on all categories, that is, all features, including the original inclusive credit prediction data and the derived inclusive credit prediction data, to obtain features after initial screening; The initial screening data conversion module is used to determine the conversion method of the features after the initial screening to confirm whether to use one of the WOE conversion method, the dummy feature conversion method and the continuous conversion method for feature conversion, and to use the best method for feature conversion for each feature after the initial screening; A feature fine screening module is used to perform in-depth screening on the initially screened features after feature conversion to obtain finely screened features; The credit default probability modeling module is used to select the logistic regression method to build a model based on the probability relationship between the finely screened features and credit default, and confirm the method used to calculate the credit default probability. The data samples obtained by the data collection module are samples of customers who have used inclusive credit services.
40. The device according to claim 39, in, The device executes the steps of the method for constructing an inclusive credit risk prediction model as described in any one of claims 29 to 38.
41. A system for building an inclusive credit risk prediction model for inclusive credit business, It is characterized in that The system includes: a memory, a processor, and a program for constructing a method for predicting inclusive credit risk, which is stored in the memory and can be run on the processor. When the program for constructing a method for predicting inclusive credit risk is executed by the processor, the steps of constructing a method for predicting inclusive credit risk are implemented as described in any one of claims 29 to 38.
Citation Information
Patent Citations
A credit risk assessment method and apparatus based on logistic regression technology
CN112686749B