Customer credit risk evaluation method integrating machine learning model and expert knowledge experience
Through the method of integrating machine learning models and expert knowledge and experience, the historical transaction data, external data and supplementary data of enterprises and customers are integrated and evaluated, which solves the problems of strong subjectivity and insufficient data integration in the existing technology, and realizes an accurate assessment of customer credit risks.
Patent Information
- Application Number
- CN202510229465.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing technology has problems such as strong subjectivity and insufficient data integration in customer credit risk assessment, making it difficult to achieve accurate assessment of customer credit risk.
Using a method that integrates machine learning models and expert knowledge and experience, we integrate, clean and process the historical transaction data, external data and supplementary data between enterprises and customers, and use machine learning algorithms such as random forests, and combine expert knowledge and experience to perform multi-dimensional ratings of customer credit risks.
A comprehensive assessment of customer credit risks has been realized, which reduces human subjectivity, improves the accuracy and reliability of the assessment, and can effectively identify risky customers and retain high-quality customers.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and specifically relates to a customer credit risk evaluation method integrating machine learning models and expert knowledge and experience. Background Art
[0002] In order to deeply explore the value of data assets and give full play to the important role of artificial intelligence algorithm models in empowering enterprise customer risk management, the present invention designs a customer credit risk rating method integrating machine learning models and expert knowledge and experience. Specifically, the present invention starts from the historical transaction data between enterprises and customers, as well as the external data of the enterprises themselves (including industrial and commercial registration information, financial information, tax arrears information, case-filing information, judgment document information, etc.) to design risk rating indicators. After integrating, cleaning, and processing the data, machine learning models and expert knowledge and experience are used to rate the risks of customers. Through this method, not only the comprehensive risk level of customers can be obtained, but also the risk matters of customers can be obtained. At the same time, applying this method to enterprise practice can achieve accurate scanning of customer risks, reduce the subjectivity and workload of manually evaluating customer quality, effectively identify risk customers while retaining high-quality customers. Summary of the Invention
[0003] In order to overcome the problems existing in the background art, the present invention provides a customer credit risk evaluation method integrating machine learning models and expert knowledge and experience. The risk rating object of this method is general enterprises with a unified social credit code starting with 91 and the transaction nature being customers. The specific rating process relies on three parts of data. One is the historical transaction data between enterprises and customers, such as the ending balance or occurrence amount data of accounting subjects such as accounts receivable, advance receipts, main business income, and other business income; the second is the supplementary data related to customers accumulated during the operation of the enterprise, including data on litigation with customers and unqualified list data; the third is public external data, including industrial and commercial registration, financial information, tax arrears information, etc. Through the integration, cleaning, and processing of the three parts of data, with the help of the mainstream machine learning algorithms in artificial intelligence models, rules are set up by integrating expert knowledge and experience to evaluate the customer credit risk. While outputting three risk levels of normal, vigilant, and cautious, the risk levels and risk matters under each rating dimension are also output.
[0004] To achieve the purpose of the present invention, the technical solution adopted is:
[0005] A customer credit risk evaluation method integrating machine learning models and expert knowledge and experience, specifically including the following steps:
[0006] S1. Data integration;
[0007] S1.1. Use the ETL tool to regularly extract the historical transaction data of enterprise customers deposited in the NC database every month. Through cooperation with other platforms, obtain external data and update it at a certain frequency.
[0008] S1.2. Manually supplement the litigation and unqualified list data related to customers monthly; integrate the historical transaction data, external data, and supplemented data to form a data warehouse based on internal and external customer data. This data warehouse is a subject-oriented, integrated, relatively stable, and historical change-reflecting data set, used to support risk analysis.
[0009] S1.3. Obtain the required multi-year accounting subject data from the data warehouse through writing SQL scripts for internal historical transaction data modeling, and obtain the public information of the enterprise from the data warehouse for external data modeling.
[0010] S2. Build risk rating indicators for different dimensions of data.
[0011] S2.1. For historical transaction data, set four indicators: accounts receivable turnover rate in the previous year and the previous two years, and accounts receivable recovery rate in the previous year and the previous two years. Among them,
[0012]
[0013] S2.2. For external data, set indicators such as registered capital, paid-in capital, situation of unpaid paid-in capital, proportion of the number of operating abnormal external investment enterprises, business status, four major financial ability indicators, whether there is operating abnormal information, whether it is seriously illegal and dishonest, whether it is listed as a person subject to enforcement for dishonesty, total amount of guarantee amount accounted for the registered capital ratio, whether it is a non-normal taxpayer for tax payment, cumulative amount involved in cases accounted for the registered capital ratio, etc.
[0014] S2.3. For the supplemented data, set indicators such as whether there is litigation with this unit and whether it is on the unqualified list of this unit.
[0015] S3. Calculate the actual values of the indicators and assign values to the indicators according to the quartile rule.
[0016] S3.1. For historical transaction data, calculate the accounts receivable turnover rate and accounts receivable recovery rate indicators. Use the quantile function in pandas to calculate the quartiles of the indicators. Divide the normal, vigilant, and cautious risk intervals according to the quartiles. When the actual value of the indicator falls into the above three risk intervals, assign the values of 0, 1, and 2 to the indicator respectively.
[0017] S3.2. For external data, directly sort the registered capital and paid-in capital and delimit quartiles, and assign risk scores of 0, 1, and 2 according to the same method as above; calculate, sort, and delimit quartiles for ratio indicators such as the situation of unpaid paid-in contributions, the proportion of abnormal operations of foreign-invested enterprises, the total amount of guarantees as a proportion of the registered capital, and the cumulative amount involved in cases as a proportion of the registered capital, and assign risk scores of 0, 1, and 2 according to the same method; for general yes / no type indicators such as whether there are administrative penalties, whether there are environmental protection penalties, and whether there is tax arrears, assign values of 1 and 0 respectively according to whether the customer triggers the relevant indicators; for important risk indicators such as whether there is serious illegal dishonesty, whether the customer is subject to enforcement for dishonesty, and whether there is major tax evasion, assign values of 2 and 0 respectively according to whether the customer triggers them.
[0018] S4. Build a learning model and generate index weights;
[0019] S4.1. For historical transaction data, set customers with accounts receivable outstanding for two consecutive years as customers on the unqualified list; set customers with the sum of risk scores of the four indicators of accounts receivable turnover in the previous year and the previous two years, and accounts receivable recovery rate in the previous year and the previous two years less than or equal to 2 as customers on the qualified list; assign values of 1 and 0 to unqualified and qualified list customers respectively, representing cautious customers and normal customers; incorporate the indicators and their risk scores designed in the historical transaction dimension into the machine learning model, and assign the industry as a control variable to the model learning dataset; use random forest to perform data modeling on the learning data, and obtain the optimized model and index weights through parameter tuning.
[0020] S4.2. For external data, basically set equal weights, but for the indicator of the cumulative amount involved in cases as a proportion of the registered capital, its index weight is set to 2 times that of other indicators, and finally the sum of the index weights of all external data is equal to 1; use the third-party tools pandas and numpy in python to build risk assessment indicators. For absolute value indicators, delimit the normal, vigilant, and cautious intervals of the indicators according to quartiles, and then obtain the normal, vigilant, and cautious grade types of the indicators according to the indicator values falling into different intervals, and assign values of 0, 1, and 2 respectively; for general yes / no type indicators, assign values of 1 and 0 respectively according to whether the indicator is triggered; for important yes / no type indicators, assign values of 2 and 0 respectively according to whether the indicator is triggered; sum the risk scores of each customer's indicators multiplied by the indicator weights to initially obtain the external data risk score; set score thresholds according to the score data distribution clustering to delimit the customer external data risk level interval, and then obtain the risk level of the customer's external data, and assign values of 0, 1, and 2 to normal, vigilant, and cautious respectively.
[0021] S5. Calculate the internal and external data risk scores of the customer and delimit the internal and external dimension ratings according to quartiles;
[0022] For internal and external data respectively, the method of multiplying the index weight by the index risk score is used to calculate the internal and external data risk scores. The internal and external data risk scores are sorted to obtain quartiles, and three risk intervals of normal, vigilant, and cautious are divided according to the quartiles. When the risk score falls into the risk interval, the risk level corresponding to the risk score is obtained, so as to obtain the internal and external dimension risk levels;
[0023] S6. Integrate the internal and external data risk levels to obtain the comprehensive customer rating;
[0024] Place the internal and external data risk levels on the horizontal and vertical coordinate axes of the first quadrant respectively to form a heat map of the customer rating nine-square grid; when the combination of the internal and external data risk levels is one of (normal, normal), (normal, vigilant), and (vigilant, normal), the comprehensive customer rating is normal; when the combination of the internal and external data risk levels is one of (vigilant, vigilant), (normal, cautious), and (cautious, normal), the comprehensive customer rating is vigilant; when the combination of the internal and external data risk levels is one of (cautious, cautious), (vigilant, cautious), and (cautious, vigilant), the comprehensive customer rating is cautious;
[0025] S7. Set strong rules for the supplementary data to demarcate cautious merchants;
[0026] On the basis of S6, revise the customer risk level according to whether the customer has litigation with the company or is on the company's unqualified list. When the customer has litigation with the company or the customer is on the company's unqualified list, some customers will be revised from normal and vigilant obtained in S6 to cautious, and the comprehensive rating will be revised with reference to the supplementary data.
[0027] Preferably, the other platforms in step S1.1 are units specializing in collecting social entity information, preferably Tianyancha, Qichacha, and Qixinbao.
[0028] Preferably, the enterprise public information in step S1.3 is the enterprise's industrial and commercial registration information, financial information, credit information, business information, tax information, and judicial litigation information.
[0029] Preferably, the specific steps of step S4.1 are as follows:
[0030] S4.1.1. Infer the black and white lists
[0031] Customers with an empty accounts receivable but a prepaid account or operating income greater than 0 are set as normal customers; for the remaining customers, calculate the accounts receivable turnover rate and accounts receivable recovery rate of different customers for four years respectively;
[0032] Based on the list of (enterprises, customers) in the previous year, the data of the following four indicators are obtained through matching, including the accounts receivable turnover rate in the previous year, the accounts receivable recovery rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable recovery rate in the previous two years;
[0033] Based on the indicator data of the customer group, calculate the quartiles Q1 and Q3 of each indicator respectively; divide the three intervals [0, Q1], (Q1, Q3), [Q3, +∞); the indicator values in these three intervals are assigned three levels of caution, vigilance, and normal, and are assigned values of 2, 1, and 0 respectively;
[0034] Infer the qualified list, and the (enterprise, customer) samples with the total indicator risk score less than or equal to 2 are designated as the qualified list samples; infer the unqualified list, and the (enterprise, customer) samples with accounts receivable outstanding for two years are designated as the unqualified list samples. Among them, accounts receivable outstanding for two years means that the beginning balance = ending balance > 0 of the enterprise's accounts receivable from the customer for two consecutive years, and the cumulative debit = cumulative credit = 0;
[0035] S4.1.2. Construct model learning data
[0036] Calculate and match the internal indicator data of the previous three years and the previous four years of the black and white lists to obtain four indicator risk scores based on the time point of the previous three years: the risk score of the accounts receivable turnover rate in the previous year, the risk score of the accounts receivable recovery rate in the previous year, the risk score of the accounts receivable turnover rate in the previous two years, and the risk score of the accounts receivable recovery rate in the previous two years;
[0037] Add industry control variables, embed the four indicator risk score data of the unqualified list and the qualified list based on the time point of the previous three years in the form of one-hot encoding, and label the unqualified list and the qualified list, where the unqualified list is assigned a value of 1 and the qualified list is assigned a value of 0 to obtain the final model learning data;
[0038] S4.1.3. Construct and train the model to obtain the internal data indicator weights
[0039] Use the third-party library sklearn of python to split the training set and the test set, and construct a random forest model. And use grid search to obtain the best model; use the feature_importance method of the random forest model in the third-party library sklearn of python to obtain the indicator weight vector ω of the best model;
[0040] S4.1.4. Calculate the overall risk of internal data
[0041] Taking the previous year's time point as a reference, four indicator data of (enterprise, customer) are obtained: the accounts receivable turnover rate in the previous year, the accounts receivable recovery rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable recovery rate in the previous two years. Then, the respective indicator risks are divided by quartiles to obtain data in the form of four indicator risk scores for the black and white lists based on the previous year's time point;
[0042] Referring to the method mentioned in S4.1.2, the industry is added as a control variable to the four indicator risk score data of the black and white lists based on the previous year's time point as the basic data for risk rating;
[0043] The basic data is extracted into an indicator data matrix A, and the risk score can be obtained as: S = A × ω T , using the third-party library numpy in python to calculate the matrix multiplication, calculating the quartiles Q1 and Q3 based on the obtained risk scores, dividing the three intervals [0, Q1], (Q1, Q3), [Q3, +∞), and assigning the risk values in these three intervals to the three levels of normal, vigilant, and cautious for the internal historical transaction data, replacing them with 0, 1, and 2 to obtain the rating result data of the internal historical transaction data.
[0044] Preferably, the risk evaluation indicators in step S4.2 are indicators such as registered capital, paid-in capital, whether the business is abnormal, whether it is illegal and dishonest, and whether it is subject to enforcement for dishonesty.
[0045] Preferably, the horizontal axis of the nine-square grid heat map in step S6 is the internal data rating result, from left to right are normal, vigilant, and cautious; the vertical axis of the nine-square grid heat map is the external data rating result, from bottom to top are normal, vigilant, and cautious.
[0046] The beneficial effects of the present invention are:
[0047] The method designed by the present invention combines internal historical transactions, customer external big data, and manually supplemented data accumulated from transaction experience, effectively evaluates the customer credit risk, and provides information support for enterprises to optimize customer cooperation and prevent customer risks. Specific embodiments
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the technical content of the embodiments described in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] A customer credit risk evaluation method integrating machine learning models and expert knowledge and experience specifically includes the following steps:
[0050] I. Data Preparation
[0051] S1: Extract historical transaction data from the NC database with the help of an ETL tool, and set up a scheduled task to update customer historical transaction data monthly; obtain external data related to customers by cooperating with units focusing on collecting social entity information such as Tianyancha, Qichacha, and Qixinbao; respectively construct data warehouses for customer internal data and external data;
[0052] S2: Obtain the required multi-year accounting subject data (4 years in this embodiment) from the internal historical transaction data warehouse through writing SQL scripts for internal historical transaction data modeling, and obtain enterprise public information such as industrial and commercial registration information, financial information, credit information, business information, tax information, and judicial litigation information from the external data warehouse for external data modeling;
[0053] II. Internal Historical Transaction Data Modeling
[0054] 1. Infer Black and White Lists
[0055] S3: Set customers with accounts receivable being empty but advance receipts or operating income greater than 0 as normal customers;
[0056] S4: For the remaining customers, calculate the accounts receivable turnover rate and accounts receivable recovery rate for four years for different customers respectively, where,
[0057]
[0058] S5: Based on the list of (enterprises, customers) in the previous year, match the data of the following four indicators, including the accounts receivable turnover rate in the previous year, the accounts receivable recovery rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable recovery rate in the previous two years; the data sample is shown in Table 1 below:
[0059] Table 1 Sample Table of Customer Historical Transaction Data Indicators
[0060]
[0061] S6: Based on the customer group index data, calculate the quartiles Q1 and Q3 for each indicator respectively; divide the three intervals [0, Q1], (Q1, Q3), [Q3, +∞); assign the indicator values in these three intervals with three levels of caution, vigilance, and normal, and use 2, 1, 0; obtain the sample data shown in Table 2 below:
[0062] Table 2 Sample Table of Customer Historical Transaction Data Indicator Risks
[0063]
[0064] S7: Inference of white list, samples of (enterprises, customers) with the total index risk score less than or equal to 2 are designated as white list samples;
[0065] S8: Inference of unqualified list, samples of (enterprises, customers) with accounts receivable on the books for two years are designated as unqualified list samples. Among them, accounts receivable on the books for two years means that the beginning balance = ending balance > 0 of the accounts receivable of the enterprise from the customer for two consecutive years, and the cumulative debit = cumulative credit = 0;
[0066] 2. Construct model learning data
[0067] S9: Calculate and match the internal index data of the previous three years and the previous four years of the black and white lists to obtain four index risk scores based on the time point of the previous three years: the risk score of accounts receivable turnover rate in the previous year, the risk score of accounts receivable recovery rate in the previous year, the risk score of accounts receivable turnover rate in the previous two years, and the risk score of accounts receivable recovery rate in the previous two years; The data example is shown in Table 3 below:
[0068] Table 3 Four index risk score table of black and white lists based on the time point of the previous three years
[0069]
[0070] S10: Add industry control variables; Embed the data in Table 3 in the form of one-hot encoding, and label the unqualified list and the qualified list, where the unqualified list is assigned a value of 1 and the qualified list is assigned a value of 0; Obtain the final model learning data, and the sample data is shown in Table 4 below:
[0071] Table 4 Black and white list model data table
[0072]
[0073] 3. Construct and train the model to obtain the internal data index weights
[0074] S11: Use the third-party library sklearn in python to split the training set and the test set, and construct a random forest model; And use grid search to obtain the best model;
[0075] S12: Use the feature_importance method of the random forest model in the third-party library sklearn in python to obtain the index weight vector ω of the best model;
[0076] 4. Calculate the overall internal data risk
[0077] S13: Based on the time point of the previous year, obtain the four indicator data of (enterprise, customer): the accounts receivable turnover rate in the previous year, the accounts receivable recovery rate in the previous year, the accounts receivable turnover rate in the two previous years, and the accounts receivable recovery rate in the two previous years; then, divide the respective indicator risks by quartiles to obtain the data in the form of Table 3;
[0078] S14: Referring to the method mentioned in S10, add the industry as a control variable to the data obtained in S13 as the basic data for risk rating;
[0079] S15: Extract the basic data into an indicator data matrix A; the risk score can be obtained as: S = A × ω T , use the third-party library numpy in python to calculate matrix multiplication; calculate the quartiles Q1 and Q3 based on the obtained risk scores, and divide them into three intervals [0, Q1], (Q1, Q3), [Q3, +∞); assign the risk values in these three intervals to the three levels of normal, vigilant, and cautious for the internal historical transaction data, and replace them with 0, 1, and 2; obtain the sample data as shown in Table 5 below:
[0080] Table 5 Rating Results Table of Internal Historical Transaction Data
[0081] Enterprise Code Customer Code Internal Rating 61040200Q 68010000P 2 93000000P 34001200P 1 93000000P 34001000P 0
[0082] III. External Data Modeling
[0083] S16: Use the third-party tools pandas and numpy in python to construct risk evaluation indicators, including registered capital, paid-in capital, whether there is an abnormal operation, whether there is illegal dishonesty, whether there is an execution for dishonesty, etc.; for absolute value indicators, divide the normal, vigilant, and cautious intervals of the indicators by quartiles, and then according to the indicator values falling into different intervals, obtain the normal, vigilant, and cautious level types of the indicators, and assign values of 0, 1, and 2 respectively; for general yes / no type indicators, assign values of 1 and 0 according to whether the indicator is triggered or not, and for important yes / no type indicators, assign values of 2 and 0 according to whether the indicator is triggered or not, and obtain the sample data as shown in Table 6 below:
[0084] Table 6 Risk Score Table of External Data Indicators
[0085]
[0086] S17: Based on the sum of the risk scores of each indicator of each customer multiplied by the indicator weights, initially obtain the external data risk score; set the score threshold according to the score data distribution for clustering to divide the external data risk level interval of the customer, and then obtain the external data risk level of the customer, and assign values of 0, 1, and 2 to normal, vigilant, and cautious respectively; an example of the external data rating result is shown in Table 7 below:
[0087] Table 7 External Data Rating Results Table
[0088] Customer Code External Data Rating 68010000P 2 34001200P 0 34001000P 0
[0089] IV. Integration of Internal and External Data Ratings
[0090] S18: After obtaining the rating results of the customer's internal and external data respectively, place the internal and external data rating results on the horizontal and vertical coordinate axes of the nine-square grid heat map; the horizontal axis is the internal data rating result, which is normal, vigilant, and cautious from left to right in sequence, and the vertical axis is the external data rating result, which is normal, vigilant, and cautious from bottom to top in sequence, forming a customer rating nine-square grid; for the three categories of (normal, normal), (normal, vigilant), and (vigilant, normal), the customer's comprehensive rating is designated as normal; for the three categories of (vigilant, vigilant), (normal, cautious), and (cautious, normal), the customer's comprehensive rating is designated as vigilant; for the three categories of (cautious, cautious), (cautious, vigilant), and (vigilant, cautious), the customer's comprehensive rating is designated as cautious; the data sample table is shown in Table 8 below:
[0091] Table 8 Internal and External Data Risk Ratings and Comprehensive Rating Table
[0092] Customer Code External Data Rating Internal Data Rating Comprehensive Rating 68010000P 2 2 2 34001200P 0 2 1 34001000P 0 1 0
[0093] V. Revise the Comprehensive Rating with Reference to Supplementary Recorded Data
[0094] S19: In the case where the customer's comprehensive rating has been obtained in S18, revise the customer rating with reference to the supplementary recorded data accumulated from the enterprise's operations. For customers who have litigation with the unit or are on the unit's unqualified list, they need to be classified as cautious customers; the revised data sample table is shown in Table 9 below:
[0095] Table 9 Revised Comprehensive Rating Table with Reference to Supplementary Recorded Data
[0096]
[0097] In the present invention, the method combines internal historical transactions, customer external big data, and manually supplementary recorded data accumulated from transaction experience, effectively evaluates the customer credit risk, and provides information support for the enterprise to optimize customer cooperation and prevent customer risks.
[0098] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A customer credit risk assessment method that integrates machine learning models and expert knowledge and experience, characterized in that: The specific steps include: S1. Data integration; S1.
1. Use ETL tools to extract the company's accumulated customer historical transaction data from the NC database on a monthly basis, and obtain external data through cooperation with other platforms and update it at a certain frequency; S1.
2. Manually record litigation and disqualified list data related to customers on a monthly basis; integrate historical transaction data, external data, and supplementary data to form a data warehouse based on internal and external data of customers. This data warehouse is a subject-oriented, integrated, relatively stable data set that reflects historical changes and is used to support risk analysis; S1.
3. Obtain the required accounting data for multiple years from the data warehouse for internal historical transaction data modeling by writing SQL scripts, and obtain corporate public information from the data warehouse for external data modeling; S2. Construct risk rating indicators for data of different dimensions; S2.
1. For historical transaction data, set four indicators: accounts receivable turnover rate in the previous year and two years, accounts receivable collection rate in the previous year and two years, among which: S2.
2. For external data, set up indicators such as registered capital, paid-in capital, unpaid paid-in capital, proportion of abnormal business operations of foreign-invested enterprises, business status, four major financial capacity indicators, whether there is abnormal business information, whether there is serious violation of law and dishonesty, whether dishonesty is subject to execution, total guarantee amount to registered capital ratio, whether it is an abnormal taxpayer, and the cumulative amount involved in the case to registered capital ratio; S2.
3. For supplementary data, set indicators of whether there is litigation with the unit and whether the unit is on the unqualified list; S3. Calculate the actual value of the indicator and assign values to the indicator according to the quartile rule; S3.
1. Based on historical transaction data, calculate the accounts receivable turnover rate and accounts receivable collection rate indicators, use the quantile function in pandas to calculate the quartiles of the indicators, and divide them into three risk intervals of normal, alert, and cautious according to the quartiles. When the actual value of the indicator falls into the above three risk intervals, assign the indicators three values of 0, 1, and 2 respectively; S3.
2. For external data, the registered capital and paid-in capital are directly sorted and quartiles are defined, and risk scores of 0, 1, and 2 are assigned in the same way as above; the ratio indicators such as the situation of unpaid paid-in capital, the proportion of abnormal business operations of foreign-invested enterprises, the total amount of guarantees as a percentage of registered capital, and the cumulative amount involved in the case as a percentage of registered capital are calculated, sorted, and quartiles are defined, and risk scores of 0, 1, and 2 are assigned in the same way; for general indicators such as whether there are administrative penalties, whether there are environmental penalties, and whether there are tax arrears, 1 and 0 are assigned respectively according to whether the customer triggers the relevant indicators; for important risk indicators such as whether there is a serious violation of the law and dishonesty, whether dishonesty is enforced, and whether there is a major tax violation, 2 and 0 are assigned respectively according to whether the customer triggers it; S4. Build a learning model and generate indicator weights; S4.
1. Based on historical transaction data, customers whose accounts receivable have been on the books for two consecutive years are set as unqualified list customers; customers whose risk scores of the four indicators of accounts receivable turnover rate in the previous year and the previous two years and accounts receivable collection rate in the previous year and the previous two years are less than or equal to 2 are set as qualified list customers; the unqualified list customers and qualified list customers are assigned values of 1 and 0, representing cautious customers and normal customers respectively; the indicators and indicator risk scores designed in the historical transaction dimension are incorporated into the machine learning model, and the industry is assigned as a control variable to the model learning data set; random forest is used to model the learning data, and the optimal model and indicator weights are obtained by adjusting the parameters; S4.
2. For external data, equal weights are basically set, but for the indicator of the ratio of the cumulative amount involved in the case to the registered capital, its indicator weight is set to 2 times that of other indicators, and finally the sum of the indicator weights of all external data is equal to 1; use python's third-party tools pandas and numpy to build risk assessment indicators. For absolute value indicators, the indicators are divided into normal, alert, and cautious intervals according to quartiles, and then the normal, alert, and cautious level types of indicators are obtained according to the indicator values falling into different intervals, and the values are assigned to 0, 1, and 2 respectively; for general yes and no indicators, the values are assigned to 1 and 0 according to whether the indicators are triggered or not; for important and no indicators, the values are assigned to 2 and 0 according to whether the indicators are triggered or not; the external data risk score is preliminarily obtained by summing the risk score of each customer's indicators × the indicator weight; the score threshold is set according to the distribution clustering of the score data to divide the customer's external data risk level interval, and then the risk level of the customer's external data is obtained, and normal, alert, and cautious are assigned to 0, 1, and 2 respectively; S5. Calculate the customer's internal and external data risk scores and divide the internal and external dimension ratings into quartiles; The method of indicator weight × indicator risk score is used to calculate the risk scores of internal and external data respectively. The risk scores of internal and external data are sorted to obtain quartiles, and the three risk intervals of normal, vigilant, and cautious are divided according to the quartiles. When the risk score falls into the risk interval, the risk level corresponding to the risk score is obtained, thereby obtaining the risk level of the internal and external dimensions; S6. Integrate internal and external data risk levels to obtain a comprehensive customer rating; The internal and external data risk levels are placed on the horizontal and vertical axes of the first quadrant, respectively, to form a nine-square heat map of customer ratings; when the internal and external data risk levels are combined into (normal, normal), (normal, vigilant), and (vigilant, normal), the customer's comprehensive rating is normal; when the internal and external data risk levels are combined into (vigilant, vigilant), (normal, cautious), and (cautious, normal), the customer's comprehensive rating is vigilant; when the internal and external data risk levels are combined into (cautious, cautious), (vigilant, cautious), and (cautious, vigilant), the customer's comprehensive rating is cautious; S7. Set strong rules for supplementary data and identify cautious merchants; On the basis of S6, the customer risk level will be revised according to whether the customer is in litigation with the unit or is on the unit's unqualified list. When the customer is in litigation with the unit or is on the unit's unqualified list, some customers will be revised from normal or alert according to S6 to cautious, and the comprehensive rating will be revised with reference to the supplementary data.
2. The customer credit risk assessment method integrating machine learning model and expert knowledge and experience according to claim 1 is characterized in that: The other platforms in step S1.1 are units that focus on collecting information on social entities, preferably Tianyancha, Qichacha, and Qixinbao.
3. The customer credit risk assessment method integrating machine learning model and expert knowledge and experience according to claim 1 is characterized in that: The enterprise public information in step S1.3 is the enterprise's business registration information, financial information, credit information, business information, tax information, and judicial litigation information.
4. The customer credit risk assessment method integrating machine learning model and expert knowledge and experience according to claim 1 is characterized in that: The specific steps of step S4.1 are: S4.1.
1. Reasoning about unqualified list and qualified list For customers whose accounts receivable are empty but whose advance payments or operating income are greater than 0, they are set as normal customers; for the remaining customers, the accounts receivable turnover rate and accounts receivable collection rate of different customers for four years are calculated respectively; Based on the list of enterprises or customers in the previous year, the following four indicators are matched to obtain data, including the accounts receivable turnover rate in the previous year, the accounts receivable collection rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable collection rate in the previous two years; Based on the customer group indicator data, the quartiles Q1 and Q3 of each indicator are calculated respectively; three intervals are divided: [0, Q1], (Q1, Q3), and [Q3, +∞); the indicator values in these three intervals are assigned three levels: cautious, vigilant, and normal, and are assigned values of 2, 1, and 0; Inferring the qualified list, the enterprise or customer samples whose total indicator risk score is less than or equal to 2 are designated as qualified list samples; inferring the unqualified list, the enterprise or customer samples whose accounts receivable have been outstanding for two years are designated as unqualified list samples, where accounts receivable outstanding for two years means that the beginning balance of the enterprise's accounts receivable to customers for two consecutive years = the ending balance > 0, and the accumulated debits = the accumulated credits = 0; S4.1.
2. Constructing model learning data Calculate and match the internal indicator data of the previous three years and the previous four years of the unqualified list and the qualified list to obtain the risk scores of four indicators based on the previous three years: the risk score of the accounts receivable turnover rate in the previous year, the risk score of the accounts receivable recovery rate in the previous year, the risk score of the accounts receivable turnover rate in the previous two years, and the risk score of the accounts receivable recovery rate in the previous two years; Add industry control variables, embed the risk score data of the four indicators of the unqualified list and qualified list based on the previous three years in the form of one-hot encoding, and label them with the unqualified list and qualified list, where the unqualified list is assigned a value of 1 and the qualified list is assigned a value of 0, and obtain the final model learning data; S4.1.
3. Build and train the model to obtain the internal data indicator weights Use the python third-party library sklearn to split the training set and test set, build a random forest model, and use grid search to get the best model; use the random forest model feature_importance method of the python third-party library sklearn to get the indicator weight vector ω of the best model; S4.1.
4. Calculate the overall risk of internal data Taking the previous year as the benchmark, we take the four indicators of the enterprise or customer: the accounts receivable turnover rate in the previous year, the accounts receivable collection rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable collection rate in the previous two years. Then, we divide the risk of each indicator by quartiles, and get the data in the form of risk scores of the four indicators, namely, the unqualified list and the qualified list, based on the previous year as the benchmark; Referring to the method mentioned in S4.1.2, the industry is added as a control variable to the risk score data of the four indicators of the unqualified list and the qualified list based on the time point of the previous year as the basic data for risk rating; By extracting the basic data into the indicator data matrix A, the risk score can be obtained as: S = A × ω T , use the python third-party library numpy to calculate matrix multiplication, calculate the quartiles Q1 and Q3 according to the obtained risk scores, and divide them into three intervals [0, Q1], (Q1, Q3), and [Q3, +∞). The risk values in these three intervals are assigned to the internal historical transaction data as normal, vigilant, and cautious, and are assigned with 0, 1, and 2 to obtain the internal historical transaction data rating result data.
5. The customer credit risk assessment method integrating machine learning model and expert knowledge and experience according to claim 1 is characterized in that: The risk assessment indicators in step S4.2 are registered capital, paid-in capital, whether the business is abnormal, whether the business is illegal or dishonest, whether the business is subject to execution for dishonesty, etc.
6. The customer credit risk assessment method integrating machine learning model and expert knowledge and experience according to claim 1 is characterized in that: The horizontal axis of the nine-square heat map in step S6 is the internal data rating result, which is normal, alert, and cautious from left to right; the vertical axis of the nine-square heat map is the external data rating result, which is normal, alert, and cautious from bottom to top.
Citation Information
Patent Citations
Small merchant credit evaluation method of fusion of expert model and machine learning model
CN107644375A
Random forest rule generator
US20240070474A1
Credit evaluation method, apparatus and device, and computer readable storage medium
WO2019080407A1