Customer credit risk assessment method integrating machine learning models and expert knowledge and experience

By integrating machine learning models and expert knowledge and experience, and combining historical transaction data, supplementary data and external data between enterprises and customers, the subjectivity and efficiency issues in customer credit risk assessment are resolved, and accurate assessment and risk identification of customer credit risk are achieved.

CN120070041BActive Publication Date: 2025-10-03HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510229465.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-10-03
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing technology in customer credit risk assessment has the problems of strong subjectivity of manual evaluation, heavy workload and difficulty in accurately identifying risky customers.

Method used

By adopting a method that combines machine learning models and expert knowledge and experience, the company integrates, cleans and processes historical transaction data, supplementary data and external data between the company and its customers, combines risk rating indicators, uses mainstream machine learning algorithms and expert knowledge rules, and outputs customer credit risk levels and risk items.

Benefits of technology

It achieves accurate assessment of customer credit risk, reduces the subjectivity of human evaluation, improves work efficiency, and can effectively identify risky customers and retain high-quality customers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005582532720000011
    Figure FDA0005582532720000011
  • Figure GDA0005582532730000031
    Figure GDA0005582532730000031
  • Figure GDA0005582532730000091
    Figure GDA0005582532730000091
Patent Text Reader

Abstract

The present invention relates to a customer credit risk assessment method that integrates machine learning models and expert knowledge and experience. The method forms a DW data warehouse by integrating historical transaction data, external data, and supplementary data, constructs risk rating indicators for data of different dimensions for the historical transaction data, external data, and supplementary data, calculates actual values ​​of the indicators and assigns values ​​to the indicators according to the quartile rule, builds a learning model based on the data warehouse and generates indicator weights, calculates risk scores for internal and external customer data, and delineates internal and external dimension ratings according to quartiles, integrates the risk levels of internal and external data to obtain a comprehensive customer rating, sets strong rules for the supplementary data, delineates cautious merchants, and revises the comprehensive rating. The method designed by the present invention combines internal historical transaction big data, external customer big data, and manually supplemented data accumulated from transaction experience to effectively evaluate customer credit risk, providing information support for enterprises to optimize customer cooperation and prevent customer risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a customer credit risk assessment method that integrates a machine learning model and expert knowledge and experience. Background Art

[0002] In order to deeply explore the value of data assets and give full play to the important role of artificial intelligence algorithm models in enabling corporate customer risk management, the present invention has designed a customer credit risk rating method that integrates machine learning models and expert knowledge and experience. Specifically, the present invention designs risk rating indicators based on the historical transaction data between the enterprise and the customer, as well as the company's own external data (including business registration information, financial information, tax arrears information, case filing information, judicial document information, etc.). After integrating, cleaning, and processing the data, the machine learning model and expert knowledge and experience are used to rate the customer's risk. This method can not only obtain the customer's comprehensive risk level, but also obtain customer risk items. At the same time, applying this method to corporate practice can achieve accurate scanning of customer risks, reduce the subjectivity and workload of manual evaluation of customer pros and cons, effectively identify risky customers while retaining high-quality customers. Summary of the Invention

[0003] In order to overcome the problems existing in the background technology, the present invention provides a customer credit risk assessment method that integrates machine learning models and expert knowledge and experience. The risk rating object of this method is a general enterprise with a unified social credit code starting with 91 and the nature of the transaction as a customer. The specific rating process relies on three parts of data. The first is the historical transaction data between the enterprise and the customer, such as the ending balance or amount data of accounting subjects such as accounts receivable, advance payments, main business income, and other business income; the second is the supplementary data related to customers accumulated during the operation of the enterprise, including data on litigation with customers and data on unqualified lists; the third is public external data, including industrial and commercial registration, financial information, tax arrears information, etc. Through the integration, cleaning and processing of the three parts of data, with the help of the mainstream machine learning algorithm in the artificial intelligence model, the expert knowledge and experience are integrated to establish rules to evaluate the customer credit risk. While outputting the three risk levels of normal, vigilant and cautious, it also outputs the risk level and risk items under each rating dimension.

[0004] To achieve the purpose of the present invention, the technical solution adopted is:

[0005] A customer credit risk assessment method that integrates machine learning models and expert knowledge and experience, specifically including the following steps:

[0006] S1. Data integration;

[0007] S1.1. Use ETL tools to extract historical customer transaction data from the NC database on a monthly basis. By collaborating with other platforms, obtain external data and update it at a regular interval.

[0008] S1.2. Manually re-enter client-related litigation and disqualification list data on a monthly basis. This data warehouse, based on internal and external client data, integrates historical transaction data, external data, and re-entered data to form a data warehouse. This data warehouse is a subject-oriented, integrated, relatively stable data set reflecting historical changes, used to support risk analysis.

[0009] S1.3. Write SQL scripts to retrieve the required multi-year accounting data from the data warehouse for internal historical transaction data modeling, and retrieve publicly available corporate information from the data warehouse for external data modeling.

[0010] S2. Construct risk rating indicators for data of different dimensions;

[0011] S2.1. For historical transaction data, set four indicators: accounts receivable turnover rate in the previous year and two years, and accounts receivable collection rate in the previous year and two years.

[0012]

[0013] S2.2. For external data, indicators include registered capital, paid-in capital, unpaid paid-in capital, the percentage of enterprises with abnormal operations, operating status, and financial capabilities, as well as the presence of abnormal operations, serious violations of law and trust, enforcement of default, the ratio of total guarantees to registered capital, abnormal taxpayer status, and the ratio of the cumulative amount involved in the case to registered capital.

[0014] S2.3. For supplementary data, set indicators for whether the data is in litigation with the unit and whether the data is on the unit's unqualified list;

[0015] S3. Calculate the actual value of the indicator and assign values ​​to the indicator according to the quartile rule;

[0016] S3.1. Calculate the accounts receivable turnover rate and accounts receivable collection rate indicators based on historical transaction data. Use the quantile function in Pandas to calculate the quartiles of these indicators. Use these quartiles to divide the indicators into three risk intervals: normal, alert, and cautious. When the actual indicator values ​​fall into these three risk intervals, assign them values ​​of 0, 1, and 2, respectively.

[0017] S3.2. For external data, directly sort registered capital and paid-in capital and assign quartiles, using the same method as above to assign risk scores of 0, 1, and 2. For ratio indicators such as the percentage of unpaid paid-in capital, the proportion of abnormal operations of outbound investment enterprises, the ratio of total guarantee amounts to registered capital, and the ratio of cumulative case amounts to registered capital, we calculate, sort, and assign quartiles, assigning risk scores of 0, 1, and 2 using the same method. For general yes / no indicators such as whether there have been administrative penalties, whether there have been environmental penalties, and whether there are tax arrears, we assign values ​​of 1 and 0, respectively, depending on whether the customer triggers the relevant indicator. For important risk indicators such as whether there has been a serious violation of law and trust, whether there has been enforcement of the breach of trust, and whether there has been a major tax violation, we assign values ​​of 2 and 0, respectively, depending on whether the customer triggers the relevant indicator.

[0018] S4. Build a learning model and generate indicator weights;

[0019] S4.1. Based on historical transaction data, customers with accounts receivable outstanding for two consecutive years are designated as ineligible. Customers whose sum of the risk scores for the four indicators, accounts receivable turnover rate and accounts receivable collection rate, in the previous year and two years is less than or equal to 2 are designated as qualified. Customers on the ineligible and qualified lists are assigned a value of 1, representing cautious customers, and 0, representing normal customers. Indicators and their risk scores designed from the historical transaction dimension are incorporated into the machine learning model, and industry is assigned as a control variable to the model learning dataset. Random forest modeling is performed on the learning data, and the optimal model and indicator weights are obtained through parameter adjustment.

[0020] S4.2. For external data, equal weights are generally set. However, for the indicator "the ratio of the cumulative amount involved in the case to registered capital," its indicator weight is set to twice that of the other indicators, so that the sum of the indicator weights for all external data equals 1. Risk assessment indicators are constructed using the third-party Python tools pandas and numpy. For absolute value indicators, the indicators are divided into normal, alert, and cautious intervals according to quartiles. Then, based on where the indicator values ​​fall into different intervals, the indicators are assigned normal, alert, and cautious levels, with values ​​of 0, 1, and 2, respectively. For general "yes" and "no" indicators, values ​​of 1 and 0 are assigned based on whether the indicator is triggered or not, respectively. For important and "no" indicators, values ​​of 2 and 0 are assigned based on whether the indicator is triggered or not, respectively. A preliminary external data risk score is derived by multiplying the risk score of each customer's indicators by the indicator weight. A score threshold is set based on the distribution and clustering of the score data to divide the customer's external data risk level intervals, thereby deriving the customer's external data risk level, with normal, alert, and cautious values ​​of 0, 1, and 2, respectively.

[0021] S5. Calculate the risk scores of the client's internal and external data and assign ratings to the internal and external dimensions according to quartiles;

[0022] The internal and external data risk scores are calculated using the indicator weight × indicator risk score method. The internal and external data risk scores are sorted to obtain quartiles, and the three risk intervals of normal, alert, and cautious are divided according to the quartiles. When the risk score falls into the risk interval, the risk level corresponding to the risk score is obtained, thereby obtaining the risk level of the internal and external dimensions;

[0023] S6. Integrate internal and external data risk levels to obtain a comprehensive customer rating;

[0024] The internal and external data risk levels are placed on the horizontal and vertical axes of the first quadrant, respectively, to form a nine-square heat map of customer ratings. When the internal and external data risk levels are combined into (normal, normal), (normal, vigilant), and (vigilant, normal), the customer's overall rating is normal. When the internal and external data risk levels are combined into (vigilant, vigilant), (normal, cautious), and (cautious, normal), the customer's overall rating is cautious. When the internal and external data risk levels are combined into (cautious, cautious), (vigilant, cautious), and (cautious, vigilant), the customer's overall rating is cautious.

[0025] S7. Set strong rules for supplementary data and identify cautious merchants;

[0026] On the basis of S6, the customer risk level will be revised according to whether the customer is in litigation with the unit or is on the unit's unqualified list. When the customer is in litigation with the unit or is on the unit's unqualified list, some customers will be revised from normal or alert according to S6 to cautious, and the comprehensive rating will be revised with reference to the supplementary data.

[0027] Preferably, the other platforms in step S1.1 are units that focus on collecting social entity information, preferably Tianyancha, Qichacha, and Qixinbao.

[0028] Preferably, the enterprise public information in step S1.3 is the enterprise's business registration information, financial information, credit information, operating information, tax information, and judicial litigation information.

[0029] Preferably, the specific steps of step S4.1 are:

[0030] S4.1.1. List of unqualified and qualified inferences

[0031] Customers whose accounts receivable are empty but whose prepayments or operating income are greater than 0 are defined as normal customers. For the remaining customers, the accounts receivable turnover rate and accounts receivable collection rate are calculated for each customer over four years.

[0032] Based on the list of (enterprises, customers) from the previous year, the following four indicators are matched to obtain data, including the accounts receivable turnover rate of the previous year, the accounts receivable collection rate of the previous year, the accounts receivable turnover rate of the previous two years, and the accounts receivable collection rate of the previous two years;

[0033] Based on the customer group indicator data, the quartiles Q1 and Q3 of each indicator are calculated. The three intervals are divided into [0, Q1], (Q1, Q3), and [Q3, +∞). The indicator values ​​in these three intervals are assigned three levels: cautious, vigilant, and normal, and are assigned values ​​of 2, 1, and 0.

[0034] For the qualified list, the samples of (enterprises, customers) with a total risk score of 2 or less are designated as qualified. For the unqualified list, the samples of (enterprises, customers) with accounts receivable outstanding for two years are designated as unqualified. Two years of accounts receivable outstanding means that the beginning balance of the enterprise's accounts receivable from customers is equal to the ending balance > 0 for two consecutive years, and the accumulated debits are equal to the accumulated credits = 0.

[0035] S4.1.2. Constructing model learning data

[0036] Calculate and match the internal indicator data of the previous three years and the previous four years on the unqualified list and the qualified list to obtain four indicator risk scores based on the previous three years: the previous year's accounts receivable turnover risk score, the previous year's accounts receivable collection rate risk score, the previous two years' accounts receivable turnover risk score, and the previous two years' accounts receivable collection rate risk score;

[0037] Industry control variables were added, and the risk score data of the four indicators, the unqualified list and the qualified list, based on the previous three years were embedded in the form of one-hot encoding. The data were then labeled as unqualified list and qualified list, with the unqualified list assigned a value of 1 and the qualified list assigned a value of 0, to obtain the final model learning data.

[0038] S4.1.3. Build and train the model to obtain internal data indicator weights

[0039] Use the Python third-party library sklearn to split the training set and test set, build a random forest model, and use grid search to find the best model. Use the feature_importance method of the random forest model of the Python third-party library sklearn to find the indicator weight vector ω of the best model.

[0040] S4.1.4. Calculate the overall risk of internal data

[0041] Using the previous year as a benchmark, we collected data on four indicators (for businesses and customers): accounts receivable turnover rate for the previous year, accounts receivable collection rate for the previous year, accounts receivable turnover rate for the previous two years, and accounts receivable collection rate for the previous two years. We then divided the risk of each indicator by quartile, generating risk scores for the four indicators, including a failed list and a qualified list, based on the previous year.

[0042] Referring to the method mentioned in S4.1.2, industry is added as a control variable to the risk score data of the four indicators of the unqualified list and qualified list based on the previous year as the basic data for risk rating;

[0043] Extracting the basic data into the indicator data matrix A, the risk score can be obtained as: S = A × ω T , use the python third-party library numpy to calculate matrix multiplication, calculate the quartiles Q1 and Q3 according to the obtained risk score, and divide it into three intervals [0,Q1], (Q1,Q3), and [Q3,+∞). The risk values ​​in these three intervals are assigned to the internal historical transaction data as normal, vigilant, and cautious, respectively, and replaced by 0, 1, and 2 to obtain the internal historical transaction data rating result data.

[0044] Preferably, the risk assessment indicators in step S4.2 are registered capital, paid-in capital, whether the business is abnormal, whether the business is illegal or dishonest, whether the business is subject to execution for dishonesty, and the like.

[0045] Preferably, the horizontal axis of the nine-square heat map in step S6 is the internal data rating result, which is normal, vigilant, and cautious from left to right; the vertical axis of the nine-square heat map is the external data rating result, which is normal, vigilant, and cautious from bottom to top.

[0046] The beneficial effects of the present invention are:

[0047] The method designed in this invention combines internal historical transactions, external big data of customers, and manually recorded data accumulated from transaction experience to effectively evaluate customer credit risk, providing information support for enterprises to optimize customer cooperation and prevent customer risks. DETAILED DESCRIPTION

[0048] The following will be combined with the technical content of the embodiments described in the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] A customer credit risk assessment method that integrates machine learning models and expert knowledge and experience, specifically including the following steps:

[0050] 1. Data Preparation

[0051] S1: Use ETL tools to extract historical transaction data from the NC database and set up scheduled tasks to update customer historical transaction data monthly. Acquire external customer data by collaborating with organizations specializing in collecting social entity information, such as Tianyancha, Qichacha, and Qixinbao. Build data warehouses for both internal and external customer data.

[0052] S2: By writing SQL scripts, the required accounting data for multiple years (four years is used in this embodiment) is obtained from the internal historical transaction data warehouse for internal historical transaction data modeling. In addition, public enterprise information such as business registration information, financial information, credit information, operating information, tax information, and judicial litigation information is obtained from the external data warehouse for external data modeling.

[0053] 2. Internal Historical Transaction Data Modeling

[0054] 1. Reasoning Unqualified List and Qualified List

[0055] S3: Set customers whose accounts receivable are empty but whose advance payments or operating income are greater than 0 as normal customers;

[0056] S4: For the remaining customers, calculate the accounts receivable turnover rate and accounts receivable collection rate for each customer for four years, where:

[0057]

[0058]

[0059] S5: Based on the list of (enterprises, customers) from the previous year, the following four indicators are matched and obtained: accounts receivable turnover rate from the previous year, accounts receivable collection rate from the previous year, accounts receivable turnover rate from the previous two years, and accounts receivable collection rate from the previous two years; a sample of the data is shown in Table 1 below:

[0060] Table 1 Sample table of customer historical transaction data indicators

[0061]

[0062] S6: Based on the customer group indicator data, calculate the quartiles Q1 and Q3 of each indicator respectively; divide the data into three intervals: [0, Q1], (Q1, Q3), and [Q3, +∞); assign the indicator values ​​in these three intervals to three levels: cautious, vigilant, and normal, and use 2, 1, and 0 as the values. The sample data shown in Table 2 is obtained:

[0063] Table 2 Customer Historical Transaction Data Risk Sample Table

[0064]

[0065] S7: Infer the qualified list, the (enterprise, customer) samples with the sum of indicator risk scores less than or equal to 2 are designated as qualified list samples;

[0066] S8: Infer the unqualified list, and designate the (enterprise, customer) samples with accounts receivable outstanding for two years as the unqualified list samples, where accounts receivable outstanding for two years means that the beginning balance of the enterprise's accounts receivable to customers is equal to the ending balance > 0 for two consecutive years, and the accumulated debits are equal to the accumulated credits = 0;

[0067] 2. Build model learning data

[0068] S9: Calculate and match the internal indicator data of the previous three years and the previous four years of the unqualified list and the qualified list to obtain the risk scores of four indicators based on the previous three years: the risk score of the accounts receivable turnover rate in the previous year, the risk score of the accounts receivable collection rate in the previous year, the risk score of the accounts receivable turnover rate in the previous two years, and the risk score of the accounts receivable collection rate in the previous two years; the data examples are shown in Table 3 below:

[0069] Table 3 Risk score table of four indicators of unqualified list and qualified list based on the time point of the previous three years

[0070]

[0071] S10: Add industry control variables; embed the data in Table 3 in the form of one-hot encoding, and label them as unqualified list and qualified list, where the unqualified list is assigned a value of 1 and the qualified list is assigned a value of 0; obtain the final model learning data, and the sample data is shown in Table 4 below:

[0072] Table 4 Model data table of unqualified list and qualified list

[0073]

[0074] 3. Build and train the model to obtain the weight of internal data indicators

[0075] S11: Use the Python third-party library sklearn to split the training set and test set, build a random forest model, and use grid search to find the best model;

[0076] S12: Use the feature_importance method of the random forest model of the Python third-party library sklearn to obtain the indicator weight vector ω of the optimal model;

[0077] 4. Calculate the overall risk of internal data

[0078] S13: Using the previous year as a benchmark, take the following four indicators (for the enterprise and customer): accounts receivable turnover rate for the previous year, accounts receivable collection rate for the previous year, accounts receivable turnover rate for the previous two years, and accounts receivable collection rate for the previous two years; then, divide the risk of each indicator by quartiles to obtain the data in Table 3.

[0079] S14: Referring to the method mentioned in S10, add the industry as a control variable to the data obtained in S13 as the basic data for risk rating;

[0080] S15: Extract the basic data into the indicator data matrix A; the risk score can be obtained as: S = A × ω T , using the Python third-party library numpy to calculate matrix multiplication; based on the obtained risk score, quartiles Q1 and Q3 are calculated and divided into three intervals [0, Q1], (Q1, Q3), and [Q3, +∞); the risk values ​​in these three intervals are assigned to the internal historical transaction data as normal, vigilant, and cautious, respectively, and replaced by 0, 1, and 2; the sample data shown in Table 5 is obtained:

[0081] Table 5 Internal historical transaction data rating results

[0082] Enterprise Code Customer Code Internal Ratings 61040200Q 68010000P 2 93000000P 34001200P 1 93000000P 34001000P 0

[0083] 3. External Data Modeling

[0084] S16: Use Python's third-party tools pandas and numpy to construct risk assessment indicators, including registered capital, paid-in capital, whether operations are abnormal, whether there is illegal or dishonest behavior, whether dishonest behavior is subject to enforcement, etc. For absolute value indicators, the indicators are divided into normal, alert, and cautious intervals according to quartiles, and then the normal, alert, and cautious levels of the indicators are obtained according to the indicator values ​​falling into different intervals, and the values ​​are assigned to 0, 1, and 2 respectively. For general yes / no indicators, the values ​​are assigned to 1 and 0 according to whether the indicators are triggered or not, respectively. For important and no indicators, the values ​​are assigned to 2 and 0 according to whether the indicators are triggered or not, respectively. The sample data is shown in Table 6 below:

[0085] Table 6 External data indicator risk score table

[0086]

[0087] S17: Based on the sum of the risk score of each customer's indicators and the indicator weight, a preliminary external data risk score is obtained. According to the distribution clustering of the score data, the score threshold is set to divide the customer's external data risk level interval, and then the risk level of the customer's external data is obtained. Normal, alert, and cautious are assigned values ​​of 0, 1, and 2 respectively. An example of the external data rating results is shown in Table 7 below:

[0088] Table 7 External data rating results

[0089] Customer Code External data rating 68010000P 2 34001200P 0 34001000P 0

[0090] IV. Integration of Internal and External Data Ratings

[0091] S18: After obtaining the rating results of the customer's internal and external data respectively, the internal and external data rating results are placed on the horizontal and vertical axes of the nine-square heat map respectively; the horizontal axis is the internal data rating result, which is normal, vigilant, and cautious from left to right, and the vertical axis is the external data rating result, which is normal, vigilant, and cautious from bottom to top, forming a customer rating nine-square grid; (normal, normal), (normal, vigilant), (vigilant, normal) are classified as normal for the three categories; (vigilant, vigilant), (normal, cautious), (cautious, normal) are classified as cautious for the three categories; (cautious, cautious), (cautious, vigilant), (vigilant, cautious) are classified as cautious for the three categories; the comprehensive rating of the customer is classified as cautious for the three categories. The data sample table is shown in Table 8 below:

[0092] Table 8 Internal and external data risk rating and comprehensive rating table

[0093] Customer Code External data rating Internal data rating Overall Rating 68010000P 2 2 2 34001200P 0 2 1 34001000P 0 1 0

[0094] 5. Revise the comprehensive rating with reference to the supplementary data

[0095] S19: If the comprehensive customer rating has been obtained in S18, the customer rating will be revised by referring to the supplementary data accumulated by the enterprise operation. Customers who have been involved in litigation with the unit or are on the unit's unqualified list need to be classified as cautious customers. The revised data sample table is shown in Table 9 below:

[0096] Table 9 Revised comprehensive rating table with reference to supplementary data In the present invention, the method combines internal historical transactions, external customer big data, and manually recorded data accumulated from transaction experience to effectively evaluate customer credit risk and provide information support for enterprises to optimize customer cooperation and prevent customer risks.

[0097] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A customer credit risk assessment method that integrates machine learning models and expert knowledge and experience, characterized by: The specific steps include: S1. Data integration; S1.

1. Use ETL tools to extract historical customer transaction data from the NC database on a monthly basis. By collaborating with other platforms, obtain external data and update it at a regular interval. S1.

2. Manually re-enter client-related litigation and disqualification list data on a monthly basis. This data warehouse, based on internal and external client data, integrates historical transaction data, external data, and re-entered data to form a data warehouse. This data warehouse is a subject-oriented, integrated, relatively stable data set reflecting historical changes, used to support risk analysis. S1.

3. Write SQL scripts to retrieve the required multi-year accounting data from the data warehouse for internal historical transaction data modeling, and retrieve publicly available corporate information from the data warehouse for external data modeling. S2. Construct risk rating indicators for data of different dimensions; S2.

1. For historical transaction data, set four indicators: accounts receivable turnover rate in the previous year and two years, and accounts receivable collection rate in the previous year and two years. S2.

2. For external data, establish indicators for registered capital, paid-in capital, unpaid paid-in capital, the percentage of enterprises with abnormal operations, operating status, and financial capabilities; whether there is abnormal operation information, whether there is serious illegality or breach of trust, whether there is a breach of trust enforcement, the ratio of total guarantee amount to registered capital, whether there is an abnormal taxpayer, and the ratio of the cumulative amount involved in the case to registered capital; S2.

3. For supplementary data, set indicators for whether the data is in litigation with the unit and whether the data is on the unit's unqualified list; S3. Calculate the actual value of the indicator and assign values ​​to the indicator according to the quartile rule; S3.

1. Calculate the accounts receivable turnover rate and accounts receivable collection rate indicators based on historical transaction data. Use the quantile function in Pandas to calculate the quartiles of these indicators. Use these quartiles to divide the indicators into three risk intervals: normal, alert, and cautious. When the actual indicator values ​​fall into these three risk intervals, assign them values ​​of 0, 1, and 2, respectively. S3.

2. For external data, directly sort registered capital and paid-in capital and assign quartiles, using the same method as above to assign risk scores of 0, 1, and 2. For indicators such as the percentage of unpaid paid-in capital, the proportion of abnormal operations of outbound investment enterprises, the ratio of total guarantees to registered capital, and the cumulative amount involved in the case to registered capital, we calculate, sort, and assign quartiles, assigning risk scores of 0, 1, and 2 using the same method. For general indicators such as whether there are administrative penalties, environmental penalties, and tax arrears, assign values ​​of 1 and 0, depending on whether the relevant indicator is triggered by the client. For key risk indicators such as whether there is serious illegality or dishonesty, whether dishonesty is subject to enforcement, and whether there is major tax violation, assign values ​​of 2 and 0, depending on whether the client triggers the trigger. S4. Build a learning model and generate indicator weights; S4.

1. Based on historical transaction data, customers with accounts receivable outstanding for two consecutive years are designated as ineligible. Customers whose sum of the risk scores for the four indicators, accounts receivable turnover rate and accounts receivable collection rate, in the previous year and two years is less than or equal to 2 are designated as qualified. Customers on the ineligible and qualified lists are assigned a value of 1, representing cautious and normal customers, respectively. Indicators and their risk scores designed from the historical transaction dimension are incorporated into the machine learning model, and industry is assigned as a control variable to the model learning dataset. Random forest modeling is performed on the learning data, and the optimal model and indicator weights are obtained through parameter adjustment. S4.

2. For external data, equal weights are generally set. However, for the indicator "the ratio of the cumulative amount involved in the case to registered capital," its indicator weight is set to twice that of the other indicators, so that the sum of the indicator weights for all external data equals 1. Risk assessment indicators are constructed using the third-party Python tools pandas and numpy. For absolute value indicators, the indicators are divided into normal, alert, and cautious intervals according to quartiles. Then, based on where the indicator values ​​fall into different intervals, the indicators are assigned normal, alert, and cautious levels, with values ​​of 0, 1, and 2, respectively. For general "yes" and "no" indicators, values ​​of 1 and 0 are assigned based on whether the indicator is triggered or not, respectively. For important and "no" indicators, values ​​of 2 and 0 are assigned based on whether the indicator is triggered or not, respectively. A preliminary external data risk score is derived by multiplying the risk score of each customer's indicators by the indicator weight. A score threshold is set based on the distribution and clustering of the score data to divide the customer's external data risk level intervals, thereby deriving the customer's external data risk level, with normal, alert, and cautious values ​​of 0, 1, and 2, respectively. S5. Calculate the risk scores of the client's internal and external data and assign ratings to the internal and external dimensions according to quartiles; The internal and external data risk scores are calculated using the indicator weight × indicator risk score method. The internal and external data risk scores are sorted to obtain quartiles, and the three risk intervals of normal, alert, and cautious are divided according to the quartiles. When the risk score falls into the risk interval, the risk level corresponding to the risk score is obtained, thereby obtaining the risk level of the internal and external dimensions; S6. Integrate internal and external data risk levels to obtain a comprehensive customer rating; The internal and external data risk levels are placed on the horizontal and vertical axes of the first quadrant, respectively, to form a nine-square heat map of customer ratings. When the internal and external data risk levels are combined into (normal, normal), (normal, vigilant), and (vigilant, normal), the customer's overall rating is normal. When the internal and external data risk levels are combined into (vigilant, vigilant), (normal, cautious), and (cautious, normal), the customer's overall rating is cautious. When the internal and external data risk levels are combined into (cautious, cautious), (vigilant, cautious), and (cautious, vigilant), the customer's overall rating is cautious. S7. Set strong rules for supplementary data and identify cautious merchants; On the basis of S6, the customer risk level will be revised according to whether the customer is in litigation with the unit or is on the unit's unqualified list. When the customer is in litigation with the unit or is on the unit's unqualified list, some customers will be revised from normal or alert according to S6 to cautious, and the comprehensive rating will be revised with reference to the supplementary data.

2. The customer credit risk assessment method integrating machine learning models and expert knowledge and experience according to claim 1 is characterized in that: The other platforms in step S1.1 are units that focus on collecting social entity information, including Tianyancha, Qichacha, and Qixinbao.

3. The customer credit risk assessment method integrating machine learning models and expert knowledge and experience according to claim 1 is characterized in that: The enterprise public information in step S1.3 is the enterprise's business registration information, financial information, credit information, business information, tax information, and judicial litigation information.

4. The customer credit risk assessment method integrating machine learning models and expert knowledge and experience according to claim 1 is characterized in that: The specific steps of step S4.1 are: S4.1.

1. Reasoning about the unqualified list and qualified list Customers whose accounts receivable are empty but whose prepayments or operating income are greater than 0 are defined as normal customers. For the remaining customers, the accounts receivable turnover rate and accounts receivable collection rate are calculated for each customer over four years. Based on the list of companies or customers in the previous year, the following four indicators are matched to obtain data, including the accounts receivable turnover rate in the previous year, the accounts receivable collection rate in the previous year, the accounts receivable turnover rate in the previous two years, and the accounts receivable collection rate in the previous two years; Based on the customer group indicator data, the quartiles Q1 and Q3 of each indicator are calculated. The three intervals are divided into [0, Q1], (Q1, Q3), and [Q3, +∞). The indicator values ​​in these three intervals are assigned three levels: cautious, vigilant, and normal, and are assigned values ​​of 2, 1, and 0. To infer the qualified list, enterprises or customer samples with a total indicator risk score of less than or equal to 2 are designated as qualified list samples; to infer the unqualified list, enterprises or customer samples with accounts receivable outstanding for two years are designated as unqualified list samples. Two years of accounts receivable outstanding means that the beginning balance of the enterprise's accounts receivable from customers is equal to the ending balance > 0 for two consecutive years, and the accumulated debits are equal to the accumulated credits = 0; S4.1.

2. Constructing model learning data Calculate and match the internal indicator data of the previous three years and the previous four years on the unqualified list and the qualified list to obtain four indicator risk scores based on the previous three years: the previous year's accounts receivable turnover risk score, the previous year's accounts receivable collection rate risk score, the previous two years' accounts receivable turnover risk score, and the previous two years' accounts receivable collection rate risk score; Industry control variables were added, and the risk score data of the four indicators, the unqualified list and the qualified list, based on the previous three years were embedded in the form of one-hot encoding. The data were then labeled as unqualified list and qualified list, with the unqualified list assigned a value of 1 and the qualified list assigned a value of 0, to obtain the final model learning data. S4.1.

3. Build and train the model to obtain internal data indicator weights Use the Python third-party library sklearn to split the training set and test set, build a random forest model, and use grid search to find the best model. Use the feature_importance method of the random forest model of the Python third-party library sklearn to find the indicator weight vector ω of the best model. S4.1.

4. Calculate the overall risk of internal data Using the previous year as a benchmark, we collected data on four indicators for the company or customer: accounts receivable turnover rate in the previous year, accounts receivable collection rate in the previous year, accounts receivable turnover rate in the previous two years, and accounts receivable collection rate in the previous two years. We then divided the risk of each indicator by quartiles, generating risk scores for the four indicators, including a failed list and a qualified list, based on the previous year. Referring to the method mentioned in S4.1.2, industry is added as a control variable to the risk score data of the four indicators of the unqualified list and the qualified list based on the time point of the previous year as the basic data for risk rating; Extracting the basic data into the indicator data matrix A, the risk score can be obtained as: S = A × ω T , use the python third-party library numpy to calculate matrix multiplication, calculate the quartiles Q1 and Q3 according to the obtained risk score, and divide it into three intervals [0,Q1], (Q1,Q3), and [Q3,+∞). The risk values ​​in these three intervals are assigned to the internal historical transaction data as normal, vigilant, and cautious, and assigned with 0, 1, and 2 to obtain the internal historical transaction data rating result data.

5. The customer credit risk assessment method integrating machine learning models and expert knowledge and experience according to claim 1 is characterized in that: The risk assessment indicators in step S4.2 are the registered capital, paid-in capital, whether the business is operating abnormally, whether the business is breaking the law and breaching trust, and whether the business is subject to execution for breach of trust.

6. The customer credit risk assessment method integrating machine learning models and expert knowledge and experience according to claim 1 is characterized in that: The horizontal axis of the nine-square heat map in step S6 is the internal data rating result, which is normal, vigilant, and cautious from left to right; the vertical axis of the nine-square heat map is the external data rating result, which is normal, vigilant, and cautious from bottom to top.

Citation Information

Patent Citations

  • Small merchant credit evaluation method of fusion of expert model and machine learning model

    CN107644375A

  • Credit evaluation method, apparatus and device, and computer readable storage medium

    WO2019080407A1