Enterprise risk assessment model training method and system based on government affair data
By using the KMeansSMOTE algorithm to expand default samples and FocalLoss corrected BPNN algorithm to build models in the credit risk rating of small and micro enterprises, data exceptions and sample imbalance are solved and the accuracy of ratings is improved.
Patent Information
- Application Number
- CN202411542410.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology has data anomalies and sample imbalance in the credit risk rating of small and micro enterprises, which leads to the model paying attention to imbalance during training, affecting the accuracy of the rating.
The enterprise risk assessment model training method based on government data is adopted, and the default samples are expanded through the KMeansSMOTE algorithm, and the enterprise risk assessment model is constructed using FocalLoss correction BPNN algorithm.
It alleviates the imbalance in model concern caused by abnormal sample data or sample imbalance during training, and improves the evaluation effect of enterprise risk assessment models.
Smart Images

Figure CN120031648A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for training an enterprise risk assessment model based on government affairs data. Background Art
[0002] As the cornerstone of national economic and social development, the healthy and stable development of small and micro enterprises is crucial to the prosperity of the overall economic environment. However, due to the relatively low transparency of small and micro enterprises in financial management and the lack of robustness in the operation process, how to accurately assess their credit risk level has become a focus of attention. At present, most institutions use the discriminant analysis method based on mathematical statistics for the credit risk rating of small and micro enterprises. Although this method has certain scientificity in theory, it has exposed some problems in actual operation. In particular, when there are outliers in the data set or the data deviates from the normal distribution, this method may produce misleading classification results, thereby affecting the accuracy of the rating. At the same time, due to the large number and wide distribution of small and micro enterprises, their data samples often show a high degree of imbalance. This imbalance causes the model to over-focus on the majority class samples during the training process, while ignoring the important characteristics of the minority class, especially the high-risk class samples, so that the model completed by the final training deviates from the actual situation when assessing risks. Summary of the invention
[0003] In view of the deficiencies in the prior art, the present invention provides a method for training an enterprise risk assessment model based on government affairs data, comprising the following steps: S1, collect basic data of sample enterprises and establish an initial credit evaluation index set, which includes but is not limited to basic enterprise index data, operating status index data, tax status index data, financial status index data, innovation capability index data and credit status index data; S2, screening and adjusting the initial credit evaluation indicator set according to the feature missing values, feature data type distribution and positive and negative sample ratio in the initial credit evaluation indicator set, and forming a test sample after marking the default features, and dividing the test sample data into a first training set, a test set and a validation set in proportion to form a risk assessment test set; S3, using the KMeansSMOTE algorithm to expand the number of default samples and perform equalization processing on the first training set in the risk assessment test set while retaining the data characteristics of the original default samples, to obtain a second training set; S4, use FocalLoss to modify the cross entropy loss function of the BPNN algorithm to build an enterprise risk assessment model and use the second training set for training, use the validation set to adjust the hyperparameters of the enterprise risk assessment model, and obtain the final enterprise risk assessment model through the test set.
[0004] Preferably, the basic indicator data of the enterprise include but are not limited to the industry category of the enterprise, the number of years since the establishment of the enterprise, the number of changes in the legal representative in the past two years, the number of changes in shareholders in the past two years, the number of shareholders, the shareholding ratio of legal person shareholders, the registered capital, and the number of employees; the operating status indicator data include but are not limited to the sales revenue in the past 12 months, the number of months in which the taxable sales revenue in the past 12 months is 0 or missing, and the maximum number of months in which the taxable sales revenue in the past 12 months is 0 or missing; the tax status indicator data include but are not limited to the month-on-month increase in the tax payable in the past three months, the year-on-year increase in the actual tax paid in the past 12 months, the actual tax paid in the past 12 months, the month-on-month increase in the actual tax paid in the past three months, and the maximum number of months in which the tax payable in the past 12 months is 0 or missing. The VAT payment for the following months is the number of the month with zero as the monthly VAT payment amount, the year-on-year VAT actual payment amount for the past 6 months, the taxpayer credit rating and taxpayer status; the financial status indicator data include but are not limited to debt repayment ability information, profitability information, and growth ability information; the innovation ability indicator data include but are not limited to enterprise qualification information and the number of intellectual property rights obtained; the credit status indicator data include but are not limited to whether the enterprise is currently included in the abnormal operation list, whether the enterprise is currently included in the serious illegal and dishonest list, the longest overdue period, the number of fines in the past 12 months, the basis for calculating the late payment fee in the past 24 months, the number of illegal and regulatory violations in the past 12 months and whether the enterprise has been hit by the dishonest debtor in the past 3 years.
[0005] Preferably, step S2 includes: eliminating variable features with a single missing rate greater than 50%; performing discretization grouping processing on the variable features, and after grouping, calculating the weight of evidence WOE for the i-th group: , is the proportion of responding customers in the group, is the proportion of non-responding customers in the group, is the amount of responding customer data in this group, is the number of unresponsive customers in this group, is the total amount of data of responding customers in this group, is the total number of unresponsive customers in the group, where responsive customers refer to positive samples and unresponsive customers refer to negative samples; Calculate the information value IV of each group of variable characteristics: ; According to the IV value of the variable in each group, the IV value of the entire variable is: ; Eliminate variables with no predictive power whose information value is lower than the set value.
[0006] Preferably, the step S3 includes: dividing the first training set into a majority-class sample data set Dmax and a minority-class sample data set Dmin, randomly selecting minority-class samples from the minority-class sample data set Dmin, and calculating the distances from it to all samples in other minority-class sample data sets through Euclidean distance to obtain its k nearest neighbor samples; Set the sampling rate according to the imbalance ratio of the samples to determine the sampling multiplier N. For the minority-class sample a, randomly select multiple samples from its k nearest neighbor samples, and assume that the selected neighbor samples are b; for each randomly selected neighbor sample b, construct a new sample with the original sample a according to the following formula: , and supplement it to the minority-class sample data set; For each majority-class sample Xmax in the majority-class data set Dmax, find its nearest minority-class sample Xmin, and for each minority-class sample Xmin in the minority-class sample data set Dmin, find its nearest majority-class sample Xmax, and compare their nearest distances d(Xmin, Xmax); Judge whether there is a sample y in the risk assessment test set such that d(Xmin, y) < d(Xmin, Xmax) or d(Xmax, y) < d(Xmin, Xmax); if there is no such sample y, then the samples Xmin and Xmax are called Tomek Links pairs; After deleting the majority class in the Tomek Links pairs, the second training set is obtained.
[0007] Preferably, the enterprise risk assessment model includes a three-layer neural network of an input layer, a hidden layer, and an output layer. Among them, the number of neural network layers is set to 10, the number of neurons in the output layer is the number m of input indicators, the number of neurons in the output layer is the number n of classification categories, and the number of neurons in the hidden layer is S: ; where the value range of a is [1, 10], the activation function used in the hidden layer is Relu, the activation function of the output layer is softmax, the dropout setting range is [0, 2], the number of training times epoch of all training data takes a value range of 100, and the batchsize of the training data division is 50; The preset network loss function of the enterprise risk assessment model uses FocalLoss, and its calculation method is as follows: ; Among them, γ is a hyperparameter used to adjust the weight of the loss of misjudged default samples in the target loss, is the real result, is the result predicted by the model.
[0008] Preferably, step S4 comprises: using the second training set to train the enterprise risk assessment model on the FocalLoss modified BP neural network, substituting the test set into the model to obtain the prediction result of the test data set; evaluating the model according to the actual situation and the prediction result, the indicators of model evaluation adopt accuracy, first error rate and second error rate, assuming that TP is the number of non-default samples correctly judged as non-default, FN is the number of non-default samples misjudged as default samples, TN is the number of default samples correctly judged as default samples, FP is the number of default samples judged as non-default samples, and the calculation formulas of each indicator are as follows: .
[0009] Preferably, it also includes: based on the results of enterprise risk assessment model training, formulating scoring card rules to form a software development kit and deploying it to the e-government cloud system, and connecting to the subject database data source; the e-government cloud system is configured to use the enterprise's name and social unified credit code as business system request parameters, parse the received request instructions and call the interface in the SDK according to actual conditions, and return the risk assessment results of the called enterprise.
[0010] The present invention also discloses an enterprise risk assessment model training system based on government data, including an indicator establishment module, a sample generation module, a sample processing module and a model training module, wherein the indicator establishment module is used to collect basic data of sample enterprises and establish an initial credit evaluation indicator set, wherein the credit evaluation indicator set includes but is not limited to basic enterprise indicator data, operating status indicator data, tax status indicator data, financial status indicator data, innovation capability indicator data and credit status indicator data; the sample generation module is used to screen and adjust the initial credit evaluation indicator set according to the feature missing values, feature data type distribution and positive and negative sample ratios in the initial credit evaluation indicator set, and to select the default feature data. The features are annotated to form a test sample, and the test sample data is divided into a first training set, a test set and a validation set in proportion to form a risk assessment test set; a sample processing module is used to use the KMeansSMOTE algorithm to expand the number of default samples and perform equalization processing on the first training set in the risk assessment test set while retaining the characteristics of the original default sample data, so as to obtain a second training set; a model training module is used to use the cross entropy loss function of the BPNN algorithm modified by FocalLoss to construct an enterprise risk assessment model and use the second training set for training, use the validation set to adjust the hyperparameters of the enterprise risk assessment model, and obtain the final enterprise risk assessment model through the test set.
[0011] The present invention also discloses an enterprise risk assessment model training device based on government data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the aforementioned methods when executing the computer program.
[0012] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the aforementioned methods are implemented.
[0013] The present invention discloses a method and system for training an enterprise risk assessment model based on government data. After establishing an initial credit evaluation index according to collected sample data, the data features in the index are screened and adjusted, and the default features are annotated to form a test sample, and a test sample set including a training set, a test set and a validation set is generated. The KMeansSMOTE algorithm is used to expand the number of default samples in the test sample while retaining the original default sample data features, and the sample data is balanced. Finally, the cross entropy loss function of the BPNN algorithm is modified by FocalLoss to construct an enterprise risk assessment model and train, verify and test it to obtain the final enterprise risk assessment model, alleviate the model attention imbalance caused by sample data anomalies or sample imbalance in the training process, and improve the evaluation effect of the enterprise risk assessment model.
[0014] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A schematic diagram of the steps of a method for training an enterprise risk assessment model based on government data disclosed in one embodiment of the present invention.
[0016] Figure 2 This is a schematic diagram of the structure of an enterprise risk assessment model training system based on government data disclosed in another embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0019] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.
[0020] Unless otherwise defined, the technical or scientific terms used herein shall have the common meanings understood by persons with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the patent application specification and claims of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one" or "a" do not indicate a quantity limitation, but indicate the existence of at least one.
[0021] As the cornerstone of national economic and social development, the healthy and stable development of small and micro enterprises is crucial to the prosperity of the overall economic environment. However, due to the relatively low transparency of small and micro enterprises in financial management and the lack of robustness in the operation process, how to accurately assess their credit risk level has become a focus of attention. At present, most institutions use the discriminant analysis method based on mathematical statistics for the credit risk rating of small and micro enterprises. Although this method has certain scientificity in theory, it has exposed some problems in actual operation. In particular, when there are outliers in the data set or the data deviates from the normal distribution, this method may produce misleading classification results, thereby affecting the accuracy of the rating. At the same time, due to the large number and wide distribution of small and micro enterprises, their data samples often show a high degree of imbalance. This imbalance not only increases the complexity of data processing, but also may cause the model to over-focus on the majority class samples during training, while ignoring the important characteristics of the minority class (especially the high-risk class) samples.
[0022] Therefore, in this embodiment, a method for training an enterprise risk assessment model based on government data is disclosed, as shown in the attached Figure 1 As shown, the following steps are included.
[0023] Step S1, collect basic data of sample enterprises and establish an initial credit evaluation index set, which includes but is not limited to basic enterprise index data, operating status index data, tax status index data, financial status index data, innovation capability index data and credit status index data.
[0024] Specifically, we first collect various basic data of small and medium-sized enterprises to establish an initial credit evaluation index set to help the model comprehensively evaluate the credit status of enterprises from multiple dimensions. The various basic data collected include but are not limited to industrial and commercial data, tax data, litigation information, water, electricity, gas data, real estate data, provident fund payment information, social security payment information and intellectual property information; the credit evaluation index set established includes basic enterprise index data, operating status index data, tax status index data, financial status index data, innovation capability index data and credit status index data.
[0025] Among them, the basic indicator data of the enterprise include but are not limited to the enterprise industry category, the company's establishment years, the number of changes in the legal representative in the past two years, the number of changes in shareholders in the past two years, the number of shareholders, the proportion of legal person shareholders' shares, registered capital, and the number of employees. It is a comprehensive score designed based on the company's face-to-face information, industrial and commercial registration information, employee social security information and other information, and is a direct manifestation of credit behavior.
[0026] The data of business condition indicators include, but are not limited to, the sales revenue in the recent 12 months, the number of months with zero or missing taxable sales revenue in the recent 12 months, and the maximum number of consecutive months with zero or missing taxable sales revenue in the recent 12 months. It is a comprehensive credit score designed based on multiple transaction behavior indicators of the enterprise, reflecting the business compliance and operating ability of the enterprise.
[0027] The data of tax payment status indicators include, but are not limited to, the month-on-month change of the payable tax amount in the recent 3 months, the year-on-year change of the actual tax paid in the recent 12 months, the actual tax paid in the recent 12 months, the month-on-month change of the actual tax paid in the recent 3 months, the number of months with zero value-added tax payment in the recent 12 months, the year-on-year change of the actual value-added tax paid in the recent 6 months, the tax credit rating, and the taxpayer status. It is a comprehensive credit score designed based on the tax payment situation of the enterprise, reflecting the tax contribution ability and status of the enterprise.
[0028] The data of financial condition indicators include, but are not limited to, solvency information, profitability information, and growth ability information. It is a comprehensive score designed by combining multiple financial indicators of the enterprise with different weights, reflecting the financial credit level of the enterprise and being an important dimension for judging the financial credit ability of the enterprise.
[0029] The data of innovation ability indicators include, but are not limited to, enterprise qualification information and the number of intellectual property rights obtained. It is a comprehensive score designed based on aspects such as enterprise team development, industry development, and intellectual property rights, reflecting the long-term development prediction of the enterprise.
[0030] The data of credit status indicators include, but are not limited to, whether it is currently listed in the business exception list, whether it is currently listed in the list of seriously illegal and dishonest enterprises, the longest overdue duration, the number of fines in the recent 12 months, the tax basis for late fees in the recent 24 months, the number of illegal and irregular records in the recent 12 months, and whether the enterprise has been listed as a person subject to enforcement for dishonesty in the recent 3 years. It is a comprehensive score designed based on the behavior performances of the enterprise such as water and electricity consumption, tax-related abnormal behaviors, and dishonesty enforcement situations, reflecting the enterprise's standardized behavior status.
[0031] Step S2: Screen and adjust the initial credit evaluation index set according to the missing feature values, the distribution of feature data types, and the positive and negative samples in the initial credit evaluation index set, mark the default features to form a test sample, and divide the test sample data into a first training set, a test set, and a validation set according to a ratio to form a risk assessment test set.
[0032] Specifically, default samples usually refer to samples of individuals or enterprises that fail to perform their obligations as agreed in loans, contracts or other financial transactions. These samples usually have some specific characteristics, such as poor historical repayment behavior, poor credit records, etc., and have important application value. Introducing default samples in the test data set can help the model identify groups with higher default risks. In this embodiment, the data in the established evaluation index data set is pre-processed, and redundant and noise features in the data are removed through feature screening and adjustment, and features closely related to credit risk assessment are retained to improve the accuracy of the model, and default feature samples are marked to enhance the model's ability to identify default behavior. At the same time, the data set is divided into a training set, a test set, and a validation set for subsequent cross-validation and model tuning to further improve the stability and reliability of the model.
[0033] In this embodiment, the process of screening and adjusting the features in the initial credit evaluation index set in step S2 specifically includes the following contents.
[0034] Step S21, remove variable features with a single missing rate greater than 50%.
[0035] Specifically, variable features with high missing rates often contain a lot of incomplete information, which may not provide useful prediction signals for the model and may cause instability and deviation in the model learning process. In the data preprocessing stage, removing variable features with a single missing rate greater than 50% can effectively reduce noise and redundancy in the data, helping to improve the overall data quality, model accuracy and generalization ability.
[0036] Step S22, discretize and group the variable features, and after grouping, calculate the weight of evidence WOE for the i-th group: , is the proportion of responding customers in the group, is the proportion of non-responding customers in the group, is the amount of responding customer data in this group, is the number of unresponsive customers in this group, is the total amount of data of responding customers in this group, It is the total data volume of non-responding customers in this group. Responding customers refer to positive samples, and non-responding customers refer to negative samples.
[0037] Specifically, the Information Value (IV) is an indicator used to represent the degree of contribution of a feature to the prediction of the target variable. By calculating the IV value of a feature, the prediction ability of each feature for the target variable can be evaluated, so as to select features that can significantly improve the prediction performance of the model. Since the calculation of the IV value is based on the Weight of Evidence (WOE), in this embodiment, the variable features are first discretized and grouped, and the proportions of responding customers (positive samples) and non-responding customers (negative samples) in the i-th group are calculated to obtain the WOE value of each group. Among them, WOE is the logarithm of the ratio difference of good and bad customers (or two categories in the target variable) under a specific grouping or segmentation, and its magnitude reflects the difference between the ratio of good and bad customers in a specific grouping and the overall ratio. The greater the difference, the greater the absolute value of the WOE value, indicating that the grouping has a stronger ability to distinguish the target variable.
[0038] Step S23, calculate the information value of each group of variable features: ; According to the IV values of the variable in each group, the IV value of the entire variable is obtained as: ; Eliminate variables with no predictive ability whose information value is lower than the set value.
[0039] Specifically, after calculating the WOE value of each group, the IV value of each group variable is calculated based on the WOE values of all groups to obtain the IV value of the entire variable. Since the higher the IV value, the stronger the predictive ability of the corresponding feature and the higher the information contribution degree, eliminating variables with no predictive ability whose IV value is lower than the set value can help screen out features that can significantly improve the prediction performance of the model. In this embodiment, the set value is taken as 0.02. Then, when IV < 0.02, it means that the corresponding feature has no predictive ability and needs to be eliminated; when 0.02 < IV < 0.1, it means that the corresponding feature has weak predictive ability; when 0.1 < IV < 0.2, it means that the corresponding feature has medium predictive ability; when 0.2 < IV, it means that the corresponding feature has strong predictive ability; Eliminate variables with no predictive ability whose information value is lower than the set value.
[0040] Step S3, for the first training set in the risk assessment test set, use the KMeansSMOTE algorithm to expand the number of default samples and perform equalization processing while retaining the characteristics of the original default sample data, and obtain the second training set.
[0041] Specifically, since the amount of default samples in real life is far less than other types of data, the constructed sample data is highly unbalanced, which greatly reduces the classification effect of the model training on the default samples. Therefore, it is necessary to expand the default samples, effectively increase the number of default samples, ensure that they reach a certain degree of balance with the majority class samples in quantity, and perform equalization processing to eliminate the bias caused by data imbalance. This embodiment uses the KMeansSMOTE algorithm to perform the above operations. The algorithm combines K-means clustering and SMOTE algorithm (synthetic minority class oversampling technology), which can effectively expand the minority class sample data and solve the data imbalance problem. Then, step S3 includes the following content.
[0042] Step S31, divide the first training set into a majority class sample data set Dmax and a minority class sample data set Dmin, randomly select a minority class sample from the minority class sample data set Dmin, and calculate its distance to all samples in other minority class sample data sets through Euclidean distance to obtain its k nearest neighbor samples.
[0043] Step S32, the sampling rate is set according to the imbalance ratio of the samples to determine the sampling multiplier N. For the minority class sample a, multiple samples are randomly selected from its k nearest neighbor samples, and it is assumed that the selected neighbor sample is b; for each randomly selected neighbor sample b, a new sample is constructed with the original sample a according to the following formula: , added to the minority class sample dataset.
[0044] Specifically, the SMOTE algorithm is a commonly used oversampling method that balances the data set by generating new minority class samples. Specifically, the SMOTE algorithm randomly selects a sample data from the minority class sample set as the starting point, and calculates the distance between the sample data and the K nearest neighbor samples around it, and then randomly selects a sample from the K nearest neighbors obtained in the previous step to calculate the difference between the selected sample and the original sample, and multiplies it by a random number (between 0 and 1), and finally adds the difference to the original sample to obtain a new synthetic sample. By repeating the above steps many times until the proportion of minority class samples reaches the set value, the expansion of minority class samples can be completed.
[0045] Furthermore, the SMOTE algorithm is usually used in combination with the Tomek Links algorithm to further solve the data imbalance problem by removing the negative example samples of the nearest neighbors. Then step S3 specifically also includes the following content.
[0046] Step S33: For each majority-class sample Xmax in the majority-class dataset Dmax, find its nearest minority-class sample Xmin, and for each minority-class sample Xmin in the minority-class sample dataset Dmin, find its nearest majority-class sample Xmax, and compare their nearest distances d(Xmin, Xmax).
[0047] Step S34: Determine whether there is a sample y in the risk assessment test set such that d(Xmin, y) < d(Xmin, Xmax) or d(Xmax, y) < d(Xmin, Xmax); if there is no such sample y, then the samples Xmin and Xmax are called a Tomek Links pair.
[0048] Step S35: After deleting the majority class in the Tomek Links pair, obtain the second training set.
[0049] Specifically, the Tomek Links algorithm is an undersampling method for dealing with class-imbalanced datasets. Combined with the SMOTE algorithm, it can be used to remove noise or duplicate samples in the dataset. The basic idea is to calculate the distances between samples to find the nearest neighbor relationships between samples. If the nearest neighbor of a sample is a sample of another class, then these two samples form a Tomek pair. For each pair of Tomek pairs, if their classes are different, then one of the samples (usually the majority-class sample) is deleted to reduce the imbalance and noise of the dataset. In this embodiment, if there is a sample y such that d(Xmin, y) < d(Xmin, Xmax), it means that for a majority-class sample Xmax in the majority-class dataset Dmax, there is a sample y closer to its nearest minority-class sample Xmin in the minority-class sample dataset Dmin; if there is a sample y such that d(Xmax, y) < d(Xmin, Xmax), it means that for the minority-class sample Xmin in the minority-class sample Xmin, there is a sample y closer to its nearest majority-class sample Xmax in the majority-class dataset Dmax; if there is no such sample y. It means that at this time Xmin and Xmax are the nearest points to each other and form a Tomek Links pair. After deleting the Tomek Links pair, the second training set after expansion and balancing processing can be obtained.
[0050] Step S4: Use the FocalLoss to correct the cross-entropy loss function of the BPNN algorithm to construct an enterprise risk assessment model and train it using the second training set, adjust the hyperparameters of the enterprise risk assessment model using the validation set, and obtain the final enterprise risk assessment model through the test set.
[0051] Specifically, after the data set is constructed, an enterprise risk assessment model is constructed based on the BPNN algorithm, and the FocalLoss function is used to correct the loss function. The enterprise risk assessment model includes a three-layer neural network of input layer, hidden layer and output layer, where the number of neural network layers is set to 10, the number of neurons in the output layer is the number of input indicators m, the number of neurons in the output layer is the number of classifications n, and the number of neurons in the hidden layer is S: ; The value range of a is [1, 10], the activation function used in the hidden layer is Relu, the activation function of the output layer is softmax, the dropout setting range is [0, 2], the number of epochs for all training data is 100, and the batchsize for training data division is 50.
[0052] The preset network loss function of the enterprise risk assessment model uses FocalLoss, which is calculated as follows: ; Among them, γ is a hyperparameter used to adjust the weight of the target loss accounted for by the misjudged default sample loss, is the actual result, and is the result predicted by the model.
[0053] After the model is built, the data in the second training set is input to train the model. The training process specifically includes the following contents.
[0054] Step S101, using the second training set to train the FocalLoss modified BP neural network to obtain an enterprise risk assessment model, substituting the test set into the model to obtain the prediction result of the test data set.
[0055] Step S102, the model is evaluated according to the actual situation and the prediction results. The indicators of model evaluation are accuracy, first error rate and second error rate. Let TP be the number of non-default samples correctly judged as non-default, FN be the number of non-default samples misjudged as default samples, TN be the number of default samples correctly judged as default samples, FP be the number of default samples judged as non-default samples, and the calculation formulas of each indicator are as follows: .
[0056] Specifically, Accuracy is used to measure the proportion of correct predictions made by the model on all samples; the first error rate (Type1-error) is used to measure the proportion of non-default samples that are mistakenly judged as default; the second error rate (Type2-error) is used to measure the proportion of default samples that are mistakenly judged as non-default. The above indicators together constitute a complete system for model evaluation. In practical applications, appropriate evaluation indicators should be selected according to specific needs and scenarios to comprehensively and accurately evaluate the performance of the model.
[0057] During the specific implementation process, the enterprise or organization formulates scoring card rules based on the results of enterprise risk assessment model training, forms a software development kit and deploys it to the e-government cloud system, and connects to the subject database data source; the e-government cloud system is configured to use the enterprise's name and social unified credit code as business system request parameters, parse the received request instructions and call the interface in the SDK according to actual conditions, and return the risk assessment results of the called enterprise.
[0058] The present embodiment discloses a method for training an enterprise risk assessment model based on government data. After establishing an initial credit evaluation index according to the collected sample data, the data features in the index are screened and adjusted, and the default features are annotated to form a test sample, and a test sample set including a training set, a test set and a validation set is generated. The KMeansSMOTE algorithm is used to expand the number of default samples in the test sample while retaining the original default sample data features, and the sample data is balanced. Finally, the cross entropy loss function of the BPNN algorithm is modified by FocalLoss to construct an enterprise risk assessment model and train, verify and test it to obtain the final enterprise risk assessment model, thereby alleviating the model attention imbalance caused by sample data anomalies or sample imbalance during the training process, and improving the evaluation effect of the enterprise risk assessment model.
[0059] In another embodiment, as shown in the attached Figure 2 As shown, a system for training an enterprise risk assessment model based on government data is also disclosed, including an indicator establishment module 1, a sample generation module 2, a sample processing module 3 and a model training module 4. The indicator establishment module 1 is used to collect basic data of sample enterprises and establish an initial credit evaluation indicator set, which includes but is not limited to basic enterprise indicator data, operating status indicator data, tax status indicator data, financial status indicator data, innovation capability indicator data and credit status indicator data. The sample generation module 2 is used to screen and adjust the initial credit evaluation indicator set according to the feature missing values, feature data type distribution and positive and negative sample ratio in the initial credit evaluation indicator set, and to form a test sample after annotating the default feature, and to form a risk assessment test set after dividing the test sample data into a first training set, a test set and a validation set in proportion. The sample processing module 3 is used to use the KMeansSMOTE algorithm for the first training set in the risk assessment test set to expand the number of default samples and perform equalization processing while retaining the original default sample data characteristics, and obtain the second training set. Model training module 4 is used to construct an enterprise risk assessment model using the cross entropy loss function of the BPNN algorithm modified by FocalLoss and train it using the second training set, adjust the hyperparameters of the enterprise risk assessment model using the validation set, and obtain the final enterprise risk assessment model through the test set.
[0060] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Similar parts between the various embodiments can be referenced to each other. For the enterprise risk assessment model training system based on government data disclosed in the embodiment, since it corresponds to the enterprise risk assessment model training method based on government data disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the aforementioned method part description.
[0061] In other embodiments, there is also provided a device for training an enterprise risk assessment model based on government data, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the various steps of the enterprise risk assessment model training method based on government data as described in the above embodiments when executing the computer program. The server may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that the schematic diagram is merely an example of a server and does not constitute a limitation on the server, and may include more or fewer components than shown in the diagram, or a combination of certain components, or different components.
[0062] If the enterprise risk assessment model training device based on government data is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned enterprise risk assessment model training method based on government data can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.
[0063] In conclusion, the above is only a preferred embodiment of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the patent of the present invention.
Claims
1. A method for training an enterprise risk assessment model based on government data, characterized in that: The steps include: S1, collect basic data of sample enterprises and establish an initial credit evaluation index set, which includes but is not limited to basic enterprise index data, operating status index data, tax status index data, financial status index data, innovation capability index data and credit status index data; S2, screening and adjusting the initial credit evaluation indicator set according to the feature missing values, feature data type distribution and positive and negative sample ratio in the initial credit evaluation indicator set, and forming a test sample after marking the default features, and dividing the test sample data into a first training set, a test set and a validation set in proportion to form a risk assessment test set; S3, using the KMeansSMOTE algorithm to expand the number of default samples and perform equalization processing on the first training set in the risk assessment test set while retaining the data characteristics of the original default samples, to obtain a second training set; S4, use FocalLoss to modify the cross entropy loss function of the BPNN algorithm to build an enterprise risk assessment model and use the second training set for training, use the validation set to adjust the hyperparameters of the enterprise risk assessment model, and obtain the final enterprise risk assessment model through the test set.
2. The enterprise risk assessment model training method based on government data according to claim 1 is characterized by: The basic indicator data of the enterprise include but are not limited to the industry category of the enterprise, the years of establishment of the enterprise, the number of changes in the legal representative in the past two years, the number of changes in shareholders in the past two years, the number of shareholders, the proportion of shares held by legal person shareholders, registered capital, and the number of employees; The business status indicator data include but are not limited to sales revenue in the past 12 months, the number of months in which taxable sales revenue in the past 12 months is 0 or missing, and the maximum number of consecutive months in which taxable sales revenue in the past 12 months is 0 or missing; The tax status indicator data include but are not limited to the month-on-month comparison of tax payable in the past three months, the year-on-year comparison of actual tax paid in the past 12 months, the actual tax paid in the past 12 months, the month-on-month comparison of actual tax paid in the past three months, the number of months with zero VAT payment in the past 12 months, the year-on-year comparison of actual VAT payment in the past six months, tax credit rating and taxpayer status; The financial status indicator data include but are not limited to debt repayment ability information, profitability information, and growth ability information; The innovation capability indicator data include but are not limited to enterprise qualification information and the number of intellectual property rights obtained; The credit status indicator data include but are not limited to whether the enterprise is currently included in the abnormal operation list, whether the enterprise is currently included in the serious illegal and dishonest list, the longest overdue period, the number of fines in the past 12 months, the basis for calculating the tax of late payment fees in the past 24 months, the number of illegal and regulatory records in the past 12 months, and whether the enterprise has been hit by the dishonest debtor in the past three years.
3. The enterprise risk assessment model training method based on government data according to claim 2 is characterized in that: The step S2 comprises: Eliminate variable features with a single missing rate greater than 50%; Discretize and group the variable features. After grouping, calculate the weight of evidence WOE for the i-th group: , is the proportion of responding customers in the group, is the proportion of non-responding customers in the group, is the amount of responding customer data in this group, is the number of unresponsive customers in this group, is the total amount of data of responding customers in this group, is the total number of unresponsive customers in the group, where responsive customers refer to positive samples and unresponsive customers refer to negative samples; Calculate the information value IV of each group of variable characteristics: ; According to the IV value of the variable in each group, the IV value of the entire variable is: ; Eliminate variables with no predictive power whose information value is lower than the set value.
4. The enterprise risk assessment model training method based on government data according to claim 3 is characterized in that: The step S3 comprises: Divide the first training set into a majority-class sample dataset Dmax and a minority-class sample dataset Dmin. Randomly select minority-class samples from the minority-class sample dataset Dmin, and calculate the distances from it to all samples in the other minority-class sample datasets through Euclidean distance to obtain its k nearest neighbor samples; The sampling rate is set according to the imbalance ratio of the sample to determine the sampling multiplier N. For the minority class sample a, multiple samples are randomly selected from its k nearest neighbor samples, and it is assumed that the selected neighbor sample is b; for each randomly selected neighbor sample b, a new sample is constructed with the original sample a according to the following formula: , added to the minority sample data set; For each majority-class sample Xmax in the majority-class dataset Dmax, find its nearest minority-class sample Xmin, and for each minority-class sample Xmin in the minority-class sample dataset Dmin, find its nearest majority-class sample Xmax, and compare their nearest distances d(Xmin, Xmax); Judge whether there is a sample y in the risk assessment test set such that d(Xmin, y) < d(Xmin, Xmax) or d(Xmax, y) < d(Xmin, Xmax); if there is no such sample y, then the samples Xmin and Xmax are called Tomek Links pairs; Delete the majority class in the Tomek Links pairs to obtain the second training set.
5. The method for training an enterprise risk assessment model based on government affairs data according to claim 4, wherein: The enterprise risk assessment model includes a three-layer neural network of input layer, hidden layer and output layer, wherein the number of neural network layers is set to 10, the number of neurons in the output layer is the number of input indicators m, the number of neurons in the output layer is the number of classifications n, and the number of neurons in the hidden layer is S: ; The value range of a is [1, 10], the activation function used in the hidden layer is Relu, the activation function of the output layer is softmax, the dropout setting range is [0, 2], the number of epochs for all training data is 100, and the batch size of the training data is 50; The preset network loss function of the enterprise risk assessment model uses Focal Loss, and its calculation method is as follows: ; Where γ is a hyperparameter used to adjust the weight of the target loss accounted for by the misjudged default sample loss. For real results, The result predicted by the model.
6. The enterprise risk assessment model training method based on government data according to claim 5 is characterized in that: The step S4 includes: Use the second training set to train an enterprise risk assessment model on a BP neural network modified by Focal Loss, substitute the test set into the model to obtain the prediction results of the test dataset; evaluate the model according to the actual situation and the prediction results. The evaluation indicators of the model adopt accuracy Accuracy, the first error rate, and the second error rate. Let TP be the number of non-default samples correctly determined as non-default, FN be the number of non-default samples misjudged as default samples, TN be the number of default samples correctly determined as default samples, and FP be the number of default samples determined as non-default samples. The calculation formulas of each indicator are as follows: 。 7. The enterprise risk assessment model training method based on government data according to claim 6 is characterized in that: It also includes: According to the training results of the enterprise risk assessment model, formulate a scoring card rule to form a software development kit and deploy it to the e-government cloud system to connect to the data source of the theme database; The e-government cloud system is configured to use the name and unified social credit code of the enterprise as business system request parameters, parse the received request instruction and call the interface in the SDK according to the actual situation, and return the risk assessment result of the called enterprise.
8. An enterprise risk assessment model training system based on government data, characterized in that: It includes: An index establishment module, used to collect basic data of sample enterprises and establish an initial credit evaluation index set, and the credit evaluation index set includes but is not limited to enterprise basic index data, business condition index data, tax payment condition index data, financial condition index data, innovation ability index data, and credit condition index data; A sample generation module, used to screen and adjust the initial credit evaluation index set according to the missing feature values, feature data type distribution, and positive and negative sample comparison in the initial credit evaluation index set, and label the default features to form test samples, and divide the test sample data into a first training set, a test set, and a validation set according to a ratio to form a risk assessment test set; A sample processing module, for obtaining a second training set by using a KMeansSMOTE algorithm to expand the number of default samples and perform equalization processing on the first training set in the risk assessment test set while retaining the data characteristics of the original default samples; The model training module is used to construct an enterprise risk assessment model using the cross entropy loss function of the BPNN algorithm modified by FocalLoss and train it using the second training set, adjust the hyperparameters of the enterprise risk assessment model using the validation set, and obtain the final enterprise risk assessment model through the test set.
9. An enterprise risk assessment model training device based on government data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.