Small and medium-sized enterprise risk identification method and device based on data amplification
By adopting data amplification technology and integrated learning methods in credit risk assessment, the problem of small amount of data and imbalance is solved, the accuracy and interpretability of the assessment are improved, and more reliable decision support is provided.
Patent Information
- Application Number
- CN202510534422.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively deal with the problem of small amount of data and imbalance in credit risk assessment, resulting in poor effectiveness of machine learning algorithms in practical applications.
Using a data amplification method, it is integrated into CTGAN through K-means clustering to generate synthetic data, and fusion of multiple individual learners is used to improve the accuracy and interpretability of credit risk assessment.
It significantly improves the prediction accuracy of credit risk assessment and the interpretability of the model, can be more effectively applied to actual business scenarios, and provides more reliable decision support.
Smart Images

Figure CN120069560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data augmentation and risk identification, and particularly to a method and device for identifying risks of small and medium-sized enterprises based on data augmentation. Background Art
[0002] Credit risk assessment is a complex issue that covers a wide range of research contents. It mainly includes: collection and processing of data, identification and quantification of risk factors, establishment of models, model verification and performance evaluation, etc. Current research focuses on the establishment of models and their accuracy, while less research is conducted on how to use credit risk assessment models to support risk management decisions.
[0003] Early risk assessment methods often relied on the subjective judgment of experts and were more inclined to qualitative analysis. Later, with the development of various statistical methods, more and more scholars began to conduct quantitative analysis of risk assessment. In practical applications, for a specific and determined problem or dataset, only by determining a targeted model can the best results be achieved. A single model often cannot effectively solve all credit problems, so the ensemble learning model came into being.
[0004] In recent years, with the development of machine learning, various machine learning algorithms have made significant progress in many fields of data analysis, including data classification and prediction. However, in practical applications, the obtained datasets have hindered the effective application of most algorithms, among which the small amount of data and data imbalance are the main influencing factors, and these two problems often occur in the research of enterprise credit risk assessment. Due to the consideration of factors such as data security by enterprises, obstacles are often encountered when collecting real data, resulting in a small amount of final sample data. Therefore, how to handle the small sample size required for machine learning and the data imbalance problem is an important content of credit assessment. Most existing sampling methods do not consider the data distribution problem when dealing with data imbalance, and often lose the potential information of the original data. Summary of the Invention
[0005] The purpose of the present invention is to propose a method and device for identifying risks of small and medium-sized enterprises based on data augmentation in view of the deficiencies of the prior art.
[0006] The purpose of the present invention is achieved through the following technical solutions: In the first aspect, the present invention provides a method for identifying risks of small and medium-sized enterprises based on data augmentation, and the method includes the following steps:
[0007] Step 1, obtain enterprise credit-related data and construct a dataset;
[0008] Step 2, integrate K-means clustering into CTGAN to obtain a data augmentation model, and use the data augmentation model to perform data augmentation processing on the dataset to obtain synthetic data;
[0009] Step 3: Based on the ensemble learning method, multiple individual learners are fused to obtain a fusion model, and the fusion model is used to perform risk assessment and prediction on small and medium-sized enterprises based on the synthetic data after data augmentation.
[0010] Step 4: Display the differential factors affecting enterprise risks through visualization means to facilitate decision-makers to make manual analysis.
[0011] Furthermore, in order to integrate K-means clustering into CTGAN, the following steps need to be taken:
[0012] (1) Use the K-means algorithm to cluster the original data D and divide the data into K clusters.
[0013] (2) For each cluster, calculate its statistical characteristics.
[0014] (3) In the generator network, associate the real data X with the cluster labels of K-means so that the clustering information is considered when generating synthetic data.
[0015] (4) Modify the loss function of the generator network so that the similarity with the K-means clusters is also considered when generating data.
[0016] Furthermore, in Step 2, find the optimal parameters of the generator network and the parameters of the discriminator network , so that the generated data is as similar as possible to the distribution of the original data under the given conditional real data X. The loss function is specifically as follows:
[0017]
[0018] where D(X, D) represents the output of the discriminator network D for the real data X, and D(X, G(Z)) represents the output for the generated data G(Z); the goal of the loss function is to maximize the ability of the discriminator network to distinguish real data and generated data, and minimize the ability of the generator network to deceive the discriminator network.
[0019] Furthermore, in Step 2, undersample the majority class samples based on OCSVM, oversample using the data augmentation model, and finally, on the basis of the obtained new dataset, use the data augmentation model again to expand the data.
[0020] Furthermore, in Step 3, calculate the feature importance based on the random forest algorithm, screen the features according to the contribution degree, and construct a risk index system.
[0021] Further, in step 3, the individual learners include Naive Bayes, Random Forest, and AdaBoost, and the model fusion method uses the Stacking algorithm.
[0022] Further, in step 4, interpretability analysis is performed based on the SHAP model. The SHAP values are calculated to represent the contribution of each data feature to the fusion model predicting the positive class and the negative class, and a visualization tool is used to present the SHAP values and the interpretation results.
[0023] In a second aspect, the present invention also provides a small and medium-sized enterprise risk identification device based on data augmentation, including a memory and one or more processors. Executable code is stored in the memory, and when the processor executes the executable code, the described small and medium-sized enterprise risk identification method based on data augmentation is implemented.
[0024] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the described small and medium-sized enterprise risk identification method based on data augmentation is implemented.
[0025] In a fourth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the described small and medium-sized enterprise risk identification method based on data augmentation is implemented.
[0026] Advantages of the present invention:
[0027] 1. Aiming at the key issues of credit risk assessment for cross-border e-commerce small and medium-sized enterprises, the present invention constructs a credit risk assessment model for small and medium-sized enterprises based on ensemble learning. First, on the basis of a real dataset, an improved CTGAN data augmentation model is constructed, and compared with conventional data processing methods, achieving good results. Second, by introducing the ensemble learning method, Naive Bayes, Random Forest, and AdaBoost models are fused into a comprehensive ensemble model to improve the accuracy and reliability of credit risk assessment. Finally, in order to help decision-makers make better decisions, it is necessary to enhance the interpretability of the model. The present invention introduces the SHAP (Shapley Additive exPlanations) framework to perform interpretability analysis of the model's prediction results from both macroscopic and microscopic perspectives. In this way, the decision-making process of the model can be understood more clearly, and more credible decision-making support can be provided for decision-makers.
[0028] 2. The credit risk assessment model based on ensemble learning proposed by the present invention has significant performance advantages in the field of cross-border small and medium-sized enterprises. Compared with traditional methods, its prediction accuracy has been significantly improved. The improved CTGAN data augmentation model has been improved to a certain extent compared with conventional methods. In addition, through the application of the SHAP framework, the interpretability of the model has been successfully improved, making it better applicable to actual business scenarios.
[0029] 3. The present invention provides a new method and perspective for the credit risk assessment of cross-border small and medium-sized enterprises, and has made significant progress in enterprise data augmentation, improving prediction performance and interpretability, providing a more reliable decision-making tool for financial institutions and investors. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 It is a schematic flow chart of a method for identifying risks of small and medium-sized enterprises based on data augmentation provided by the present invention.
[0032] Figure 2 It is a flow chart of the improved K-means CTGAN.
[0033] Figure 3 It is a schematic diagram for analyzing the positive and negative impacts of features.
[0034] Figure 4 It is a schematic diagram of the global importance calculated by SHAP values.
[0035] Figure 5 It is a schematic diagram for SHAP feature interaction analysis, where (a) represents the interaction between the contract period and accounts receivable, and (b) represents the interaction between the transaction cycle and the registered capital.
[0036] Figure 6 It is a structural diagram of a device for identifying risks of small and medium-sized enterprises based on data augmentation according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0038] As Figure 1As shown in the figure, a risk identification method for small and medium-sized enterprises based on data augmentation provided by the present invention specifically includes the following steps:
[0039] Step 1: Obtain real data from cross-border e-commerce small and medium-sized enterprises to construct a data set. Enterprise credit risk assessment belongs to a binary classification problem in machine learning, and the target variable is the risk level (0 or 1), where 0 represents no risk or low risk, and 1 represents high risk.
[0040] Step 2: Perform data augmentation processing on the data set: As Figure 2 shown in the figure, the present invention combines clustering and conditional adversarial networks to improve the quality and diversity of tabular data generation. Design a generative model K-means CTGAN (Conditional Tabular Generative Adversarial Network) for generating structured data, which combines the ideas of K-means clustering and generative adversarial networks (GANs). The model aims to simulate the multidimensional tabular data distribution under given conditions in order to generate synthetic data with similar statistical characteristics.
[0041] The goal of the model is to find the optimal parameters of the generator network and the parameters of the discriminator network , so that the generated data is as similar as possible to the distribution of the original data under the given conditional real data X. This can be achieved through the following loss function
[0042]
[0043] where D(X, D) represents the output of the discriminator network D for the real data X, and D(X, G(Z)) represents the output for the generated data G(Z); the goal of the loss function is to maximize the ability of the discriminator to distinguish real data and generated data, and minimize the ability of the generator to deceive the discriminator.
[0044] To integrate K-means clustering into CTGAN, the following steps need to be taken:
[0045] (1) Use the K-means algorithm to cluster the original data D and divide the data into K clusters.
[0046] (2) For each cluster, calculate its statistical characteristics (such as mean and covariance matrix).
[0047] (3) In the generator network, associate the condition X with the cluster labels of K-means so that clustering information is considered when generating synthetic data.
[0048] (4) Modify the loss function of the generator network so that the similarity to the K-means clusters is also considered when generating data.
[0049] In this way, the generated data can better simulate the distribution characteristics of the original data under specific conditions while retaining the clustering information. The final generated data will be closer to the statistical characteristics of the original data and also consider the structure of the K-means clustering.
[0050] The CTGAN model based on K-means combines K-means clustering and generative adversarial networks to generate structured data that meets specific conditions. The design of this model enables the generated data to better retain the distribution characteristics and clustering structure of the original data.
[0051] To better solve the problems of data imbalance and insufficient data volume, the present invention, based on the K-means CTGAN model, combines OCSVM (One-Class Support Vector Machine) to undersample the majority class samples, uses K-means CTGAN for oversampling, and finally, based on the obtained new dataset, uses K-means CTGAN again for data augmentation. OCSVM is a machine learning model for anomaly detection, usually used to handle class imbalance problems. Its core idea is to identify abnormal samples by constructing a hypersphere that encloses normal samples. The objective function of OCSVM is as follows:
[0052]
[0053] where, is a weight vector used to define the hyperplane that separates normal samples from abnormal samples. is the distance from the hyperplane to the origin, which can be regarded as the tolerance of abnormal samples. is a set of slack variables used to allow some normal samples to fall inside the hypersphere. represents projecting the input sample onto the mapping in the high-dimensional feature space. is a user-defined parameter that controls the degree of tolerance. is a slack variable used to handle the misclassification of samples. For each sample , if , then , representing the distance of the sample to the hypersphere boundary; if , then , representing that the sample is inside the hypersphere.
[0054] The present invention uses a dataset containing a certain amount of real tabular data for training the model. In the data preprocessing stage, operations such as missing value handling, data cleaning, data transformation, and feature selection are performed on the original data to improve the quality and consistency of the data, and the number of samples in each dataset is controlled at 1000. Then, the K-means clustering algorithm is used to cluster the tabular data, grouping similar data together to avoid the impact of overlapping data categories on data generation.
[0055] The model of the conditional adversarial network consists of a generator and a discriminator. The generator is responsible for generating new tabular data, while the discriminator is responsible for judging the difference between the generated data and the real data. Here, the framework of the conditional adversarial network is used, and corresponding objective functions are defined, including the loss functions of the generator and the discriminator and the balance parameter between them. In addition, the numerical interval and the conditions associated with the feature values are set.
[0056] By using an optimization algorithm for training and conducting experiments in a GPU-accelerated environment, the quality and diversity of the generated tabular data are evaluated here. Quantitative evaluations are carried out using metrics such as accuracy, recall, and F1-score. The data generated by the K-means CTGAN model has shown significant improvements in both prediction accuracy and precision, recall, and F1-score. On the basis of solving the data imbalance problem, combined with data generation, the prediction accuracy is improved.
[0057] Step 3: Synthetic data evaluation: When generating data, the distribution information of the data should be considered simultaneously, and the distribution information of the original data should be preserved as much as possible. The indicators for evaluating the quality of the distribution information of the tabular synthetic data adopted in the present invention include the following two aspects, as shown in Table 1:
[0058] Table 1 Indicators for the quality of tabular synthetic data
[0059]
[0060] Specifically, the degree of simulation can be further divided into correlation similarity and statistical similarity. The former can describe whether the correlation between the characteristics of the synthetic data is consistent with that of the real data, measure the correlation between a pair of numerical columns, and calculate the similarity degree of the correlation between the real data and the synthetic data. The latter describes whether the statistical indicators of the synthetic data are consistent with those of the real data, and measures the similarity between the real data and the synthetic data by comparing the summary statistical indicators (including mean, median, and standard deviation). The formula is as follows:
[0061]
[0062] Among them, It represents the application of statistical functions to a set of sample data, including mean function, variance function, standard deviation function, etc., and the continuous variable coverage is used to measure whether the synthetic data covers all the value ranges existing in the real data. The definition is as follows:
[0063] If and s represent the same column in the real and synthetic data respectively, then this metric calculates the closeness of the minimum and maximum values of s to the minimum and maximum values in r according to the following formula.
[0064]
[0065] The discrete variable coverage is used to measure whether a synthetic data column covers all the possible categories that appear in the real data column. The definition is as follows:
[0066] First, calculate the number of categories existing in a certain column r in the real data , and then calculate the number of categories existing in the corresponding column s in the synthetic data . It returns the proportion of real categories in the synthetic data, as shown in the following formula:
[0067]
[0068] The correlations between the characteristics of the synthetic data are highly consistent with those of the real data, showing a high degree of similarity. According to the statistical similarity, the statistical indicators of the synthetic data are highly consistent with those of the real data. And according to the continuous and discrete variable coverages, the generated data can almost cover all the numerical situations of the real data. Thus, it can be seen that the generated data has good effectiveness.
[0069] Step 4: Construct a risk index system. Based on the synthetic data after data augmentation, use the random forest algorithm to calculate the feature importance, then sort the obtained features, and screen out the features with greater contribution degrees. These features contain more information in the sample data and play a key role in the correct evaluation.
[0070] Use Naive Bayes, random forest, and the iterative algorithm AdaBoost as individual learners based on the ensemble learning method, and perform the fusion of learners based on the model fusion Stacking algorithm to conduct risk assessment and prediction on the synthetic data; the ensemble learning method can effectively improve the prediction effect under the enterprise risk assessment data. The reason is that the fusion of multiple individual learners can, to a certain extent, alleviate the problem of local optimal solutions. The Stacking fusion model has good generalization ability in the enterprise risk assessment scenario and can stably evaluate the risk level of the enterprise.
[0071] Step 5: Interpretability Analysis Based on the SHAP Model: Through visualization means, more intuitively display the differential factors affecting the risk levels of different enterprises, facilitating decision-makers to conduct manual analysis.
[0072] SHAP (Shapley Additive exPlanations) is a powerful framework for interpreting the prediction results of machine learning models. It provides an interpretable explanation of model predictions based on the theoretical foundation of game theory. The core idea of SHAP values is that for any given prediction result, we hope to determine the "contribution" of each feature to this result, that is, in the game, if a feature is removed from the model, how the result will change. The calculation formula of SHAP values is as follows:
[0073]
[0074] where, represents the SHAP value of feature , represents the prediction result of the model given the feature set , represents the set of all features, represents the subset that contains feature , , represents the size of the set . It can be seen that the core idea of SHAP values is for each feature , by adding it to different feature sets , calculating the changes in the model's prediction results in different situations, and then performing a weighted sum over all possible feature sets , where the weights are determined according to the size of the feature set and the size of the total feature set.
[0075] An important property of SHAP values is "fairness", which ensures that the SHAP value of each feature is allocated according to its contribution to the result. Specifically, if a feature contributes more to the result, then its SHAP value is higher. Therefore, the prediction results of the model can be interpreted based on SHAP values to understand the relative importance of each feature for this result. Enterprise risk assessment belongs to a binary classification problem. For binary classification problems, SHAP values can represent the contribution of each feature to the model predicting the positive class and the contribution to the model predicting the negative class. And visualization tools can be used to present SHAP values and interpret the results. The SHAP framework provides a variety of visualization methods, including SHAP value plots, summary plots, and force plots, etc., which can better understand the model's prediction results and the impact of features.
[0076] Example: The present invention will be described with the application scenario of small and medium-sized cross-border e-commerce enterprises. Real data from small and medium-sized cross-border e-commerce enterprises is obtained to construct a data set. The data set contains a total of 878 samples, all of which are obtained through on-site visits and research. Features such as company names and business descriptions that are useless for risk assessment are removed. The data set contains 14 features including historical transaction volume, transaction cycle, accounts receivable, total export volume, contract duration, number of transactions, number of overdue times, registered capital, legal litigation, tax credit score, recruitment status, business risk, intellectual property, and risk level. Enterprise credit risk assessment belongs to a binary classification problem in machine learning. The target variable is the risk level (0 or 1), where 0 indicates no risk or low risk, and 1 indicates high risk. Statistical analysis of the target variable in the data set shows that there are 706 enterprises with a risk level of 0, accounting for 80.4% of the total number of samples, and 172 enterprises with a risk level of 1, accounting for 19.6% of the total number of samples.
[0077] After data collection is completed, data preprocessing is required because in the original data, problems such as incomplete data, inconsistent field types, and non-uniform dimensions often occur. Therefore, in order to obtain a clean, consistent, and reliable data set for better evaluation model construction and analysis, the present invention performs the following preprocessing steps:
[0078] (1) Data cleaning: Data cleaning aims to detect and correct errors, inconsistencies, and outliers in the data set. It includes handling missing values, removing duplicates, and handling outliers. The total number of samples is reduced from 878 to 797.
[0079] (2) Data transformation: Standardization is used to transform all features into the same scale. In the enterprise credit risk assessment of the present invention, mean standardization is adopted. Because mean standardization retains the original distribution information of the data, scales the data to a standard normal distribution centered on the mean, and does not change the original distribution shape of the data.
[0080] The data of small and medium-sized cross-border e-commerce enterprises collected is amplified. For machine learning algorithms, the processed 797 data points are relatively few, and further data generation is required. Here, a data amplification model is applied to the preprocessed small and medium-sized enterprise credit data, expanding the sample size to 1200, and increasing the ratio of low-risk to high-risk data from the original 4:1 to 7:3 to reduce the misclassification probability of the evaluation model.
[0081] After calculating the feature importance using the random forest algorithm based on the synthetic data after data augmentation, the obtained features are sorted. The contract term and accounts receivable contribute more to the performance of the evaluation model. These features contain more information in the sample data and play a key role in the correct evaluation. While the business risk and recruitment information contain less sample data information and contribute little to the evaluation model, and can be considered as noise data or irrelevant features.
[0082] The model after Stacking ensemble learning fusion is used to conduct risk assessment prediction on the selected features, and based on the SHAP model, an interpretable feature importance analysis is carried out on 13 features. As Figure 3 shown, for the influence direction and influence strength of different features on the final result, the abscissa is the SHAP value, and each point in the figure represents a sample. The red points indicate that the current feature value of the sample point is larger, and the blue points indicate that the current feature value of the sample point is smaller. When the accounts receivable value is larger, the SHAP value is greater than 0, meaning that the higher the value of this feature, the more significant the positive impact on the evaluation model. Conversely, the higher the accounts receivable value, the greater the probability that the model evaluates as high risk. Most of the intellectual property feature values are concentrated around 0, and the SHAP value is between -2 and 1.5, indicating that intellectual property has little impact on the enterprise's risk level. However, it can be seen that for most points with higher intellectual property values, their SHAP values are less than 0. Therefore, it can be inferred that the higher the number of intellectual properties of an enterprise, the more it tends to be a low-risk enterprise. As Figure 4 shown, the average absolute value of the SHAP value of each feature is taken to obtain the bar chart as shown. Among them, the abscissa is the feature importance, and the ordinate is the name of each feature, and finally they are arranged in descending order according to the size. Among them, accounts receivable has the greatest impact on the result of the evaluation model globally, followed by the contract term and the number of overdue times. In addition, the transaction cycle, the number of transactions, the total export volume, the historical trading volume, the registered capital, intellectual property, and legal litigation also have a greater impact on whether the enterprise risk level is high risk. While the business risk, the current recruitment situation, and the tax credit score have less impact on the actual evaluation result. Combining Figure 3 it can be seen that enterprises with larger accounts receivable, shorter contract terms, more overdue times, less total export volume, and fewer transaction times are more inclined to high risk and thus require more attention.
[0083] The present invention also studies the feature interaction effects of the enterprise credit risk data set. As Figure 5 shown, the abscissa in the figure is the value of the theme feature, and the ordinate is the SHAP value. The redder the color, the larger the value. Conversely, the bluer the color, the smaller the value of this feature. From Figure 5In (a), the interaction between the contract term and accounts receivable can be seen. When the contract term is less than 1 year, the probability of high risk for the enterprise is relatively high. Generally speaking, the longer the contract term, the higher the accounts receivable will be. Figure 5 In (b), the interaction effect between the transaction cycle and the registered capital is shown. The results show that there is a positive correlation between the transaction cycle and the registered capital. That is to say, enterprises with a high registered capital usually have a longer transaction cycle. However, as the transaction cycle increases, the risk level of the enterprise gradually tends to be high risk.
[0084] To further improve the interpretability of the enterprise credit risk assessment model, from the micro level, that is, from the perspective of a single sample, the high-risk probability of the enterprise is analyzed. In the present invention, factors such as accounts receivable, contract term, number of transactions, and number of overdue times have an important decisive impact on whether an enterprise will be evaluated as high risk. Therefore, when evaluating an enterprise loan, these aspects should be focused on.
[0085] Corresponding to the foregoing embodiment of a method for identifying risks of small and medium-sized enterprises based on data augmentation, the present invention also provides an embodiment of a device for identifying risks of small and medium-sized enterprises based on data augmentation.
[0086] See Figure 6 , an embodiment of a device for identifying risks of small and medium-sized enterprises based on data augmentation provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement a method for identifying risks of small and medium-sized enterprises based on data augmentation in the foregoing embodiment.
[0087] An embodiment of a device for identifying risks of small and medium-sized enterprises based on data augmentation provided by the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where a device for identifying risks of small and medium-sized enterprises based on data augmentation provided by the present invention is located. Except for Figure 6 the processor, memory, network interface, and non-volatile memory shown, generally, according to the actual functions of the any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0088] For the implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.
[0089] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0090] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for identifying risks of small and medium-sized enterprises based on data augmentation in the above embodiments.
[0091] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.
[0092] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for identifying risks of small and medium-sized enterprises based on data augmentation described above.
[0093] The above embodiments are used to explain the present invention, rather than limiting the present invention. Any modifications and changes made within the spirit and scope of the protection of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A method for identifying risks of small and medium-sized enterprises based on data augmentation, characterized in that: The method comprises the following steps: Step 1: Obtain enterprise credit-related data and build a data set; Step 2: Integrate K-means clustering into CTGAN to obtain a data augmentation model, and use the data augmentation model to perform data augmentation processing on the data set to obtain synthetic data; Step 3: Based on the ensemble learning method, multiple individual learners are integrated to obtain a fusion model, and the fusion model is used to perform risk assessment and prediction on small and medium-sized enterprises based on the synthetic data after data amplification; Step 4: Use visualization to display the differential factors that affect corporate risks, making it easier for decision makers to make manual analysis.
2. According to the method for identifying risks of small and medium-sized enterprises based on data augmentation in claim 1, it is characterized in that: In step 2, in order to integrate K-means clustering into CTGAN, the following steps need to be taken: (1) Use the K-means algorithm to cluster the original data D and divide the data into K clusters; (2) For each cluster, calculate its statistical characteristics; (3) In the generator network, the real data X is associated with the cluster labels of K-means so that the clustering information is considered when generating synthetic data; (4) Modify the loss function of the generator network so that the similarity with the K-means clusters is also considered when generating data.
3. According to the method for identifying risks of small and medium-sized enterprises based on data amplification in claim 1, it is characterized in that: In step 2, find the optimal parameters of the generator network and the parameters of the discriminator network , so that the distribution of the generated data is as similar as possible to the distribution of the original data under the given condition real data X. The loss function is as follows: ; Where D(X, D) represents the output of the discriminator network D for the real data X, and D(X, G(Z)) represents the output for the generated data G(Z); the goal of the loss function is to maximize the ability of the discriminator network to distinguish between real data and generated data, and minimize the ability of the generator network to deceive the discriminator network.
4. The method for identifying risks of small and medium-sized enterprises based on data augmentation according to claim 1 is characterized in that: In step 2, the majority class samples are undersampled based on OCSVM, and oversampled using the data augmentation model. Finally, based on the new data set obtained, the data augmentation model is used again to expand the data.
5. The method for identifying risks of small and medium-sized enterprises based on data augmentation according to claim 1 is characterized in that: In step 3, the feature importance is calculated based on the random forest algorithm, the features are screened according to their contribution, and a risk indicator system is constructed.
6. The method for identifying risks of small and medium-sized enterprises based on data augmentation according to claim 1 is characterized in that: In step 3, the individual learners include Naive Bayes, Random Forest and AdaBoost, and the model fusion method uses the Stacking algorithm.
7. The method for identifying risks of small and medium-sized enterprises based on data augmentation according to claim 1 is characterized in that: In step 4, an interpretability analysis is performed based on the SHAP model. The SHAP value is calculated to represent the contribution of each data feature to the fusion model prediction of the positive category and the contribution to the fusion model prediction of the negative category. A visualization tool is used to present the SHAP value and the explanation results.
8. A device for identifying risks of small and medium-sized enterprises based on data amplification, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a method for identifying risks of small and medium-sized enterprises based on data amplification as described in any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a method for identifying risks of small and medium-sized enterprises based on data amplification as described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, a method for identifying risks of small and medium-sized enterprises based on data amplification as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Medical abnormity violation big data risk early warning method based on unsupervised machine learning and integrated learning
CN117764741A
Method for generating repeated violation person user portraits on unbalanced data based on DBSCAN-cGAN-XGBoost model
CN118211087A
Credit risk assessment system for small and medium-sized enterprises in production and manufacturing industry
CN119250962A