A synthetic oversampling method and system based on safe gradient distribution
By introducing gradient contribution to divide gradient intervals and synthesizing pseudo-sample samples into the synthetic oversampling method, the problem of imbalance in pseudo-sample distribution in high-dimensional samples is solved, and the performance and recognition accuracy of the classification model are improved.
Patent Information
- Application Number
- CN202411523296.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-10-30
AI Technical Summary
Existing synthetic oversampling methods are difficult to accurately measure sample spatial distance in high-dimensional samples, resulting in imbalance in pseudo-sample distribution and affecting the recognition accuracy of the model.
By introducing gradient contribution as a classification indicator for a few types of samples, the samples are divided into different gradient intervals, and pseudo-samples are synthesized within the safe gradient interval to avoid relying on high-dimensional spatial distance measurements.
It realizes that the category balance of the data set is improved without relying on spatial characteristics and the performance of the classification model is improved, especially in the high-dimensional sample scenarios of financial risk control data.
Smart Images

Figure CN119046740B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of imbalanced classification of artificial intelligence / data mining technology, and in particular relates to a synthetic oversampling method and system based on secure gradient distribution. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In the field of financial risk control, supervised learning tasks of binary classification usually face the problem of class imbalance, that is, the number of samples in one category in the test data set of financial risk control is often much larger than the number of samples in the other category, where the category with more samples is called the majority class (negative class) and the other category is called the minority class (positive class). Taking default prediction and fraudulent transaction detection in the field of credit risk control in financial risk control as an example, the class imbalance of the credit default prediction data set is reflected in the fact that the number of non-default samples is far greater than the default samples, and the class imbalance of the fraud consumption detection data set is reflected in the fact that the number of normal consumption samples is far greater than the number of fraud consumption samples. It can be seen that the data set categories of these task scenarios in financial risk control are usually seriously unbalanced, and the accurate identification of abnormal samples is of great significance to the scenario tasks.
[0004] Specifically, the characteristic of supervised machine learning is that the more samples there are in a certain category, the more the model can learn the pattern of the samples, that is, the model's ability to recognize each category depends on the contribution of each category's samples to model optimization. However, under the condition of class imbalance, the absolute advantage of the majority class samples in quantity makes their contribution to the model relatively large, resulting in a good fitting result for the majority class data; the opposite is true for minority class samples. Although the overall recognition accuracy of the model will be very high, the recognition performance of minority class samples is very low. This result contradicts the goal of the scene task, and the failure to correctly identify minority class samples may have serious consequences. As a result, the oversampling technology of minority class sample data has become the first research direction to solve the problem of class imbalance. It synthesizes a certain number of minority class pseudo samples to make the number of majority class and minority class samples roughly equal, thereby balancing their contributions to the model.
[0005] At present, there are research results on synthetic oversampling methods to solve the problem of class imbalance in practical applications and academic research, such as SMOTE and its improved methods: NaNSMOTE, ENaNSMOTE, NLDAO, WRND, etc.; synthetic oversampling methods based on clustering: SMOTE-COF, AWTDO, etc.; synthetic oversampling methods based on noise filtering: GMF-SMOTE, etc. It can be seen that although the existing work and research have proposed effective improvement measures for synthetic oversampling methods of imbalanced data sets, the above improvements are all based on feature space and clustering interpolation. This improvement still has the problem of dependence on the spatial distribution of samples. In addition, the spatial distance of high-dimensional samples is difficult to measure, and in many cases the distance measurement of high-dimensional space is invalid; the method based on adaptive interpolation pays too much attention to the synthetic weight of a certain sample; the method based on filtering interpolation will lose the precious minority class samples, which are already few in number, and make the synthetic samples lose diversity.
[0006] Therefore, the inventors found that the existing synthetic oversampling method still has some technical problems, such as:
[0007] (1) Existing methods need to rely on high-dimensional space distance measurement or sample density when selecting root samples, auxiliary samples, and determining sampling weights. Since high-dimensional distance measurement is difficult, this method is not applicable to high-dimensional samples. Financial risk control data are all high-dimensional samples (more than dozens of dimensions), so the quality of samples synthesized by existing methods for financial risk control data sampling is poor.
[0008] (2) The distribution of pseudo samples synthesized in existing methods is not conducive to the balanced recognition of different categories of samples by the model, making it easy for the model to overfit the positive samples and then lose the recognition accuracy of negative samples, which is specifically manifested in low indicators such as F1-Score and MCC. Summary of the invention
[0009] In order to overcome the above-mentioned shortcomings of the prior art, the present invention provides a synthetic oversampling method and system based on safe gradient distribution, which can avoid the error accumulation of noise samples and is independent of spatial features, thereby ensuring the category balance of the data set and improving the performance of the classification model.
[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0011] A first aspect of the present invention provides a synthetic oversampling method based on safe gradient distribution.
[0012] A synthetic oversampling method based on secure gradient distribution, comprising:
[0013] Get the initial sample data set, use the cross entropy gradient of the sample to determine the gradient contribution of the sample, divide the gradient contribution in the range of 0 to 1 into multiple intervals and set a safety gradient threshold, and take the interval with a gradient contribution less than the set safety gradient threshold as the safety gradient interval;
[0014] All minority class samples are assigned to different gradient intervals according to their gradient contributions and safe gradient distribution is calculated; samples within the safe gradient interval are used as root samples, the gradient right neighbors of the root samples are used as auxiliary samples, and the number of sample synthesis is determined based on the safe gradient distribution approximation strategy;
[0015] Linear interpolation method is used to synthesize pseudo samples for each safety gradient interval to achieve sample synthesis oversampling.
[0016] Furthermore, the gradient contribution is obtained by derivation of a binary cross entropy loss function with respect to the sample.
[0017] Furthermore, the interval whose gradient contribution is greater than the safety gradient threshold is taken as the dangerous gradient interval; at the same time, the safe gradient interval closest to the safety gradient threshold is taken as the critical safe gradient interval, and the dangerous gradient interval closest to the safety gradient threshold is taken as the critical dangerous gradient interval.
[0018] Furthermore, the security gradient distribution calculation is performed, specifically: based on the allocation results of different gradient intervals of the samples, the number of positive samples in each security gradient interval is counted, that is, the security gradient distribution of the positive samples is calculated.
[0019] Furthermore, the gradient right neighbor of a sample is a sample whose gradient contribution is greater than that of the sample but whose gradient contribution difference with that of the sample is the smallest.
[0020] Furthermore, before and after the pseudo samples are synthesized, the ratio of the number of samples in each safety gradient interval to the number of samples in all safety gradient intervals remains consistent.
[0021] Furthermore, the synthesis method of synthesizing pseudo samples for each safety gradient interval using the linear interpolation method is as follows:
[0022] ;
[0023] in, represents the synthesized pseudo sample, represents the root sample, Represents a random number between 0 and 1. Represents auxiliary samples.
[0024] A second aspect of the present invention provides a synthetic oversampling system based on a safe gradient distribution.
[0025] A synthetic oversampling system based on a secure gradient distribution, comprising:
[0026] The gradient interval division module is configured to: determine the gradient contribution of the sample by the cross entropy gradient of the sample, divide the gradient contribution in the range of 0 to 1 into multiple intervals and set a safety gradient threshold, and take the interval whose gradient contribution is less than the set safety gradient threshold as the safety gradient interval;
[0027] The gradient interval allocation module is configured to: allocate all minority class samples to different gradient intervals according to gradient contribution and perform safe gradient distribution calculation; use samples in the safe gradient interval as root samples, use the gradient right neighbors of the root samples as auxiliary samples, and determine the number of sample synthesis based on the safe gradient distribution approximation strategy;
[0028] The sample synthesis module is configured to: synthesize pseudo samples for each safety gradient interval using a linear interpolation method to achieve sample synthesis oversampling.
[0029] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in a synthetic oversampling method based on a secure gradient distribution as described in the first aspect of the present invention.
[0030] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a synthetic oversampling method based on a secure gradient distribution as described in the first aspect of the present invention are implemented.
[0031] One or more of the above technical solutions have the following beneficial effects:
[0032] The present invention creatively proposes a new indicator for classifying minority class samples in the synthetic oversampling method - gradient contribution, and proposes different gradient interval concepts. Therefore, the present invention does not need to calculate high-dimensional space distances and does not rely on spatial features, which can avoid the problem of high-dimensional space distance measurement failure, and then ensure the category balance of the data set, which is suitable for high-dimensional samples of financial risk control data.
[0033] The present invention uses samples within the safety gradient interval as root samples, which can effectively avoid the error accumulation caused by noise samples with large gradient contributions to a certain extent, and does not eliminate minority class samples. At the same time, the number of synthesized samples in each gradient interval is calculated based on the safety gradient distribution approximation strategy. Therefore, the present invention can improve the accuracy of minority class samples without sacrificing the accuracy of majority class samples, which is specifically reflected in the fact that indicators such as F1-Score, G-Mean, MCC, and KS are higher than existing methods. The significance of this in the field of credit default is that it can accurately identify defaulting customers without misidentifying non-defaulting users, which is in line with the requirements of modern risk control and wide credit.
[0034] The present invention combines the maximum classification interval principle of the classification model, so the optimized data set can move the decision boundary of the classification model toward the majority class in a soft way, thereby further improving the recognition accuracy of minority class samples.
[0035] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0037] Figure 1 This is a flow chart of a synthetic oversampling method based on security gradient distribution in Embodiment 1 of the present invention.
[0038] Figure 2 Schematic diagram of the gradient interval in the first embodiment of the present invention.
[0039] Figure 3 It is a schematic diagram of the synthesis process of the synthetic oversampling method based on spatial distance in the prior art.
[0040] Figure 4 Schematic diagram of the synthesis process of the synthetic oversampling method based on security gradient distribution in Example 1 of the present invention. DETAILED DESCRIPTION
[0041] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0042] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0043] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0044] The overall idea proposed by the present invention is as follows: The present invention aims at the binary classification task of the unbalanced data set of financial risk control data, and creatively proposes a new synthetic oversampling method based on security gradient distribution. By proposing new minority class sample classification indicators, root sample selection strategies, auxiliary sample selection strategies based on the concept of right nearest neighbors, synthetic sample quantity calculation principles and other concepts and methods, the present synthetic oversampling methods solve the two major problems of reliance on spatial features and error accumulation of noise samples. By synthesizing minority class samples, the category balance of the data set is ensured, and the performance of the classification model is improved.
[0045] Embodiment 1
[0046] This embodiment discloses a synthetic oversampling method based on safe gradient distribution.
[0047] like Figure 1 As shown, a synthetic oversampling method based on secure gradient distribution includes:
[0048] Step S1, obtain an initial sample data set, use the cross entropy gradient of the sample to determine the gradient contribution of the sample, divide the gradient contribution in the range of 0 to 1 into multiple intervals and set a safety gradient threshold, and take the interval whose gradient contribution is less than the set safety gradient threshold as the safety gradient interval;
[0049] Step S2: All minority class samples are assigned to different gradient intervals according to their gradient contributions and a safe gradient distribution calculation is performed;
[0050] Step S3, taking samples within the safety gradient interval as root samples, taking the gradient right neighbor of the root sample as auxiliary samples, and determining the number of sample synthesis based on the safety gradient distribution approximation strategy;
[0051] Step S4: Synthesize pseudo samples for each safety gradient interval using a linear interpolation method to achieve sample synthesis oversampling.
[0052] Through the above process, the error accumulation of noise samples in the synthetic oversampling process can be avoided and does not rely on spatial features, thereby ensuring the category balance of the data set and improving the performance of the classification model. In order to facilitate the understanding of the present invention, the implementation process of the present invention is further described in detail below:
[0053] In step S1, the whole process can be implemented by the following steps:
[0054] Step S1-1, obtaining an initial data set.
[0055] In this embodiment, the data processing in the field of credit default in financial risk control is taken as an example: first, the labeled data in the field of credit default is obtained as sample data, and an initial data set is formed based on the obtained sample data. Specifically, the labeling of credit default data includes credit default and non-credit default. Classification is performed, and the data set with a larger number of samples corresponding to the labeled type is regarded as the majority class sample; correspondingly, the data set with a smaller number of samples corresponding to the labeled type is regarded as the minority class sample.
[0056] Specifically, using credit card consumption record data as the source of credit default data, based on the actual use of credit cards, obtain credit card credit data that has been marked as having credit defaults and not having credit defaults as sample data, and form an initial data set based on several such sample data obtained ; Preferably, samples with credit default can be marked as 1, and samples without credit default can be marked as 0. The data source of credit card consumption records and the labeling of samples can be provided by banks and other work units with relevant functions. Finally, the data set with a larger number of samples corresponding to the labeling type is regarded as the majority class sample; correspondingly, the data set with a smaller number of samples corresponding to the labeling type is regarded as the minority class sample. In actual situations, the samples with credit default are much smaller than the samples without credit default. Therefore, the result generally obtained is to regard the sample set composed of the sample data with credit default as the minority class sample, and the sample set composed of the sample data without credit default as the majority class sample.
[0057] Step S1-2, using the cross entropy gradient of the sample to determine the gradient contribution of the sample, dividing the gradient contribution in the range of 0 to 1 into multiple intervals and setting a safety gradient threshold, and taking the interval whose gradient contribution is less than the set safety gradient threshold as the safety gradient interval.
[0058] Specifically, the present invention proposes a new classification index for minority class samples in the synthetic oversampling method, namely gradient contribution. Specifically, the cross entropy gradient of the sample is defined as gradient contribution, which is denoted as .
[0059] Furthermore, the gradient contribution The sample is categorized by the binary cross entropy loss function. x The derivative shows that the cross entropy loss function for the two-class classification is expressed as: .in, Represents the label of the sample, which is 0 or 1 in the case of binary classification; Refers to the probability that the model predicts the sample as a minority class sample (i.e., a sample whose label is marked as 1). Correspondingly, It is expressed as the probability that the model predicts the sample as the majority class sample (i.e., the sample whose label is marked as 0). It reflects the gradient contribution of the sample to the model, that is, the difficulty of fitting the sample. The larger it is, the harder it is to fit the sample and the closer it is to the decision boundary in space. Specifically, the gradient contribution The calculation method is as follows:
[0060] ;
[0061] Furthermore, the present invention uses the classification index of the above-mentioned minority class samples to The gradient within the range is divided into intervals and set the safety gradient threshold . When dividing the safety gradient interval, the interval whose gradient contribution is less than the safety gradient threshold is defined as the safe gradient interval (Safe Gradient Interval, SGI), and the interval whose gradient contribution is greater than the safety gradient threshold is defined as the dangerous gradient interval (Danger Gradient Interval, DGI). The safe gradient interval closest to the safety gradient threshold is defined as the critical safe gradient interval (Critical Safe Gradient Interval, CSGI), and the dangerous gradient interval closest to the safety gradient threshold is defined as the critical dangerous gradient interval (Critical Danger Gradient Interval, CDGI). The concepts of each gradient interval are first proposed and defined by the present invention. Specifically, the form of each defined interval is as follows Figure 2 shown.
[0062] In step S2, all minority class samples are assigned to different gradient intervals according to their gradient contributions and the safe gradient distribution is calculated, which can be achieved through the following process:
[0063] Step S2-1, sample allocation. Based on the above classification indicators and gradient interval definitions, the sample allocation is based on the gradient contribution. All minority class samples are assigned to different gradient intervals. Specifically: each gradient interval is Indicates that the interval index is an integer Assumptions Indicates minority class samples, Represents the interval, then the gradient range of each gradient interval can be expressed as: , further, the gradient interval to which each minority class sample belongs is expressed as follows:
[0064] ;
[0065] Step S2-2: Calculate the security gradient distribution. First, define the number of majority class samples in the class imbalanced dataset as , the number of minority class samples is ; Then, the number of samples required to be synthesized is Based on the distribution results of different gradient intervals of samples, the number of positive samples in each safety gradient interval is counted, that is, the safety gradient distribution of positive samples is calculated. Assume that the number of safety gradient intervals divided according to the safety gradient threshold is , security gradient distribution It means that, then, the security gradient distribution The statistical result is a dimension vector; let Indicates The number of samples in the safety gradient interval, then the safety gradient distribution It is expressed as: .
[0066] Furthermore, the security gradient distribution approximation strategy is derived from the cosine similarity metric of two vectors. Specifically, assuming that the security gradient distribution of the positive class sample after synthesis is express, The form of the positive sample security gradient distribution before synthesis Consistent, too dimensional vector, as shown below, where After synthesis, The number of samples in the interval, Indicates the parameters that need to be solved.
[0067] ;
[0068] Cosine similarity is used to measure and The similarity of and The similarity is the greatest, that is When , the approximate principle of safe gradient distribution can be satisfied, that is:
[0069] ;
[0070] Based on the above formula, the number of samples in each safety gradient interval after synthesis can be solved ,Right now:
[0071] ;
[0072] The number of samples required to be synthesized in each safety gradient interval It can be solved by the following formula:
[0073] ;
[0074] .
[0075] In step S3, the samples in the safety gradient interval are used as root samples, the gradient right neighbors of the root samples are used as auxiliary samples, and the number of sample synthesis is determined based on the safety gradient distribution approximation strategy, which can be specifically achieved through the following process:
[0076] Step S3-1, select the root sample. The present invention uses the samples within the safety gradient interval as the root sample. , expressed as This root sample selection strategy is based on the performance of the sample during the classification model training process. It does not depend on the spatial characteristics of the sample, and overcomes the shortcomings of the failure of the high-dimensional distance metric relied on by the existing methods. At the same time, selecting samples within the safe gradient interval as root samples can also reduce the error accumulation caused by synthesizing pseudo samples with some noise samples with high gradient contributions as root samples.
[0077] Step S3-2, select auxiliary samples. First, the present invention proposes a new concept of neighbors that is independent of the sample space characteristics, namely, the gradient right neighbor. Specifically, the gradient right neighbor of a sample refers to a sample whose gradient contribution is greater than that of the sample but whose gradient contribution difference with that of the sample is the smallest. Assume that the gradient right neighbor is expressed as ,sample The gradient contribution of ,sample The gradient contribution of the right neighbor of the gradient is expressed as ; Then, the sample The gradient right neighbor of is expressed as: The present invention selects the gradient right neighbor of the root sample as the auxiliary sample. The gradient right neighbor is a new neighbor concept proposed by the present invention, which is different from the one that depends on spatial features. Compared with existing neighbor selection strategies such as nearest neighbor and natural neighbor, the present invention determines neighbor samples according to the gradient contribution of the sample to the model. Therefore, selecting the gradient right nearest neighbor as the auxiliary sample not only gets rid of the dependence on spatial features, but also ensures that the synthesized pseudo sample and the current root sample fall in the same gradient interval.
[0078] Step S3-3, determine the number of synthesis. The present invention proposes a new synthesis number determination strategy, namely, safe gradient distribution approximation, whose premise is that the gradient distribution of minority class samples in real samples is consistent with the gradient distribution of minority class samples in the training set; therefore, the purpose of the safe gradient distribution approximation strategy of the present invention is to ensure that the distribution of minority class samples after synthetic oversampling and the original minority class samples in the safe gradient interval remains similar.
[0079] Specifically, the number of samples that need to be synthesized in each safety gradient interval is , that is, before and after the pseudo samples are synthesized, the ratio of the number of samples in each security gradient interval to the number of samples in all security gradient intervals remains the same. represents the distribution of synthesized minority samples in all safety gradient intervals, After synthesis, The number of positive samples in the safety gradient interval, then the safety gradient distribution approximation strategy can be expressed as .
[0080] In step S4, a linear interpolation method is used to synthesize pseudo samples for each safety gradient interval to achieve sample synthesis oversampling. Specifically, according to the synthesis information determined in steps S3-1 to S3-3, a linear interpolation method is used to synthesize pseudo samples for each safety gradient interval as follows:
[0081] ;
[0082] in, represents the synthesized pseudo sample, represents the root sample, Represents a random number between 0 and 1. Represents auxiliary samples.
[0083] Compared with the prior art, the present invention introduces the concept of gradient contribution, which can characterize the difficulty of sample fitting to a certain extent. In space, it represents the distance of the sample from the decision boundary, that is, the more difficult the sample is to fit, the closer it is to the decision boundary, without relying on spatial distance.
[0084] Furthermore, if Figure 3 FIG. 1 is a schematic diagram of the synthesis process of the synthetic oversampling method (two-dimensional) based on spatial distance in the prior art, assuming that the spatial distance metric is valid. Figure 3 The oversampling method shown on the left is implemented by determining auxiliary samples based on spatial distance. When synthesizing a pseudo sample for the sample with the highest position in the dotted circle (that is, when the sample is the root sample), assuming that the number of search neighbors is 5, it can only find auxiliary samples within the dotted line range. Figure 3 The oversampling method shown on the right is based on synthetic oversampling of spatial distance clustering. When the sample distribution within the dotted line is sparse, this method will synthesize a large number of pseudo samples for the subcluster, which will force the decision boundary to move towards the majority class, thereby reducing the recognition accuracy of the majority class samples. The most important thing is that the distance measurement in high-dimensional space is invalid, that is, the assumption at the beginning of this paragraph is not valid under high-dimensional conditions.
[0085] like Figure 4The figure shows a schematic diagram of the synthesis process of the synthetic oversampling method based on the safe gradient distribution of the present invention. It is not difficult to see that the samples in the safe gradient interval are root samples, which reduces the error accumulation of noise samples, and the pseudo samples synthesized by the gradient right neighbor as auxiliary samples always fall closer to the decision boundary than the root sample and not far from the gradient right neighbor; the safe gradient distribution approximation principle ensures that the number of minority class samples in each gradient interval in the data set is consistent with the distribution of all minority class samples in each safe gradient interval in reality; that is, under the premise of class balance, the proportion of minority class samples in each gradient interval to the total samples is restored. Combined with the maximum classification interval principle of the classification model, the decision boundary can be moved to the majority class samples in a soft way, ensuring that the algorithm of the present invention will not significantly reduce the recognition accuracy of the majority class samples.
[0086] In general, the present invention is a new synthetic oversampling method for unbalanced classification tasks. Experiments have shown that the method of the present invention can not only improve the recognition accuracy of minority class samples, but also does not sacrifice the recognition accuracy of majority class samples, making the classification model's recognition ability for samples of different categories more balanced.
[0087] To further verify the significant improvement of the synthetic oversampling method based on secure gradient distribution provided by the present invention over the prior art, this embodiment compares the performance improvement of the method of the present invention and other eight synthetic oversampling methods on six indicators of logistic regression (LR) and multi-layer perceptron (MLP) classifiers on four unbalanced financial data sets. Specifically, the experimental results are shown in Table 1:
[0088] Table 1 Performance comparison based on different classifiers
[0089]
[0090] From Table 1, we can see that: First, for dataset No. 1, the GDSO method makes all the indicators of the MLP classifier achieve the optimal value, among which the F1-Score, MCC and KS indicators have a more obvious improvement effect. At the same time, the F1-Score, Recall, G-Mean, and MCC indicators of the LR classifier have all achieved the optimal value, and the AUC indicator is only 0.08 percentage points different from the optimal KM-SMOTE method, and the KS indicator is 1.23 percentage points different from the optimal SMOTEN method. On dataset No. 2, the F1-Score, G-Mean, and MCC indicators of the two classifiers are all optimal. At the same time, the AUC and KS indicators of the MLP classifier are also optimal. In addition to the KM-SMOTE method, the Recall value is also optimal, but the F1-Score value of GDSO is 2.85 percentage points higher than that of the KM-SMOTE method. The Recall value of the LR classifier is the best except for the SMOTE-ENN method, but the F1-Score value of GDSO is 2.8 percentage points higher than that of SMOTE-ENN. The F1-Score, G-Mean, MCC, and KS indicators of the two classifiers are all optimal on the No. 3 dataset, and the AUC indicator of the LR classifier is also optimal. Similar to the performance on the No. 2 dataset, the Recall value of the LR classifier is the best except for the SMOTE-ENN method, but the F1-Score value of GDSO is 3.07 percentage points higher than that of SMOTE-ENN. The Recall value of the MLP classifier is the highest except for SMOTE, but the F1-Score value of the GDSO method is 8.54 percentage points higher than that of SMOTE. On the No. 4 dataset, the Recall, G-Mean, and MCC of the MLP classifier are all optimal. When the safety gradient threshold of GDSO is 0.5, the F1-Score, MCC, and KS indicators of the LR classifier are all optimal. While achieving a higher Recall value, the GDSO method significantly improves the F1-Score value, which is 10.46 percentage points higher than the other optimal KM-SMOTE. When the safety gradient threshold is 0.9, the AUC and MCC indicators are optimal. At this time, GDSO achieves a higher Recall, but the F1-Score value is still higher than other methods that approximate the Recall value.
[0091] The experimental results prove the advantages of the method of the present invention, that is, the method of the present invention can not only improve the recognition accuracy of minority class samples to be consistent with or exceed other algorithms, but also obtain higher F1-Score and MCC index values than other algorithms, that is, without losing the recognition accuracy of majority class samples. Therefore, for the credit default prediction scenario under financial risk control data, the present invention can well identify default and non-default users at the same time, which meets the requirements of modern risk control and wide credit, and is conducive to the incremental development of credit business.
[0092] Embodiment 2
[0093] This embodiment discloses a synthetic oversampling system based on a safe gradient distribution, including:
[0094] The gradient interval division module is configured to: determine the gradient contribution of the sample by the cross entropy gradient of the sample, divide the gradient contribution in the range of 0 to 1 into multiple intervals and set a safety gradient threshold, and take the interval whose gradient contribution is less than the set safety gradient threshold as the safety gradient interval;
[0095] The gradient interval allocation module is configured to: allocate all minority class samples to different gradient intervals according to gradient contribution and perform safe gradient distribution calculation; use samples in the safe gradient interval as root samples, use the gradient right neighbors of the root samples as auxiliary samples, and determine the number of sample synthesis based on the safe gradient distribution approximation strategy;
[0096] The sample synthesis module is configured to: synthesize pseudo samples for each safety gradient interval using a linear interpolation method to achieve sample synthesis oversampling.
[0097] Embodiment 3
[0098] The purpose of this embodiment is to provide a computer-readable storage medium.
[0099] A computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps in a synthetic oversampling method based on a secure gradient distribution as described in the first embodiment of the present disclosure.
[0100] Embodiment 4
[0101] The purpose of this embodiment is to provide an electronic device.
[0102] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in a synthetic oversampling method based on a secure gradient distribution as described in the first embodiment of the present disclosure are implemented.
[0103] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0104] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0105] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A synthetic oversampling method based on safe gradient distribution, characterized in that: include: Obtaining an initial sample data set, specifically: using credit card consumption record data as a source of credit default data, obtaining credit card credit data that has been marked as having credit default or not having credit default as sample data based on actual credit card usage, and forming an initial sample data set based on a number of such sample data obtained; wherein, samples that have credit default are marked as 1, and samples that have not credit default are marked as 0; The initial sample data set is divided into categories according to different annotation types. Specifically, the sample data set with a larger number of samples corresponding to the annotation type is regarded as the majority class sample, that is, the samples without credit default are regarded as the majority class sample; the sample data set with a smaller number of samples corresponding to the annotation type is regarded as the minority class sample, that is, the samples with credit default are regarded as the minority class sample; The cross entropy gradient of the sample is used to determine the gradient contribution of the sample. The gradient contribution in the range of 0 to 1 is divided into multiple intervals and a safety gradient threshold is set. The interval with a gradient contribution less than the set safety gradient threshold is taken as the safety gradient interval. All minority class samples are assigned to different gradient intervals according to their gradient contributions and safe gradient distribution is calculated; samples within the safe gradient interval are used as root samples, the gradient right neighbors of the root samples are used as auxiliary samples, and the number of sample synthesis is determined based on the safe gradient distribution approximation strategy; The linear interpolation method is used to synthesize pseudo samples for each security gradient interval to achieve sample synthesis oversampling, so as to improve the recognition accuracy of minority samples that have credit defaults.
2. A synthetic oversampling method based on security gradient distribution as claimed in claim 1, characterized in that: The gradient contribution is obtained by derivation of the binary cross entropy loss function on the sample.
3. The synthetic oversampling method based on security gradient distribution according to claim 1, characterized in that: The interval whose gradient contribution is greater than the safety gradient threshold is taken as the dangerous gradient interval; at the same time, the safe gradient interval closest to the safety gradient threshold is taken as the critical safe gradient interval, and the dangerous gradient interval closest to the safety gradient threshold is taken as the critical dangerous gradient interval.
4. The synthetic oversampling method based on security gradient distribution according to claim 1, characterized in that: The security gradient distribution calculation is specifically: based on the allocation results of different gradient intervals of the samples, the number of positive samples in each security gradient interval is counted, that is, the security gradient distribution of the positive samples is calculated.
5. The synthetic oversampling method based on security gradient distribution according to claim 1, characterized in that: The gradient right neighbor of a sample is the sample whose gradient contribution is greater than that of the sample but whose gradient contribution difference with that of the sample is the smallest.
6. The synthetic oversampling method based on security gradient distribution according to claim 1, characterized in that: Before and after synthesizing pseudo samples, the ratio of the number of samples in each safety gradient interval to the number of samples in all safety gradient intervals remains consistent.
7. The synthetic oversampling method based on security gradient distribution according to claim 1, characterized in that: The synthesis method of synthesizing pseudo samples for each safety gradient interval using linear interpolation method is as follows: ; in, represents the synthesized pseudo sample, represents the root sample, Represents a random number between 0 and 1. Represents auxiliary samples.
8. A synthetic oversampling system based on a safety gradient distribution, characterized in that: include: The gradient interval division module is configured to: obtain an initial sample data set, specifically: use credit card consumption record data as the source of credit default data, obtain credit card credit data that has been marked as having credit default and not having credit default as sample data based on the actual use of the credit card, and form an initial sample data set based on several such sample data obtained; wherein, samples with credit default are marked as 1, and samples without credit default are marked as 0; classify the initial sample data set according to different annotation types, specifically: use a sample data set with a larger number of samples corresponding to the annotation type as the majority class samples, that is, use samples without credit default as the majority class samples; use a sample data set with a smaller number of samples corresponding to the annotation type as the minority class samples, that is, use samples with credit default as the minority class samples; determine the gradient contribution of the sample through the cross entropy gradient of the sample, divide the gradient contribution in the range of 0 to 1 into multiple intervals and set a safety gradient threshold, and use the interval with a gradient contribution less than the set safety gradient threshold as the safety gradient interval; The gradient interval allocation module is configured to: allocate all minority class samples to different gradient intervals according to gradient contribution and perform safe gradient distribution calculation; use samples in the safe gradient interval as root samples, use the gradient right neighbors of the root samples as auxiliary samples, and determine the number of sample synthesis based on the safe gradient distribution approximation strategy; The sample synthesis module is configured to: synthesize pseudo samples for each safety gradient interval using a linear interpolation method to achieve sample synthesis oversampling to improve the recognition accuracy of minority class samples that have credit defaults.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the synthetic oversampling method based on secure gradient distribution as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the synthetic oversampling method based on security gradient distribution as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Minor class sample self-paced synthesis algorithm based on contribution value grade
CN115795357A
Oversampling method and system for differential evolution of highly unbalanced data set
CN115878999A
Cited By
Credit default prediction method and system based on knowledge graph and extreme gradient lifting
CN121213219A