Hybrid resampling method for unbalanced data

By dividing and clustering the unbalanced data set, synthesized samples are generated and the majority class sample set is optimized, and neighborhood graphs are constructed for fusion, which solves the bias problem in traditional methods and improves the recognition rate of minority categories and the accuracy of multi-label classification.

CN120508846APending Publication Date: 2025-08-19CHONGQING UNIV OF EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510589948.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

When the prior art deals with unbalanced data sets, traditional classification methods are prone to introduce bias, resulting in extremely low recognition rates for a few categories and increasing the complexity of decision-making boundaries of multi-label classification, and existing methods have limitations.

Method used

By dividing the unbalanced data set, performing clustering of minority samples, generating synthetic samples and calculating the allocation weight of synthetic samples, combining the density optimization of most categories of sample sets, constructing neighborhood graphs for fusion, and generating a mixed resampling sample set.

Benefits of technology

The generated samples conform to the local structure of a few classes of distribution, maintain the direction and density in the feature space, improve the learning effect of the classifier, and enhance the applicability to large-scale complex data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508846A_ABST
    Figure CN120508846A_ABST
Patent Text Reader

Abstract

The invention provides a hybrid resampling method for unbalanced data, and relates to the technical field of machine learning, and the method comprises the steps: dividing an unbalanced data set, and obtaining a majority class sample set and a minority class sample set; performing clustering processing on the minority class sample set to obtain a resampled minority class sample set; based on each sample in the resampling minority class sample set, calculating by using an interpolation method to obtain a synthetic sample and a corresponding synthetic sample distribution weight; optimizing the majority class sample set by calculating the density of each sample in the majority class sample set to obtain a resampled majority class sample set; and based on the synthetic sample distribution weight, fusing the synthetic sample, the resampling minority class sample set and the resampling majority class sample set to obtain a mixed resampling sample set, and completing the mixed resampling of the unbalanced data. The problem that unbalanced data are difficult to accurately sample is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of machine learning technology, and in particular to a hybrid resampling method for imbalanced data. Background Art

[0002] In the field of machine learning (ML), various tasks encounter significant obstacles when faced with datasets exhibiting high class imbalance. These tasks range from detecting fraudulent transactions to identifying faults. In these cases, the primary goal is often to minimize false negatives, as the minority class holds greater significance. However, applying traditional machine learning classifiers to imbalanced datasets introduces a bias toward the more common class, limiting the predictive power of the minority class. This phenomenon is widely recognized in the machine learning community as the challenge of imbalanced learning.

[0003] Traditional classification methods face challenges when dealing with imbalanced datasets, often resulting in extremely poor recognition rates for the minority class. Unfortunately, the cost of misclassification of the minority class is often higher than that of the majority class. This degradation in classification performance primarily stems from complex distributional characteristics. These characteristics include small disjuncts, class overlap, rare instances, and outliers in the minority class.

[0004] Multi-label classification involves distinguishing between more than two classes, which poses considerable challenges due to the increased complexity of defining the decision boundary. In contrast, binary classification, involving only two classes, typically allows for simpler and easier-to-define decision boundaries. An effective approach to multi-label classification is to decompose it into multiple binary classification tasks using techniques such as One-Vs-One (OVO) and One-Vs-All (OVA). OVO methods address the multi-class problem by creating multiple binary classifiers, each trained to distinguish between a specific pair of classes. The overall predicted class is obtained by combining the outputs from the individual base classifiers. OVA methods involve training a separate binary classifier for each class, where each classifier is designed to distinguish its target class from all other classes. A positive prediction from any base classifier results in the corresponding class being assigned as the final output. Although many methods have been proposed in the literature for handling imbalanced data, many of them still have certain limitations. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a hybrid resampling method for unbalanced data, which solves the problem that unbalanced data is difficult to sample accurately.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a hybrid resampling method for unbalanced data, comprising: S1: Divide the imbalanced data set into majority class sample sets and minority class sample sets; S2: performing clustering processing on the minority class sample set to obtain a resampled minority class sample set; S3: Based on each sample in the resampled minority class sample set, an interpolation method is used to calculate and obtain a synthetic sample and the corresponding synthetic sample allocation weight; S4: Optimize the majority class sample set by calculating the density of each sample in the majority class sample set to obtain the resampled majority class sample set; S5: Based on the weights assigned to the synthetic samples, the synthetic samples, the resampled minority class sample set, and the resampled majority class sample set are fused to obtain a mixed resampled sample set, thereby completing the mixed resampling of the imbalanced data.

[0007] Furthermore, the S2 includes: Analyze the minority class sample set to obtain the number of clusters; Based on the number of clusters, the intra-cluster sum of squares of each cluster is obtained by calculation; Clustering the minority class sample set using the intra-cluster sum of squares to obtain a clustering result; A neighborhood graph is constructed for each cluster in the clustering result to obtain a resampled minority class sample set.

[0008] Furthermore, the expression of the intra-cluster sum of squares is: ; in, represents the within-cluster sum of squares, represents the number of clusters, Indicates the number of samples in the minority class samples, Indicates the clusters, Indicates the The centroids of the clusters, represents the L2 norm.

[0009] Furthermore, the expression of the synthetic sample is: ; in, represents a synthetic sample, Indicates the number of samples in the minority class samples, represents random weights, express Neighbor samples of express The neighbor sample set of .

[0010] Furthermore, the expression for assigning weights to the synthetic samples is: ; in, represents the weight of synthetic sample allocation, Indicates the clusters, Indicates the number of samples in the minority class samples, represents a synthetic sample.

[0011] Furthermore, the S4 includes: By calculating the majority class sample set, the density of each sample in the majority class sample set is obtained; Based on the density of each sample in the majority class sample set, samples with a density lower than a density threshold are deleted to obtain a resampled majority class sample set.

[0012] Furthermore, the density of each sample in the majority class sample set is expressed as: ; in, represents the density of each sample in the majority class sample set, Indicates the majority class sample set samples, express Neighbor samples of represents the majority class sample set, Indicates the conditional separator, Indicates the density threshold.

[0013] The beneficial effects of the present invention are: a hybrid resampling method for imbalanced data. (1) By constructing a neighborhood graph, the generated samples are ensured to conform to the local structure of the minority class distribution, and synthetic samples are generated using interpolation between adjacent points, maintaining the direction and density of the minority class in the feature space, thereby obtaining more accurate sampling results; (2) synthetic samples are inserted into the hybrid resampling sample set using synthetic sample assignment weights, and adaptive weights are assigned to the synthetic samples based on their similarity to existing instances, ensuring that the classifier learns from more informative and representative synthetic data, further enhancing the applicability to large-scale complex datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein: Figure 1 This is an exemplary flow chart of a hybrid resampling method for unbalanced data according to some embodiments of this specification. DETAILED DESCRIPTION

[0015] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0016] Example 1 Figure 1 FIG. 1 is an exemplary flow chart of a hybrid resampling method for unbalanced data according to some embodiments of this specification. Figure 1 As shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.

[0017] S1: Divide the imbalanced data set into majority class sample sets and minority class sample sets.

[0018] An imbalanced dataset is a dataset where the data is unevenly distributed. For example, an imbalanced dataset can include an imbalanced dataset of forest fire data, etc.

[0019] In some embodiments, the imbalanced dataset can be represented as: ; in, represents an unbalanced dataset, Indicates the imbalanced dataset data, express Tags, Indicates the amount of data in the imbalanced dataset; , represents the majority class sample set, represents the minority class sample set.

[0020] The majority class sample set is a set of data in the imbalanced dataset that is labeled as the majority class sample set.

[0021] The minority class sample set is a set of data labeled as the minority class sample set in the imbalanced dataset.

[0022] In some embodiments, the processor may use the majority class sample set and the minority class sample set to perform calculations to obtain an imbalance rate of the imbalanced data set. When the imbalance rate is greater than an imbalance threshold, a mixed resampling method of the imbalanced data is performed to obtain a mixed resampled sample set.

[0023] In some embodiments, the imbalance rate of the imbalanced dataset may be expressed as: ; in, Indicates the imbalance rate of the imbalanced dataset.

[0024] S2: performing clustering processing on the minority class sample set to obtain a resampled minority class sample set.

[0025] The resampled minority class sample set is a neighborhood graph reflecting the connection relationship between each sample in the minority class sample set. For example, the resampled minority class sample set may include a constructed neighborhood graph having a point set and an edge set.

[0026] In some embodiments, the processor can implement S2 based on the following steps: analyzing the minority class sample set to obtain the number of clusters; based on the number of clusters, obtaining the intra-cluster sum of squares of each cluster by calculation; using the intra-cluster sum of squares to cluster the minority class sample set to obtain a clustering result; constructing a neighborhood graph for each cluster in the clustering result to obtain a resampled minority class sample set.

[0027] The number of clusters is the number of clusters in which the minority class samples are clustered.

[0028] In some embodiments, the processor may perform k-means clustering on the minority class sample set to obtain the number of clusters.

[0029] The intra-cluster sum of squares of each cluster is the data reflecting the distance from each sample in the minority class sample set to the centroid.

[0030] In some embodiments, the expression for the within-cluster sum of squares may be: ; in, represents the within-cluster sum of squares, represents the number of clusters, Indicates the number of samples in the minority class samples, Indicates the clusters, Indicates the The centroids of the clusters, represents the L2 norm.

[0031] The clustering result is the result that reflects the clustering of the minority class sample set. For example, the expression of the clustering result can be: ;in, represents the first cluster, represents the second cluster, Indicates the Clusters.

[0032] In some embodiments, the processor may calculate a coherence threshold of samples within each cluster in the clustering result to obtain a resampled minority class sample set.

[0033] In some embodiments, the expression for resampling the minority class sample set may be: ; in, represents the resampled minority class sample set, represents a point set, represents an edge set; , , , represents the coherence threshold.

[0034] S3: Based on each sample in the resampled minority class sample set, an interpolation method is used to calculate and obtain the synthetic sample and the corresponding synthetic sample allocation weight.

[0035] Synthetic samples are samples used to interpolate the constructed neighborhood graph.

[0036] In some embodiments, the expression of the synthetic sample may be: ; in, represents a synthetic sample, Indicates the number of samples in the minority class samples, represents random weights, express Neighbor samples of express The neighbor sample set of .

[0037] The synthetic sample allocation weight is the weight data of the synthetic sample allocated to the constructed neighborhood graph.

[0038] In some embodiments, the expression for assigning weights to synthetic samples may be: ; in, represents the weight of synthetic sample allocation, Indicates the clusters, Indicates the number of samples in the minority class samples, represents a synthetic sample.

[0039] In some embodiments, the processor may use the synthetic sample allocation weights to insert the synthetic samples into the mixed resampling sample set to complete the mixed resampling of the unbalanced data.

[0040] S4: Optimize the majority class sample set by calculating the density of each sample in the majority class sample set to obtain the resampled majority class sample set.

[0041] The resampled majority class sample set is the majority class sample set after resampling.

[0042] In some embodiments, the processor can obtain the density of each sample in the majority class sample set by calculating the majority class sample set; based on the density of each sample in the majority class sample set, delete samples with density lower than a density threshold to obtain a resampled majority class sample set.

[0043] In some embodiments, the density of each sample in the majority class sample set may be expressed as: ; in, represents the density of each sample in the majority class sample set, Indicates the majority class sample set samples, express Neighbor samples of represents the majority class sample set, Indicates the density threshold.

[0044] S5: Based on the weights assigned to the synthetic samples, the synthetic samples, the resampled minority class sample set, and the resampled majority class sample set are fused to obtain a mixed resampled sample set, thereby completing the mixed resampling of the imbalanced data.

[0045] The mixed resampled sample set is a mixed sample set obtained by fusing a resampled minority class sample set with a resampled majority class sample set. For example, the mixed resampled sample set may include a mixed resampled sample set of forest fires.

[0046] In some embodiments, the mixed resampled sample set can be expressed as: ; in, represents a mixed resampled sample set, represents a synthetic sample, represents the resampled majority class sample set, Represents the resampled minority class sample set.

[0047] In some embodiments of the present specification, a hybrid resampling method for imbalanced data is provided. (1) By constructing a neighborhood graph, the generated samples are ensured to conform to the local structure of the minority class distribution. Interpolation between adjacent points is used to generate synthetic samples, maintaining the direction and density of the minority class in the feature space, and obtaining more accurate sampling results. (2) Synthetic samples are inserted into the hybrid resampling sample set using synthetic sample assignment weights. Adaptive weights are assigned to the synthetic samples based on their similarity to existing instances, ensuring that the classifier learns from more informative and representative synthetic data, further enhancing its applicability to large-scale complex datasets.

[0048] Example 2 Most raw remote sensing images in existing forest fire datasets suffer from the problem of extremely small fire areas. The proportion of fire pixels can range from 0.2% to 10%. As a result, after image processing, the ratio of positive and negative samples will be severely unbalanced.

[0049] In some embodiments, a hybrid resampling method for imbalanced data includes: S1: Divide the unbalanced forest fire dataset into a majority class sample set and a minority class sample set; S2: performing clustering processing on the forest fire minority class sample set to obtain a resampled minority class sample set; S3: Based on each sample in the resampled minority class sample set, an interpolation method is used to calculate and obtain a synthetic sample and the corresponding synthetic sample allocation weight; S4: Optimize the forest fire majority class sample set by calculating the density of each sample in the forest fire majority class sample set to obtain the resampled majority class sample set; S5: Based on the weights assigned to the synthetic samples, the synthetic samples, the resampled minority class sample set and the resampled majority class sample set are fused to obtain a forest fire mixed resampled sample set, thereby completing the mixed resampling of the unbalanced data.

[0050] In some embodiments, the S2 includes: Analyze the forest fire minority sample set to obtain the number of clusters; Based on the number of clusters, the intra-cluster sum of squares of each cluster is obtained by calculation; Clustering the forest fire minority sample set using the intra-cluster sum of squares to obtain a clustering result; A neighborhood graph is constructed for each cluster in the clustering result to obtain a resampled minority class sample set.

[0051] In some embodiments, the expression for the intra-cluster sum of squares is: ; in, represents the within-cluster sum of squares, represents the number of clusters, Indicates the number of forest fire minority samples in the samples, Indicates the clusters, Indicates the The centroids of the clusters, represents the L2 norm.

[0052] In some embodiments, the expression of the synthetic sample is: ; in, represents a synthetic sample, Indicates the number of forest fire minority samples in the samples, represents random weights, express Neighbor samples of express The neighbor sample set of .

[0053] In some embodiments, the expression for assigning weights to the synthetic samples is: ; in, represents the weight of synthetic sample allocation, Indicates the clusters, Indicates the number of forest fire minority samples in the samples, represents a synthetic sample.

[0054] In some embodiments, the S4 includes: By calculating the forest fire majority class sample set, the density of each sample in the forest fire majority class sample set is obtained; Based on the density of each sample in the forest fire majority class sample set, samples with a density lower than a density threshold are deleted to obtain a resampled majority class sample set.

[0055] In some embodiments, the density of each sample in the majority class sample set is expressed as: ; in, represents the density of each sample in the forest fire majority class sample set, Indicates the number of samples in the majority class of forest fires samples, express Neighbor samples of represents the majority class sample set of forest fires, Indicates the conditional separator, Indicates the density threshold.

[0056] In some embodiments of this specification, a hybrid resampling method for unbalanced data is provided. An unbalanced forest fire dataset is partitioned to obtain a hybrid resampled forest fire sample set. By utilizing this hybrid resampling method, a substantial balance is maintained between positive and negative samples in the dataset, improving the performance and reliability of the model in forest fire detection and classification tasks.

Claims

1. A hybrid resampling method for unbalanced data, characterized in that: include: S1: Divide the imbalanced data set into majority class sample sets and minority class sample sets; S2: performing clustering processing on the minority class sample set to obtain a resampled minority class sample set; S3: Based on each sample in the resampled minority class sample set, an interpolation method is used to calculate and obtain a synthetic sample and the corresponding synthetic sample allocation weight; S4: Optimize the majority class sample set by calculating the density of each sample in the majority class sample set to obtain the resampled majority class sample set; S5: Based on the weights assigned to the synthetic samples, the synthetic samples, the resampled minority class sample set, and the resampled majority class sample set are fused to obtain a mixed resampled sample set, thereby completing the mixed resampling of the imbalanced data.

2. The hybrid resampling method for unbalanced data according to claim 1, characterized in that: The S2 includes: Analyze the minority class sample set to obtain the number of clusters; Based on the number of clusters, the intra-cluster sum of squares of each cluster is obtained by calculation; Clustering the minority class sample set using the intra-cluster sum of squares to obtain a clustering result; A neighborhood graph is constructed for each cluster in the clustering result to obtain a resampled minority class sample set.

3. The hybrid resampling method for unbalanced data according to claim 2, characterized in that: The expression of the intra-cluster sum of squares is: ; in, represents the within-cluster sum of squares, represents the number of clusters, Indicates the number of samples in the minority class samples, Indicates the clusters, Indicates the The centroids of the clusters, represents the L2 norm.

4. The hybrid resampling method for unbalanced data according to claim 1, wherein: The expression of the synthetic sample is: ; in, represents a synthetic sample, Indicates the number of samples in the minority class samples, represents random weights, express Neighbor samples of express The neighbor sample set of .

5. The hybrid resampling method for unbalanced data according to claim 4, characterized in that: The expression of the synthetic sample allocation weight is: ; in, represents the weight of synthetic sample allocation, Indicates the clusters, Indicates the number of samples in the minority class samples, represents a synthetic sample.

6. The hybrid resampling method for unbalanced data according to claim 1, wherein: The S4 includes: By calculating the majority class sample set, the density of each sample in the majority class sample set is obtained; Based on the density of each sample in the majority class sample set, samples with a density lower than a density threshold are deleted to obtain a resampled majority class sample set.

7. The hybrid resampling method for unbalanced data according to claim 6, characterized in that: The density expression of each sample in the majority class sample set is: ; in, represents the density of each sample in the majority class sample set, Indicates the majority class sample set samples, express Neighbor samples of represents the majority class sample set, Indicates the conditional separator, Indicates the density threshold.