Unbalanced time sequence classification enhancement method based on minority class label merging
By adopting the dual enhancement strategy and label mapping mechanism of minority class label merging in unbalanced time series classification, the problems of noise introduction, category proportion imbalance and classification boundary ambiguity are solved, and the recognition accuracy of minority classes is significantly improved.
Patent Information
- Application Number
- CN202510258524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
AI Technical Summary
The existing unbalanced time series classification methods have problems such as noise introduction, category proportion imbalance and classification boundary ambiguity, resulting in a decrease in the accuracy of recognition of minority classes.
An unbalanced time series classification enhancement method based on the merger of labels of minority classes is proposed. Through dual enhancement strategies (sample enhancement and label enhancement) and tag mapping mechanism, the proportion of minority classes in the tag space is improved, and classification boundaries are optimized through joint tag learning.
It effectively improves the characterization ability of a few types of samples, improves the accuracy of unbalanced time series classification, avoids noise interference, and ensures the clarity of classification boundaries.
Smart Images

Figure FSA0000300009580000011 
Figure FSA0000300009580000012 
Figure HSA0000300009600000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and time series data processing, and specifically relates to an enhanced method for classifying imbalanced time series based on the merging of minority class labels. This method effectively improves the representation ability of minority class samples in the classification model through a double-enhancement strategy of minority class merging and a label mapping mechanism, and at the same time combines joint label learning to optimize the clarity of the classification boundary, and is applicable to time series data classification tasks in fields such as financial risk control, medical diagnosis, and industrial equipment monitoring. Background Art
[0002] In time series classification tasks, class imbalance is a common challenge. Since the number of minority class samples is much lower than that of the majority class, traditional classification models tend to favor the learning of majority class features, resulting in a significant decrease in the recognition accuracy of the minority class. Existing solutions mainly include the following two categories:
[0003] Traditional data augmentation methods include oversampling and undersampling methods, but both of them have limitations. Oversampling increases the number of minority class samples by replicating minority class samples (such as random oversampling) or generating synthetic samples (such as the SMOTE algorithm). However, simple replication of samples is prone to overfitting, and synthetic samples may introduce noise or destroy the original time series features (such as destroying the time domain continuity or frequency domain stability). Undersampling balances the class ratio by randomly deleting majority class samples, but it will cause information loss. Especially when there are complex sub-patterns within the majority class, key features may be removed. Although the joint augmented label learning method JobDA generates augmented samples through time series distortion methods and assigns independent self-supervised labels to each augmentation operation, jointly learning the original label and the self-supervised label. However, this method does not adjust the class ratio, and the distortion operation may change the original data distribution.
[0004] Generally speaking, there are three problems to be solved in the existing enhancement methods for imbalanced time series classification, namely the noise introduction problem, the class ratio imbalance problem, and the classification boundary fuzziness problem. For the noise introduction problem, the reason is that synthetic samples or data distortion operations destroy the original time series features, such as the time shift of sensor data leading to a fuzzy classification boundary. For the class ratio imbalance problem, existing methods only alleviate the surface imbalance by adjusting the sample quantity, do not enhance the weight of the minority class at the label space level, and do not change the label distribution. For the classification boundary fuzziness problem, it is because the augmented samples share the same label as the original samples, resulting in the diffusion of the distribution of samples of the same class, such as traditional oversampling strategies such as oversampling and undersampling. Summary of the Invention
[0005] Objective of the Invention: To solve the above problems existing in the unbalanced time series data enhancement method, the present invention proposes an unbalanced time series classification enhancement method based on minority class label merging. Its core objective is to retain the original time series characteristics through repetitive sample enhancement, avoid noise interference, and use the label mapping mechanism to increase the proportion of minority classes in the label space. Finally, the classification boundary is optimized through joint label learning to improve the recognition accuracy of the model for minority classes.
[0006] Technical Solution: There is provided an unbalanced time series classification enhancement method based on minority class label merging, characterized in that this method can expand both samples and labels through double augmentation to obtain a larger-scale data set. And for the minority classes, their original and enhanced labels are merged to increase the proportion of minority class labels in the potential sample space. Then, through joint label learning, the deep learning model can learn more features of the minority classes during training, improving the accuracy of unbalanced time series classification, including the following steps:
[0007] S1. Double augmentation processing, including sample augmentation and label augmentation;
[0008] Sample augmentation is to generate enhanced samples from the original time series data. Let the original sample be (L is the time step), after repeating A times, an enhanced sample set T' = {t' i,1 , …, t' i,2 , …, t' i,A} is generated, and the final training set is T aug = T ∪ T'. Using the sample repetition operation can not change the time domain and frequency domain characteristics and avoid the noise introduced by traditional distortion or interpolation. Label augmentation is to assign independent labels to each enhanced sample. The label of the original sample is y i ∈ {1, 2, …, C}, and the label of the enhanced sample is extended to y' i ∈ {C + 1, C + 2, …, C × (A + 1)}, forming a joint label space.
[0009] S2. Label mapping mechanism, increasing the proportion of minority class labels through minority class label merging and then mapping to training labels:
[0010] Define the C 2 categories with the least number of samples in the data set as minority classes, and the remaining C 1 categories as the number of majority classes. Minority class label merging is to map the labels of the A - time enhanced samples of the minority classes back to the original label indices, so that the number of minority class labels is amplified to A + 1 times. While the majority class labels are retained, and the enhanced labels of the majority classes remain independent, forming C 1 × (A + 1) labels;
[0011] Let the original number of classes be C = C 1+C 2 , after A times of enhancement, the label space expands from C to C×(A + 1). Through mapping, the final label space is adjusted to: C final = C 1 ×(A + 1)+C 2 . At this time, the proportion of the minority class in the label space increases from to to achieve implicit balance.
[0012] S3. Joint label learning. Specifically, use the original labels of the majority class to train the original samples of the majority class, the enhanced labels of the majority class to train the enhanced samples of the majority class, and the original labels of the minority class to train the original samples and enhanced samples of the minority class;
[0013] Perform probability fusion on the majority class samples, and the model output is the probability weighted sum of the original label and all enhanced labels. Let the probability of sample x under the enhanced label k be Then its final classification probability is: For the minority class samples, directly use the probability of the merged original label for classification.
[0014] S4. Optimize the classification model through the cross-entropy loss function, and the loss function is defined as: where y i,j is the one-hot encoding of sample i on label j, and p i,j is the model prediction probability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the overall framework diagram of the method described in the present invention.
[0016] Figure 2 is the flowchart of the implementation steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] The present invention provides an imbalance time series classification enhancement method based on minority class label merging. To make the objectives, advantages and technical solutions of the embodiments of the present invention clearer, the following will be combined with Figure 1 the method framework diagram shown and Figure 2 the step flowchart shown to describe the technical solutions in the embodiments of the present invention completely and clearly.
[0018] S1. Perform sample enhancement on the data set;
[0019] Sample enhancement generates enhanced samples by transforming the original time series to expand the scale of the training set and improve the generalization ability of the model. Different sample enhancement methods can be used for sample enhancement. These include: Dropout, frequency domain Dropout, jitter, amplitude distortion, scaling, translation, time distortion, sample repetition.
[0020] S2. Perform label augmentation on the dataset;
[0021] In the label augmentation stage, different labels are assigned to the original samples and the augmented samples to make the classification boundary clearer and avoid introducing noise. During the sample augmentation process, the present invention expands the number of sample categories in the dataset to C×(A + 1) through label augmentation, including C original categories and C×A augmented categories. After label augmentation, a label mapping mechanism is used to perform different subsequent operations on the labels of the majority class and the minority class. Traditional data augmentation methods usually do not specifically handle labels but directly assign the same class labels to samples from the same source. This results in the number of class labels in the dataset remaining unchanged. When the data patterns of the original samples and the augmented samples are significantly different, this method will introduce noise, thus expanding the data distribution in the potential space of the model and affecting the classifier performance.
[0022] S3. Perform label mapping on the original labels and the augmented labels;
[0023] Label mapping is the process of converting the augmented labels into training labels. This involves merging the augmented minority class labels and then mapping them to indices, which are subsequently used as training labels. Suppose the total number of classes in the time series dataset is C = C 1 +C 2 , where C 1 represents the number of majority classes and C 2 represents the number of minority classes. After double augmentation, due to label augmentation, the original time series obtains additional majority classes and minority classes, and the number of classes remains C 1 and C 2 . The present invention performs minority class label merging on the augmented labels, merging the augmented labels of the minority classes into their original labels while keeping the majority class labels unchanged. After this process, there are three types of sample categories: the original minority class samples with the number of classes C 1 , the original majority class samples with the number of classes C 2 , and the majority class augmented samples with the number of classes A×C 1 . At the same time, the number of classes in the augmented sample space is reduced from C×(A + 1) to C 1 ×(A + 1)+C 2 . After merging the minority class labels, the number of minority class labels in the sample space becomes A + 1 times the original number. For the majority class, since it is divided into original labels and augmented labels, there are multiple different labels in the majority class sample space, and each label maintains its original number. This increases the proportion of the minority group relative to the majority group. Finally, the merged labels are mapped to a set of consecutive index values and used as training labels.
[0024] S4. Use the original labels and the augmented labels for joint label training.
[0025] After the double augmentation and label mapping operations, a new sample space and labels are obtained. The present invention does not expand the distribution of the majority class samples, but assigns independent labels to each augmented sample of the majority class, thus forming a compact clustering. Since the number of minority class samples is limited, after the label merging, the number of majority class samples in the sample space is significantly larger than that of the minority class. This avoids the problem of expanding the minority class samples, and at the same time, the deep learning model can capture more details of the minority class. The present invention performs A augmentations on the original time series. Therefore, from the perspective of self-supervised learning, the number of self-supervised classes is A. The core of the joint label learning is to extract information from the joint data distribution composed of the original labels and the self-supervised labels, thus avoiding the boundary expansion problem caused by augmentation. In the augmentation strategy proposed by the present invention, the joint label learning is only applicable to the majority class. By introducing the augmented labels, the majority class can learn multiple clusters, while the minority class benefits from the additional original labels to obtain more comprehensive information. The present invention will use a classifier (joint classifier) that learns with joint labels, and this classifier is trained on the augmented training set T aug which contains An time series samples and AC 1 +C 2 class labels. Let g(·; ω) denote the joint classifier, where ω represents the weights of the classifier. For the training time series t in T aug , the conditional distribution of each joint label can be defined as follows:
[0026] z = g(t; ω)
[0027] P(AC 1 +C 2 |t) = softmax(z)
[0028] where denotes the output vector generated by the joint classifier. P(AC 1 +C 2 |t) represents the conditional label distribution of the input time series t.
[0029] When predicting the category, it is only necessary to classify the time series into the original C categories, but the classifier is trained to classify into AC 1 +C 2 categories. To predict the original labels, the present invention sums the joint label probabilities of all categories belonging to the same original category. Given a time series x, then the conditional distribution of x over the AC 1 +C 2 category is:
[0030] p = softmax(g(x; ω))
[0031] where denotes the joint probability distribution of AC 1 + C 2 class labels, denotes the probability distribution of the majority class of AC 1 and denotes the probability distribution of the minority class of C 2 Meanwhile, g(x; ω) is a classifier parameterized by ω. The conditional distribution over each original class label is defined as:
[0032] p = (C|x) = {p 1 , …, p i , …, p C}
[0033] When distinguishing by majority and minority classes, it can be denoted as:
[0034]
[0035] where
[0036] When predicting the class, the classification results generated by different majority-class augmented labels are probability-weighted, while the minority class remains unchanged. The final classification result of the majority class is the sum of its original samples and all augmented samples. Probability weighting generates a compact distribution for the majority class, thus ensuring accurate classification of the majority class. Meanwhile, as the number of minority-class samples increases, this method can better distinguish the minority class.
Claims
1. A method for improving the classification of unbalanced time series based on merging minority class labels, characterized in that: The following steps are involved: S1. Use the time series perturbation method to enhance the samples of the original time series data set, generate enhanced time series samples, and obtain a larger data set; S2, enhance the original labels and assign new labels to the original samples and enhanced samples; S3, classify the enhanced labels according to the category imbalance ratio; S4. The classification model is trained based on the mapped labels. The weighted sum of the probability distribution of the original labels of the majority class and the enhanced labels is taken as the final classification result. The minority class is directly classified using the merged labels.
2. According to the method for improving the classification of unbalanced time series based on merging minority class labels according to claim 1, it is characterized in that: The sample enhancement in step S1 generates enhanced samples by repeating the original time series data, retaining the time-frequency characteristics of the original data, and the number of times of sample enhancement is 1 to 3 times.
3. According to the method for improving the classification of unbalanced time series based on merging minority class labels according to claim 1, it is characterized in that: The label enhancement in step S2 assigns an independent label to each original sample and its corresponding enhanced sample.
4. According to claim 1, the method for improving the classification of unbalanced time series based on merging minority class labels is characterized in that: Step S3 includes the following operations: (1) Calculate the category distribution of the statistical data set and define the C2 categories with the smallest sample size as the minority class and the rest as the majority class; (2) For minority class samples, their enhanced labels are merged into the original labels to form a unified minority class label; (3) For the majority class samples, retain their original labels and enhanced labels as independent categories; (4) Map the majority and minority class labels to training labels.
5. According to claim 1, the method for improving the classification of unbalanced time series based on merging minority class labels is characterized in that: Step S4 includes the following operations: (1) Use the joint classifier on the enhanced training set T aug Training is performed on , g(·; ω) represents the joint classifier, and ω represents the weight of the classifier. For the time series t, the conditional distribution of each joint label can be defined as follows: z=g(t;ω) P(AC1+C2|t)=softmax(z) Where z represents the output vector generated by the joint classifier. P represents the conditional label distribution of the input time series t; (2) For the majority class samples, the probability distribution output by the classification model is the weighted sum of the probabilities of its original label and all enhanced labels. Its calculation formula can be expressed as: (3) For minority class samples, the classification model directly outputs the combined original label probability, and its calculation formula can be expressed as: The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be within the protection scope of the present invention.