Multi-label time sequence classification method based on spatial-temporal characteristics and dynamic loss

Through a multi-task learning strategy that combines a multi-task learning strategy with multi-scale packet convolutional network, channel calibration mechanism and cascade feature multiplexing network combined with dynamic loss ratio weighting, the problems of insufficient feature extraction and category imbalance in multi-label timing classification are solved, and classification accuracy and recognition ability of a few categories are improved.

CN120579094APending Publication Date: 2025-09-02LUDONG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510740264.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the dynamic interaction between time series variables in multi-label timing classification, which has high computational cost of high-dimensional data and fixed convolution kernel parameters leading to information imbalance, and there is label imbalance in multi-label classification, which affects model performance.

Method used

By constructing a multi-scale packet convolutional network to extract timing features, combining channel calibration mechanism and cascade feature multiplexing network for spatiotemporal features fusion, and adopting a multi-task learning strategy with dynamic loss ratio weighting to solve the problem of category imbalance.

Benefits of technology

It significantly improves the classification accuracy of multi-label timing data, effectively solves the problem of label imbalance, improves the recognition ability of a few categories, and realizes efficient spatio-temporal feature extraction and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579094A_ABST
    Figure CN120579094A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-label time sequence classification method based on spatio-temporal characteristics and dynamic loss. The method comprises three core modules including a data preprocessing module, a spatio-temporal characteristic learning module and a multi-task optimization module. The data preprocessing module adopts filtering and standardization processing to improve the input quality; the spatio-temporal feature learning module firstly constructs a multi-scale packet convolutional network to extract the time sequence feature of each channel, then adaptively adjusts the channel weight by constructing a channel calibration mechanism to enhance key information, and then constructs a cascade feature multiplexing network to extract spatial correlation features to realize deep fusion of spatio-temporal features, so as to improve the time sequence feature of each channel. And finally, the extracted spatial-temporal features are used for classification. The multi-task optimization module establishes a dynamic loss proportion weighted multi-task learning strategy, updates historical loss through an index moving average mechanism, and solves the problem of class imbalance through loss inverse proportion weight. Experimental results show that the classification accuracy of the method on the multi-label time sequence data set is remarkably improved, and the method has high practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of time series analysis, and in particular to a multi-label time series classification method based on spatiotemporal features and dynamic loss. Background Art

[0002] Time series classification is an important research area in computer technology, widely used in fields such as physiological signal analysis, industrial equipment monitoring, and financial data forecasting. Multi-label time series classification is a key branch of time series classification, characterized by the fact that each sample may correspond to multiple class labels. Taking electrocardiograms as an example, each ECG may simultaneously contain multiple anomalies, such as atrial fibrillation and premature ventricular contractions. This requires classification models to not only identify multiple anomalies but also handle high-dimensional outputs and complex loss functions. The complexity of this problem places higher demands on classification methods.

[0003] Traditional methods employ a paradigm that separates feature engineering from classifiers. They extract features from the time-frequency domain, combine them with feature selection for dimensionality reduction, and then use classifiers such as support vector machines or random forests for multi-label classification. These methods suffer from two significant drawbacks: First, manually designed features struggle to capture the dynamic interactions between time series variables; second, when dealing with high-dimensional time series data, the computational cost is high, and it is difficult to improve overall classification performance through end-to-end optimization. Deep learning-based methods significantly improve the feature extraction process through end-to-end architectures. Hybrid networks based on convolutional neural networks and recurrent neural networks, such as the Transformer architecture, are typical solutions. Hybrid networks extract local spatial features through convolutional layers and capture temporal dependencies through recurrent layers, resulting in excellent performance when processing data with both local patterns and temporal dynamics. The Transformer architecture utilizes a self-attention mechanism to establish global temporal associations, overcoming the limitations of traditional recurrent networks in modeling long-range dependencies. Its parallel computing capabilities significantly improve training efficiency.

[0004] However, existing research still faces a series of technical bottlenecks in processing multi-label time series tasks. First, each variable in the time series represents a different data dimension, which has both certain independence and possible interdependence. Traditional methods usually assume that the relationship between variables is equal and fail to explicitly model this balance between independence and dependence. It is difficult to simultaneously capture the dynamics of the time series and the spatial topological structure, thus affecting model performance. Second, fixed convolution kernel parameters lead to uneven weight distribution of key feature channels. The information of certain channels in high-dimensional data may contribute little to the classification or even introduce noise, which has a negative impact on the results. Third, the problem of label imbalance is common in multi-label classification tasks. Especially in medical data, the number of samples of minority class labels is significantly insufficient, resulting in the model's weak learning ability for these categories. Traditional loss functions cannot effectively handle the imbalance between labels, making the classification performance of minority class labels difficult to meet actual needs. Breaking through these technical bottlenecks is crucial to improving the performance of multi-label time series classification. Summary of the Invention

[0005] To address the problems existing in the multi-label time series classification task, the present invention proposes a multi-label time series classification method based on spatiotemporal features and dynamic loss, which can accurately classify multi-label time series.

[0006] The present invention is achieved through the following technical solutions:

[0007] S1. Build a data preprocessing module to preprocess the original multi-label time series data, remove noise and dimensional differences in the data, build standardized input, and improve data quality;

[0008] S2. Construct a spatiotemporal feature learning module. For preprocessed data, first, a multi-scale grouped convolutional network is constructed to extract the temporal features of each channel, capturing feature changes at different time scales while retaining the independence of the temporal features within each channel. Then, based on the global information of the temporal features, a calibration matrix is ​​calculated to perform weighted calibration on the channel features, enhancing key temporal feature information and suppressing redundant information. A cascaded feature reuse network is then used to extract spatial correlation features, achieving dense fusion of spatiotemporal features and further enriching feature representation. Finally, the spatiotemporal features are input into the classification layer for dimensionality reduction and regularization, and multi-label categories are output through a fully connected layer and activation function.

[0009] S3. Build a multi-task optimization module. By establishing a multi-task learning strategy with dynamic loss ratio weighting and combining iterative optimization of model parameters with the back-propagation algorithm, we can solve the category imbalance problem in multi-label classification and improve the recognition ability of minority categories.

[0010] Furthermore, the step S1 specifically includes the following sub-steps:

[0011] S11. Read the time series feature dataset and perform length normalization. If the time series length is less than the preset length, zero padding is performed at the end of the sequence to make up the length. If the time series length exceeds the preset value, only the portion before the preset length is retained.

[0012] S12, using a fourth-order Butterworth bandpass filter to reduce noise on the raw data and dynamically adjust the cutoff frequency according to the different signal spectrum characteristics;

[0013] S13, performing Z-value standardization on the filtered data so that the data of each channel satisfies a normal distribution with a mean of 0 and a variance of 1;

[0014] S14, remove unlabeled samples and convert multi-label into one-hot encoding matrix;

[0015] S15. Convert the multi-label data into a one-hot encoding matrix and divide it into training set, validation set and test set.

[0016] Furthermore, the step S2 specifically includes the following sub-steps:

[0017] S21. Construct a multi-scale group convolutional network to extract temporal features:

[0018] (1) Automatically determine the number of groups G based on the number of channels C of the input data, so that G = C, and divide the time series data into C independent analysis subgroups according to the channel dimension, where each subgroup corresponds to the complete time series data of a single channel. The input is denoted as X, and its dimensions are (B, C, T), where B is the batch size, C is the number of channels, and T is the time step;

[0019] (2) Data for each subgroup , respectively input two parallel large kernel convolution branches, and their outputs are spliced ​​across branches in the channel dimension, recorded as :

[0020]

[0021]

[0022]

[0023] and is a convolution kernel with two different sizes 、 The calculated intermediate characteristic state, It is a splicing operation;

[0024] (3) The output of each subgroup is input into two new parallel large-kernel convolution branches again, and the outputs of the two branches are fused across branches twice in the channel dimension to further refine the feature information;

[0025] (4) The local features of the above outputs are optimized by sequentially using two convolutional small kernels, thereby generating the temporal features of each channel, which have multi-scale characteristics and channel independence;

[0026] The size of the large kernel convolution is in the range of 19×1 to 29×1, which is used to extract long-range trend features. The size of the small kernel convolution is in the range of 3×1 to 8×1, which is used to enhance instantaneous state recognition. The convolution kernel size is positively correlated with the length of the input sequence, and the optimal size of the convolution kernel can be dynamically determined according to the characteristics of the input time series data. Each subgroup shares the initialization weight, but the feature extraction process is carried out independently, and there is no cross-group information interaction between the subgroups.

[0027] S22. Build a channel calibration mechanism to adaptively adjust channel weights to enhance key information:

[0028] (1) Perform global average pooling on each channel of the temporal feature to generate a channel global vector :

[0029]

[0030] in Indicates the time series characteristics in The channel in The value of the time step, is the length of the time step;

[0031] (2) Based on the channel global vector, a bottleneck structure consisting of two layers of convolution kernel size 1×1 is constructed to learn the dependency between channels. The first layer of convolution compresses the number of channels to 1 / 3 of the original value. After applying the ReLU activation function to capture the nonlinear association, the second layer of convolution restores the original channel dimension, thereby generating the channel attention coefficient;

[0032] (3) The channel attention coefficient is activated by the Sigmoid function to generate a calibration matrix;

[0033] (4) Let the calibration matrix be , broadcast it to get the same dimension as the input, and perform element-by-element multiplication with the input time series features to enhance the time series features:

[0034]

[0035] S23. Construct a cascade feature reuse network to extract spatial correlation features and perform dense fusion of spatiotemporal features:

[0036] (1) Rearrange the dimensions of the time series features with (B, C, T) dimensions, where C is the channel dimension, T is the time dimension, and B is the batch dimension. Expand the channel dimension and the time dimension into a two-dimensional feature plane of (H, W), where H×W = C×T, and add a new channel dimension of size 1 to form a four-dimensional tensor of (B, 1, H, W) structure;

[0037] (2) Configure the first-level cascade feature reuse network for processing and extract shallow spatial correlation features. The structure contains 5 two-dimensional convolution units in series. Suppose the input , where the output of the intermediate layer l is It can be expressed as:

[0038] {H}_{l}=\sigma \left ( {{W}_{l}\cdot \left [ {{H}_{0},{H}_{1},...,{H}_{l-1}} \right ]+{b}_{l}} \right )

[0039] where \left [ {{H}_{0},{H}_{1},...,{H}_{l-1}} \right ] Represents the splicing operation of features, and Respectively The weights and biases of the layer convolution, It is a nonlinear activation function. The output of each two-dimensional convolutional unit is passed to all subsequent convolutional units through layer-by-layer connections, building a dense feedback path from the end to the initial layer. This allows the gradient to flow back through multiple paths during backpropagation, alleviating the gradient attenuation problem in deep networks. A downsampling layer with a stride of 2 and a kernel size of 2×2 is set at the end to reduce the spatial dimension of the features and reduce computational complexity.

[0040] (3) Construct a second-layer cascade feature reuse network to deepen the network depth and perform dense fusion of deep spatiotemporal features. The structure of the second-layer cascade feature reuse network is consistent with that of the first-layer cascade feature reuse network.

[0041] S24. Input the spatiotemporal features into the classification layer and output the multi-label prediction category:

[0042] (1) Perform global average pooling on the spatiotemporal features, aggregate contextual information along the H×W spatial dimension, and compress the four-dimensional tensor of dimension (B, C, H, W) to (B, C, 1, 1);

[0043] (2) Construct a random shielding layer of neurons with an inactivation rate of 0.5 to suppress overfitting during training;

[0044] (3) Reconstruct the four-dimensional tensor of dimension (B, C, 1, 1) into a two-dimensional tensor (B, C), implement a nonlinear transformation from feature space to label space through a fully connected layer, and output the predicted probability distribution of the number of categories;

[0045] (4) A configurable discrimination threshold T = 0.5 is used to implement a binary decision on the predicted probability. The corresponding label is activated only when the category probability exceeds the threshold, and finally a one-hot encoded category prediction result is generated.

[0046] Furthermore, the step S3 specifically includes the following sub-steps:

[0047] S31. During training, assume that the model outputs a predicted value z of dimension (B, T), where B is the batch size and T is the number of tasks (number of categories). The true label y is of dimension (B, T), representing the target value of each sample. First, the binary cross entropy loss function is used to calculate the binary cross entropy loss for each sample on each task:

[0048] {L}_{i,j}=-\left [ {{y}_{i,j}\cdot log\left ( {\sigma \left ( {{z}_{i,j}} \right )} \right )+(1-{y}_{i,j})\cdot log(1-\sigma ({z}_{i,j}))} \right ]

[0049] in is the Sigmoid function, Indicates the The sample in The true labels on the tasks, is the corresponding predicted value;

[0050] S32. Average the loss of each task in the batch dimension and calculate the average loss of all tasks:

[0051]

[0052] in Each element in represents the average loss value of a task;

[0053] S33. Construct an exponential moving average mechanism to dynamically update the historical loss of each task to smooth out loss fluctuations:

[0054]

[0055] in is the current time point, It is the current loss, is the attenuation factor;

[0056] S34. Calculate dynamic weights based on the inverse proportional relationship of historical losses, so that tasks with larger losses receive relatively smaller weights, and vice versa:

[0057]

[0058] in is a numerical stability constant with a value of 1e-8 to prevent division by zero errors;

[0059] S35. Normalize the dynamic weights to ensure that the sum of the task weights is 1 and maintain the stability of the weight distribution:

[0060]

[0061] S36. Calculate the weighted total loss based on the normalized dynamic weights:

[0062]

[0063] S37. After each training round, the model performance indicators are calculated on the validation set. Based on the calculated total loss, the parameter gradients are calculated using the backpropagation algorithm, and the model parameters are updated using the Adam optimizer.

[0064] Compared with the prior art, the present invention has the following advantages:

[0065] 1. The present invention combines a multi-scale group convolutional network, a channel calibration mechanism, a cascaded feature reuse network, and a dynamic loss ratio weighted multi-task learning strategy to achieve efficient and comprehensive spatiotemporal feature extraction, effectively solve the label imbalance problem, and significantly improve the classification accuracy of multi-label time series data.

[0066] 2. This paper constructs a multi-scale grouped convolutional network that divides the number of channels into multiple subgroups and performs multi-scale feature extraction on each subgroup. Subsequently, a channel calibration mechanism is used to automatically filter key features. This method can retain the unique temporal characteristics within each channel while effectively filtering out irrelevant or redundant information, thereby effectively learning temporal features.

[0067] 3. After learning the temporal features, the cascaded feature reuse network constructed by the present invention can explicitly capture the spatial correlation features between channels. The convolution kernel of its two-dimensional convolution unit can slide and learn the global correlation between channels, so as to better understand the interaction between channels. In addition, the network structure allows low-level features to be directly transmitted to deep layers, which can minimize the loss of internal information of a single channel and help to achieve a more refined fusion of spatiotemporal features.

[0068] 4. The dynamic loss ratio weighted multi-task learning strategy constructed by the present invention can systematically adjust the weights of each label task so that the model will not tend to learn the dominant category, but will learn all categories more evenly, improving the recognition ability of minority categories, thereby solving the problem of category imbalance in multi-label classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 :Overall flow chart of the proposed method

[0070] Figure 2 :Channel calibration mechanism structure diagram

[0071] Figure 3 : Cascade feature reuse network structure diagram DETAILED DESCRIPTION

[0072] In order to more intuitively demonstrate the application effects of the present invention in different scenarios, seven representative real-world time series data sets were selected for experimental verification. These data sets cover a variety of fields, including speech recognition, disease detection, network traffic monitoring, and traffic flow analysis, and can fully reflect the performance and advantages of the present invention in processing different types of time series data. The basic information of the data sets is shown in Table 1. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. The component configurations of the embodiments of the present invention generally shown here can be arranged according to the characteristics of the task.

[0073]

[0074] This paper proposes a multi-label time series classification method based on spatiotemporal features and dynamic loss, aiming to solve the problem of low classification accuracy caused by insufficient feature extraction and class imbalance in the existing technology when processing multi-label time series data. The overall flow chart of the method is shown in Figure 1 , which is achieved through the following steps:

[0075] S1. Build a data preprocessing module to preprocess the original data set, remove noise and dimensional differences in the data, and improve data quality. The specific steps are as follows:

[0076] S11. Since the time series lengths in different data sets are inconsistent, the time series lengths of all time series are first preset to 5000 time points. If the time series length is less than 5000 time points, the end of the sequence is padded with 0; if the time series length exceeds 5000 time points, only the first 5000 time points of the sequence are retained and the excess are discarded;

[0077] S12. Use a fourth-order Butterworth bandpass filter to perform noise reduction on the raw data. According to the spectral characteristics of the signal, the frequency range is set to 0.5 to 50 Hz to remove high-frequency noise and low-frequency interference and retain the effective components of the signal;

[0078] S13, performing Z-value normalization on the filtered electrocardiogram data so that the data of each channel satisfies a normal distribution with a mean of 0 and a variance of 1, thereby eliminating the dimensional differences between the data of different channels;

[0079] S14: Remove unlabeled samples, convert labels into one-hot encoded vectors, and randomly split the training, validation, and test sets in an 8:1:1 ratio to provide an appropriate sample distribution for model training and evaluation. The resulting data dimensions are (N, C, 5000), where N is the number of samples, 5000 is the time step, and C is the number of channels. Since the number of channels varies across datasets, the number of channels can be adjusted.

[0080] S2. Construct a spatiotemporal feature learning module. This module implements spatiotemporal feature learning by sequentially connecting a multi-scale grouped convolutional network, a channel calibration mechanism, and a cascaded feature reuse network. The specific steps are as follows:

[0081] S21. Construct a multi-scale group convolutional network to extract temporal features:

[0082] (1) Divide the output of the data preprocessing module into C subgroups according to the channel, so that each subgroup contains only the time series data of a single channel;

[0083] (2) For each subgroup, large kernel convolution with kernel size of 27×1 and 25×1 is used for processing, and the output is spliced ​​and fused across branches in the channel dimension. This process aims to capture the feature changes of each channel at different time scales and extract long-range trend features;

[0084] (3) The above output is processed again with large kernel convolution of kernel size 21×1 and 19×1, and the output is fused across branches twice in the channel dimension to further refine the feature information;

[0085] (4) The local features of the above outputs are optimized by sequentially using two concatenated small kernel convolutions with kernel sizes of 5×1 and 7×1 to obtain the temporal features of each channel.

[0086] S22. Build a channel calibration mechanism to perform weighted calibration on timing features, and adaptively adjust channel weights to enhance key timing feature information. The specific steps are as follows:

[0087] (1) If Figure 2As shown in Figure 1, firstly, global average pooling is performed on the temporal features to obtain the channel global vector of each channel, which is used to represent the overall features of each channel.

[0088] (2) Based on the channel global vector, a bottleneck structure consisting of two layers of convolution kernel size 1×1 is constructed to learn the dependency between channels. The first layer of convolution compresses the number of channels to 1 / 3 of the original value, and after applying the ReLU activation function to capture the nonlinear association, the second layer of convolution restores the original channel dimension, thereby generating the channel attention coefficient ;

[0089] (3) Perform Sigmoid activation on the attention coefficient to generate a calibration matrix;

[0090] (4) The calibration matrix is ​​broadcasted to obtain the same dimension as the input, and element-by-element multiplication is performed with the original time series features to enhance the time series features.

[0091] S23. Construct a cascade feature reuse network to extract spatial correlation features and perform dense fusion of spatiotemporal features. The specific steps are as follows:

[0092] (1) Rearrange the dimensions of the time series features with (B, C, T) dimensions, where C is the channel dimension, T is the time dimension, and B is the batch dimension. Expand the channel dimension and the time dimension into a two-dimensional feature plane of (H, W), where H×W = C×T, and add a new channel dimension of size 1 to form a four-dimensional tensor of (B, 1, H, W) structure;

[0093] (2) If Figure 3 As shown in the figure, the reorganized tensor extracts spatial correlation features through the first set of cascaded feature reuse networks. The cascaded feature reuse network contains 5 series-connected two-dimensional convolutional units. Each layer receives the output of all previous layers as input, thereby achieving feature reuse and direct propagation of gradients, capturing the spatial relationship and potential patterns between channels;

[0094] (3) A downsampling layer with a step size of 2 and a kernel size of 2×2 is set at the end of the cascade feature reuse network to reduce the spatial dimension of the feature and reduce the computational complexity;

[0095] (4) Then, the spatiotemporal features are further densely fused through the second set of cascaded feature reuse networks, which have the same configuration as the first set of cascaded feature reuse networks.

[0096] S24: Input the spatiotemporal features into the classification layer and output the multi-label prediction category. The specific steps are as follows:

[0097] (1) Perform global average pooling on the spatiotemporal features, aggregate contextual information along the H×W spatial dimension, and compress the four-dimensional tensor of dimension (B, C, H, W) to (B, C, 1, 1);

[0098] (2) Setting a random shielding layer of neurons with an inactivation rate of 0.5 to suppress overfitting during training;

[0099] (3) The four-dimensional tensor of dimension (B, C, 1, 1) is reconstructed into a two-dimensional tensor (B, C). The nonlinear transformation from feature space to label space is implemented through the fully connected layer, and the output dimension is the predicted probability distribution of the number of categories;

[0100] (4) A configurable discrimination threshold T = 0.5 is used to implement element-level binarization decision on the predicted probability vector. The corresponding label is activated when and only when the category probability exceeds the threshold, and finally a one-hot encoded category prediction result is generated.

[0101] S3. Build a multi-task optimization module and use a dynamic loss ratio weighted multi-task learning strategy to train the network, solve the category imbalance problem in multi-label classification, and improve the recognition ability of minority categories. The specific steps are as follows:

[0102] S31. During training, set the batch size to 16 and use the binary cross entropy loss function to calculate the binary cross entropy loss of each sample on each task to quantify the prediction error of the model on each task.

[0103] S32. Average the losses of each task in the batch dimension and calculate the average loss of all tasks;

[0104] S33. Construct an exponential moving average mechanism to dynamically update the historical loss of each task to smooth the loss fluctuations during training;

[0105] S34. Calculate dynamic weights based on the inverse proportional relationship of historical losses, so that high-loss tasks receive more attention and vice versa.

[0106] S35. Normalize the dynamic weights to ensure that the sum of the weights of each task is 1 and maintain the stability of the weight distribution;

[0107] S36. Calculate the weighted total loss based on the normalized dynamic weights, use the backpropagation algorithm to calculate the parameter gradients, and use the Adam optimizer to update the model parameters. Experimental results

[0108] The experiment systematically traversed multiple hyperparameter combinations using a grid search method and ultimately determined the optimal hyperparameter configuration. The detailed hyperparameter configuration is shown in Table 2. The entire training process was completed on an NVIDIA Tesla V100 GPU, and all code was implemented in Python and the PyTorch framework.

[0109]

[0110] The experimental evaluation metrics used were accuracy (ACC), area under the receiver operating characteristic curve (AUC), precision (Pr), recall (Re), and F1 score (F1). Detailed experimental results are summarized in Table 3. The proposed method performed well on multiple datasets, achieving high levels of AUC, ACC, and F1. For example, the ArabicDigits dataset achieved an AUC of 100 and an ACC of 99.93; the UWave dataset achieved an AUC of 99.79 and an ACC of 99.22. Even on complex datasets such as PTB-XL, the method maintained good performance in some categories, achieving an AUC of 98.70 and an ACC of 97.05 for rhythm. Overall, the proposed method demonstrated high accuracy and stability, effectively demonstrating its versatility across a wide range of time series domains.

[0111]

[0112] The present invention was compared with existing methods in five classification tasks on the PTB-XL dataset, with experimental results shown in Table 4. In the diagnosis superclass classification task, the present invention achieved an AUC of 93.14% and an ACC of 88.85%, respectively, improving by 0.98% and 0.59% over the state-of-the-art ECGTransForm. In the diagnosis subclass classification task, the AUC of 92.86% surpassed ECGNet's 92.83%, while the ACC reached a leading level of 96.51%. In particular, the present invention achieved breakthrough progress in the form classification task, with an AUC of 88.10% and an ACC of 94.82% improving on ResNet-101 by 4.06% and 3.76%, respectively, demonstrating superior recognition of signal morphological features. In rhythm classification, the present invention achieved an AUC of 98.70%, a 3.34 percentage point improvement over the next-best method, ECGNet, and an ACC of 97.05%, the current state-of-the-art. A comprehensive comparison of the performance of various methods across five task categories revealed that the proposed method achieved the best results across four AUC metrics and three ACC metrics. Specifically, in the clinically important dimensions of form and rhythm, the proposed method achieved an average improvement of 3.7% in AUC and 3.2% in ACC over existing methods, demonstrating its superior accuracy in multi-label time series classification.

[0113] Ablation experiments

[0114] To validate the contributions of each core module to the final performance, we designed and conducted systematic ablation experiments on the PTB-XL multi-label ECG dataset. Specifically, we first constructed a base model consisting solely of a multi-scale grouped convolutional network, designated Method A. This model was then supplemented with a channel calibration mechanism to form Method B. A cascaded feature reuse network was then added to form Method C. Finally, Method C was integrated with a multi-task learning strategy weighted by dynamic loss scaling to form the complete Method D. The experimental results are shown in Table 5. With the gradual introduction of each technical component, the model demonstrated stable and significant performance improvements across all classification tasks. In the diagnosis superclass classification task, the AUC and ACC indicators gradually improved from 90.23% and 86.50% of method A to 93.14% and 88.85% of method D; in the diagnosis subclass task, the AUC and ACC increased from 89.12% and 95.20% to 92.86% and 96.51% respectively; in the diagnosis class task, the AUC increased from 84.56% to 88.73%, and the ACC increased from 96.50% to 97.20%; the AUC of the form class and rhythm class increased from 80.10% and 92.30% to 88.10% and 98.70% respectively, and the ACC also increased from 89.20% and 94.50% to 94.82% and 97.05%. These results fully demonstrate the gain effect of each module. The channel calibration mechanism enhances the model's ability to focus on key temporal features, the cascaded feature reuse network further explores the spatial dependencies between channels, and the dynamic loss weighting strategy significantly alleviates the category imbalance problem in multi-label learning and improves the recognition ability of minority class labels.

[0115] In summary, the multi-label time series classification method based on spatiotemporal features and dynamic loss proposed in the present invention effectively solves the problems of insufficient feature extraction and category imbalance in the existing technology in multi-label time series classification tasks, significantly improves the classification accuracy and generalization ability, and provides an efficient and reliable solution for practical applications in related fields. It has important scientific value and broad application prospects.

[0116] References

[0117] [1] Liu K, Liu T, Wen D, et al. SRTNet: Scanning, Reading, andThinking Network for myocardial infarction detection and localization [J].Expert Systems with Applications, 2024, 240.

[0118] [2] El-Ghaish H, Eldele E. ECGTransForm: Empowering adaptive ECGarrhythmia classification framework with bidirectional transformer [J].Biomedical Signal Processing and Control, 2024, 89.

[0119] [3] Murugesan B, Ravichandran V, Ram K, et al. Ecgnet: Deep networkfor arrhythmia classification; proceedings of the 2018 IEEE InternationalSymposium on Medical Measurements and Applications (MeMeA), F, 2018 [C].IEEE.

[0120] [4] He T, Zhang Z, Zhang H, et al. Bag of tricks for imageclassification with convolutional neural networks; proceedings of theProceedings of the IEEE / CVF conference on computer vision and patternrecognition, F, 2019 [C]。

Claims

1. A multi-label time series classification method based on spatiotemporal features and dynamic loss, characterized by: The following steps are involved: S1. Build a data preprocessing module to perform data preprocessing, including: S11. Read the time series data set and perform length normalization. If the time series length is less than a preset length, zero padding is performed at the end of the sequence to make up the length. If the time series length exceeds a preset value, only the portion before the preset length is retained. S12, using a fourth-order Butterworth bandpass filter for denoising; S13, performing Z value normalization processing on the filtered data in the channel dimension; S14, convert the multi-label into a one-hot encoding matrix and divide it into training set, validation set and test set; S2. Build a spatiotemporal feature learning module, which implements spatiotemporal feature learning by sequentially connecting a multi-scale grouped convolutional network, a channel calibration mechanism, and a cascaded feature reuse network, including: S21. Construct a multi-scale group convolutional network to extract temporal features: (1) Automatically determine the number of groups G based on the number of channels C of the input data, so that G = C, and divide the time series data into C independent analysis subgroups according to the channel dimension, where each subgroup corresponds to the complete time series data of a single channel; (2) The data of each subgroup is input into two parallel large-kernel convolution branches, and their outputs are spliced ​​and fused across branches in the channel dimension; (3) The output of each subgroup is input into two new parallel large-kernel convolution branches again, and the outputs of the two branches are concatenated and fused twice across branches in the channel dimension; (4) Two series-connected small-kernel convolutions are used to optimize the local features of the output after the secondary cross-branch splicing and fusion to obtain the temporal features of each channel; S22. Build a channel calibration mechanism to adaptively adjust channel weights to enhance key information: (1) Perform global average pooling on each channel of the temporal feature to generate a channel global vector; (2) Based on the channel global vector, a bottleneck structure consisting of two layers of convolution is constructed to learn the inter-channel dependency. The first layer of convolution compresses the number of channels to 1 / 3 of the original value. After applying the ReLU activation function to capture the nonlinear association, the second layer of convolution restores the original channel dimension, thereby generating the channel attention coefficient. (3) The channel attention coefficient is activated by the Sigmoid function to generate a calibration matrix; (4) Perform element-by-element multiplication of the calibration matrix and the original time series features to enhance the time series features; S23. Construct a cascade feature reuse network to extract spatial correlation features and perform dense fusion of spatiotemporal features: (1) Rearrange the dimensions of the time series features with (B, C, T) dimensions, where C is the channel dimension, T is the time dimension, and B is the batch dimension. Expand the channel dimension and the time dimension into a two-dimensional feature plane of (H, W), where H×W = C×T, and add a new channel dimension of size 1 to form a four-dimensional tensor of (B, 1, H, W) structure; (2) Constructing a first-level cascade feature reuse network to extract shallow spatial correlation features; (3) Constructing a second-layer cascade feature reuse network to deepen the network depth and perform dense fusion of deep spatiotemporal features; S24. Input the spatiotemporal features into the classification layer and output the multi-label prediction category: (1) Perform global average pooling on the spatiotemporal features, aggregate contextual information along the H×W spatial dimension, and compress the four-dimensional tensor of dimension (B, C, H, W) to (B, C, 1, 1); (2) Constructing a random shielding layer of neurons to suppress overfitting during training; (3) Reconstruct the four-dimensional tensor of dimension (B, C, 1, 1) into a two-dimensional tensor (B, C), implement a nonlinear transformation from feature space to label space through a fully connected layer, and output the predicted probability distribution of the number of categories; (4) Configure the discrimination threshold and implement a binary decision on the predicted probability. The corresponding label is activated only when the class probability exceeds the threshold, and finally a one-hot encoded class prediction result is generated. S3. Build a multi-task optimization module to solve the class imbalance problem by constructing a multi-task learning strategy with dynamic loss ratio weighting and improve the ability to identify minority classes, including: S31. Calculate the binary cross entropy loss for each task to quantify the prediction error of the model on each task. S32. Average the losses in the batch dimension and calculate the average loss of each task; S33, using the exponential moving average mechanism to update the historical loss to alleviate the sharp fluctuations in task loss during training; S34. Calculate dynamic weights based on the inverse proportional relationship of historical losses, so that tasks with larger losses receive relatively smaller weights, and vice versa. S35. Normalize the dynamic weights and calculate the weighted total loss; S36. Use this weighted total loss for backpropagation to update the network parameters.

2. The multi-label time series classification method based on spatiotemporal features and dynamic loss according to claim 1 is characterized in that: In step S21, both the large kernel convolution and the small kernel convolution are one-dimensional, and the convolution kernel size is positively correlated with the length of the input sequence; each subgroup shares the initialization weight, but the feature extraction process is performed independently, and no cross-group information exchange is performed between the subgroups.

3. The multi-label time series classification method based on spatiotemporal features and dynamic loss according to claim 1 is characterized in that: The cascaded feature reuse network in step S23 includes multiple two-dimensional convolutional units in series, wherein the output of each two-dimensional convolutional unit is transmitted to all subsequent convolutional units in a layer-by-layer connection manner in the channel dimension, thereby constructing a dense feedback path from the end to the initial layer, so that the gradient can flow back through multiple paths during back propagation; and downsampling layers are respectively configured at the end of the network to implement spatial dimension downsampling of the feature map.

Citation Information

Cited By

  • Multi-domain feature fusion network and bearing fault diagnosis method and system

    CN121278492A