Gtn time series classification method fused with mamba module

By integrating the GTN time series classification method with the Mamba module, the MAGTN model was constructed, which solved the problem of integrating multi-source heterogeneous data, enhanced the feature extraction and long-term dependency capture of multivariate time series, and improved the classification accuracy and recall rate of the earthquake early warning system.

CN120541608BActive Publication Date: 2025-11-11TAIYUAN UNIVERSITY OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510637657.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-11-11
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Existing earthquake early warning systems struggle to effectively integrate multi-source heterogeneous data, fail to capture the nonlinear spatial correlation characteristics between multi-sensor time series data, lack the ability to capture weak, low-frequency signals of long-period geological activity during earthquake gestation, and traditional methods neglect the correlation between multivariate time series data, resulting in insufficient ability of models to capture short-term fluctuations and changes in the sequence.

Method used

We adopt the GTN time series classification method that integrates Mamba modules. By constructing a MAGTN model that coordinates the attention mechanism with the Mamba module, and combining hierarchical attention mechanism and depthwise separable convolution, we can capture multi-scale features and long-term dependencies. We use the AdamW optimizer to guide training and improve the feature extraction and training stability of the model.

Benefits of technology

It improves the classification accuracy and recall of multivariate time series data, achieves efficient processing of complex time series data, and enhances the model's classification performance on multivariate time series datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541608B_ABST
    Figure CN120541608B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of deep learning technology, specifically relating to a GTN time series classification method integrating Mamba modules. The method includes the following steps: Based on the GTN model, a time series classification model MAGTN is constructed using an attention mechanism and a collaborative Mamba module; a hierarchical attention mechanism is employed to capture multi-scale features in the time series; causal convolutions are performed using depthwise separable convolutions, and dilated convolutions are used to fully capture long-term dependencies in the time series, combined with normalization layers and residual connections to improve training stability; the AdamW optimizer is used to guide model training and improve convergence speed. This invention, based on the GTN model, constructs a time series classification model that integrates a hierarchical attention mechanism and Mamba modules, and uses convolutional layers for local convolution operations, enabling the model to classify high-dimensional multivariate time series data, indirectly improving the efficiency of the GTN model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning technology, specifically relating to a GTN time series classification method that integrates the Mamba module. Background Technology

[0002] In the field of earthquake monitoring, forecasting, and early warning, multivariate time series data exhibit extremely complex dynamic characteristics. Traditional earthquake early warning systems face the challenge of effectively integrating multi-source heterogeneous data, such as seismograph waveform data, crustal deformation monitoring data, underground fluid monitoring data, and electromagnetic monitoring data recorded by earthquake monitoring networks, among other multi-dimensional time series data. Existing models based on fixed thresholds or single neural networks have high difficulty in identifying anomalies in multi-dimensional earthquake precursor signals, often exhibiting the following shortcomings: 1) failure to effectively capture the nonlinear spatial correlation characteristics between multi-sensor time series data; 2) lack of continuous ability to capture weak low-frequency signals formed by long-period geological activities during earthquake gestation, leading to difficulty in identifying anomalous data.

[0003] These shortcomings affect the model's performance and effectiveness. Furthermore, multivariate time series data typically contains multiple related sensors or variables with strong correlations. Traditional methods often neglect these correlations during feature extraction, potentially inputting redundant and repetitive information as independent features into the model. This not only increases computational complexity but can also lead to overfitting. Local temporal dependencies often exist between time steps in multivariate time series data, particularly short-term dependencies, which may be more important than long-term dependencies. However, many traditional feature extraction methods (e.g., convolutional or simple recurrent neural networks) perform poorly in modeling these local dependencies, often focusing too much on global features while ignoring short-term local changes and patterns in the sequence. This results in the model's inadequate ability to capture short-term fluctuations and changes in the sequence. Summary of the Invention

[0004] To address the technical problem that existing models are insufficient in capturing short-term fluctuations and changes in sequences, this invention provides a GTN time series classification method that integrates the Mamba module.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] The GTN time series classification method incorporating the Mamba module includes the following steps:

[0007] S1. Based on the GTN model, construct a time series classification model MAGTN that coordinates the attention mechanism and the Mamba module;

[0008] S2. Employ a hierarchical attention mechanism to capture multi-scale features in time series and enhance the model's ability to model local and global time dependencies.

[0009] S3. Use depthwise separable convolution for causal convolution, utilize dilated convolution to fully capture long-term dependencies in time series, and combine normalization layers and residual connections to improve training stability.

[0010] S4. The AdamW optimizer is used to guide model training, which improves the convergence speed and optimizes the model performance while avoiding gradient explosion and gradient vanishing.

[0011] The main framework of the GTN model in S1 is a GTN structure, which adopts a dual-branch model to process the temporal and spatial features in the time series simultaneously. The results are then fused through the Gated mechanism and the classification results are output. To enhance the feature extraction capability of the model, a processing mechanism that coordinates the Mamba module and the attention mechanism is constructed. Through multi-scale information extraction and feature fusion, the model's ability to process complex time series data is strengthened.

[0012] The method for constructing the collaborative processing mechanism of the Mamba module and the attention mechanism in S1 is as follows: the input multivariate time series data is processed by the Mamba module before entering the Encoder, and then processed by the attention mechanism module. The weighted output features of the attention mechanism are then input into the Mamba module. Since spatial features are usually less complex than temporal features, it is sufficient to add the Mamba module directly after the attention mechanism to capture spatial dependencies. To reduce the computational burden, the Mamba module is only added at the Encoder of the spatial channel.

[0013] The method for capturing multi-scale features in time series in S2 is as follows:

[0014] Information is extracted from the input data step by step through multi-layer attention. The attention mechanism of each layer is weighted and aggregated according to the output of the previous layer. Let the input time series be X, the input X is mapped to multiple query Q matrix, key K matrix and value V matrix. The similarity between the query Q matrix and key K matrix is ​​calculated by dot product to obtain the attention weight score. The value V matrix is ​​weighted and summed to obtain the attention output. The input feature x is processed by multi-head attention in each layer and x is passed to the next layer. After each attention mechanism except the last layer, a linear layer and the ReLU activation function are applied to transform the output.

[0015] The method of applying a linear layer and the ReLU activation function to transform the output is as follows:

[0016] X (l+1) =Dropout(σ(W)(l) X (l) +MHA(X (l) )),l=1,…,L-1

[0017] Among them, X (l) This represents the input features of the l-th layer, initially the original input X; MHA() represents multi-head self-attention computation, W (l) It is the linear transformation matrix between layers, σ represents the ReLU activation function, and Dropout is used to prevent overfitting.

[0018] The method in S3 for fully capturing the long-term dependencies of time series using dilated convolution is as follows: after inputting time series data, the shape of the data is adjusted, causal convolution is performed through depthwise separable convolution to extract local features in the time series, and channel information is further fused by pointwise convolution. Residual connections and normalization operations are used to maintain the stability of information flow. Then, the GELU activation function is applied to increase the nonlinearity of the network, enabling the model to learn complex time patterns more effectively. Finally, regularization is performed through random deactivation layers.

[0019] The method for using the AdamW optimizer to guide model training in S4 is as follows:

[0020] AdamW combines the Adam algorithm with weight decay. The Adam algorithm calculates the first and second moments of the gradient and adaptively adjusts the learning rate of each parameter to enable the model to converge quickly in the early stages of training and remain stable in the later stages. Since the model integrates multiple modules and has a complex structure, when the AdamW optimizer guides the model training, it helps the model train better by leveraging the adaptive learning rate and weight decay characteristics, avoiding the optimization problems common in complex networks.

[0021] Compared with the prior art, the beneficial effects of this invention are:

[0022] This invention proposes a Time Series Classification Method (MAGTN) that integrates the Mamba module. It combines the advantages of Transformer and Mamba, addressing issues such as redundant feature selection and insufficient local dependency modeling in time series classification to some extent. On 10 publicly available multivariate time series datasets, this invention achieves an average classification accuracy of 95.09%, a macro-average recall of 0.9445, a macro-average precision of 0.9517, and an F1 score of 0.9458, outperforming most models across all four evaluation metrics. Attached Figure Description

[0023] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0024] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0025] Figure 1 This is a diagram of the overall architecture of the MAGTN of the present invention;

[0026] Figure 2 This is a structural diagram of hierarchical attention in the MAGTN model of this invention;

[0027] Figure 3 This is a structural diagram of depthwise separable convolution in the MAGTN model of this invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. These descriptions are only for further illustrating the features and advantages of the present invention, and not for limiting the claims of the present invention. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0030] This embodiment is implemented using the PyTorch deep learning framework. It provides a GTN time series classification method incorporating the Mamba module, specifically including the following steps:

[0031] 1. Data Preparation

[0032] The data samples in this embodiment come from a dataset compiled by Mustafa Baydogan, one of the most widely used datasets for multivariate time series classification. Ten datasets were selected for experiments. These datasets have different categories, lengths, and number of channels, covering different fields and application scenarios, such as human activity recognition and health monitoring. Most datasets contain multidimensional time series data collected from sensors (such as accelerometers, gyroscopes, heart rate sensors, etc.) and are designed to train and evaluate different machine learning and deep learning models through classification tasks on this data. These datasets include sequence data at multiple time steps, each time step recording multidimensional features from different sensors, used to capture and analyze dynamic change patterns in the time series.

[0033] 2. Model Building

[0034] The main framework of the constructed MAGTN model is a GTN structure, and the specific network structure is as follows: Figure 1 As shown, the model processes the temporal and spatial features of the time series using two different encoder modules. The input data is mapped to a unified feature dimension through a linear layer. Then, local features from the time series are extracted using depthwise separable convolutional layers. The specific network structure of the depthwise separable convolutional layer is shown below. Figure 2 As shown in the diagram. Then, a Mamba module was added to the time channel to further optimize the representation of time series features. The entire process includes multiple encoder layers to extract global temporal dependencies and handle long-short-term dependencies in the time series. Finally, the model fuses features from different encoders through a gating mechanism to generate the final output. The encoder layer first processes the data through a hierarchical attention module to capture global dependencies in the time series. The specific network structure of the hierarchical attention module is shown below. Figure 3 As shown. Next, the feature representation is further optimized using the Mamba module to enhance the model's feature extraction capability. Residual connections and layer normalization are applied after each submodule. Subsequently, the data is further processed through a feedforward neural network to enhance the model's nonlinear transformation capability, and finally, the normalized features are output.

[0035] 3. Model Training

[0036] In the MAGTN network model constructed using the training set, cross-entropy loss is used to calculate the average error between correct and incorrect classifications, measuring the difference between the model's classification probability distribution and the true labels. The loss formula is defined as follows:

[0037]

[0038] Where N represents the number of categories in the sample. i It is a one-hot encoded representation of the true distribution. It is the probability distribution of the model output, that is, the predicted value after Softmax normalization.

[0039] 4. Test Results

[0040] The training process mainly involves loading multiple datasets and using MAGTN for multivariate time series classification. Each dataset is first split into training and test sets, and the data is loaded in batches using DataLoader. During training, the AdamW optimizer and cross-entropy loss function are used, and the model parameters are updated through backpropagation. The number of training iterations is fixed at 150, the batch size is 3, and the learning rate is fixed at 1e-4.

[0041] 5. Model Evaluation

[0042] Accuracy, recall, precision, and F1 score are calculated using the classification results and the true labels to evaluate the model's performance.

[0043] Accuracy is one of the most basic classification metrics, representing the proportion of samples correctly predicted by the model out of the total sample size. It is calculated by dividing the number of correctly predicted samples by the total number of samples. The formula is as follows:

[0044]

[0045] Specifically, TP represents the number of samples correctly classified as positive. TN represents the number of samples correctly classified as negative. FP represents the number of negative samples misclassified as positive. FN represents the number of positive samples misclassified as negative.

[0046] Macro-average recall is calculated by averaging the recall rates of all classes after calculating the recall rate for each class individually. Recall measures the comprehensiveness of the model for each class, that is, the proportion of samples that are actually positive that the model correctly predicts as positive. The formula is as follows:

[0047]

[0048] Where C is the total number of categories, TP i and FN i These are the true examples and false negative examples of the i-th class, respectively.

[0049] Macro-average precision is calculated by averaging the precision of all classes after calculating the precision for each class individually. Precision measures the accuracy of the model's predictions for each class; that is, the proportion of samples predicted as positive that are actually positive. The formula is as follows:

[0050]

[0051] Where C is the total number of categories, TP iand FP i These are the true examples and false negative examples of the i-th class, respectively.

[0052] The F1 score is the harmonic mean of precision and recall, providing a trade-off between the two. A higher F1 score indicates that the model performs well in both precision and recall when handling positive and negative samples. The formula is as follows:

[0053]

[0054] Precision represents the proportion of samples predicted as positive that are actually positive, while Recall represents the proportion of samples that were actually positive that were correctly predicted as positive by the model.

[0055] Table 1. Test results of accuracy of different models on the dataset.

[0056]

[0057] Table 2. Test results of macro-average recall for different models on the dataset.

[0058]

[0059] Table 3. Test results of macro-average precision of different models on the dataset.

[0060]

[0061] Table 4. Test results of F1 scores for different models on the dataset.

[0062]

[0063] As shown in Tables 1-4, the evaluation of multivariate time series classification was conducted on ten datasets. Precision, recall, accuracy, and F1 score were calculated using the predicted results and the true labels. The results were compared with six different classification models: encoder, CNN, MLP, TLENet, MCCDCNN, and GTN. The best metric in the table is highlighted in bold.

[0064] The MAGTN in this embodiment performs well on most datasets, especially on datasets such as AUSLAN, CMUsubject16, CharacterTrajectories, JapaneseVowels, and KickVsPunch, where it significantly outperforms other models.

[0065] The above description only illustrates the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A GTN time series classification method integrating Mamba modules, characterized in that, Includes the following steps: S1. Based on the GTN model, a time series classification model MAGTN is constructed that integrates an attention mechanism and a Mamba module. The main framework of the GTN model in S1 is a GTN structure, employing a dual-branch model to simultaneously process temporal and spatial features in the time series. To enhance the model's feature extraction capability, a collaborative processing mechanism between the Mamba module and the attention mechanism is constructed. Through multi-scale information extraction and feature fusion, the model's ability to process complex time series data is strengthened. The method for constructing the collaborative processing mechanism between the Mamba module and the attention mechanism in S1 is as follows: the input multivariate time series data is processed by the Mamba module before entering the Encoder, and then processed by the attention mechanism module. The weighted output features of the attention mechanism are then input into the Mamba module. To reduce computational burden, the Mamba module is only added to the Encoder in the spatial channel. S2. A hierarchical attention mechanism is adopted to capture multi-scale features in time series and enhance the model's ability to model local and global time dependencies. Information in the input data is extracted step by step through multi-layer attention. Each layer's attention mechanism is weighted and aggregated based on the output of the previous layer. Let the input time series be X, and map the input X to multiple query Q matrices, key K matrices, and value V matrices. The similarity between the query Q matrix and the key K matrix is ​​calculated by dot product to obtain the attention weight score. The value V matrix is ​​weighted and summed to obtain the attention output. The input feature x is processed by the multi-head attention mechanism of each layer and is passed to the next layer. After each attention mechanism except the last layer, a linear layer and the ReLU activation function are applied to transform the output. S3. Use depthwise separable convolution for causal convolution, utilize dilated convolution to fully capture long-term dependencies in time series, and combine normalization layers and residual connections to improve training stability. S4. The AdamW optimizer is used to guide model training, which improves the convergence speed and optimizes the model performance while avoiding gradient explosion and gradient vanishing.

2. The GTN time series classification method incorporating the Mamba module according to claim 1, characterized in that, The method of applying a linear layer and the ReLU activation function to transform the output is as follows: Among them, X (l) The input features of the l-th layer are represented by the original input X; MHA() represents multi-head self-attention computation, W (l) It is the linear transformation matrix between layers. σ This represents the ReLU activation function, and Dropout is used to prevent overfitting.

3. The GTN time series classification method integrating Mamba modules according to claim 1, characterized in that, The method in S3 for fully capturing the long-term dependencies of time series using dilated convolution is as follows: after inputting time series data, the shape of the data is adjusted, causal convolution is performed through depthwise separable convolution to extract local features in the time series, and channel information is further fused by pointwise convolution. Residual connections and normalization operations are used to maintain the stability of information flow. Then, the GELU activation function is applied to increase the nonlinearity of the network, enabling the model to learn complex time patterns more effectively. Finally, regularization is performed through random deactivation layers.

4. The GTN time series classification method integrating the Mamba module according to claim 1, characterized in that, The method for using the AdamW optimizer to guide model training in S4 is as follows: AdamW combines the Adam algorithm with weight decay. The Adam algorithm calculates the first and second moments of the gradient and adaptively adjusts the learning rate of each parameter to enable the model to converge quickly in the early stages of training and remain stable in the later stages. Since the model integrates multiple modules and has a complex structure, when the AdamW optimizer guides the model training, it helps the model train better by leveraging the adaptive learning rate and weight decay characteristics, avoiding the optimization problems common in complex networks.

Citation Information

Patent Citations

  • A convolutional echo state network time sequence classification method based on a multi-head self-attention mechanism

    CN109919205A

  • Multi-modal attention deep learning method for enhancing virus recognition in metagenome data

    CN118918954A