Multivariate time sequence classification method and device fusing convolutional neural network and Transform
By combining the convolutional neural network and the Transformer model, the multi-scale convolution and channel fusion convolution modules are used to extract local features, and the multi-head attention modules capture global features in combination with feature dimensions and time dimensions, the shortcomings of local and global feature extraction in the multivariate time series classification are solved, and more efficient classification performance is achieved.
Patent Information
- Application Number
- CN202510766377.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing time series classification methods are difficult to effectively capture local and global features in multivariate time series data, especially traditional CNN models have shortcomings in local feature extraction, while Transformer models may weaken time dependence during global feature capture.
Fusion convolutional neural network and Transformer model, local features are extracted through multi-scale convolution modules and channel fusion convolution modules, and global features are captured through feature dimension multi-head attention and time dimension multi-head attention modules, combining learnable position coding to enhance time information processing.
It significantly improves the accuracy of multivariate time series classification, can better adapt to complex time series data, overcomes the limitations of single model in feature extraction and feature modeling, and improves the model's expression ability and classification performance.
Smart Images

Figure CN120336972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series data processing, and more particularly to a multivariate time series classification method and device that integrates a convolutional neural network and a Transformer. Background Art
[0002] In real life, time series data exists widely, such as network traffic, electrocardiograms, sound data, and sensor data. Time series analysis studies these data that change over time and is widely used in multiple fields such as activity recognition, electrocardiogram classification, and network anomaly detection. According to the dimension of the data, time series can be divided into two categories: univariate and multivariate. In practical applications, multivariate time series are more common because they can simultaneously reflect the mutual relationship between multiple variables and the complexity of system behavior. Time series classification (TSC) is a core task in time series analysis, aiming to classify according to the change pattern of the data. Multivariate time series classification (MTSC) is more complex because it not only has to process the characteristics of each variable changing over time but also needs to consider the mutual dependence between multiple variables. Since multivariate time series data usually exhibits complex spatio-temporal dependencies, traditional classification methods often have difficulty effectively capturing this information. Therefore, the research on MTSC is particularly important, especially in fields such as intelligent healthcare, financial forecasting, and industrial monitoring, where it can provide more accurate data analysis and decision support, promoting the development of various industries.
[0003] In the MTSC task, existing methods are mainly divided into two categories: one is the method based on traditional machine learning, and the other is the method based on deep learning. The former relies on manual feature engineering and domain knowledge, and improves the classification performance by artificially designing features. However, it has limited effects when dealing with high-dimensional and complex data sets and is difficult to capture the deep structural features of the data. The latter automatically learns multi-level feature representations from the original time series through neural networks, reduces the dependence on manual feature engineering, and can better adapt to complex and high-dimensional data, especially showing more superiority in large-scale complex data scenarios. In the MTSC method based on deep learning, the convolutional neural network (CNN) is widely used because it can effectively extract local features of time series through local convolution operations. Models represented by the Inception network further improve the understanding and classification performance of complex patterns by using multi-scale convolutional kernels in parallel. However, due to the limitations of convolution operations, CNN-based models have deficiencies in capturing global features. In contrast, the Transformer model can effectively capture global features in time series through the attention mechanism, but its processing method of time information (such as position encoding) may weaken the temporal dependence of time series. In addition, existing Transformer models usually only focus on a single dimension in feature extraction in the time dimension and channel dimension, and it is difficult to efficiently capture both local and global features of time series simultaneously. These limitations provide directions for further optimization of model design in the MTSC task. Summary of the Invention
[0004] The object of the present invention is to propose a multi-variate time series classification method and device that combines a convolutional neural network and a Transformer in view of the limitations in local feature extraction and global feature capture in the prior art. The present invention first applies multi-scale two-dimensional convolution to each channel of the input time series sample through a multi-scale convolution module and a channel fusion convolution module, extracts local features using convolution kernels with different time steps, and at the same time extracts features of each channel and fuses cross-channel information through the channel fusion convolution module to generate a high-dimensional feature representation containing multi-scale local features and channel information. Subsequently, the improved Transformer model effectively retains time information by introducing learnable position encoding, and combines two types of multi-head attention mechanisms, namely multi-head attention in the feature dimension and multi-head attention in the time dimension, to capture the correlation between features and the dependence between time steps respectively. The present invention gives full play to the advantages of CNN in local feature extraction and the ability of Transformer in global feature capture, effectively makes up for the deficiencies of a single model in dealing with complex time series data, and significantly improves the classification performance and the expressive ability of the model.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A multi - variable time - series classification method integrating convolutional neural network and Transformer, comprising the following steps: Step 1: Construct a multi - variable time - series classification model, which includes a multi - scale convolutional module, a channel - fusion convolutional module, a feature - dimension multi - head attention module, and a time - dimension multi - head attention module. Step 2: Pass the time - series sample set data sequentially through the multi - scale convolutional module and the channel - fusion convolutional module for multi - scale convolution and channel - fusion convolution operations to obtain fused features , Step 3: Remove the data - type dimension from the fused features and perform position - encoding operations to obtain position - encoded fused features. The position - encoded fused features are input into the feature - dimension multi - head attention module to obtain feature - dimension residual - connection features . Exchange the channel dimension and the time dimension of the feature - dimension residual - connection features to obtain residual - connection features . Input the residual - connection features into the time - dimension multi - head attention module to obtain time - dimension residual - connection features . Exchange the time dimension and the channel dimension of the time - dimension residual - connection features , perform layer normalization, then input into the feed - forward neural network layer for non - linear transformation and feature extraction, and apply layer normalization again to obtain feature - time - dimension abstract features ; Step 4: Perform average pooling on the time dimension of the feature - time - dimension abstract features , remove the time dimension, then perform linear transformation and normalization to obtain the predicted classes of each time - series sample; Step 5: Train the multi - variable time - series classification model based on minimizing the loss function to obtain the optimal parameters of the multi - variable time - series classification model; Step 6: Use the trained multi - variable time - series classification model to predict the time - series sample set data to be recognized.
[0006] The above - mentioned multi - scale convolution and channel - fusion convolution operations include the following steps: Step 2.1: Let the time - series sample set data be , the dimensions of the time - series sample set data include the total number of samples dimension, the data - type dimension, and the time dimension. The time - series sample set data X contains B time - series samples, each time - series sample contains D data types, different data types correspond to different types of monitoring data of the monitoring object, and each data type includes T time - point data. Add a channel dimension to the time - series sample set data X to obtain the set data , represents the real number field; Step 2.2: There are a total of types of scale convolutions used in the multi-scale convolution module. represents the convolution operation result of the e-th scale convolution module on the set of data . Concatenate the convolution operation results of each scale convolution module in the channel dimension to obtain the multi-scale locally sensitive feature y. After batch normalization of the multi-scale locally sensitive feature , use the GELU activation function to obtain the enhanced multi-scale locally sensitive feature .; Step 2.3: Channel fusion convolution module There are groups of convolution units. The size of the convolution kernel of the channel fusion convolution module is . The number of output channels of the channel fusion convolution module is , and the number of input channels is . is the number of output channels of the e-th scale convolution module. Then, use the channel fusion convolution module to perform a convolution operation on the multi-scale locally sensitive feature to obtain the fused feature . After batch normalization of the fused feature , use the GELU activation function to obtain the fused feature .
[0007] As described above, when the e-th scale convolution module performs convolution on the set of data , pad 0 at the front and back of the time dimension of the set of data so that the time scales of the output features of each scale convolution module remain unchanged.
[0008] As described above, the position encoding fused feature in step 3 is obtained based on the following steps: Perform dimensional compression on the data type dimension of the fused feature to obtain the fused feature . Set the learnable position encoding as . Duplicate the position encoding times to expand it into the position encoding . Add the fused feature to the position encoding to obtain the position encoding fused feature .
[0009] The above-mentioned feature dimension residual connection feature is obtained based on the following steps: Fuse the positional encoding feature Input it into the multi-head attention module for feature dimensions, calculate the query, key, and value of the th attention head of the multi-head attention module for feature dimensions, and further calculate the operation result of the th attention head, the number of subspace dimensions , is the floor function, is the number of attention heads. Concatenate the operation results of each attention head in the subspace dimension to obtain the concatenation result , and perform a linear transformation on the concatenation result through the weight matrix to obtain the linear transformation result . Then, perform a residual connection on the positional encoding fusion feature and the linear transformation result to obtain the feature dimension residual connection feature .
[0010] The above-mentioned time dimension residual connection feature is obtained based on the following steps: Exchange the channel dimension and the time dimension of the feature dimension residual connection feature to obtain the residual connection feature . Input the residual connection feature into the multi-head attention module for time dimensions, calculate the query, key, and value of the th attention head of the multi-head attention module for time dimensions, and further calculate the operation result of the th attention head, the number of subspace dimensions , is the floor function, is the number of attention heads. Concatenate the operation results of each attention head in the subspace dimension to obtain the concatenation result , and perform a linear transformation on the concatenation result through the weight matrix to obtain the linear transformation result . Then, perform a residual connection on the residual connection feature and the linear transformation result to obtain the feature dimension residual connection feature .
[0011] The above-mentioned loss function is the cross-entropy loss function.
[0012] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above classification method are implemented.
[0013] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the above classification method are implemented.
[0014] A computer program product includes a computer program. When the computer program is executed by a processor, the steps of the above classification method are implemented.
[0015] Compared with the prior art, the present invention has the following advantages and effects: 1. Through the multi-scale convolution module, channel fusion convolution module, feature dimension multi-head attention module, and time dimension multi-head attention module, the present invention can simultaneously extract local and global features of time series, overcoming the limitations of traditional CNN and Transformer models when dealing with local or global features separately.
[0016] 2. The present invention extracts rich local features through multi-scale convolution operations and enhances the fusion of cross-channel information through a channel fusion mechanism, improving the performance of the model in complex time series data.
[0017] 3. Compared with traditional models, the present invention further improves the ability to understand time series data through the feature dimension multi-head attention mechanism and the time dimension multi-head attention mechanism, especially in modeling the dependence relationship between time and feature dimensions, better adapting to the multi-channel time series classification task and effectively improving the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic structural diagram of the multivariate time series classification model (MSCFormer) of the present invention.
[0019] Figure 2 is a visualization diagram of the multi-scale feature channel fusion module (multi-scale convolution module and channel fusion convolution module). DETAILED IMPLEMENTATION MANNER
[0020] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.
[0021] Embodiment 1
[0022] Time series data refers to a series of data points recorded in chronological order. Each data point is usually composed of a continuous numerical value and changes over time. A time series can be regarded as consisting of multiple data points arranged in sequence. The total number of data points represents the length of the series, and each data point itself contains multiple numerical values, which respectively reflect the values on different variables or channels at that moment. The goal of the time series classification task is to enable the computer to automatically identify and determine the category to which a time series data belongs. To achieve this goal, it is usually necessary to prepare a dataset containing a large amount of labeled time series data and use a classification algorithm to train a model that can correctly classify newly input time series using this data. Specifically, such a dataset contains multiple time series data and their corresponding class labels, and each series has a label indicating its belonging category. The number of records in the dataset represents the total number of all time series, and these class labels come from a predefined set of categories. The core of the multivariate time series classification task is to accurately correspond each input time series data to its true belonging category.
[0023] The basic framework of this method is as shown in the appendix Figure 1 as follows.
[0024] Step 1: Construct a multivariate time series classification model (MSCFormer). The multivariate time series classification model (MSCFormer) includes a multi-scale convolution module, a channel fusion convolution module, a feature dimension multi-head attention module, and a time dimension multi-head attention module. Step 2: Pass the time series sample set data sequentially through the multi-scale convolution module and the channel fusion convolution module for multi-scale convolution and channel fusion convolution operations to obtain fused features .
[0025] The purpose of this step is to enhance the feature extraction ability of time series data through multi-scale convolution and channel fusion convolution operations. Multi-scale convolution extracts local features at multiple time scales by using convolution kernels of different scales, and can capture the long-term and short-term change patterns and complex local dependencies in time series data. Then, the channel fusion convolution effectively fuses the features from different channels, thereby reducing redundant information and enhancing the diversity and high-dimensional expression of features.
[0026] Step 2.1: Let the time series sample set data be , the dimensions of the time - series sample set data include the total number of samples dimension (B), the data type dimension (D), and the time dimension (T). The time - series sample set data X contains B time - series samples, and B corresponds to the number of different monitoring objects. Each time - series sample contains D data types, and different data types correspond to different types of monitoring data of the monitoring object. Each data type includes T time - point data. Denote the i - th time - series sample as , represents the real - number field, B represents the number of input time - series samples, D is the number of data types of time - point data, and T is the total number of time - point data. Add a channel dimension to the time - series sample set data X to obtain the set data for subsequent steps of feature extraction.
[0027] In this embodiment, the time - series samples are the time - point data detected by the monitoring device worn by the monitored person at each time point during a certain activity period. The number of time - series samples is B, that is, the total number of monitored persons is B; the number of time - point data is T, that is, the total number of time points is T; the number of data types D of time - point data is 6, that is, the time - point data includes three - dimensional accelerometer data (3 dimensions) and three - dimensional gyroscope data (3 dimensions).
[0028] Step 2.2: A total of types of scale convolutions are used in the multi - scale convolution module, which are respectively denoted as the first - scale convolution module , the second - scale convolution module , …, the E - th - scale convolution module , the first - scale convolution module , the second - scale convolution module , …, the E - th - scale convolution module The scales of the first - scale convolution module , the second - scale convolution module , …, the E - th - scale are respectively the first scale , denote the e - th - scale convolution module as , where e ∈ {1, 2, …, E}. Among them, the size of the convolution kernel of the e - th - scale convolution module is , the number of output channels of the e - th - scale convolution module is , and the number of input channels is 1. Use the first - scale convolution module , the second - scale convolution module , …, the E - th - scale convolution module to perform convolution operations on the set data (1)
[0029] Among them, Represents a convolution operation, with the input channels being , and the output channels being , Represents padding during the convolution operation, that is, by padding 0 before and after the time dimension (the last dimension) of the set data , ensuring that the time scale of the input set data and the time scale of the output features of each scale convolution module remain unchanged. Represents the convolution operation result of the e-th scale convolution module on the set data . After that, the convolution operation results are concatenated along the channel dimension (the second dimension) to obtain the multi-scale locally sensitive feature y: (2)
[0030] Among them, is the concatenation function, and the obtained is a set of multi-scale locally sensitive features. After that, the multi-scale locally sensitive features are batch-normalized (Batch Normalization, BN) and then the GELU activation function is used to further improve the stability and non-linear expression ability of the model, obtaining the enhanced multi-scale locally sensitive feature .
[0031] Step 2.3, Channel Fusion Convolution Module Suppose there are groups of convolution units. The size of the convolution kernel of the channel fusion convolution module is , the number of output channels of the channel fusion convolution module is , and the number of input channels is . After that, the channel fusion convolution module performs a convolution operation on the multi-scale locally sensitive feature , effectively fusing the information of different channels: (3)
[0032] Among them, represents the channel fusion convolution module, with the input channels being , and the output channels being . represents the fused feature output by the channel fusion convolution module. Similarly, after batch-normalizing and then using the GELU activation function to further improve the stability and non-linear expression ability of the model, the enhanced fused feature output by the fused convolution module . The operation in Step 2 can not only extract the deep features of each channel, but also capture the interaction information between channels, thus generating a more comprehensive and high-dimensional feature representation. This method helps to reduce the redundant information between features and improve the generalization ability and prediction performance of the model. Figure 2 Visualizes the change process of Step 2 when B = 1, E = 2, M = 2, and N = 3.
[0033] Step 3: Strengthen the features using the improved Transformer model. The improved Transformer model includes a multi-head attention module in the feature dimension and a multi-head attention module in the time dimension. Specifically, it includes: taking the fused features removing the data type dimension and performing a positional encoding operation to obtain the position-encoded fused features, and inputting the position-encoded fused features into the multi-head attention module in the feature dimension to obtain the residual connection features in the feature dimension , taking the residual connection features in the feature dimension exchanging the channel dimension and the time dimension to obtain the residual connection features , taking the residual connection features inputting them into the multi-head attention module in the time dimension to obtain the residual connection features in the time dimension , taking the residual connection features in the time dimension exchanging the time dimension and the channel dimension, then performing layer normalization and inputting them into the feed-forward neural network layer for non-linear transformation and feature extraction, and applying layer normalization again to obtain the abstract features in the feature time dimension ; The purpose of this step is to improve the model's ability to capture the time-dependent relationships in the time series through the improved Transformer module. The learnable positional encoding can dynamically adjust the position information between time steps, thereby enhancing the sensitivity to time dependence and improving the prediction performance. The multi-head attention module in the feature dimension and the multi-head attention module in the time dimension respectively promote the information interaction between features and between time steps, and capture complex temporal dependence relationships. These operations can better understand the spatio-temporal patterns of the time series, and improve the expressiveness and robustness of the feature representation. The specific process is as follows: Step 3.1. Perform dimensional compression (squeeze) on the data type dimension (the third dimension) of the fused features to obtain the fused features , set the learnable positional encoding as , where the positional encoding is a learnable parameter, and the positional encoding is the self-learnable parameter of the Transformer module. Then, copy the positional encoding times to expand it into the positional encoding , and take the fused features Add it to the position encoding so that the features are integrated with the position encoding information, obtaining the position encoding fused features .
[0034] Step 3.2: Use the position encoding fused features from Step 3.1 as the input features and input them into the multi-head attention module for feature dimensions, where is the number of output channels of the channel fusion convolutional module in Step 2.3. Let (4)
[0035] where are the query, key, and value of the multi-head attention module for feature dimensions respectively. It is assumed that the multi-head attention module for feature dimensions has attention heads, and map the query, key, and value to different subspaces: (5)
[0036] where represents the serial number of the attention head, are respectively the learnable linear transformation matrices of the th attention head, , and are respectively the query, key, and value of the th attention head. The number of dimensions of the subspace , is the floor function. Perform the attention operation on each subspace: (6)
[0037] where represents the matrix transpose operator, represents the dimension of the subspace corresponding to the attention head in the multi-head attention module for feature dimensions. represents the normalized exponential function, is the operation result of the th attention head. Concatenate the operation results of each attention head in the subspace dimension (the dimension where is located): (7)
[0038] where is the concatenation function, is the concatenation result of the multi-head attention module for feature dimensions. Next, perform a linear transformation on the concatenation result using the weight matrix to obtain the linear transformation result , then perform a residual connection on the input features and the linear transformation result of the feature dimension multi-head attention module to obtain the feature dimension residual connection feature , to prevent gradient explosion and increase the stability of training. Through this operation, the high-dimensional features of different channels can interact with each other to generate richer and more comprehensive feature representations, while enhancing the model's overall understanding of time series data.
[0039] Step 3.3, swap the channel dimension (the second dimension) and the time dimension (the third dimension) of the feature dimension residual connection feature to obtain the residual connection feature , use the residual connection feature as the input feature of the time dimension multi-head attention module, and calculate the query , key and value of the time dimension multi-head attention module. (8) where are the query, key, and value of the time dimension multi-head attention module respectively. It is assumed that the time dimension multi-head attention module has attention heads, and map the query, key, and value to different subspaces: (9)
[0040] where represents the serial number of the attention head, are respectively the learnable linear transformation matrices of the -th attention head, , and are respectively the query, key, and value of the -th attention head. The subspace dimension number , is the floor function. Perform the attention operation on each subspace: (10) where represents the matrix transpose operator, represents the subspace dimension corresponding to the attention head in the time dimension multi-head attention module. represents the normalized exponential function, is the operation result of the -th attention head. Concatenate the operation results of each attention head in the subspace dimension ( the dimension where (11) where is a splicing function, which is the splicing result of the multi-head attention module in the time dimension. Next, through the weight matrix perform a linear transformation on the splicing result to obtain the linear transformation result , and then perform a residual connection on the input features (residual connection features ) of the multi-head attention module in the time dimension and the linear transformation result to obtain the residual connection features in the time dimension . The attention mechanism at this stage can identify and understand the mutual influence between each time step in the high-dimensional features, so as to more accurately model the time dependence and improve the performance of the module.
[0041] Step 3.4: Next, swap the time dimension (the second dimension) and the channel dimension (the third dimension) of the residual connection features in the time dimension to obtain the residual connection features in the time dimension . Then, apply layer normalization (Layer Normalization, LN) to the residual connection features in the time dimension to stabilize the training process and promote convergence. Subsequently, perform a non-linear transformation and feature extraction through a feed-forward neural network layer (Feed-Forward Network, FFN), and apply layer normalization again to obtain the abstract features in the feature time dimension . The feed-forward neural network layer further improves the ability of feature representation, ensuring that the features processed by the attention mechanism have better interpretability and robustness.
[0042] Step 4: Perform average pooling on the time dimension of the abstract features in the feature time dimension to obtain the feature . Remove the time dimension of the feature to obtain the final feature . Perform a linear transformation on the final feature through the weight matrix to obtain the linearly transformed feature . Perform normalization on the linearly transformed feature to obtain the predicted category of each time series sample.
[0043] The purpose of this step is to extract key information from the high-dimensional features after model feature extraction and use this information for the final classification task.
[0044] Step 4.1: First, input the abstract features in the feature time dimension of Step 3.4 . For the abstract features in the feature time dimension Perform average pooling on the time dimension (the third dimension) to obtain features . This step compresses the information in the time dimension and retains important global features. Then, remove (squeeze) the time dimension of the features to obtain the final features . At this time, the temporal features of each time series sample are represented as N-dimensional vectors for input into the fully connected layer. Next, perform a linear transformation on the final features using the weight matrix to obtain the linearly transformed features . This step maps the N-dimensional feature space to the G-dimensional classification space, where G is the total number of categories in the classification space, generating the raw scores for each classification category. Finally, perform normalization on the linearly transformed features to obtain the prediction probability that the -th time series sample belongs to category , satisfying . Denote as the probability distribution of the -th time series sample over all categories. The category corresponding to the maximum value in the probability distribution is the final predicted category of the -th time series sample.
[0045] Step 5: When training the multivariate time series classification model (MSCFormer), the loss function is used to measure the difference between the model's prediction results and the true labels, thereby guiding the optimization process of the model. To ensure that the model can effectively perform multivariate time series classification, this step selects the cross-entropy loss function as the loss function and trains the multivariate time series classification model (MSCFormer) based on minimizing the loss function to obtain the optimal parameters of the multivariate time series classification model (MSCFormer).
[0046] Step 6: After completing the training of the multivariate time series classification model (MSCFormer), it is necessary to use the trained multivariate time series classification model (MSCFormer) to predict the set of time series samples to be recognized (multivariate time series data) to evaluate its classification performance.
[0047] In the prediction stage, first, a batch of time series data to be classified needs to be prepared as time series samples, and a set of time series sample data to be recognized is constructed. Next, the trained multivariate time series classification model is loaded and placed in the evaluation mode to avoid gradient updates or parameter changes during the inference process. Subsequently, the set of time series sample data to be recognized is input into the multivariate time series classification model. After receiving the set of time series sample data to be recognized, the multivariate time series classification model will perform forward propagation based on the multi-scale feature representations and temporal dependencies learned during training, extract features layer by layer and perform calculations, and finally output the class probability distributions of each time series sample in the set of time series sample data to be recognized. For each time series sample, the class with the highest probability is selected as the final prediction result, thereby determining the final predicted class to which the time series sample belongs. Finally, the time series samples of the entire batch are summarized to obtain the final predicted classes of all time series samples.
[0048] In summary, the method of the present invention (MSCFormer) extracts multi-scale local features and channel information of the input data through multi-scale convolution and channel fusion convolution operations, and performs global feature modeling with the help of the feature dimension multi-head attention module and the time dimension multi-head attention module. The synergistic effect of these steps makes the model more adaptable to time series data, and finally realizes the efficient classification of multi-channel time series. This combination method gives full play to the advantages of CNN and Transformer, and successfully overcomes the limitations of a single model in feature extraction and feature modeling.
[0049] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0050] Next, the advantages of the method of the present invention are verified through experiments: I. Data Description To test the effectiveness of the present invention, it is applied to the time series data classification experiment herein. The time series datasets used in the experiment are all UEA time series classification datasets, which are widely used datasets for multivariate time series classification. The datasets used in the experiment are shown in Table 1.
[0051] Table 1 Dataset Details Table
[0052] II. Description of Benchmark Methods To verify the superiority of the present invention, the present invention (MSCFormer model) is compared with the following 7 advanced models, and these models are briefly introduced below: 1. FEAT: A general framework for multivariate time series representation learning.
[0053] 2. DKN: A dense knowledge-aware network based on Transformer for MTSC.
[0054] 3. MICOS: A hybrid supervised contrastive learning model for MSTC.
[0055] 4. ShapeNet: A novel neural network structure.
[0056] 5. Conv-GLU: A convolutional network with gated linear unit kernels.
[0057] 6. DA-NET: A neural network with dual attention.
[0058] 7. ConvTran: A Transformer model based on improved positional encoding.
[0059] III. Description of Experimental Settings A total of 5 hyperparameters are involved in the experiment, which are respectively: 1. Learning rate; 2. Four different scale combinations of scale convolutions are used in the experiment, and the different scales are , and the corresponding number of output channels ; 3. The number of output channels of the channel fusion convolution kernel ; 4. The number of heads of the multi-head attention mechanism in the feature dimension , and the number of heads of the multi-head attention mechanism in the time dimension ; According to the results of existing research, the number of heads of the multi-head attention in the experiment and are directly set to 8. The number of output channels of each convolution kernel is set to the number of output channels of the channel fusion convolution kernel, that is, let , the learning rate is set to 0.001, and the optimal values of the remaining parameters are searched by the method of grid search. Their search ranges are respectively: different scale combinations The search range is ; The search range of the number of output channels of the channel fusion convolution kernel is .
[0060] IV. Results of Algorithm Performance Comparison Using the overall classification accuracy as the evaluation metric, the experimental results are shown in Table 2 below: Table 2 Comparison Results Table
[0061] The comparison experimental results are shown in Table 2. The MSCFormer of the present invention is overall superior to all baseline models in terms of classification performance. Among the 27 datasets, the MSCFormer of the present invention performs best on 15 datasets, significantly leading the baseline models, with an average accuracy of 80%, which is 4.1% higher than the best-performing baseline model. On most datasets, the MSCFormer of the present invention performs very well. For example, on datasets such as AWR, AF, CT, ER, BM, etc., the MSCFormer of the present invention reaches or exceeds the highest accuracy. Even on some datasets with relatively poor performance, the MSCFormer of the present invention also exceeds most of the baseline methods, further demonstrating its significant advantages in handling traditional time series tasks. In summary, the MSCFormer of the present invention performs better than the existing baseline models on multiple datasets, proving its powerful ability in processing complex time series data.
[0062] Example 2
[0063] This example provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method embodiments are implemented.
[0064] Example 3
[0065] This example provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method embodiments are implemented.
[0066] Example 4
[0067] This example provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above method embodiments are implemented.
[0068] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A multi - variable time - series classification method integrating convolutional neural network and Transformer, comprising the following steps: Step 1: Construct a multi - variable time - series classification model, where the multi - variable time - series classification model includes a multi - scale convolutional module, a channel - fusion convolutional module, a feature - dimension multi - head attention module, and a time - dimension multi - head attention module. Step 2: Collect the time series sample data Multi-scale convolution and channel fusion convolution operations are performed in turn through the multi-scale convolution module and the channel fusion convolution module to obtain fusion features , Step 3: The fused feature removes the data type dimension and performs a positional encoding operation to obtain a position-encoded fused feature, which is input into the feature-dimension multi-head attention module to obtain a feature-dimension residual connection feature . The feature-dimension residual connection feature has its channel dimension and time dimension swapped to obtain a residual connection feature . The residual connection feature is input into the time-dimension multi-head attention module to obtain a time-dimension residual connection feature . The time-dimension residual connection feature has its time dimension and channel dimension swapped, followed by layer normalization and then input into a feed-forward neural network layer for non-linear transformation and feature extraction, and layer normalization is applied again to obtain an abstract feature in the feature time dimension ; Step 4: Abstract features in the characteristic time dimension After performing average pooling on the time dimension and removing the time dimension, perform a linear transformation and normalization to obtain the predicted categories of each time series sample; Step 5: Train the multi - variable time - series classification model based on minimizing the loss function to obtain the optimal parameters of the multi - variable time - series classification model. Step 6: Use the trained multi - variable time - series classification model to predict the data of the time - series sample set to be recognized.
2. The multivariate time series classification method integrating a convolutional neural network and a Transformer according to claim 1, wherein The multi - scale convolution and channel - fusion convolution operations include the following steps: Step 2.1: Let the time series sample set data be , the dimensions of the time series sample set data include the total number of samples dimension, the data type dimension, and the time dimension. The time series sample set data X contains B time series samples, each time series sample contains D data types, different data types correspond to different types of monitoring data of the monitoring object, each data type includes T time point data, add a channel dimension to the time series sample set data X to obtain the set data , represents the real number field; Step 2.
2. There are types of scale convolutions used in the multi-scale convolution module. represents the convolution operation result of the e-th scale convolution module on the set data . The convolution operation results of each scale convolution module are concatenated in the channel dimension to obtain the multi-scale locally sensitive feature y. After batch normalization of the multi-scale locally sensitive feature , the GELU activation function is used to obtain the enhanced multi-scale locally sensitive feature . Step 2.3, Channel Fusion Convolution Module There are groups of convolution units. The convolution kernel size of the channel fusion convolution module is , the output channel number of the channel fusion convolution module is , the input channel number is , is the output channel number of the e-th scale convolution module. Then, the channel fusion convolution module performs a convolution operation on the multi-scale locally sensitive features to obtain the fused features . Then, the fused features are batch-normalized and then the GELU activation function is used to obtain the fused features .
3. The multi-variate time series classification method integrating a convolutional neural network and a Transformer according to claim 2, wherein The e-th scale convolution module When performing convolution on the set data , zeros are padded before and after the time dimension of the set data so that the time scales of the output features of each scale convolution module remain unchanged.
4. The multivariate time series classification method integrating a convolutional neural network and a Transformer according to claim 2, characterized in that The position - encoding fusion feature in Step 3 is obtained based on the following steps: Perform dimensionality compression on the data type dimension of the fused feature to obtain the fused feature , set the learnable position encoding as , and copy the position encoding for times to expand it into the position encoding . Then add the fused feature and the position encoding to obtain the position encoding fused feature .
5. The multi-variate time series classification method integrating a convolutional neural network and a Transformer according to claim 4, characterized in that, The residual connection feature of the feature dimension is obtained based on the following steps: Fused position encoding feature Input feature dimension multi-head attention module, calculate the query, key, and value of the th attention head of the feature dimension multi-head attention module, and further calculate the operation result of the th attention head , number of subspace dimensions , is floor function, is the number of attention heads, concatenate the operation results of each attention head in the subspace dimension to obtain the concatenated result , perform a linear transformation on the concatenated result through the weight matrix to get the linear transformation result , then fuse the position encoding feature and the linear transformation result with a residual connection to obtain the feature dimension residual connection feature .
6. The multivariate time series classification method integrating a convolutional neural network and a Transformer according to claim 5, characterized in that, The time dimension residual connection feature is obtained based on the following steps: Connect the feature dimension residual to the feature Exchange the channel dimension and the time dimension to obtain the residual connection feature , and input the residual connection feature into the multi-head attention module in the time dimension, calculate the query, key, and value of the th attention head of the multi-head attention module in the time dimension, and further calculate the operation result of the th attention head , the number of subspace dimensions , is the floor function, is the number of attention heads. Concatenate the operation results of each attention head in the subspace dimension to obtain the concatenation result . Perform a linear transformation on the concatenation result through the weight matrix to obtain the linear transformation result . Then, perform a residual connection on the residual connection feature and the linear transformation result to obtain the feature dimension residual connection feature .
7. The multivariate time series classification method integrating a convolutional neural network and a Transformer according to claim 1, characterized in that The loss function is a cross - entropy loss function.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the classification method described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the classification method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the classification method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Time series data anomaly detection method based on channel fusion self-attention mechanism
CN117034175A
Short temporary rainfall prediction method based on multi-scale attention and convolution fusion
CN118298222A
Upsampling of compressed financial time-series data using a jointly trained Vector Quantized Variational Autoencoder neural network
US12229679B1
Systems and methods for convolutional neural network and transformer-based time series modeling
US20250045566A1
Cited By
Time series data end-to-end classification method and device, equipment and medium
CN120744426A
Generative statistical semantic guidance processing method oriented to time sequence classification
CN122286526A