Sequence memory guided multi-time sequence anomaly detection method
The reconstruction model trained in two stages utilizes separable convolutional networks and memory modules to extract sequence features, solving the problem of capturing normal patterns in multi-time series anomaly detection, improving detection accuracy and adapting to concept drift, and achieving better reconstruction results.
Patent Information
- Application Number
- CN202510935679.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies struggle to effectively capture normal patterns in multi-time series anomaly detection, especially in the presence of noise or anomalies, and also face concept drift issues, resulting in low detection accuracy.
A two-stage training reconstruction model is adopted, which extracts sequence features through a separable convolutional network and a memory module, uses memory prototypes for attention calculation and sparsity optimization, and combines anomaly thresholds for anomaly detection.
It improves the accuracy of anomaly detection across multiple time series, adapts to concept drift, reduces the impact of noise and anomalies, and enhances reconstruction results.
Smart Images

Figure CN120873659A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of time series and data mining, and specifically to a sequence memory-guided method for detecting anomalies in multiple time series. Background Technology
[0002] Multi-time series anomaly detection is a widely in-demand area within the time series field. Time series data are prevalent in various systems and possess universality. The health of a system can be determined by analyzing its key performance indicators (KPIs). Multi-time series anomaly detection analyzes historical data of these KPIs to determine whether anomalies have occurred within a given timeframe. This type of anomaly detection requirement is widely applied in real-world scenarios. For example, in industrial equipment monitoring, users want machines to notify maintenance personnel for intervention when anomalies occur, such as noise or overheating, to reduce losses due to equipment failure. Furthermore, multi-time series anomaly detection models capable of real-time data processing can be used in edge devices such as human health monitoring equipment and traffic flow detection equipment.
[0003] Traditional time series anomaly detection methods primarily rely on statistical methods such as the 3-δ rule or quantiles. These methods directly treat the current state of the time series data as a feature, using the distribution of historical states as the standard feature distribution. If the current state falls within the low-probability range of the historical state set, it is considered an anomalous. Statistical methods presuppose a normal distribution for time series data, such as a Gaussian distribution, and they work well when the actual distribution is the same as or close to the presupposed distribution. However, most real-world time series data not only do not conform to the presupposed distribution but may also change over time, a phenomenon known as concept drift. Therefore, traditional statistical methods are not entirely effective when directly applied to time series anomaly detection. In recent years, with the rapid development of deep learning, researchers have proposed numerous deep learning methods for multi-time series anomaly detection, using recurrent neural networks (RNNs), temporal convolutional networks (TCNs), variational autoencoders (VAEs), and Transformers to detect anomalies in multiple time series. Some works reconstruct the input sequence to obtain a reconstructed sequence and determine anomalies based on the error between the reconstructed and input sequences. Others directly generate anomaly scores based on the latent representations or prediction errors output by the model. For example, autoencoder-based methods capture the inherent patterns of normal data through compression and reconstruction processes, triggering alarms when anomalous data deviates from the normal distribution, causing a sharp increase in reconstruction errors. However, these methods still face many challenges: First, complex time dependencies (such as long-period, multi-scale patterns) may make it difficult for models to accurately reconstruct normal patterns; second, the concept drift problem has not been systematically addressed, making it difficult for statically trained models to adapt to dynamically changing normal behavior; third, the modeling of nonlinear interactions between variables in multivariate time series is insufficient, easily overlooking causal anomalies across channels.
[0004] The challenge of anomaly detection across multiple time series lies in capturing normal patterns. In real-world scenarios, normal patterns in multiple time series not only mix with unlabeled anomalous patterns but may also change over time, exhibiting concept drift. The key challenge is to enable the model to capture normal patterns in noisy or anomalous time series to ensure detection accuracy, and to capture the altered data distribution after concept drift occurs. Therefore, an effective method is needed to address these issues. Summary of the Invention
[0005] The purpose of this invention is to provide a sequence memory-guided method for detecting anomalies in multiple time series, so as to improve the accuracy of the task.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A sequence memory-guided multi-time series anomaly detection method is proposed. The sequence to be detected, X, is input into a two-stage reconstruction model obtained through two-stage training to obtain a reconstructed sequence, Y. The error between the last time node of the reconstructed sequence Y and the sequence to be detected, X, is used as the reconstruction error. The reconstruction error is compared with an anomaly threshold τ. If the reconstruction error is greater than or equal to the anomaly threshold τ, the current time point is determined to be an anomaly time point. If the reconstruction error is less than the anomaly threshold τ, the current time point is determined to be a normal time point. The second phase of training is as follows: First-stage training: For a training set of multiple time series with shape N × T, length T, and dimension N; during the first-stage training, firstly, k memory prototypes in the reconstruction model are randomly initialized, and then the reconstruction model is initialized; the multiple time series of the training set are input into the initialized reconstruction model for training to obtain the reconstruction sequence, thus obtaining the first-stage reconstruction model. During the training process, the mean squared error loss function, i.e., formula (1), is used as the optimization objective, where Y i To reconstruct the sequence, and X i The input sequence;
[0007] Second stage training: After obtaining the first stage model, the validation set is used to perform sliding window splitting to obtain the sampling sequence; the sampling sequence is input into the first stage model to obtain the sampling reconstruction sequence; the sampling reconstruction sequence is subjected to k-means clustering with the number of memory prototypes k as the number of clusters to obtain k cluster centers; the k cluster centers are used as the initial memory prototypes to initialize the reconstruction model; the multiple time series of the training set are input into the first stage reconstruction model for training to obtain the reconstruction sequence, and the second stage reconstruction model is obtained. The training process is optimized with formula (2) as the objective, where A is the attention relationship between sequence features and prototypes, and α is the proportion of sparsity of the subsequent target, i.e., the memory prototype. .
[0008] The reconstructed model includes a separable convolutional network and a memory module; the model training in both the first and second training phases includes the following steps: S1. Overlapping linear coding: The time series of the training set is divided into several equal-length overlapping sequences using a fragmentation operation. These overlapping sequences are then encoded using a shared linear layer to obtain several sequence codes. S2. Extract sequence features: Input the sequence encoding into a separable convolutional neural network to extract different features from the sequence encoding, thus obtaining sequence features; S3. Obtain prototype features: Use the sequence features obtained in S2 as the query input memory module, perform attention calculation with the prototype items in the memory module, and obtain the prototype similarity ratio after Softmax normalization. Then, perform calculation and superposition of the similarity ratio with the prototype to obtain the prototype features. S4. Obtain enhanced features and reconstruct the sequence: Connect the sequence features with the prototype features to obtain enhanced features. Then, use a linear layer as a weak decoder, input the enhanced features, reconstruct the fragment sequence, and restore all fragment sequences to obtain the reconstructed sequence.
[0009] In S1, the fragmentation operation is based on a sliding window operation, using two hyperparameters, window length and stride, to divide the input multi-time series into finer-grained small window forms S = {S1, S2, ...., S}. P Meanwhile, the shared linear layer used can express the features of different fragment sequences based on the encoding of the fragment sequence; after the fragmentation operation, the original shape (B, N, W) of the input sequence is transformed into (B, N, P, d), where B represents the size of each batch, N is the dimension of the input sequence, W is the window length of the input sequence, P is the number of fragments finally obtained by the fragmentation operation, and d is the output dimension of the shared linear layer.
[0010] S2 is specifically as follows: S2.1 Input sequence encoding, use depthwise convolution to extract features within the sequence segment. The kernel size of the depthwise convolution is fixed and greater than 1 to obtain the internal relationship of the segment sequence encoding, which is the internal feature encoding of the sequence. S2.2. Input the sequence internal feature code obtained in S2.1, use point-direction convolution with a kernel size of 1 to extract features between channels, and obtain channel convolution codes of the same shape; S2.3. Input the channel convolutional encoding obtained in S2.2. After dimension swapping, use the same point-direction convolution to extract features between sequence segment encodings to obtain segment convolutional encoding. Then restore the dimensions. The resulting features are the sequence features. The shape of the final sequence features is consistent with the fragmented input features, which is (B, N, P, d).
[0011] S3 is specifically as follows: S3.1 Perform matrix multiplication between the sequence features and the prototype terms in the memory module, i.e., A = E c @ M′, where E c Let M' be the sequence feature, M' be the transpose of all prototype terms in the sequence memory, and @ denote matrix multiplication; finally, we obtain the attention relationship between the sequence feature and the prototype, which has the shape (B, N, P, k), where k is the number of prototype terms. S3.2. Perform Softmax normalization on the attention relationship to obtain the prototype similarity ratio between 0 and 1; S3.3, Perform matrix multiplication between the prototype similarity ratio and the prototype in the memory module, i.e., E t = A@M, obtain the prototype feature E t The shape of the prototype features is consistent with that of the sequence features.
[0012] During the two-stage training, the reconstruction model only back-propagates the relevant gradients of the separable convolutional network and automatically updates the reconstruction model. The sequence memory prototype is manually updated based on attention and sequence features. The manual update uses the sequence features E obtained in S2. c The prototype M in the sequential memory term is first obtained through matrix multiplication E. c @ M′ obtains the attention relationship A between sequence features and prototypes; then, A is Softmax normalized to obtain the sequence feature-prototype similarity ratio; next, matrix multiplication A′@ Ec is performed between the similarity ratio and the sequence features to obtain the prototype modification term M. n Finally, the modified prototype and the original prototype are weighted and combined using an expert network and impulse to obtain the updated prototype, as shown in the following formula:
[0013]
[0014] In this context, U and V are both expert networks, M is the original prototype, Mn is the prototype modification term, and β controls the rate of prototype update.
[0015] The abnormal threshold τ is obtained as follows: After completing the two-stage training, the training set multiple time series are input into the two-stage reconstruction model again. The reconstruction error of the last time node of each sequence is taken as the basic abnormal score set. Then, the abnormal ratio is estimated by using the prior knowledge of the real scene task corresponding to the abnormal detection, and the abnormal threshold τ is obtained.
[0016] By adopting the above scheme, this invention introduces a two-stage training strategy to obtain the initial information of the memory terms. In the first stage, the model mainly focuses on optimizing reconstruction capabilities, so the reconstruction error of the directly reconstructed sequence and the input sequence are used as optimization targets. Before the second stage of training, a certain amount of validation set data is input into the model trained in the first stage, and the current sequence features are obtained after separable convolution. These sequence features are clustered, and the cluster centers are used as the initial memory terms for the second-stage model for training. The second-stage training uses reconstruction error and sparsity in the memory terms as joint targets for training optimization. In summary, this invention fragments multiple time series, extracts short-term features of the fragmented time series from different perspectives, and uses historical sequence features from the sequence memory terms to enhance and reconstruct the current sequence features, thereby reducing the noise and anomaly effects in the sequence features, achieving better reconstruction results, and thus improving the accuracy of multi-time series anomaly detection tasks. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the principle of the present invention; Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the following embodiments will be used in conjunction with the accompanying drawings to further illustrate the invention. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. Rather, the invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the invention as defined by the claims. Furthermore, to provide the public with a better understanding of the invention, some specific details are described in detail below. Those skilled in the art will fully understand the invention even without these detailed descriptions.
[0019] like Figure 1 and Figure 2 As shown, this invention discloses a multi-time series anomaly detection method guided by sequence memory. It uses the error between the last time node of the reconstructed sequence Y and the sequence to be detected X as the reconstruction error. The reconstruction error is compared with the anomaly threshold τ. If the reconstruction error is greater than or equal to the anomaly threshold τ, the current time point is determined to be an abnormal time point. If the reconstruction error is less than the anomaly threshold τ, the current time point is determined to be a normal time point.
[0020] The second phase of training is as follows: First-stage training: For a training set of multiple time series with shape N × T, length T, and dimension N; during the first-stage training, firstly, k memory prototypes in the reconstruction model are randomly initialized, and then the reconstruction model is initialized; the multiple time series of the training set are input into the initialized reconstruction model, and the reconstruction sequence is obtained through training through the following S1-S4, thus obtaining the first-stage reconstruction model. During the training process, the mean squared error loss function, i.e., formula (1), is used as the optimization objective, where Y i To reconstruct the sequence, and X i The input sequence; .
[0021] Second stage training: After obtaining the first stage model, the validation set is used to perform sliding window splitting to obtain the sampling sequence; the sampling sequence is input into the first stage model to obtain the sampling reconstruction sequence; the sampling reconstruction sequence is clustered with k-means clustering with the number of memory prototypes k as the number of clusters to obtain k cluster centers; the k cluster centers are used as the initial memory prototypes to initialize the reconstruction model; the multiple time series of the training set are input into the first stage reconstruction model, and the reconstruction sequence is obtained through the following S1-S4 to obtain the second stage reconstruction model. The training process is optimized with formula (2) as the objective, where A is the attention relationship between sequence features and prototypes, and α is the proportion of sparsity of the subsequent target, i.e., the memory prototype. .
[0022] The reconstruction model includes a separable convolutional network and a memory module, and its training process S1-S4 is as follows: S1. Overlapping linear coding: The input time series is divided into several equal-length overlapping sequences using a fragmentation operation. These overlapping sequences are then encoded using a shared linear layer to obtain several sequence codes.
[0023] Fragmentation is based on sliding window operations, using two hyperparameters, window length and stride, to divide the input multi-time series into finer-grained small window forms S = {S1, S2, ..., S...}. P The shared linear layer used simultaneously can uniformly express the features of different segment sequences based on the encoding of the segment sequence. After the fragmentation operation, the original shape (B, N, W) of the input sequence is transformed into (B, N, P, d), where B represents the size of each batch, N is the dimension of the input sequence, W is the window length of the input sequence, P is the number of segments obtained after the fragmentation operation, and d is the output dimension of the shared linear layer.
[0024] S2. Extract sequence features: Input the sequence encoding into three convolutional neural networks to extract different features from the sequence encoding to obtain sequence features.
[0025] We extract features from three different channels of the input segment sequence using different operations in two separable convolutions: depthwise convolution and pointwise convolution. The depthwise convolution is responsible for the output channels of the shared linear layer obtained in S1, mainly processing features within the segment sequence. There are two pointwise convolutions, each with a different first dimension to fuse the intermediate input sequence dimension and the inter-segment dimension, and then use different groups of convolutions to extract features from different channels.
[0026] S2.1 Input sequence encoding, use depthwise convolution to extract features from the internal features of the sequence segment, and obtain internal feature encodings of the same shape; The kernel size of the depthwise convolution is fixed and greater than 1 to obtain the internal relationship of the segment sequence encoding, which is the internal feature encoding of the sequence.
[0027] S2.2. Input the sequence internal feature code obtained in S2.1, use point-direction convolution with a kernel size of 1 to extract features between channels, and obtain channel convolution codes of the same shape; S2.3. Input the channel convolutional encoding obtained in S2.2. After dimension swapping, use the same point direction convolution to extract features between sequence segment encodings to obtain segment convolutional encoding. Then restore the dimension. The resulting features are the sequence features. The shape of the final sequence features is consistent with the fragmented input features, which is (B, N, P, d). S3. Obtain prototype features: Use the sequence features obtained in S2 as the query input memory module, perform attention calculation with the prototype items in the memory module, and obtain the prototype similarity ratio after Softmax normalization. Then, perform calculation and superposition of the similarity ratio with the prototype to obtain the prototype features. S3.1 Perform matrix multiplication between the sequence features and the prototype terms in the memory module, i.e., A = E c @ M′, where E c Let M' be the sequence feature, M′ be the transpose of all prototype terms in the sequence memory, and @ denote matrix multiplication. The final result is the attention relationship between the sequence feature and the prototypes, with shape (B, N, P, k), where k is the number of prototype terms. S3.2. Perform Softmax normalization on the attention relationship to obtain the prototype similarity ratio between 0 and 1; S3.3, Perform matrix multiplication between the prototype similarity ratio and the prototype in the memory module, i.e., E t = A@M, obtain the prototype feature, the shape of the prototype feature is consistent with the sequence feature; S4. Obtain enhanced features and reconstruct the sequence: Concatenate the sequence features with the original features to obtain the enhanced features, i.e., E = Concat(E c Et The Concat function indicates that the two are directly concatenated in the last dimension, so the shape of the final enhanced feature is (B, N, P, 2*d). Then, a linear layer is used as a weak decoder. The enhanced feature is input, the fragment sequence is reconstructed, and all fragment sequences are restored to obtain the reconstructed sequence.
[0028] For the anomaly threshold τ, after the two-stage model is trained, the training set multiple time series are input into the model again. The reconstruction error of the last time node of each series is taken as the basic anomaly score set. Then, the anomaly ratio is estimated by using the prior knowledge of the real scene task corresponding to the anomaly detection, and the threshold τ is obtained.
[0029] This anomaly detection method manually updates the intrinsic parameters of the memory prototype during model training and real-time model adjustment. When updating parameters according to the objective function, gradient backpropagation is not applied to the intrinsic parameters of the memory prototype. The specific processing method is as follows: During training, only the relevant gradients of the separable convolutional network are backpropagated and automatically updated, while the sequence memory prototype is manually updated based on attention and sequence features. The manual update uses the sequence features E obtained in S2. c The prototype M in the sequential memory term is first obtained through matrix multiplication E. c @ M′ obtains the attention relationship A between sequence features and prototypes; then A is Softmax normalized to obtain the sequence feature-prototype similarity ratio; then the similarity ratio is used to perform matrix multiplication A′@ Ec to obtain the prototype modification term M. n Finally, the modified prototype and the original prototype are weighted and combined using an expert network and impulse to obtain the updated prototype, as shown in the following formula:
[0030]
[0031] In this context, U and V are both expert networks, M is the original prototype, Mn is the prototype modification term, and β controls the rate of prototype update.
[0032] This invention employs a two-stage training strategy to unsupervised train a multi-time-series reconstruction model, obtaining sparse sequence memory initial terms to store features from different historical sequences. A two-stage objective function is used to optimize and balance the model's reconstruction capability and the sparsity of the sequence memory terms. When the sequence to be detected is input into the model, it is first segmented into equal-length sequence encoding segments. Then, separable convolution is used to capture short-term dependencies at different angles within the sequence encoding segments, resulting in sequence features storing these short-term dependencies. Next, historical sequence features similar to the sequence features are queried from the sequence memory terms and exported, then concatenated with the sequence features to obtain enhanced features. Finally, a simple reconstruction layer is used to reconstruct and restore the enhanced features, yielding the reconstructed sequence and reconstruction error. Anomaly detection uses the reconstruction error at the last time point in the input and reconstructed sequences as an anomaly score, which is compared with an anomaly threshold calculated from the anomaly scores in the training set to determine whether an anomaly has occurred at the current time point.
[0033] In summary, this invention proposes a simple and efficient method for multi-time series anomaly detection. This method fragments multiple time series, extracts short-term features from these fragments from different perspectives, and then uses historical sequence features from the sequence memory to enhance and reconstruct the current sequence features. This reduces noise and anomaly effects in the sequence features, resulting in better reconstruction and thus improving the accuracy of multi-time series anomaly detection.
[0034] In addition, the method fully considers the problems that may occur in time series in real-world scenarios when designing the model, and can deal with the concept drift phenomenon by periodically updating the two-stage reconstruction model online.
[0035] In multi-time series anomaly detection tasks, the effectiveness of this invention is demonstrated by testing it on three widely recognized multi-time series anomaly detection datasets: PSM, SMD, and SWAT. This invention is also compared with numerous standard time series anomaly detection methods, such as Isolation Forest, using Pred... A Recall A and F1 A The effectiveness of the method of this invention is compared using three retrieval metrics to demonstrate its superiority, as shown in Table 1. The method of this invention is called MSCN.
[0036] Table 1
[0037] In Table 1, Pred A Recall represents the ratio of the length of the interval detected as abnormal to the length of the actual abnormal interval. A F1 represents the ratio of the length of the detected actual abnormal intervals to the total length of the actual abnormal intervals. AIt is the average of the two, used to measure the effectiveness of multi-time series anomaly detection methods.
[0038] The leftmost column of Table 1 represents the method names, while the top row, PSM, SMD, and SWAT, represents three real-world multi-time-series anomaly detection task datasets. Specifically, IFOrest is a traditional partitioning-based method; OCSVM is a traditional classification-based method; VAE is a widely used deep reconstruction method; DeepSVDD is a deep classification-based method; AnomalyTransformer is a Transformer-based reconstruction method; ModernTCN is a feature extractor based on temporal convolutional networks, which can be used as a downstream task for comparison in anomaly detection; and MSCN is a sequence memory-guided multi-time-series anomaly detection method (this invention).
[0039] Table 1 presents experimental results for anomaly detection on three recognized real-world multi-time-series anomaly detection datasets, using the Pred report. A Recall A and F1 A Three criteria are used to demonstrate the accuracy of different methods in multi-time-series anomaly detection tasks. Table 1 shows that this invention surpasses other anomaly detection methods in terms of accuracy.
[0040] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.
Claims
1. A multi-time series anomaly detection method guided by sequence memory, characterized in that: The sequence X to be detected is input into the two-stage reconstruction model obtained after two-stage training to obtain the reconstructed sequence Y. The error between the last time node of the reconstructed sequence Y and the sequence X to be detected is taken as the reconstruction error. The reconstruction error is compared with the abnormal threshold τ. If the reconstruction error is greater than or equal to the abnormal threshold τ, the current time point is determined to be an abnormal time point. If the reconstruction error is less than the abnormal threshold τ, the current time point is determined to be a normal time point. The second phase of training is as follows: First-stage training: For a training set of multiple time series with shape N × T, length T, and dimension N; during the first-stage training, firstly, k memory prototypes in the reconstruction model are randomly initialized, and then the reconstruction model is initialized; the multiple time series of the training set are input into the initialized reconstruction model for training to obtain the reconstruction sequence, thus obtaining the first-stage reconstruction model. During the training process, the mean squared error loss function, i.e., formula (1), is used as the optimization objective, where Y i To reconstruct the sequence, and X i The input sequence; Second stage training: After obtaining the first stage model, the validation set is used to perform sliding window splitting to obtain the sampling sequence; the sampling sequence is input into the first stage model to obtain the sampling reconstruction sequence; the sampling reconstruction sequence is subjected to k-means clustering with the number of memory prototypes k as the number of clusters to obtain k cluster centers; the k cluster centers are used as the initial memory prototypes to initialize the reconstruction model; the multiple time series of the training set are input into the first stage reconstruction model for training to obtain the reconstruction sequence, and the second stage reconstruction model is obtained. The training process is optimized with formula (2) as the objective, where A is the attention relationship between sequence features and prototypes, and α is the proportion of sparsity of the subsequent target, i.e., the memory prototype. 。 2. The multi-time series anomaly detection method guided by sequence memory according to claim 1, characterized in that: The reconstructed model includes a separable convolutional network and a memory module; the model training in both the first and second training phases includes the following steps: S1. Overlapping linear coding: The time series of the training set is divided into several equal-length overlapping sequences using a fragmentation operation. These overlapping sequences are then encoded using a shared linear layer to obtain several sequence codes. S2. Extract sequence features: Input the sequence encoding into a separable convolutional neural network to extract different features from the sequence encoding, thus obtaining sequence features; S3. Obtain prototype features: Use the sequence features obtained in S2 as the query input memory module, perform attention calculation with the prototype items in the memory module, and obtain the prototype similarity ratio after Softmax normalization. Then, perform calculation and superposition of the similarity ratio with the prototype to obtain the prototype features. S4. Obtain enhanced features and reconstruct the sequence: Connect the sequence features with the prototype features to obtain enhanced features. Then, use a linear layer as a weak decoder, input the enhanced features, reconstruct the fragment sequence, and restore all fragment sequences to obtain the reconstructed sequence.
3. The multi-time series anomaly detection method guided by sequence memory according to claim 2, characterized in that: In S1, the fragmentation operation is based on a sliding window operation, using two hyperparameters, window length and step size, to divide the input multi-time series into finer-grained small window forms S={S1, S2, ...., S...}. P Meanwhile, the shared linear layer used can express the features of different fragment sequences based on the encoding of the fragment sequence; after the fragmentation operation, the original shape (B, N, W) of the input sequence is transformed into (B, N, P, d), where B represents the size of each batch, N is the dimension of the input sequence, W is the window length of the input sequence, P is the number of fragments finally obtained by the fragmentation operation, and d is the output dimension of the shared linear layer.
4. The multi-time series anomaly detection method guided by sequence memory according to claim 3, characterized in that: S2 is specifically as follows: S2.1 Input sequence encoding, use depthwise convolution to extract features within the sequence segment. The kernel size of the depthwise convolution is fixed and greater than 1 to obtain the internal relationship of the segment sequence encoding, which is the internal feature encoding of the sequence. S2.
2. Input the sequence internal feature code obtained in S2.1, use point-direction convolution with a kernel size of 1 to extract features between channels, and obtain channel convolution codes of the same shape; S2.
3. Input the channel convolutional encoding obtained in S2.
2. After dimension swapping, use the same point-direction convolution to extract features between sequence segment encodings to obtain segment convolutional encoding. Then restore the dimensions. The resulting features are the sequence features. The shape of the final sequence features is consistent with the fragmented input features, which is (B, N, P, d).
5. The multi-time series anomaly detection method guided by sequence memory according to claim 3, characterized in that: S3 is specifically as follows: S3.1 Perform matrix multiplication between the sequence features and the prototype terms in the memory module, i.e., A = E c @ M′, where E c Let M' be the sequence feature, M' be the transpose of all prototype terms in the sequence memory, and @ denote matrix multiplication; finally, we obtain the attention relationship between the sequence feature and the prototype, which has the shape (B, N, P, k), where k is the number of prototype terms. S3.
2. Perform Softmax normalization on the attention relationship to obtain the prototype similarity ratio between 0 and 1; S3.3, Perform matrix multiplication between the prototype similarity ratio and the prototype in the memory module, i.e., E t = A@M, obtain the prototype feature E t The shape of the prototype features is consistent with that of the sequence features.
6. The multi-time series anomaly detection method guided by sequence memory according to claim 2, characterized in that: During the two-stage training, the reconstruction model only back-propagates the relevant gradients of the separable convolutional network and automatically updates the reconstruction model. The sequence memory prototype is manually updated based on attention and sequence features. The manual update uses the sequence features E obtained in S2. c The prototype M in the sequential memory term is first obtained through matrix multiplication E. c @ M′ obtains the attention relationship A between sequence features and prototypes; then, A is Softmax normalized to obtain the sequence feature-prototype similarity ratio; next, matrix multiplication A′@ Ec is performed between the similarity ratio and the sequence features to obtain the prototype modification term M. n Finally, the modified prototype and the original prototype are weighted and combined using an expert network and impulse to obtain the updated prototype, as shown in the following formula: In this context, U and V are both expert networks, M is the original prototype, Mn is the prototype modification term, and β controls the rate of prototype update.
7. The multi-time series anomaly detection method guided by sequence memory according to claim 1, characterized in that: The abnormal threshold τ is obtained as follows: After completing the two-stage training, the training set multiple time series are input into the two-stage reconstruction model again. The reconstruction error of the last time node of each sequence is taken as the basic abnormal score set. Then, the abnormal ratio is estimated by using the prior knowledge of the real scene task corresponding to the abnormal detection, and the abnormal threshold τ is obtained.