Time sequence anomaly detection method based on multi-cycle mode guidance
By introducing multi-period mode in time series anomaly detection, the problem of insufficient feature extraction in the existing methods is solved, and the detection accuracy and reconstruction effect are significantly improved.
Patent Information
- Application Number
- CN202510021344.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing time series anomaly detection methods rely on a single feature or pattern representation, resulting in insufficient feature extraction and limiting the reconstruction effect, especially when dealing with complex nonlinear relationships and high-dimensional data.
The time series anomaly detection method based on multi-period mode guidance is adopted. By deeply mining and capturing complex periodic mode information in the time series, these mode information is integrated into the feature space of the reconstruction model, and the reconstruction ability of the model is enhanced.
The accuracy of time series abnormality detection is significantly improved, and by expanding the reconstruction error of abnormal data, it makes normal and abnormal characteristics easier to distinguish.
Smart Images

Figure CN120067635A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series data anomaly detection, and particularly relates to a time series anomaly detection method guided by multi-period patterns. Background Art
[0002] With the rapid development of interconnected devices and sensors in the Internet of Things and industrial systems, a large amount of time series data is collected and used for analysis. The mining and analysis of this data can effectively identify abnormal behaviors of sensors, thereby preventing potential system failures. This process is called multivariate time series anomaly detection. Its core goal is to detect abnormal points that deviate from the normal pattern. Currently, time series anomaly detection technology has been widely applied in many fields, such as credit card fraud detection, industrial fault diagnosis, and medical pathology analysis.
[0003] In the past two decades, time series anomaly detection has been widely studied. Early methods were mainly based on traditional machine learning and statistical techniques, including clustering, distance analysis, density analysis, reconstruction, and prediction methods. However, these methods have obvious limitations in dealing with complex non-linear relationships and high-dimensional data. In recent years, time series anomaly detection methods based on deep learning have gradually replaced traditional methods due to their advantages in accuracy and effectiveness. Deep learning methods are mainly divided into two categories: prediction and reconstruction. However, due to the rapid change and unpredictability of time series, prediction-based models perform poorly in detecting anomalies, especially when the number of prediction time points increases, the error accumulates significantly. Therefore, reconstruction-based deep learning methods have become the research focus.
[0004] Reconstruction-based methods have achieved certain results by learning the latent representation of normal time series and detecting anomalies by analyzing the differences between the original sequence and the reconstructed sequence. However, existing methods rely on single feature or pattern representation of the input sequence, resulting in insufficient feature extraction, which limits the reconstruction effect. To solve this problem, it is particularly important to mine multi-period patterns in time series data. Multi-period patterns can more accurately describe the normal features of the sequence and be effectively integrated into the reconstruction model, thereby enhancing the reconstruction ability of the model and making it easier to distinguish normal and abnormal features. Summary of the Invention
[0005] The present invention aims to provide a time series anomaly detection method guided by multi-period patterns. This method analyzes time series from the perspective of multi-period patterns, deeply mines and captures complex periodic pattern information in time series, and integrates these unique pattern information into the feature space of the reconstruction model. While enhancing the model's reconstruction ability, this method effectively expands the reconstruction error of abnormal data, thereby significantly improving the accuracy of time series anomaly detection.
[0006] To solve the above technical problems, a time series anomaly detection method based on multi-period pattern guidance proposed by the present invention includes the following steps:
[0007] Step 1: Collect multivariate time series data to be detected for anomalies;
[0008] Step 2: Perform downsampling on the multivariate time series data collected in Step 1 to reduce the data scale and retain the main information; perform periodic analysis on the downsampled data, extract the k frequencies with the highest amplitudes, and calculate k periods based on the k frequencies; use the k periods to reshape the two-dimensional time series into a three-dimensional representation through Padding and Reshape operations, generating k three-dimensional time series representations; input the k three-dimensional time series representations into the Inception convolutional network to extract features. After the feature extraction is completed, reshape the k three-dimensional features back into k two-dimensional features through the Reshape operation for subsequent work;
[0009] Step 3: Perform feature sampling on the k two-dimensional feature representations reshaped in Step 2, adopt a two-stage sampling strategy, and sample the features within and between periods respectively, obtaining intra-period feature samples and inter-period feature samples; use the samples obtained from the above feature sampling as positive and negative examples for unsupervised contrastive learning to train the Inception convolutional network and optimize the feature spaces of different periods; finally, the features output by the Inception convolutional network represent k pattern information, and the k pattern information represents the feature representations corresponding to k periods of the multivariate time series data input in Step 1;
[0010] Step 4: Perform sliding window sampling on the multivariate time series data collected in Step 1, and input the feature representation of each window sample data into the encoder of the reconstruction model; fuse the k pattern information obtained in Step 3 into the feature representation of each window sample data in the form of similarity weights to enhance the reconstruction ability; obtain the final output through the decoder. The output is the reconstruction of the sliding window sampling samples of the input multivariate time series data. Use the mean absolute error between the reconstruction output of the decoder and the input of the encoder as the final anomaly score, set a threshold, compare the anomaly score with the threshold, and determine whether the multivariate time series is abnormal to complete the detection task.
[0011] Further, in the time series anomaly detection method of the present invention:
[0012] In Step 1, the multivariate time series data is industrial system data or Internet company server data.
[0013] The specific content of Step 2 includes:
[0014] The method of the downsampling process in Step 2-1) includes one of random downsampling, mean downsampling, and median downsampling; preferably, the original data samples are downsampled by taking the median, and the values at every 10 time points are combined into one measurement, retaining key information while reducing the data dimension.
[0015] The method of calculating k periods based on k frequencies in Step 2-2) is to use the fast Fourier transform method to extract the k frequencies with the highest amplitudes and calculate the Top-k main periods accordingly. The specific process is as follows:
[0016] Convert the downsampled data X ∈ R (Td×N) into the frequency domain representation, where Td represents the length of the downsampled time series, N represents the number of variables, and the conversion formula is as follows:
[0017]
[0018] In Equation (1), Amp(·) represents the calculation of amplitude, FFT(·) represents the fast Fourier transform, represents the calculation of the average value of N dimensions, and A represents the calculated amplitude of each frequency;
[0019] Considering the sparsity in the frequency domain and avoiding noise from high frequencies, only the first k amplitude values are selected to obtain the most significant frequencies:
[0020] {f 1 ,f 2 ,…,f k}=argTop-k(A) (2)
[0021] In Equation (2), argTop-k(·) represents selecting the k frequency values with the highest amplitudes from all amplitudes;
[0022] The cycle lengths {p 1 ,p 2 ,…,p k} are calculated from the selected frequencies, and the formula is as follows:
[0023]
[0024] In Equation (3), p k represents the k cycles corresponding to the kth frequency;
[0025] In Step 2-3), using the k cycles, the two-dimensional time series is reshaped into a three-dimensional representation through Padding and Reshape operations to generate k three-dimensional time series representations:
[0026]
[0027] In Equation (4), Padding(·) is an operation of zero-padding the time series in the time dimension to make it adaptable operation, p i and f i represent the number of rows and columns of the tensor respectively, represents the three-dimensional time series representation of the i-th period;
[0028] The k generated three-dimensional time series representations X 3D are respectively input into the Inception convolutional network to extract features. The Inception convolutional network contains multi-scale two-dimensional convolutional kernels and is one of the currently recognized visual backbone networks with excellent performance. It can efficiently capture the feature relationships within and between periods in the three-dimensional time series. This design can effectively improve the model's ability to extract complex periodic features and provide higher accuracy and robustness for subsequent anomaly detection. After feature extraction is completed, the feature representations are reshaped back into k two-dimensional forms:
[0029]
[0030] In Equation (5), represents the two-dimensional representation of the features of the i-th period, and Conv represents the Inception convolution.
[0031] The above design can effectively improve the model's ability to extract complex periodic features and provide higher accuracy and robustness for subsequent anomaly detection.
[0032] The specific content of Step 3 includes:
[0033] Step 3-1) Perform feature sampling on the reshaped k two-dimensional feature representations. A two-stage sampling strategy is used to sample the features within and between periods. This strategy is based on the following two key observations:
[0034] Sampling within a period: Continuously sample time windows within the same period to obtain intra-period feature samples from the same periodic pattern. Even if these windows span different states, their feature compositions are still highly similar. These samples are defined as intra-period samples, and their representations in the embedding space should be close to each other;
[0035] Sampling between periods: Calculate the average value of the samples in the continuous time windows of each period, representing the window samples of different periods. The time window data from different periods includes inter-period feature samples from different periodic patterns. These samples reflect significantly different pattern characteristics and are called inter-period samples. The representations of these features in the embedding space should be kept significantly distinct.
[0036] Step 3-2) Perform unsupervised training on the feature samples obtained in Step 3-1) to optimize the feature spaces of different cycles. The aim is to ensure that the feature representations within a cycle are close to each other, while enhancing the differences in feature representations between cycles. This method can help the model learn more distinct feature pattern representations, thereby improving the performance of anomaly detection. The specific form of the loss function is as follows:
[0037] Intra-cycle feature loss L intra , aiming to maximize the similarity between feature samples within the same cycle so that they are close to each other in the embedding space:
[0038]
[0039] In Equation (6), represents the average similarity, σ represents the sigmoid activation function, and o m is the feature representation of the m-th sample in the i-th cycle, representing the calculation of the similarity between different samples;
[0040] Inter-cycle feature loss L inter , aiming to minimize the similarity between feature samples between different cycles so that they are far from each other in the embedding space:
[0041]
[0042] In Equation (7), represents the average similarity, represents the average of consecutive samples in different cycles;
[0043] Add the intra-cycle loss L intra and the inter-cycle loss L inter as the final loss function LSE. This loss function can ensure that the features within a cycle are close to each other while making the features between cycles as far apart as possible, enabling the model to learn the pattern information unique to each cycle:
[0044] LSE = L intra + L inter (8)
[0045] Perform unsupervised training on the Inception convolutional network using the loss function LSE;
[0046] The trained Inception convolutional network outputs k pieces of pattern information, which represent the feature representations corresponding to k cycles of the input multivariate time series.
[0047] In step 4, features are extracted from the input data by the reconstruction model, and the similarities between these features and the pre-learned multi-period patterns are calculated. These similarities are used as weights to guide the fusion process of the multi-period patterns and the extracted features.
[0048] The encoder and decoder of the reconstruction model are respectively represented as:
[0049] Weights=Encoder(X in )×(Pattern) T (9)
[0050] X o =Decoder(Weights×Pattern+X c ) (10)
[0051] In equations (9) to (10), Encoder and Decoder respectively represent the encoder and decoder of the reconstruction model, X in represents the input data, X c represents the feature representation in the embedding space, Pattern represents the extracted multi-period pattern information, Weights represents the weights between each time point of X c and the multi-period pattern, and X o represents the output of the reconstruction model;
[0052] The anomaly score for each time point is Score:
[0053] Score =|X in - X o | (11)
[0054] Points with an anomaly score greater than the threshold are considered anomaly points, and the threshold is set by selecting the 90th percentile of all anomaly scores.
[0055] Compared with the prior art, the present invention has the following advantages:
[0056] (1) The time series anomaly detection method provided by the present invention starts from the perspective of multi-period patterns and proposes an effective multi-period pattern extraction method. The method includes two main modules: a period extraction module and a pattern representation module. The first module extracts periodic features from multi-variable time series. The second module obtains features within the same period and between different periods. And through the contrast loss function, it is ensured that the features within the same period are closely combined, while the features of different periods are clearly separated. This process facilitates the extraction of unique pattern information for each period.
[0057] (2) The time series anomaly detection method provided by the present invention is designed as a pluggable method, which can integrate multi-period pattern information into the feature space of any reconstruction model, and proposes a multi-period pattern fusion strategy that can adaptively adjust the feature representation. For inputs highly similar to the normal pattern, the fusion process retains more features of the normal pattern. For abnormal data, due to its difference from the normal pattern, the fusion process generates a large deviation. This design amplifies the difference in reconstruction error between normal data and abnormal data, improving the sensitivity and accuracy of anomaly detection. Description of the Drawings
[0058] Figure 1 It is a flowchart of the time series anomaly detection method based on multi-period pattern guidance of the present invention.
[0059] Figure 2 It is a schematic structural diagram of the time series anomaly detection model based on multi-period pattern guidance of the present invention. Detailed Embodiments
[0060] A time series anomaly detection method based on multi-period pattern guidance proposed by the present invention aims to achieve accurate and efficient multivariate time series anomaly detection. Its design concept is: using multi-period patterns to guide the reconstruction process of the model, aiming to improve the feature extraction ability of complex time series. By extracting periodic features from multivariate time series, using unsupervised loss to capture pattern information of different periods, and fusing this pattern information into the feature space of the reconstruction model, the performance of time series anomaly detection is significantly enhanced.
[0061] To more clearly elaborate and explain the purpose, measure solutions, and key points of the present invention, the method proposed by the present invention will be introduced in detail below with reference to the accompanying drawings. However, the following embodiments are by no means restrictive of the present invention.
[0062] Figure 1 The step flow of the method is shown, mainly including:
[0063] The collected multivariate time series data is downsampled and then subjected to periodic analysis, converting the two-dimensional time series into a three-dimensional time series representation, inputting it into the Inception convolutional network to extract features, reshaping the extracted three-dimensional features back into a two-dimensional feature representation through a Reshape operation, and performing feature sampling. Through an unsupervised loss function, the Inception convolutional network is trained to optimize the feature space of different periods, ensuring that the features within the same period are aggregated together and the features between different periods are separated from each other, thereby extracting unique pattern information for each period; the pattern information output by the Inception convolutional network represents the corresponding feature representation of the input multivariate time series data;
[0064] Perform sliding window sampling on the collected multivariate time series data, and input each window sample data into the encoder of the reconstruction model to extract the feature representation of each window sample data;
[0065] Fuse the pattern information output by the Inception convolutional network into the feature representation of each window sample data in the form of similarity weights, which can effectively improve the reconstruction accuracy of the model for normal samples and the anomaly detection ability; After the decoder reconstructs the sliding window sampling samples of the input multivariate time series data, use the mean absolute error between the reconstruction output of the decoder and the input of the encoder as the final anomaly score, and set a threshold for anomaly detection.
[0066] Figure 2 Fig. shows the model structure of the time series anomaly detection method based on multi-period pattern guidance of the present invention. The main technical problem to be solved is: in the multivariate time series anomaly detection problem, given a historical multivariate time series X = {x 1 , x 2 ,... x t ,... x T} ∈ R T×N , where x t is the value of the sensor at time point t, T represents the length of the time series, N represents the number of variables. In unsupervised anomaly detection, assume that the training sample X train is completely composed of normal samples, while the test sample X test contains both normal and abnormal data. The labels of the test set are represented as y = {y 1 , …, y T} ∈ {0, 1}, where y t = 0 indicates a normal point at time point t, y t = 1 indicates an abnormal point at time point t. Finally, the reconstruction error between the reconstructed sequence and the input sequence is used as the anomaly score for each time point. If the anomaly score of a point is greater than the decision threshold, then that point is regarded as an abnormal time point.
[0067] As Figure 1 and Figure 2 shown, the specific steps of the time series anomaly detection method based on multi-period pattern guidance of the present invention are as follows:
[0068] Step 1: Collect and prepare the multivariate time series data to be detected for anomalies;
[0069] Step 2: Downsample the multivariate time series data. Use the period acquisition module to perform period decomposition on the multivariate time series, and convert the original two-dimensional time series into a three-dimensional time series based on the extracted periods. It consists of three core operations: downsampling, period extraction, and dimension reshaping.
[0070] 2-1) The downsampling operation is specifically as follows:
[0071] Traditional methods use a sliding window to sample the input data and extract local periods. This method not only limits the acquisition of local period information but also increases the amount of data processed by the model. The present invention extracts the period of the entire sequence through a global downsampling method, which can obtain more accurate and complete global period information. Compared with local window sampling, global downsampling can extract period features more comprehensively, significantly accelerate the training speed, and improve the model performance.
[0072] Specifically, the median downsampling method is adopted: take one median value from every 10 data points, and reduce the values of the original 10 time points to this median value.
[0073] X d = Downsample(X train ,r)
[0074] where X d is the result after downsampling, Downsample(·) is the median downsampling, r is the downsampling factor, and in this method, it is set to 10.
[0075] 2-2) The specific operation of period extraction is as follows:
[0076] Considering the sparsity in the frequency domain and better avoiding noise from high frequencies, only the first k amplitude values are selected and the most significant frequencies are obtained. As is well known, in the fast Fourier transform, the frequencies with larger amplitudes represent the main periods, which helps to capture more accurate period characteristics.
[0077] First, calculate the frequency domain representation of the time series:
[0078]
[0079] where Amp(·) represents the calculation of amplitude, FFT(·) represents the fast Fourier transform, represents the calculation of the average value of N dimensions. A represents the calculated amplitude of each frequency. Since the fast Fourier output has symmetry, only half (the independent information part) is taken to represent the spectrum, and the other half is the redundant conjugate symmetric part. Therefore, A ∈ R L ,{L=(T / 2)+1}.
[0080] Select the k frequency values with the highest amplitudes from all amplitudes:
[0081] {f 1 ,f 2 ,…,f k} = argTop-k(A)
[0082] Calculate k periods based on the frequency:
[0083] k ∈ {1, …, k}; p k represents the period corresponding to the k-th frequency. In this embodiment, k is set to 5 periods.
[0084] 2-3) The specific operation of dimensional reshaping is as follows:
[0085] The changing characteristics of the time series can be divided into two categories: intra-period changes and inter-period changes. Intra-period changes reflect the short-term time patterns of a single period, while inter-period changes reflect the long-term trends across periods. Traditional one-dimensional time series are difficult to simultaneously and explicitly present these two types of changes, which limits the comprehensive capture of the complex dynamics of the time series. To overcome this limitation, the present invention extends the time change analysis to a multi-dimensional space.
[0086] Taking a univariate time series as an example, we reshape the one-dimensional time series into a two-dimensional tensor. Specifically, each column represents the time points within a time period, and each row contains the time points at the same phase in different time periods. Through this transformation, we have successfully broken through the representation bottleneck of the one-dimensional space and achieved the unified presentation of intra-period and inter-period changes. Generalizing this method to multi-variate time series, the two-dimensional representation can be further reshaped into a three-dimensional representation.
[0087]
[0088] where Padding(·) is an operation of zero-padding the time series in the time dimension to make it adapt to the operation, p i and f i represent the number of rows and columns of the tensor respectively.
[0089] Step three: Use Inception convolution to learn the feature representation of each period sequence. By sampling the intra-period and inter-period features and using an unsupervised loss function to learn the feature patterns corresponding to each period. It mainly consists of three core operations: Incption convolution operation, two-stage sampling operation, and unsupervised loss calculation operation.
[0090] 3-1) The specific operation of Incption convolution is:
[0091] The Inception network is currently recognized as one of the vision backbone networks with excellent performance. It can efficiently capture the intra-cycle and inter-cycle feature relationships in three-dimensional time series. The network consists of two consecutive inception blocks, connected by the GELU activation function to capture multi-scale time features. Each inception block contains six parallel 2D convolutional layers with gradually increasing kernel sizes, enabling the model to capture local details and global patterns. The number of output channels for all convolutional layers remains unchanged to balance the contributions of features at different scales. After feature extraction, the feature representation is reshaped back into a two-dimensional form, which can effectively improve the model's ability to extract complex periodic features and provide higher accuracy and robustness for subsequent anomaly detection.
[0092]
[0093] Among them represents the two-dimensional representation of the features of the i-th cycle extracted by the convolutional network, where d model represents the output feature dimension, and this parameter is set to 128 in this embodiment.
[0094] 3-2) The specific operation of the two-stage sampling operation is as follows:
[0095] The purpose of the two-stage sampling strategy is to sample the intra-cycle and inter-cycle features. This strategy is based on the following two key observations:
[0096] Intra-cycle sampling: In consecutive time windows within the same cycle, the feature samples are from the same periodic pattern. Even if these windows span different states, their feature compositions are still highly similar. These samples are defined as intra-cycle samples, and their representations in the embedding space should be close to each other.
[0097] Inter-cycle sampling: Time windows from different cycles contain feature samples from different periodic patterns. These samples reflect significantly different pattern characteristics and are called inter-cycle samples. The representations of these features in the embedding space should remain significantly distinguishable.
[0098] Specifically, the process of intra-cycle sampling is as follows: First, randomly select n consecutive windows on the feature representation of each cycle. Each window is regarded as an intra-cycle sample. The position t of the first window ∼ U(0, T - w - n) is randomly drawn from a uniform distribution, where w is the window length. Then, move forward n - 1 times along the time axis with a step size of 1 to obtain n feature samples within each cycle. The n feature samples from different cycles then form the inter-cycle feature samples. In this embodiment, n is set to 30 and w is set to 64.
[0099] 3-3) The specific operation of the unsupervised loss is as follows:
[0100] The unsupervised loss consists of two parts, namely intra-cycle loss and inter-cycle loss. This loss function aims to ensure that the feature representations within a cycle are close to each other while enhancing the difference in feature representations between cycles. This can not only guarantee the similarity of features within a cycle but also make the features between cycles as far apart as possible, which is beneficial for the model to learn more distinct feature patterns, thereby improving the anomaly detection performance. The specific form of the loss function is as follows:
[0101] LSE = L intra +L inter
[0102] Among them, the intra-cycle feature loss L intra aims to maximize the similarity between feature samples within the same cycle so that they are close to each other in the embedding space:
[0103]
[0104] where represents the average similarity, σ represents the sigmoid activation function, and o m is the feature representation of the m-th window in the i-th cycle, represents calculating the similarity between different windows.
[0105] Among them, the inter-cycle feature loss L inter aims to minimize the similarity between feature samples between different cycles so that they are far from each other in the embedding space:
[0106]
[0107] where represents the average similarity. It should be noted that the negative sign before c i indicates that this loss function needs to be minimized. Given that calculating the embedding similarity of all cycle samples is too expensive, the present invention adopts an optimization strategy: using the similarity between different groups of embedding centers, and calculating the embedding center through the following formula:
[0108]
[0109] This embodiment adopts this loss function to keep the feature representations within a cycle close to each other while keeping the feature representations between cycles different. This method helps the model learn more discriminative feature pattern representations.
[0110] The Inception convolutional network is trained unsupervised using the loss function LSE; the k pattern information output by the trained Inception convolutional network represents the feature representations corresponding to k cycles of the input multivariate time series.
[0111] Step 4: Calculate the similarity between the features of the input data extracted by the reconstruction model and the multi-period patterns using a pattern fusion method. Based on the similarity as the weight, effectively integrate the multi-period patterns into the feature space of the reconstruction model, thereby retaining more normal pattern features and improving the reconstruction accuracy. Specifically:
[0112] Perform sliding window sampling on the input time series to obtain the input X of the reconstruction model in ∈R L×N , where L represents the length of the sliding window. The feature representation of X' is extracted by the encoder of the reconstruction model. Next, we calculate the similarity between these extracted features and the pre-learned normal patterns. These similarities are used as weights to guide the fusion of the normal patterns and the extracted features. The specific steps are as follows:
[0113] Weights = Encoder(X in )×(Pattern) T
[0114] X o = Decoder(Weights×Pattern + X c )
[0115] where Encoder and Decoder represent the encoder and decoder of the reconstruction model respectively, X in represents the input data, X c represents the feature representation in the embedding space of the reconstruction model, Pattern represents the extracted multi-period pattern information, Weights represents the weight between each time point of X c and the multi-period pattern, and X o represents the output of the reconstruction model.
[0116] The key of this fusion mechanism lies in its ability to adaptively adjust the feature representation. For inputs highly similar to the normal patterns, the fusion process retains more features of the normal patterns. For abnormal data, due to its differences from the normal patterns, the fusion process generates a large deviation. This design amplifies the difference in the reconstruction errors between normal and abnormal data, improving the sensitivity and accuracy of anomaly detection.
[0117] In this embodiment, the reconstruction error is selected as the final anomaly score, and a threshold is set for anomaly detection:
[0118] Score = |X - X O |
[0119] In summary, the time series anomaly detection method described in the present invention analyzes the time series from the perspective of multi-period patterns, deeply mines and captures the complex periodic pattern information in the time series, and integrates these unique pattern information into the feature space of the reconstruction model. While enhancing the model's reconstruction ability, this method effectively expands the reconstruction error of abnormal data, thus significantly improving the accuracy of time series anomaly detection.
[0120] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many improvements and changes without departing from the purpose of the present invention, and all of these fall within the scope of protection of the present invention.
Claims
1. A time series anomaly detection method based on multi-period pattern guidance, characterized in that: The following steps are involved: Step 1: Collect multivariate time series data for anomaly detection; Step 2: Downsampling the multivariate time series data collected in step 1, performing periodic analysis on the downsampled data, extracting k frequencies with the highest amplitudes, and calculating k periods based on the k frequencies; using the k periods, reshaping the two-dimensional time series into a three-dimensional representation through Padding and Reshape operations to generate k three-dimensional time series representations; inputting the k three-dimensional time series representations into the Inception convolutional network to extract features, and after the feature extraction is completed, reshaping the extracted k three-dimensional features back into k two-dimensional features through the Reshape operation; Step 3: Perform feature sampling on the k two-dimensional feature representations reshaped in step 2, including intra-cycle feature samples and inter-cycle feature samples obtained by intra-cycle sampling and inter-cycle sampling; use the samples obtained by the above feature sampling as positive and negative examples for unsupervised comparative learning, train the Inception convolutional network, and optimize the feature space of different cycles; finally, the features output by the Inception convolutional network represent k pattern information, which represents the feature representation corresponding to the k cycles of the multivariate time series data input in step 1; Step 4: Perform sliding window sampling processing on the multivariate time series data collected in step 1, and input each window sample data into the encoder of the reconstruction model to extract the feature representation of each window sample data; The k pattern information obtained in step 3 is fused into the feature representation of each window sample data in the form of similarity weight; The final output is obtained after the decoder, which is a reconstruction of the sliding window sampling samples of the input multivariate time series data. The mean absolute error between the decoder's reconstructed output and the encoder's input is used as the final anomaly score, and a threshold is set for anomaly detection.
2. The time series anomaly detection method according to claim 1, characterized in that: In step one, the multivariate time series data is industrial system data or Internet company server data.
3. The time series anomaly detection method according to claim 1, characterized in that: The specific contents of step 2 include: The downsampling method in step 2-1) includes one of random downsampling, mean downsampling and median downsampling; Step 2-2) The method of calculating k cycles based on k frequencies is to use the fast Fourier transform method to convert the downsampled data X∈R (Td×N) Converted into frequency domain representation, where Td represents the length of the time series after downsampling, and N represents the number of variables. The conversion formula is as follows: In formula (1), Amp(·) represents the calculation of amplitude, FFT(·) represents fast Fourier transform, represents the calculation of the average value of N dimensions, and A represents the calculated amplitude of each frequency; Select the first k amplitude values and get the most significant frequency: {f1,f2,…,f k }=argTop-k(A) (2) In formula (2), argTop-k(·) represents selecting k frequency values with the highest amplitude from all amplitudes; The length of the cycle {p1,p2,…,p k} Calculated by the selected frequency, the formula is as follows: In formula (3), p k represents the k cycles corresponding to the kth frequency; Step 2-3) Using the k cycles, the two-dimensional time series is reshaped into a three-dimensional representation through Padding and Reshape operations to generate k three-dimensional time series representations: In formula (4), Padding(·) is a zero-filling operation on the time series in the time dimension to make it fit Operation, p i and f i Represent the number of rows and columns of the tensor, respectively. represents the three-dimensional time series representation of the i-th period; The generated k three-dimensional time series are represented as X 3D They are respectively input into the Inception convolutional network to extract features. The Inception convolutional network contains multi-scale two-dimensional convolution kernels. After the feature extraction is completed, the feature representation is reshaped back to k two-dimensional forms: In formula (5), It represents the two-dimensional representation of the features of the i-th cycle, and Conv represents Inception convolution.
4. The time series anomaly detection method according to claim 1, characterized in that: The specific contents of step three include: Step 3-1) performs feature sampling on the reshaped k two-dimensional feature representations, including: Intra-cycle sampling: continuous time window sampling is performed within the same cycle to obtain intra-cycle feature samples from the same cycle pattern; Inter-cycle sampling: Calculate the average value of continuous time window samples in each cycle, representing window samples of different cycles, including time window data from different cycles containing inter-cycle feature samples derived from different cycle patterns; Step 3-2) performs unsupervised training on the feature samples obtained in step 3-1) to optimize the feature space of different cycles; the specific steps are as follows: The feature loss L within the cycle intra , aims to maximize the similarity between feature samples in the same cycle so that they are close to each other in the embedding space: In formula (6), represents the average similarity, σ represents the sigmoid activation function, o m is the feature representation of the mth sample in the ith period, Indicates calculating the similarity of different samples; The feature loss between cycles L inter , which aims to minimize the similarity between feature samples in different cycles so that they are far away from each other in the embedding space: In formula (7), represents the average similarity, Represents the average value of consecutive samples of different periods; The loss L in the period intra The loss between cycles L inter Add them together as the final loss function LSE: LSE=L intra +L inter (8) Use the loss function LSE to perform unsupervised training on the Inception convolutional network; Step 3-3) The trained Inception convolutional network outputs k pattern information, which represents the feature representation corresponding to the k periods of the input multivariate time series.
5. The time series anomaly detection method according to claim 1, characterized in that: In step 4, the encoder and decoder of the reconstruction model are expressed as: Weights=Encoder(X in )×(Pattern) T (9) X o =Decoder(Weights×Pattern+X c ) (10) In equations (9) and (10), Encoder and Decoder represent the encoder and decoder of the reconstruction model respectively, and X in represents the input data, X c represents the feature representation in the embedding space, Pattern represents the extracted multi-cycle pattern information, and Weights represents X c The weight between each time point and the multi-period pattern, X o represents the output of the reconstruction model; The anomaly score at each time point is Score: Score=|X in -X o | (11) A point with an anomaly score greater than a threshold is considered an anomaly point, and the threshold is set by selecting the 90% quantile of all anomaly scores.
Citation Information
Patent Citations
Cloud network end resource multi-dimensional time sequence anomaly detection method based on multi-scale decoding
CN115169430A
Method for detecting time series data anomalies generated by infrastructure equipment in network
CN116738302A
Electrocardio atrial fibrillation detection method and device based on self-supervised learning, equipment and medium
CN117679042A
Time sequence anomaly detection method and system based on differential multi-resolution decomposition
CN118094443A
Unsupervised multivariate time sequence anomaly detection method based on auto-encoder
CN118820922A