Multivariable time sequence anomaly detection method based on space-time diagram learning
Through the space-time graph learning method, combining the correlation learning layer and multiple neural network modules, the space-time dependence relationship of multivariate time series is captured, which solves the problems of blurred boundaries and limited detection capabilities of traditional methods in abnormal detection, and achieves higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510213759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional methods have difficulty accurately defining the boundaries between abnormal behavior and non-abnormal behavior in multivariable time series data, and their detection capabilities in large-scale data sets are limited.
The method based on space-time graph learning is adopted to capture paired correlations through the correlation learning layer, and a spatio-time graph neural network is constructed by combining long and short-term memory networks, time convolution networks and graph convolution networks to capture the time and space dependencies in the multivariate time series, and output the exception scores using the PCA exception scorer.
The accuracy of multivariate time series spatiotemporal modeling is improved, the accuracy and robustness of anomaly detection are improved, and the effectiveness and universality of anomaly detection are enhanced.
Smart Images

Figure CN120145252A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a multivariate time series anomaly detection method based on spatiotemporal graph learning, and belongs to the technical field of time series anomaly detection. Background Art
[0002] Cyber-Physical Systems (CPSs) are a new generation of intelligent systems that integrate computing, communication, and control. They are widely used in medical monitoring, intelligent transportation, smart grids, and other fields. With the rapid development of CPSs, the amount of time series data has grown exponentially. In practical applications, the time series data generated by multiple devices or sensors form a complex multivariate time series. However, anomalies in time series usually indicate a state that does not match the expected pattern, which may represent a failure and service unavailability. [1] .
[0003] In order to effectively monitor large-scale systems and provide early warnings, anomaly detection technology is essential, which is usually achieved by mining the potential information in multivariate time series to identify data points that significantly deviate from other data points. However, methods based on statistical models lack the ability to capture nonlinear spatiotemporal relationships. The application of traditional machine learning methods greatly improves efficiency and accuracy compared to statistical model methods, but cannot capture the complex structure in the data and cannot be extended to large-scale data sets for detection. For the above traditional methods, the boundary between abnormal behavior and non-abnormal behavior in time series data is fuzzy and not precisely defined. However, the introduction of deep learning methods has broken the limitations of traditional methods to a certain extent. It can automatically learn hierarchical discriminant features from time series data. Later, the introduction of dual-channel deep learning methods based on graph neural networks can explicitly learn and capture the pairwise correlation of variables in multivariate time series, providing more accurate spatiotemporal modeling and more robust performance when dealing with more variables and more complex anomaly detection tasks.
[0004] The technology of this invention comes from the enterprise commissioned R&D project (KKK0202403218); Yunnan Province Major Science and Technology Special Plan (202302AD080002); Yunnan Province Computer Technology Application Key Laboratory Open Fund (CB22144S073A); "Xingdian Talent Support Plan" Industrial Innovation Talent Project (Yunnan Development and Reform Personnel
[2019] No. 1096); Yunnan Provincial Department of Education Postgraduate Research Fund (2024Y129). Summary of the invention
[0005] The technical problem solved by the present invention is as follows: The present invention provides a multivariate time series anomaly detection method based on spatio-temporal graph learning to solve the problem that the boundary between abnormal behavior and non-abnormal behavior in time series data is fuzzy and not precisely defined. The present invention improves the accuracy of spatio-temporal modeling of multivariate time series, enhances the accuracy and robustness of anomaly detection, and thus improves the effectiveness and universality of anomaly detection.
[0006] The technical solution of the present invention is: A multivariate time series anomaly detection method based on spatio-temporal graph learning, the method comprising:
[0007] Using a correlation learning layer to capture pairwise correlations and obtain a graph adjacency matrix reflecting the correlation relationship between variables; then, inputting the graph adjacency matrix into a spatio-temporal graph neural network composed of a long short-term memory network, a temporal convolutional network, and a graph convolutional network, and the spatio-temporal graph neural network captures the temporal and spatial dependencies in the multivariate time series and accurately models the spatio-temporal relationship; finally, using a PCA-based anomaly score to output an anomaly score.
[0008] Further, the method comprises the following steps:
[0009] Step1. Perform randomly initialized node embeddings on each multivariate time series, and adaptively learn the graph adjacency matrix through the correlation learning layer in an end-to-end manner;
[0010] Step2. Input the sliding window into the hidden space through a convolutional layer and take the residual channel as the output, and then input the residual into the subsequent spatio-temporal graph neural network;
[0011] Step3. Input the graph adjacency matrix obtained in Step1 and the residual obtained in Step2 into a spatio-temporal graph neural network composed of three neural network modules, namely, a long short-term memory network LSTM, a temporal convolutional network TCN, and a graph convolutional network GCN, to capture spatio-temporal dependencies and merge them through skip connections to obtain encapsulated spatio-temporal hidden features;
[0012] Step4. The prediction head further projects the spatio-temporal hidden features into a single-step prediction result;
[0013] Step5. Input the current prediction result and all observation-prediction pairs calculated before the time stamp, and use a PCA-based scorer to calculate the normalized prediction error and output the anomaly metric score in real time.
[0014] Furthermore, the multivariate time series in Step1 can be used for anomaly detection by adopting the time series in the WADI dataset or the SMD dataset, or the multivariate time series data composed of the CPU usage rate, memory usage rate, database query response time, fault log, server load, and request count during the operation of the intelligent hospital system.
[0015] Furthermore, Step1 includes:
[0016] Design a correlation learning layer to adaptively learn the graph adjacency matrix X, where nodes and edges represent variables and variable connectivity; the calculation process of the graph adjacency matrix X is as follows:
[0017]
[0018] where are two randomly initialized node embedding matrices, are two sets of trainable parameters, the hyperparameter α represents the non-linear activation saturation rate, and both tanh() and ReLU() represent activation functions;
[0019] Introduce a hyperparameter k to control the sparsity of the graph adjacency matrix X, that is, further mask the elements that are zero except for the top-k nearest neighbors of each node in the learned graph adjacency matrix X; for the i-th row in the graph adjacency matrix X, it is processed as follows:
[0020]
[0021] where argmax(, k) returns the first k largest indices in the input vector, i represents the i-th row of the graph adjacency matrix X, X[i, :] represents the elements of the i-th row of the graph adjacency matrix X, topk represents the k nodes closest to each node, -topk represents the nodes other than the k nodes closest to each node, and X[i, -topk] represents the elements of the i-th row of the graph adjacency matrix X other than its k closest elements.
[0022] Furthermore, Step3 includes:
[0023] Design a graph convolutional network GCN to capture the spatial correlation between variables; for the given graph adjacency matrix X and the initial input state G in , the definition of the given graph convolutional network GCN is as follows:
[0024]
[0025] where k represents the depth of graph propagation, G 0 = G in , represents the normalized graph adjacency matrix, and and Let \(I\) denote the identity matrix, and \(X_{ij}\) ij denote the element in the \(i\)-th row and \(j\)-th column of the graph adjacency matrix \(X\). and is an intermediate bridge for normalizing the graph adjacency matrix \(X\) to and has no practical meaning; use the parameter \(\beta\) to control the amount of information retained to avoid over-smoothing; assign different weights \(\Theta_k\) k to the hidden state \(G_k\) of the nodes, where \(k\in\{0,\ldots,K\}\); use different graph adjacency matrices \(X\) and \(\hat{X}\) k to separately store the incoming and outgoing information of the nodes, \(G_k^l\) T denote the state at the \(k\)-th depth of graph propagation, and \(\hat{G}_k^l\) k denote the hidden state at the \(k\)-th depth of graph propagation, and \(G^L\) k+1 denote the complete output of the graph convolutional network GCN, and \(\hat{\Theta}_k\) out denote the product of the hidden state at the \(k\)-th depth of graph propagation and the weight assigned to this graph propagation state; k \(\Theta_k\) k denote the product of the hidden state at the \(k\)-th depth of graph propagation and the weight assigned to this graph propagation state;
[0026] When the multivariate time series flows out of the LSTM, the output \(h_t\) out ; the output \(h_t\) out is expressed as:
[0027] \(h_t\) out \(=\) out \(o_t\) out \(\cdot\tanh(C_t)\)
[0028] where \(h_t\) out denote the output of the last hidden layer of the LSTM, \(o_t\) out denote the final output of the output gate, and \(C_t\) out denote the final cell state;
[0029] In the process of using the entire LSTM-TCN network to capture time dependencies, the output sequence of the LSTM serves as the input sequence of the TCN and flows into the TCN in a residual input manner; for the TCN, the specific definition is as follows:
[0030] \(T^l\) l+1 \(=\Gamma(T^{l - 1},P^l)\) l \(+\) l+1 \(TCN(T^{l - 1},\Phi^l)\), \(l\in\{0,\ldots,L\}\) l \(+\) l \(TCN(T^{l - 1},\Phi^l)\), \(l\in\{0,\ldots,L\}\)
[0031] where the output \(T^l\) out \(=\) L \(T^{l - 1}\), and \(TCN(\cdot,\Phi^l)\) l is the \(\Phi^l\) of the \(l\)-th layer lThe parameterized convolution function, Γ(T l ,P l+1 ) indicates that along T l The sequence length axis starts from T l Intercept the last P l+1 elements, where
[0032] P l+1 =P l -q l ×(k-1) and P 1 =Q-k+1, Q, q, k represent the size of the receptive field, the dilation factor and the size of the convolution kernel respectively, k also represents the depth of graph propagation, l refers to the lth layer of temporal convolution, P l+1 It represents the number of elements intercepted by the l-th layer of temporal convolution, T l+1 It represents the state of the lth layer of temporal convolution, T l represents the hidden state of the lth layer of temporal convolution, Φ l is the parameterization of the l-th layer of temporal convolution;
[0033] The gating mechanism is used to guide the information flow. The process of using the gating mechanism to guide the information flow is expressed as:
[0034]
[0035] Among them, f C (·) and f G (·) represents time-filtered convolution and gated convolution respectively, ⊙ represents the product of elements, and They represent the parameterization of the l-th layer of filter convolution and gated convolution respectively; the definitions of filter convolution and gated convolution are as follows:
[0036]
[0037] Among them, ★Δ represents the dilated convolution operation; and It is set to consist of multiple convolutional filters with width h∈{2,3,6,7} and composition, and represents the gated convolutional unit, and h represents the width of the convolution;
[0038] Based on the completion of the above temporal and spatial dependency modeling work, a spatiotemporal graph neural network is obtained; the processing process of the spatiotemporal graph neural network is expressed as:
[0039]
[0040] Among them, Θ represents the weight, T0 Refers to the initial input state of the temporal convolution.
[0041] Furthermore, Step4 includes:
[0042] Use the final output state T of the spatio-temporal graph neural network out To achieve single-step prediction through a multi-layer perceptron, and obtain the result of single-step prediction The result of single-step prediction Is expressed as:
[0043]
[0044] Among them, MLP() represents the prediction process of the multi-layer perceptron, and T out Is the final output state of the spatio-temporal graph neural network, and W mlp Refers to using a multi-layer perceptron for prediction.
[0045] Furthermore, Step5 includes:
[0046] For any univariate time series Calculate the absolute prediction error of the current timestamp Z Then normalize the absolute prediction error value; the normalization process is expressed as:
[0047]
[0048] Among them, And Are respectively the median and interquartile range of the error values in the sliding window Of, and W a Is the length of the sliding window, Represents the prediction result of the univariate time series Of, Represents the normalized result;
[0049] After normalizing each univariate time series, a multivariate time series normalized error vector of the current timestamp is obtained Then use PCA to summarize the normalized errors to obtain the anomaly score.
[0050] Furthermore, Step5 also includes:
[0051] Calculate the normalized error in the validation set And by exploring the validation mean vector Covariance matrix and the relationship between the orthogonal eigenvectors J to fit the normalized error; where J consists of H orthogonal eigenvectors related to the H largest eigenvalues in the diagonal matrix Λ, and this diagonal matrix Λ can be derived from the formula E v = JΛJ -1 To adapt to PCA, the normalized error at the current timestamp is reconstructed through the following definition:
[0052]
[0053] In the above equation, first, the normalized error at the current timestamp is subtracted by the average validation error to achieve zero-centering, and then these results are projected using the validation eigenvector J; in addition, only the first L principal components are retained; finally, the zero-centering inference is restored by adding the validation average error to the L-dimensional reconstructed normalized error after dimensionality reduction; is the reconstructed normalized error;
[0054] To calculate the final anomaly score at the current timestamp, the reconstructed normalized error is used to take the L1 distance between the denoised normalized error and the original normalized error as the final anomaly score X(Z); the final anomaly score X(Z) is expressed as:
[0055]
[0056] The present invention also provides a multivariate time series anomaly detection system based on spatio-temporal graph learning, and the system includes: a module for executing the multivariate time series anomaly detection method based on spatio-temporal graph learning described above.
[0057] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the multivariate time series anomaly detection method based on spatio-temporal graph learning.
[0058] The beneficial effects of the present invention are:
[0059] 1. Compared with traditional methods, the present invention can adaptively learn and capture the pairwise correlations between variables, and accurately model the spatio-temporal dependence relationships;
[0060] 2. The present invention can capture the spatial dependence information of multi-hop neighbors, which can better help the model master and understand the global data information, rather than only focusing on local abnormal perturbations. When facing more variables and more complex anomaly detection tasks, the accuracy of multi-variable time series spatio-temporal modeling is improved, and the accuracy and robustness of anomaly detection are enhanced, thereby enhancing the effectiveness and universality of anomaly detection.
[0061] 3. Experimental results show that the performance has been improved both in terms of overall detection performance and early detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is the flowchart of the steps of the present invention;
[0063] Figure 2 is a bar chart comparing the overall detection performance of the present invention and other baseline methods on the WADI dataset;
[0064] Figure 3 is a bar chart comparing the overall detection performance of the present invention and other baseline methods on the SMD dataset;
[0065] Figure 4 is a line chart comparing the early detection performance of the present invention and other baseline methods under six different delay constraints on the WADI dataset;
[0066] Figure 5 is a line chart comparing the early detection performance of the present invention and other baseline methods under six different delay constraints on the SMD dataset. DETAILED DESCRIPTION OF THE INVENTION
[0067] Example 1: As Figures 1 - 5 shown, a multi-variable time series anomaly detection method based on spatio-temporal graph learning includes: using a correlation learning layer (CL) to capture pairwise correlations and obtain a graph adjacency matrix reflecting the correlation relationship between variables; then, inputting the graph adjacency matrix into a spatio-temporal graph neural network (STGNN) composed of a long short-term memory network (LSTM), a temporal convolutional network (TCN), and a graph convolutional network (GCN). The spatio-temporal graph neural network captures rich temporal and spatial dependence relationships in multi-variable time series and accurately models the spatio-temporal relationships; finally, using a PCA-based anomaly score to output an anomaly score; improving the accuracy of multi-variable time series spatio-temporal modeling, enhancing the accuracy and robustness of anomaly detection, thereby enhancing the effectiveness and universality of anomaly detection.
[0068] The multivariate time series can adopt the time series in the WADI dataset and the SMD dataset; for applications in the intelligent hospital system, the method of the present invention can be used to perform anomaly detection on the multivariate time series data composed of the CPU usage rate, memory usage rate, database query response time, fault logs, server load, and number of requests, etc. during the operation of the system, and then analyze and predict potential attack behaviors and the occurrence of other abnormal activities in the system, so as to provide early warning for system security and ensure the security of the intelligent hospital system.
[0069] Further, the method includes the following steps:
[0070] Step1. Perform randomly initialized node embeddings on each multivariate time series, and adaptively learn the graph adjacency matrix through the correlation learning layer in an end-to-end manner;
[0071] Further, Step1 includes:
[0072] In order to model the pairwise dependencies between multivariate time series, a correlation learning layer is designed to adaptively learn the graph adjacency matrix X, where nodes and edges represent variables and variable connectivity; the calculation process of the graph adjacency matrix X is as follows:
[0073]
[0074] where are two randomly initialized node embedding matrices, are two sets of trainable parameters, the hyperparameter α represents the non-linear activation saturation rate, and both tanh() and ReLU() represent activation functions;
[0075] In order to facilitate optimization and reduce the computational complexity, a hyperparameter k is introduced to control the sparsity of the graph adjacency matrix X, that is, by further masking the elements that are zero except for the top-k nearest neighbors of each node in the learned graph adjacency matrix X; for the i-th row in the graph adjacency matrix X, it is processed as follows:
[0076]
[0077] where argmax(, k) returns the first k largest indices in the input vector, i represents the i-th row of the graph adjacency matrix X, X[i, :] represents the elements of the i-th row of the graph adjacency matrix X, topk represents the k closest nodes to each node, -topk represents the nodes other than the k closest nodes to each node, and X[i, -topk] represents the elements of the i-th row of the graph adjacency matrix X other than its k closest elements.
[0078] Step 2: Input the sliding window into the latent space through a 1×1 convolutional layer, and use the residual channels as the output. Then, input the residual into the subsequent spatio-temporal graph neural network;
[0079] Step 3: Feed the graph adjacency matrix obtained in Step 1 and the residual obtained in Step 2 through a spatio-temporal graph neural network composed of three neural network modules: the long short-term memory network LSTM, the temporal convolutional network TCN, and the graph convolutional network GCN, to capture spatio-temporal dependencies and merge them through skip connections to obtain encapsulated spatio-temporal hidden features;
[0080] Furthermore, Step 3 includes:
[0081] Design a graph convolutional network GCN to capture the spatial correlation between variables; for the given graph adjacency matrix X and the initial input state G in , the definition of the graph convolutional network GCN is given as follows:
[0082]
[0083] where k represents the depth of graph propagation, G 0 = G in , represents the normalized graph adjacency matrix, and and I represents the identity matrix, X ij represents the element in the i-th row and j-th column of the graph adjacency matrix X, and are used as an intermediate bridge to normalize the graph adjacency matrix X to , which has no practical meaning; use the parameter β to control the amount of information retained to avoid over-smoothing; assign different weights Θ k to the hidden state G k , where k ∈ {0,..., K}; use different graph adjacency matrices X and X T to separately store the inflow and outflow information of the nodes, G k represents the state at graph propagation depth k, G k+1 represents the hidden state at graph propagation depth k, G out represents the complete output of the graph convolutional network GCN, G k Θ k represents the product of the hidden state at graph propagation depth k and the weight assigned in this graph propagation state;
[0084] When the multivariate time series flows out of the LSTM, the output is h out ; the output h out is expressed as:
[0085] h out= o out ·tanh(C out )
[0086] where h out represents the output of the last hidden layer of the LSTM, o out represents the final output of the output gate, and C out represents the final cell state;
[0087] During the process of using the entire LSTM-TCN network to capture time dependencies, the output sequence of the LSTM serves as the input sequence of the TCN and flows into the TCN in the form of residual input; for the TCN, the specific definition is as follows:
[0088] T l+1 = Γ(T l , P l+1 ) + TCN(T l , Φ l ), l ∈ {0,..., L}
[0089] where the output T out of the TCN = T L , and TCN(·, Φ l ) is the convolutional function parameterized by the l-th layer Φ l , Γ(T l , P l+1 ) represents the truncation function that intercepts the last P l elements from T l along the sequence length axis of T l+1 , where
[0090] P l+1 = P l - q l × (k - 1) and P 1 = Q - k + 1, where Q, q, and k represent the receptive field size, dilation factor, and convolutional kernel size respectively, k also represents the depth of graph propagation, l refers to the l-th layer of temporal convolution, P l+1 represents the number of elements intercepted by the l-th layer of temporal convolution, T l+1 represents the state of the l-th layer of temporal convolution, T l represents the hidden state of the l-th layer of temporal convolution, and Φ l is the parameterization of the l-th layer of temporal convolution;
[0091] The gating mechanism is adopted to guide the information flow, and the process of adopting the gating mechanism to guide the information flow is expressed as:
[0092]
[0093] where f C(·) and f G (·) represent temporal filtering convolution and gated convolution respectively, and ⊙ represents element-wise product. and represent the parameterizations of the l-th layer of filtering convolution and gated convolution respectively; the definitions of temporal filtering convolution and gated convolution are as follows:
[0094]
[0095] where, ★Δ represents dilated convolution operation; and are set to be composed of multiple convolution filters with width h ∈ {2, 3, 6, 7} and ; and represent gated convolutional units, and h represents the width of convolution;
[0096] Based on the completion of the above work on modeling temporal and spatial dependencies, a spatio-temporal graph neural network is obtained; the processing process of the spatio-temporal graph neural network is expressed as:
[0097]
[0098] where, Φ represents weights, and T 0 refers to the initial input state of temporal convolution.
[0099] Step4. The prediction head further projects the spatio-temporal hidden features into a single-step prediction result;
[0100] Furthermore, Step4 includes:
[0101] Using a multi-layer perceptron to achieve single-step prediction on the final output state T of the spatio-temporal graph neural network, and obtaining the result of single-step prediction out ; The result of single-step prediction is expressed as:
[0102]
[0103] where, MLP() represents the prediction process of the multi-layer perceptron, T out is the final output state of the spatio-temporal graph neural network, and W mlp refers to using the multi-layer perceptron for prediction.
[0104] Furthermore, Step5 includes:
[0105] For any univariate time series calculate the absolute prediction error at the current timestamp Z and then normalize the absolute prediction error value; the normalization process is expressed as:
[0106]
[0107] Among them, and are the median and interquartile range of the error values in the sliding window respectively, and W is the length of the sliding window, a and represents the prediction result of the univariate time series ; represents the normalization result;
[0108] After normalizing each univariate time series, a normalized error vector of the multivariate time series at the current timestamp is obtained Then, PCA is used to aggregate the normalized errors to obtain the anomaly score.
[0109] Step 5: Input the current prediction result and all the observation-prediction pairs calculated before the timestamp, and use the PCA-based scorer to calculate the normalized prediction error, and output the anomaly metric score in real time.
[0110] Furthermore, Step 5 further includes:
[0111] Calculating the normalized error in the validation set and fitting the normalized error by exploring the relationship between the validation mean vector covariance matrix and the orthogonal eigenvector J; where J is composed of H orthogonal eigenvectors related to the H largest eigenvalues in the diagonal matrix Λ, and this diagonal matrix Λ can be derived from the formula E v = JΛJ -1 ; To adapt to PCA, the normalized error at the current timestamp is reconstructed through the following definition:
[0112]
[0113] In the above equation, first, the normalized error at the current timestamp is subtracted by the average validation error to achieve zero-centering, and then these results are projected by the validation eigenvector J; in addition, only the first L principal components are retained; finally, the zero-centering inference is restored by adding the validation average error to the L-dimensional reconstructed normalized error after dimensionality reduction; is the reconstructed normalized error; in this process, L is set to the number of components required to achieve a symmetric mean absolute percentage error (sMAPE) less than 10% on the validation set.
[0114] To calculate the final anomaly score of the current timestamp, the reconstructed normalized error is utilized The distance L1 between the denoised normalized error and the original normalized error is used as the final anomaly score X(Z); the final anomaly score X(Z) is expressed as:
[0115]
[0116] The present invention also provides a multivariate time series anomaly detection system based on spatio-temporal graph learning, and the system includes: a module for executing the multivariate time series anomaly detection method based on spatio-temporal graph learning as described above.
[0117] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the multivariate time series anomaly detection method based on spatio-temporal graph learning.
[0118] The effectiveness of the present invention is illustrated by comparative experiments below.
[0119] (1) Verification of overall detection performance;
[0120] The present invention classifies whether there is an anomaly by taking each timestamp as an independent observation value, and evaluates the performance of the model by the scores of the areas under the receiver operating characteristic (ROC-AUC) and precision-recall curve (PRC-AUC). Generally, the closer the scores of ROC-AUC and PRC-AUC are to 1, the better the model's ability to distinguish between anomalous and non-anomalous time points, and the better the overall detection performance of the model.
[0121] As Figure 2 shown, for the WADI dataset, the ROC-AUC and PRC-AUC scores of the method LTG of the present invention are the highest among all methods, which are 0.8542 and 0.5450 respectively. This indicates that the method LTG of the present invention shows stronger robustness and superior detection performance when dealing with the multivariate time series anomaly detection task.
[0122] As Figure 3 shown, for the average performance of 12 servers in the SMD dataset, the ROC-AUC and PRC-AUC scores of the method LTG of the present invention are the highest, which are 0.8506 and 0.4254 respectively. This indicates that the method LTG of the present invention shows the best ability and accuracy in distinguishing anomalous samples.
[0123] (2) Verification of early detection performance;
[0124] In this invention, the detection of adjacent abnormal segments is regarded as a true positive if and only if an abnormal point is correctly detected and its timestamp is at most δ steps after the first abnormality in the adjacent abnormal segment. For early anomaly detection, generally, the earlier an anomaly is detected, the more the negative impact brought by the anomaly can be reduced. Therefore, this invention evaluates the early detection ability of the model under delay constraints of 0, 1, 5, 10, 20, and 30 minutes and reports the best F1 score under each delay constraint condition.
[0125] As Figure 4 shown, on the WADI dataset, the method LTG of this invention (represented by the red line) is always at the top of the graph and achieves the highest F1 score under all delay constraints. This shows that the detection performance of the method LTG of this invention is significantly better than other methods. For the WADI dataset, the F1 scores of all methods increase as the delay constraint increases, indicating that the performance of early anomaly detection gradually improves as the delay constraint increases.
[0126] As Figure 5 shown, on the SMD dataset, the method LTG of this invention (represented by the red line) is always at the top of the graph and obtains the highest F1 score under all delay constraints. As the delay time increases, the F1 score of the method LTG of this invention steadily rises from about 0.5 initially to about 0.85 at 30 minutes. For the SMD dataset, as the delay time increases, most methods show varying degrees of improvement in the F1 score, especially the performance improvement is particularly obvious after the 10 - minute mark.
[0127] In the verification of the dataset by the method LTG of this invention, the WADI dataset is collected and open - sourced by the iTrust institution of the Singapore University of Technology and Design (SUTD), and it has a complete and realistic water treatment, storage, and distribution network. The training set of WADI is a two - week normal operation scenario, while the test set is a two - day attack scenario; the SMD dataset is a real - world server machine dataset collected by a large Internet company, which contains time - series data of servers, each server has 38 multivariate variables, and is divided into training sets and test sets of equal size.
[0128] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above - mentioned implementation manners. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A multivariate time series anomaly detection method based on spatiotemporal graph learning, characterized by: include: A correlation learning layer is used to capture pairwise correlations and obtain a graph adjacency matrix that reflects the correlation between variables. Then, the graph adjacency matrix is input into a spatiotemporal graph neural network composed of a long short-term memory network, a temporal convolutional network, and a graph convolutional network. The spatiotemporal graph neural network captures the temporal and spatial dependencies in multivariate time series and accurately models the spatiotemporal relationships. Finally, a PCA-based anomaly scorer is used to output the anomaly score.
2. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 1, characterized in that: The method comprises the following steps: Step 1: For each multivariate time series, randomly initialized node embedding is performed, and the graph adjacency matrix is adaptively learned through the correlation learning layer in an end-to-end manner; Step 2: The sliding window input is put into the latent space through the convolution layer, and the residual channel is used as the output, and then the residual is input into the subsequent spatiotemporal graph neural network; Step 3, the graph adjacency matrix obtained in Step 1 and the residual input obtained in Step 2 are sequentially passed through the spatiotemporal graph neural network composed of three neural network modules: long short-term memory network LSTM, temporal convolution network TCN and graph convolution network GCN to capture the spatiotemporal dependency and merge them through skip connections to obtain encapsulated spatiotemporal hidden features; Step 4, the prediction head further projects the spatiotemporal hidden features into single-step prediction results; Step 5: Input the current prediction result and all observation-prediction pairs calculated before the timestamp, and use the PCA-based scorer to calculate the normalized prediction error and output the anomaly indicator score in real time.
3. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 2 is characterized in that: The multivariate time series in Step 1 uses the time series in the WADI dataset, the SMD dataset, or the multivariate time series data consisting of CPU usage, memory usage, database query response time, fault log, server load, and number of requests when the smart hospital system is running for anomaly detection.
4. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 2 is characterized in that: The Step 1 includes: A relevance learning layer is designed to adaptively learn the graph adjacency matrix X, where nodes and edges represent variables and variable connectivity; the calculation process of the graph adjacency matrix X is as follows: in are two randomly initialized node embedding matrices, There are two sets of trainable parameters. The hyperparameter α represents the nonlinear activation saturation rate. Both tanh() and ReLU() represent activation functions. A hyperparameter k is introduced to control the sparsity of the graph adjacency matrix X, that is, by further masking the zero elements in the learned graph adjacency matrix X except for the top-k nearest neighbors of each node; for the i-th row in the graph adjacency matrix X, it is processed as follows: where argmax(,k) returns the top k largest indices in the input vector, i represents the i-th row of the graph adjacency matrix X, X[i,:] represents the element in the i-th row of the graph adjacency matrix X, topk represents each node and its k closest nodes, -topk represents each node other than its k closest nodes, and X[i,-topk] represents the element in the i-th row of the graph adjacency matrix X other than its k closest elements.
5. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 2, characterized in that: The Step 3 includes: Design a graph convolutional network GCN to capture the spatial correlation between variables; given a graph adjacency matrix X and an initial input state G in , the definition of the graph convolutional network GCN given is as follows: Among them, k represents the depth of graph propagation, G 0 =G in , represents the normalized graph adjacency matrix, and and I represents the identity matrix, X ij represents the element in row i and column j of the graph adjacency matrix X, and Used to normalize the graph adjacency matrix X as It is an intermediate bridge of the node and has no practical significance. The parameter β is used to control the amount of information retained to avoid over-smoothing. k Assign different weights Θ k , where k∈{0,...,K}; using different graph adjacency matrices X and X T To save the inflow and outflow information of the node respectively, G k represents the state when the graph propagation depth is k, G k+1 It represents the hidden state when the graph propagation depth is k, G out Represents the complete output of the graph convolutional network GCN, G k Θ k Represents the product of the hidden state when the graph propagation depth is k and the weight assigned to the graph propagation state; When a multivariate time series flows out of the LSTM, the output h out ; Output h out Expressed as: h out =o out ·tanh(C out ) Among them, h out represents the output of the last hidden layer of LSTM, o out Represents the final output of the output gate, C out represents the final cell state; In the process of using the entire LSTM-TCN network to capture time dependencies, the output sequence of LSTM is used as the input sequence of TCN and flows into TCN in the form of residual input; for TCN, the specific definition is as follows: T l+1 =Γ(T l ,P l+1 )+TCN(T l ,Φ l ),l∈{0,...,L} Among them, the output of TCN is out =T L ,TCN(·,Φ l ) is the lth layer Φ l The parameterized convolution function, Γ(T l ,P l+1 ) indicates that along T l The sequence length axis starts from T l Intercept the last P l+1 elements, where P l+1 =P l -q l ×(k-1) and P 1 =Q-k+1, Q, q, k represent the size of the receptive field, the dilation factor and the size of the convolution kernel respectively, k also represents the depth of graph propagation, l refers to the lth layer of temporal convolution, P l+1 It represents the number of elements intercepted by the l-th layer of temporal convolution, T l+1 It represents the state of the lth layer of temporal convolution, T l represents the hidden state of the lth layer of temporal convolution, Φ l is the parameterization of the l-th layer of temporal convolution; The gating mechanism is used to guide the information flow. The process of using the gating mechanism to guide the information flow is expressed as: Among them, f C (·) and f G (·) represents time-filtered convolution and gated convolution respectively, ⊙ represents the product of elements, and They represent the parameterization of the l-th layer of filter convolution and gated convolution respectively; the definitions of filter convolution and gated convolution are as follows: Among them, ★Δ represents the dilated convolution operation; and It is set to consist of multiple convolutional filters with width h∈{2,3,6,7} and composition, and represents the gated convolutional unit, and h represents the width of the convolution; Based on the completion of the above temporal and spatial dependency modeling work, a spatiotemporal graph neural network is obtained; the processing process of the spatiotemporal graph neural network is expressed as: Among them, Θ represents the weight, T 0 Refers to the initial input state of the temporal convolution.
6. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 2, characterized in that: The Step 4 includes: The final output state T of the spatiotemporal graph neural network out Through the multi-level perceptron to achieve single-step prediction, get the result of single-step prediction Results of one-step forecast It is expressed as: Among them, MLP() represents the prediction process of the multi-level perceptron, T out is the final output state of the spatiotemporal graph neural network, W mlp Refers to the use of multi-level perceptron prediction.
7. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 2, characterized in that: The Step 5 includes: For any univariate time series Calculate the absolute prediction error at the current timestamp Z Then the absolute prediction error value is normalized; the normalization process is expressed as: in, and are the error values in the sliding window respectively. The median and interquartile range, W a is the length of the sliding window, Represents a univariate time series The prediction results, represents the normalized result; After normalizing each univariate time series, a multivariate time series normalized error vector of the current timestamp is obtained. Then PCA is used to summarize the normalized errors to obtain the anomaly score.
8. The multivariate time series anomaly detection method based on spatiotemporal graph learning according to claim 6, characterized in that: The Step 5 further includes: Calculate the normalized error in the validation set And by exploring and verifying the average vector Covariance matrix and the relationship between the orthogonal eigenvectors J to fit the normalized error; where J is composed of H orthogonal eigenvectors related to the H largest eigenvalues in the diagonal matrix Λ, this diagonal matrix Λ can be expressed by formula E v =JΛJ -1 It is derived that in order to adapt PCA, the normalized error of the current timestamp is reconstructed by the following definition: In the above equation, first let the normalized error of the current timestamp be Subtract the mean validation error To achieve zero centering, the validation feature vector J is used to project these results. In addition, only the first L principal components are retained. Finally, the zero centering inference is restored by adding the validation mean error to the normalized error of the L-dimensional reconstruction after dimensionality reduction. is the normalized error after reconstruction; To calculate the final anomaly score at the current timestamp, the reconstructed normalized error is used The normalized error after denoising and the original normalized error The distance L1 between them is taken as the final anomaly score X(Z); the final anomaly score X(Z) is expressed as:
9. A multivariate time series anomaly detection system based on spatiotemporal graph learning, characterized in that: The system comprises: a module for executing a multivariate time series anomaly detection method based on spatiotemporal graph learning as claimed in any one of claims 1 to 8.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, a multivariate time series anomaly detection method based on spatiotemporal graph learning as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Substation wireless access point anomaly detection method, system and device and medium
CN121099365A
Multivariable time sequence anomaly detection method based on spatio-temporal feature fusion
CN121412886A
A multivariate time series anomaly detection method based on spatiotemporal feature fusion
CN121412886B