Construction method of multidimensional time sequence anomaly detection system

By building a multi-dimensional time series data anomaly detection system based on graph neural network and association differences, the problem of high-dimensional time series data high-quality calculation costs and insufficient accuracy is solved, and efficient and accurate anomaly detection is achieved, which is suitable for distributed environments.

CN120277571APending Publication Date: 2025-07-08HOHAI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510325689.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When processing high-dimensional time series data, the prior art has high computational cost, insufficient accuracy, and it is difficult to capture complex anomaly patterns. Especially in a large-scale data flow environment, it is difficult to locate and trace the abnormality.

Method used

A multi-dimensional time series data anomaly detection system based on graph neural network and association differences is adopted. By constructing an association graph, the graph structure is optimized using the Gumbel-Softmax sampling method, combining the improved graph attention network to learn node features, and timing prediction is performed through multi-layer stacked full connection layers, local and global connections are established, anomaly scores are calculated, and real-time detection is achieved using a distributed computing framework.

Benefits of technology

It significantly reduces the computing overhead in large-scale data flow environments, improves the accuracy and stability of abnormal detection, and is suitable for real-time abnormal detection in distributed environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277571A_ABST
    Figure CN120277571A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a multi-dimensional time sequence anomaly detection system. The method comprises the following steps: acquiring multi-dimensional time sequence data; preprocessing the obtained multi-dimensional time sequence data; constructing a depth anomaly detection algorithm, constructing an association graph of the multi-dimensional time series data by using cosine similarity and a TopK strategy, learning the association graph through a graph neural network, extracting space and time features of the multi-dimensional time series data, and calculating an anomaly score of each data point; and constructing an anomaly detection system suitable for real-time multi-dimensional time series data based on Apache Flink. Partitioning and efficient detection of multi-dimensional time series data are completed through a real-time stream processing module, and an anomaly positioning map and an anomaly propagation map are generated. The method has the advantages of high efficiency, accuracy, real-time performance and the like on the anomaly detection of the multi-dimensional time series data, can effectively deal with a large amount of multi-dimensional time series data generated in the operation process of a complex system, and timely discovers and locates the anomaly condition so as to support quick response of system managers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a multi-dimensional time series data anomaly detection system based on a graph neural network and correlation difference, and relates to the technical field of time series data anomaly detection. Background Art

[0002] With the wide application of various sensors and monitoring devices, a large amount of high-dimensional time series data is generated for monitoring and analysis. When dealing with high-dimensional time series data, traditional anomaly detection methods often face challenges such as high computational cost and insufficient accuracy. The traditional threshold setting method has errors and is difficult to capture complex anomaly patterns. Although the anomaly detection method based on deep learning can improve the detection accuracy, in the large-scale data stream environment, the computational efficiency is still not high enough. In addition, due to the complex correlation between data measurement points, the positioning and root cause analysis of anomalies also become a difficult problem. Therefore, in a distributed environment, how to implement an efficient and accurate multi-dimensional time series data anomaly detection system is still an important topic that needs to be deeply explored and studied at present. Summary of the Invention

[0003] Object of the Invention: Aiming at the problems and deficiencies in the prior art, the present invention provides a method for constructing a multi-dimensional time series data anomaly detection system based on a graph neural network and correlation difference, and performs efficient and accurate anomaly detection on multi-dimensional time series data.

[0004] Technical Solution: A method for constructing a multi-dimensional time series data anomaly detection system based on a graph neural network and correlation difference, comprising the steps of:

[0005] Step 1: Obtain multi-dimensional time series data, where the multi-dimensional time series data includes device status, sensor readings, and other monitoring information;

[0006] Step 2: Preprocess the obtained multi-dimensional time series data, which are maximum-minimum normalization, data encoding, and position encoding in sequence;

[0007] Step 3: Construct a deep anomaly detection algorithm based on graph neural network and correlation difference. In the spatial dimension, construct an association graph by using cosine similarity and TopK strategy for the variable features of multi-dimensional time series data, and optimize the association graph structure by using the Gumbel-Softmax sampling method; secondly, use the improved graph attention network to learn the features of nodes in the association graph; finally, combine the variable features with the graph node features, and adopt the way of multi-layer stacked fully connected layers to realize time series prediction, and calculate the anomaly score by comparing the error between the predicted value and the measured value. In the time dimension, two independent model branches are used to establish the local connection and global connection of multi-dimensional time series data respectively. The deep anomaly detection model is trained based on the similarity between the two branches, and normal and abnormal data are distinguished by comparing the differences between the two;

[0008] Step 4: Based on the distributed computing framework technology and the idea of association partitioning, combine the method in Step 3 to construct a real-time anomaly detection system for multi-dimensional time series data, and realize the visualization of anomaly detection through the Web front end.

[0009] Preferably, in Step 2, the maximum-minimum normalization method is used to preprocess the multi-dimensional time series data, and the calculation formula is as (1):

[0010]

[0011] where x is the original multi-dimensional time series data, min represents the minimum observed value of the current feature in the whole data set, and max represents the maximum observed value of the current feature in the whole data set.

[0012] Preferably, in Step 2, data encoding is used to extract the features of multi-dimensional time series data. The embedded vector after data encoding is mainly used as the input of the deep anomaly detection model in the time dimension and the spatial dimension, and the calculation formula is as (2):

[0013] DE=1DCNN(data size ,model in ) (2)

[0014] where data size is the original size, model in is the input dimension, and the features of multi-dimensional time series data are extracted by using a one-dimensional convolutional neural network.

[0015] Preferably, in Step 2, encoding information is introduced to reflect the time series characteristics of multi-dimensional time series data. By using sine and cosine functions with different frequencies, additional encoding information is added to each element input into the model to capture the order and interval characteristics of elements in the multi-dimensional time series, thereby enhancing the model's ability to understand time features, and the calculation formulas are as (3) and (4):

[0016]

[0017] Among them, i is the dimension index, taking values in {0, 1, …, d - 1}, d is the input dimension of the model, pos is the position of the data point, taking values in {0, 1, …, w - 1}, and w is the window length of the multi-dimensional time series data.

[0018] Preferably, in step 3, the relationship between multi-dimensional time series data variables is captured by constructing an association graph, including the following steps:

[0019] (1) Use a directed graph to model the relationship between multi-dimensional time series data variables. The nodes in the graph represent different variables, and the edges represent the dependency relationships between different variables. An adjacency matrix A is introduced to represent the directed graph.

[0020] (2) A candidate relationship set C is introduced i , to determine the relationship between multi-dimensional time series data variables, and relevant variables are selected according to prior information. The expression formula is as in (5):

[0021]

[0022] Among them, N represents the total number of multi-dimensional time series data variables in the model.

[0023] (3) Calculate the correlation between different variables of multi-dimensional time series data through cosine similarity. The calculation formula is as in (6):

[0024]

[0025] Among them, r i ∈R d represents the feature vector of the encoded variable i, sim represents the cosine similarity calculation, performs a normalized dot product on all multi-dimensional time series data in the candidate set, and selects the top k data variables with the highest similarity through the TopK strategy. The value of k is selected according to the sparsity of the association matrix required by the user, and thus the association matrix A of the multi-dimensional time series data is constructed.

[0026] (4) Improve the graph structure based on the Gumbel-Softmax sampling method through the following steps:

[0027] First, generate the edge sampling probability s between node i and node j ij , and combine it with the Gumbel noise g k to adjust the probability. The calculation formulas are as in (7) and (8):

[0028] g k =-log(-log(u k)) (7)

[0029]

[0030] Among them, log represents the natural logarithm, and u k is randomly sampled from a uniform distribution, that is, u k ~Uniform(0,1).

[0031] Secondly, calculate the probability c of the existence of an edge ij , and the calculation formula is as (9):

[0032]

[0033] Among them, τ is the temperature parameter, and the smaller τ is, the sparser the graph structure is.

[0034] Preferably, in step 3, use an improved graph attention network to learn the features of nodes in the graph, including the following steps:

[0035] (1) By concatenating the components of each node with the node feature r i to fuse the feature information, the calculation formula is as (10):

[0036]

[0037] Among them, W is a weight matrix, is the input data of node i at time t.

[0038] (2) Calculate the attention weight α i,j , which is used to measure the attention score contributed by node i to node j, and the calculation formulas are as (11) and (12):

[0039]

[0040] Among them, W e is the edge weight matrix, e ij is the information of the edge between nodes, respectively represent the features of node i and node j at time t, a T is the weight vector used to calculate the attention score, represents the neighbor nodes of node i.

[0041] (3) For each graph node, complete the extraction of graph features by aggregating the information of neighbor nodes, and the calculation formula is as (13):

[0042]

[0043] Among them, α i,j is the attention weight, and W is the weight matrix, and are the input data of nodes i and j at time t, and the aggregated features are used for node representation.

[0044] (4) The final graph feature output z is obtained through the residual connection structure and layer normalization (t) , and the calculation formulas are as (14) and (15):

[0045] z′ = LayerNorm(r + h (t) ) (14)

[0046] z (t) = LayerNorm(FNN(z′)) + z′ (15)

[0047] where LayerNorm represents layer normalization, FNN is a feed-forward neural network, z′ is the feature representation after layer normalization, r is the static feature of the node, and h (t) represents the feature aggregated by the graph attention layer.

[0048] Preferably, in step 3, the variable features are combined with the graph node features, and a multi-layer stacked fully connected layer is used to achieve time series prediction. The error between the predicted value and the measured value is calculated to obtain the anomaly score, including the following steps:

[0049] (1) A multi-layer stacked fully connected layer is used for prediction, and the calculation formula is as (16):

[0050]

[0051] where {r1, r2…r N} are data features, and is the graph feature obtained through the graph neural network, represents the prediction result for the specified time t.

[0052] (2) The mean square error is used as the loss function to measure the difference between the predicted output and the observed data x (t) , and the calculation formula is as (17):

[0053]

[0054] where T represents the total length of the entire time series, and w represents the size of the sliding window.

[0055] (3) The anomaly score a i is used to measure the degree of deviation of the data point from the normal mode, and the calculation formula is as (18):

[0056]

[0057] Among them, is the predicted value, and x (t) is the measured value.

[0058] (4) Standardize the anomaly scores a i = {a1, a2, …, a N} of multiple data variables, and the calculation formula is as in (19):

[0059]

[0060] Among them, μ is the mean value, and σ is the standard deviation.

[0061] (5) To further calculate the overall anomaly score AS(t), only consider the maximum one-quarter anomaly score at time t, and the calculation formula is as in (20):

[0062]

[0063] Among them, N represents the number of multi-dimensional time series data variables, and sort{a′ i} represents the anomaly score sequence at time t after descending order sorting, and take the first one-quarter of the anomaly scores.

[0064] (6) To reduce the influence of spike noise, use the exponential moving average method to smooth the anomaly scores, and the calculation formula is as in (21):

[0065]

[0066] Among them, α is the smoothing coefficient, and the smoothed anomaly score AS′(t) is used for the final anomaly detection result.

[0067] Preferably, in step 3, establish the local connection and global connection of multi-dimensional time series data, including the following steps:

[0068] (1) By introducing the global self-attention mechanism, calculate the correlation between each data point in the multi-dimensional time series data and all other data points within the time window, and the calculation formulas are as in (22) and (23):

[0069] Q = XW q , K = XW k (22)

[0070]

[0071] Among them, Q and K represent the query and key respectively, W q and W k are the corresponding parameter matrices, and d model is the model input dimension.

[0072] (2) Calculate the correlation between the data point at the current time step and its surrounding data points through the local self-attention mechanism. Define a sliding window of length ω, which includes the current time point and its adjacent ω - 1 data points. The calculation formula is as shown in (23).

[0073] Preferably, in step 3, identify the intermediate process anomalies in the long time series anomalies through the reconstruction structure. Use the output parameters of the global self-attention to construct the reconstruction structure, and implement the reconstruction of multi-dimensional time series data through the Transformer structure. The reconstruction structure introduces the parameters and outputs based on the existing correlation difference module, which can not only simplify the model parameters but also guide the construction of the global correlation of multi-dimensional time series data. By combining the reconstruction structure and the correlation difference structure, further amplify the difference between the anomaly points and the normal points, and improve the detectability of the anomaly points.

[0074] Preferably, in step 3, calculate the correlation difference loss and calculate the anomaly score, including the following steps:

[0075] (1) Calculate the distribution gap between the local correlation and the global correlation using the stacked KL divergence. The calculation formula is as shown in (24):

[0076] Loss ass (S,P) = [KL(S i ,P i ) + KL(P i ,S i )] i=1,…,w (24)

[0077] Among them, S represents the global correlation, P represents the local correlation, w is the window length, KL is the KL divergence, and the calculation formula of the KL divergence is as shown in (25):

[0078]

[0079] Among them, P represents the probability distribution of the local correlation, Q represents the probability distribution of the global correlation, and log represents the natural logarithm.

[0080] (2) Reconstruct the multi-dimensional time series data through the reconstruction layer. The calculation formula is as shown in (26):

[0081]

[0082] Among them, X is the original input data, is the reconstructed value.

[0083] (3) Realize the reconstruction loss by fusing the reconstruction difference and the correlation difference. In the minimization stage, the main goal is to reduce the reconstruction difference and the correlation difference. The calculation formula is as shown in (27):

[0084]

[0085] In the maximization stage, the training objective is to reduce the reconstruction difference and increase the correlation difference, and the calculation formula is as shown in (28):

[0086]

[0087] where, ‖·‖ represents the norm, X is the original input data, is the reconstructed value, λ is the control coefficient, stopgrad represents stopping the gradient, S represents the global correlation, and P represents the local correlation.

[0088] (4) Calculate the correlation difference score. The smaller the correlation difference, the more likely it is an anomaly. Therefore, the difference is negated and output through the softmax function. The calculation formulas are as shown in (29) and (30):

[0089] L score = [-KL(S i , stopgrad(P i )) - KL(P i , stopgrad(S i ))] i=1,…,w (29)

[0090] Score ass = softmax(L score ) (30)

[0091] where, S i represents the global correlation distribution at the i-th position, P i represents the local correlation distribution at the i-th position, stopgrad represents stopping the gradient, and w is the sequence length of the input multi-dimensional time series data.

[0092] (5) Calculate the anomaly score at time t according to the fusion of the reconstruction difference and the correlation difference, and adopt two calculation methods AS mul and AS add , to adapt to different scenarios. The calculation formulas are as shown in (31) and (32):

[0093]

[0094] where, AS mul is used for scenarios that require marking the start and end of anomalies, and AS add is universal and applicable to detecting various types of anomalies.

[0095] Preferably, in step 3, after normalizing the anomaly scores of the two, weighted merging is used to obtain the final anomaly score, and the calculation formula is as in (33):

[0096] Anomaly Score=uAS spatial +(1 - u)AS temporal (33)

[0097] Among them, AS spatial represents the anomaly score obtained by the spatial dimension model, AS temporal represents the anomaly score obtained by the time dimension model, and the parameter u is a weight parameter used to reflect the emphasis on the two models.

[0098] Preferably, in step 3, the maximum anomaly score is used as the anomaly threshold, which specifically includes: traversing the validation set to calculate the anomaly score of each window, selecting the maximum score as the threshold, and judging the abnormal data through this threshold.

[0099] Preferably, in step 3, the training process of the deep anomaly detection model is optimized by using a dynamic learning rate and early stopping technology, the dynamic learning rate is adjusted using step decay, and overfitting is prevented through early stopping technology. The calculation formula is as in (34):

[0100] lr new =lr old α (epoch-1) / / k (34)

[0101] Among them, lr new is the updated learning rate, lr old is the learning rate before adjustment, α is the decay factor, and k is the decay interval.

[0102] Preferably, in step 4, based on the idea of associated partitioning, the multi-dimensional time series data is spatially partitioned through an association graph and spectral clustering, reducing the computational overhead and improving the anomaly detection efficiency.

[0103] Preferably, in step 4, the distributed computing framework implemented based on Apache Flink completes the partitioning and efficient detection of multi-dimensional time series data through real-time stream processing.

[0104] Preferably, in step 4, the core construction of the real-time anomaly detection system includes a data storage module, a data partitioning module, an anomaly detection module, an anomaly reasoning module, and a Web front-end module. First, the data storage module is mainly responsible for storing real-time data, model-related data, association graph data, and historical detection results. In this module, the time series database TDengine is used to store real-time monitoring data and historical test results, the relational database MySQL is used to store the file addresses of the models, as well as the data partitions corresponding to the data measurement points, and Neo4J is used to store the association graph structure of multi-dimensional time series data. Secondly, the data partitioning module is mainly responsible for classifying the original multi-dimensional time series data. Thirdly, the anomaly detection module constructs an anomaly detection algorithm for multi-dimensional time series data based on graph neural networks and association differences. At the same time, using the support of Apache Flink for data streams and distribution, it performs accurate real-time anomaly detection on the multi-dimensional time series data within the partition, and at the same time solves the problems of data merging and missing value processing, and uses the DJL framework to implement the Java program to load the deep anomaly detection model. Finally, the anomaly reasoning module provides the location of anomaly data and the search for the anomaly propagation chain by constructing the association graph structure of the data measurement points in the system. In addition, the visualization of anomaly detection is realized through a Web front-end, providing an interactive tool for users.

[0105] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the method for constructing the multi-dimensional time series data anomaly detection system as described above.

[0106] The beneficial effects of the present invention are as follows: The present invention is beneficial to improving the anomaly detection efficiency of high-dimensional time series data, is applicable to real-time anomaly detection in a distributed environment, significantly reduces the computational overhead in a large-scale data stream environment, and improves the accuracy and stability of anomaly detection. Description of the Drawings

[0107] Figure 1 It is a schematic flowchart of the construction method according to an embodiment of the present invention;

[0108] Figure 2 It is a schematic diagram of the real-time anomaly detection system for multi-dimensional time series data according to an embodiment of the present invention. Detailed Embodiments

[0109] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification made by those skilled in the art to the present invention all fall within the scope defined by the appended claims of this application.

[0110] A construction method for an abnormal detection system of multi-dimensional time series data based on graph neural network and association difference, comprising the steps:

[0111] Step 1: Obtain multi-dimensional time series data, which includes device status, sensor readings, and other monitoring information;

[0112] Step 2: Preprocess the obtained data, including maximum-minimum normalization, data encoding, and positional encoding in sequence;

[0113] The maximum-minimum normalization method is used to preprocess the multi-dimensional time series data. Data encoding extracts data features through a one-dimensional convolutional neural network, and at the same time, encoding information is introduced to reflect the time series characteristics of the multi-dimensional time series data. By using sine and cosine functions of different frequencies, additional encoding information is added to each element input into the deep anomaly detection model to capture the order and interval characteristics of the elements in the multi-dimensional time series, thereby enhancing the model's ability to understand time features.

[0114] Step 3: Construct an association graph of multi-dimensional time series data using cosine similarity and TopK strategy, and optimize the graph structure using the Gumbel-Softmax sampling method; Use an improved graph attention network to learn the features of the nodes in the graph, combine the variable features with the graph node features, and calculate the anomaly score of each data point;

[0115] Construct an association graph to capture the relationships between variables in multi-dimensional time series data, including the following steps:

[0116] The first step is to use a directed graph to model the relationships between variables in multi-dimensional time series data. The nodes in the graph represent different variables, and the edges represent the dependencies between variables. The adjacency matrix A is used to represent the directed graph.

[0117] The second step is to introduce a candidate relationship set c i , and the formula is as shown in (5):

[0118]

[0119] where N represents the total number of variables in the multi-dimensional time series data in the model.

[0120] The third step is to calculate the correlation between different variables of multi-dimensional time series data through cosine similarity and select the top k data variables with the highest similarity through the TopK strategy. The calculation formula is as shown in (6):

[0121]

[0122] where r i ∈R dThe eigenvector representing the encoded variable i, and sim represents the cosine similarity calculation.

[0123] Fourthly, optimize the graph structure based on the Gumbel-Softmax sampling method through the following steps:

[0124] Firstly, generate the edge sampling probability s between node i and node j ij , combined with the Gumbel noise g k Adjust the probability, and the calculation formulas are as in (7) and (8):

[0125] g k =-log(-log(u k )) (7)

[0126]

[0127] where, log represents the natural logarithm, and u k is randomly sampled from a uniform distribution, that is, u k ~Uniform(0,1).

[0128] Secondly, calculate the probability c of the existence of the edge ij , and the calculation formula is as in (9):

[0129]

[0130] where, τ is the temperature parameter, and the smaller τ is, the sparser the graph structure is.

[0131] Use the improved graph attention network to learn the features of the nodes in the graph, including the following steps:

[0132] Firstly, by concatenating the components of each node with the node feature r i to fuse the feature information, and the calculation formula is as in (10):

[0133]

[0134] where, W is a weight matrix, is the input data of node i at time t.

[0135] Secondly, calculate the attention weight α i,j , and the calculation formulas are as in (11) and (12):

[0136]

[0137] where, W e is the edge weight matrix, and e ij is the information of the edge between nodes, respectively represent the features of node i and node j at time t, a T is the weight vector used to calculate the attention score, represents the neighbor nodes of node i.

[0138] Thirdly, aggregate the information of neighbor nodes to complete the extraction of graph features, and the calculation formula is as (13):

[0139]

[0140] where, α i,j is the attention weight, W is the weight matrix, and are the input data of node i and node j at time t, and the aggregated features are used for node representation.

[0141] Fourthly, perform the final graph feature output z (t) through the residual connection structure and layer normalization, and the calculation formulas are as (14) and (15):

[0142] z′ = LayerNorm(r + h (t) ) (14)

[0143] z (t) = LayerNorm(FNN(z′)) + z′ (15)

[0144] where, LayerNorm represents layer normalization, FNN is the feed-forward neural network, z′ is the feature representation after layer normalization, r is the static feature of the node, and h (t) represents the feature aggregated through the graph attention layer.

[0145] Combine the variable features with the graph node features, and calculate the anomaly score by comparing the errors between the predicted value and the measured value, including the following steps:

[0146] First, use the multi-layer stacked fully connected layer method for prediction, and the calculation formula is as (16):

[0147]

[0148] where, {r1, r2…r N} are the data features, and is the graph feature obtained through the graph neural network, represents the prediction result for the specified time t.

[0149] Second, calculate the loss function to measure the difference between the predicted output and the observed data x (t) The difference is calculated as (17):

[0150]

[0151] Among them, T represents the total length of the entire time series, and w represents the size of the sliding window.

[0152] The third step is to calculate the anomaly score a i , and the calculation formula is as in (18):

[0153]

[0154] Among them, is the predicted value, and x (t) is the measured value.

[0155] The fourth step is to standardize the anomaly scores a i ={a1, a2,..., a N} of multiple multi-dimensional time series data variables, and the calculation formula is as in (19):

[0156]

[0157] Among them, μ is the mean value, and σ is the standard deviation.

[0158] The fifth step is to calculate the overall anomaly score AS(t), only considering the maximum one-quarter anomaly score at time t, and the calculation formula is as in (20):

[0159]

[0160] Among them, N represents the number of multi-dimensional time series data variables, and sort{a′ i} represents the anomaly score sequence at time t after descending order, and the first one-quarter anomaly scores are taken.

[0161] The sixth step is to smooth the anomaly scores using the exponential moving average method, and the calculation formula is as in (21):

[0162]

[0163] Among them, α is the smoothing coefficient, and the smoothed anomaly score AS′(t) is used for the final anomaly detection result.

[0164] Establish local and global connections for multi-dimensional time series data, calculate the associated difference loss, and calculate the anomaly score, including the following steps:

[0165] The first step is to calculate the correlation between each data point in the multi-dimensional time series data and all other data points within the time window, and the calculation formulas are as in (22) and (23):

[0166] Q = XW q , K = XWk (22)

[0167]

[0168] Among them, Q and K represent the query and the key respectively, and W q and W k are the corresponding parameter matrices, and d model is the model input dimension.

[0169] Second, calculate the correlation between the data point at the current time step and its surrounding data points, and the calculation formula is as shown in (23).

[0170] Third, use the stacked KL divergence to calculate the distribution gap between the local correlation and the global correlation, and the calculation formulas are as shown in (24) and (25):

[0171] Loss ass (S, P) = [KL(S i , P i ) + KL(P i , S i )] i=1,…,w (24)

[0172] Among them, S represents the global correlation, P represents the local correlation, w is the window length, KL is the KL divergence, and the calculation formula of the KL divergence is as shown in (25):

[0173]

[0174] Among them, P represents the probability distribution of the local correlation, Q represents the probability distribution of the global correlation, and log represents the natural logarithm.

[0175] Fourth, reconstruct the multi-dimensional time series data through the reconstruction layer, and the calculation formula is as shown in (26):

[0176]

[0177] Among them, X is the original input data, is the reconstruction value.

[0178] Fifth, achieve the reconstruction loss by fusing the reconstruction difference and the correlation difference. In the minimization stage, the main goal is to reduce the reconstruction difference and the correlation difference, and the calculation formula is as shown in (27):

[0179]

[0180] In the maximization stage, the training objective is to reduce the reconstruction difference and increase the correlation difference, and the calculation formula is as shown in (28):

[0181]

[0182] where, ‖·‖ represents the norm, X is the original input data, is the reconstructed value, λ is the control coefficient, stopgrad represents stopping the gradient, S represents the global correlation, and P represents the local correlation.

[0183] Step 6: Calculate the correlation difference score, and the calculation formulas are as shown in (29) and (30):

[0184] L score = [-KL(S i , stopgrad(P i )) - KL( i , stopgrad(S i ))] i=1,…,w (29)

[0185] Score ass = softmax(L score ) (30)

[0186] where, S i represents the global correlation distribution at the i-th position, P i represents the local correlation distribution at the i-th position, stopgrad represents stopping the gradient, and w is the sequence length of the input multi-dimensional time series data.

[0187] Step 7: Calculate the anomaly score at time t according to the reconstruction difference and the correlation difference fusion. Two calculation methods, AS mul and AS add , are used to adapt to different scenarios. The calculation formulas are as shown in (31) and (32):

[0188]

[0189] where, AS mul is used for scenarios that require marking the start and end of anomalies, and AS add is universal and applicable to detecting various types of anomalies; after normalizing the anomaly scores of both, weighted merging is used to obtain the final anomaly score, and the maximum anomaly score is used as the anomaly threshold. The calculation formula is as shown in (33):

[0190] Anomaly Score = uAS spatial + (1 - u)AS temporal (33)

[0191] where, AS spatial represents the anomaly score obtained from the spatial dimension model, AS temporal represents the anomaly score obtained from the time dimension model, and the parameter u is a weight parameter used to reflect the emphasis on the two models;

[0192] Step 4: Combine the deep anomaly detection algorithm in Step 3 to build an anomaly detection system applicable to real-time multi-dimensional time series data based on Apache Flink, including: a data storage module, a data partitioning module, an anomaly detection module, an anomaly reasoning module, and a Web front-end module.

[0193] Build an anomaly detection system based on Apache Flink. The system includes a data storage module, a data partitioning module, an anomaly detection module, an anomaly reasoning module, and a Web front-end module. Among them, the data storage module is used for the differential storage of multi-dimensional time series data. The time series database TDengine is used to store real-time monitoring data and historical test results, the relational database MySQL is used to store the file address of the model, and the data partitioning corresponding to the data measurement points. Neo4J is used to store the associated graph structure of multi-dimensional time series data; the data partitioning module is used to classify the original multi-dimensional time series data; the anomaly detection module is used to detect anomalies in real-time stream data and solve problems such as data merging and missing value processing. The DJL framework is used to implement the Java program to load the deep anomaly detection model; the anomaly reasoning module is used to locate abnormal data and find the abnormal propagation chain; finally, the visualization of anomaly detection is realized through the Web front-end.

[0194] Using the detection method and application system of this embodiment can provide relatively efficient anomaly detection for multi-dimensional time series data in a distributed environment, which is beneficial to reducing the computational overhead in a large-scale data stream environment and improving the accuracy and robustness of anomaly detection.

[0195] Obviously, those skilled in the art should understand that each step of the method for constructing the multi-dimensional time series data anomaly detection system of the above embodiment of the present invention or each module of the multi-dimensional time series data anomaly detection system can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A construction method for an abnormal detection system of multi-dimensional time series data based on graph neural network and association difference, characterized in that, Including the steps: Step 1: Obtain multi-dimensional time series data, which includes device status, sensor readings, and other monitoring information; Step 2: Preprocess the obtained multi-dimensional time series data, including maximum-minimum normalization, data encoding, and positional encoding in sequence; Step 3: Construct a deep anomaly detection algorithm based on graph neural network and association difference; In the spatial dimension, construct an association graph by using cosine similarity and TopK strategy for the variable features of multi-dimensional time series data, and use the Gumbel-Softmax sampling method to optimize the association graph structure. Secondly, use an improved graph attention network to learn the features of nodes in the association graph. Finally, combine the variable features with the graph node features, and use a multi-layer stacked fully connected layer to achieve time series prediction, and calculate the anomaly score by comparing the error between the predicted value and the measured value. In the time dimension, use two independent model branches to establish local and global connections of multi-dimensional time series data respectively. The deep anomaly detection model is trained based on the similarity between the two branches, and normal and abnormal data are distinguished by comparing the differences between them; Step 4: Based on the distributed computing framework technology and the idea of association partitioning, combine the method in Step 3 to construct a real-time anomaly detection system for multi-dimensional time series data, and realize the visualization of anomaly detection through the Web front end.

2. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and association difference according to claim 1, characterized in that Extract features from multi-dimensional time series data through a one-dimensional convolutional neural network, and use the embedded vector after data encoding of multi-dimensional time series data as the input of the model in both the time dimension and the spatial dimension. In the spatial dimension, the input size of data encoding is the time window length, which is used to obtain the feature representation of the variables of multi-dimensional time series data. In the time dimension, the input size of data encoding is the dimension of the data, which is used to map the data to a fixed-size vector space.

3. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and association difference according to claim 1, characterized in that In Step 2, introduce encoding information to reflect the time series characteristics of multi-dimensional time series data; add additional encoding information to each element input to the model through sine and cosine functions with different frequencies to capture the order and interval characteristics of elements in multi-dimensional time series, thereby improving the model's ability to understand time features. The calculation formulas are as shown in (3) and (4): where i is the dimension index, taking values in {0, 1, …, d - 1}, d is the input dimension of the model, pos is the position of the data point, taking values in {0, 1, …, w - 1}, and w is the window length of multi-dimensional time series data.

4. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and correlation difference according to claim 1, characterized in that, In Step 3, capture the relationship between variables of multi-dimensional time series data by constructing an association graph, including the following steps: (1) Use a directed graph to model the relationship between variables of multi-dimensional time series data. Nodes in the graph represent different variables, and edges represent the dependence relationship between different variables. Introduce an adjacency matrix A to represent the directed graph; (2) Introduce the candidate relationship set C i , to determine the relationships between the multi-dimensional time series data variables, and select relevant variables according to prior information; among them, N represents the total number of multi-dimensional time series data variables in the model; (3) Calculate the correlation between different variables of multi-dimensional time series data through cosine similarity. The calculation formula is as shown in (6): where r i ∈R d represents the feature vector of the encoded variable i, sim represents the cosine similarity calculation, performs a normalized dot product on all multi-dimensional time series data in the candidate set, and selects the top k data variables with the highest similarity through the TopK strategy; the value of k is selected according to the sparsity of the correlation matrix required by the user, thereby constructing the correlation matrix A of the multi-dimensional time series data; (4) Improve the graph structure based on the Gumbel-Softmax sampling method through the following steps: First, generate the edge sampling probability s between node i and node j ij , and combine it with the Gumbel noise g k to adjust the probability. The calculation formulas are as shown in (7) and (8): g k = -log(-log(u k )) (7) where log denotes the natural logarithm, and u k is randomly sampled from a uniform distribution, i.e., u k ~ Uniform(0,1); Secondly, calculate the probability c of the existence of an edge ij , and the calculation formula is as shown in (9): where τ is the temperature parameter, and the smaller τ is, the sparser the graph structure is.

5. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and association difference as claimed in claim 1, wherein, In step 3, an improved graph attention network is used to learn the features of nodes in the graph, including the following steps: (1) By concatenating the components of each node with the node feature r i to fuse the feature information, the calculation formula is as shown in (10): Among them, W is a weight matrix, is the input data of node i at time t; (2) Calculate the attention weight α i,j , which is used to measure the attention score of node i's contribution to node j. The calculation formulas are as shown in (11) and (12): Among them, W e is the edge weight matrix, and e ij is the information of the edge between nodes, respectively represent the features of node i and node j at time t, and a T is the weight vector used to calculate the attention score, represents the neighbor nodes of node i; (3) For each graph node, extract the graph features by aggregating the information of neighbor nodes, and the calculation formula is as shown in (13): where α i,j is the attention weight, W is the weight matrix, and are the input data of nodes i and j at time t, and the aggregated features are used for node representation; (4) The final graph feature output z is obtained through the residual connection structure and layer normalization (t) , and the calculation formulas are as shown in (14) and (15): z′ = LayerNorm(r + h (t) ) (14)z (t) = LayerNorm(FNN(z′)) + z′ (15) Among them, LayerNorm represents layer normalization, FNN is a feed-forward neural network, z′ is the feature representation after layer normalization, r is the static feature of the node, and h (t) represents the feature aggregated by the graph attention layer.

6. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and correlation difference as described in claim 1, characterized in that, In step 3, combine the variable features with the graph node features, and use a multi-layer stacked fully connected layer to achieve time series prediction. Calculate the anomaly score by comparing the error between the predicted value and the measured value, including the following steps: (1) Use a multi-layer stacked fully connected layer for prediction, and the calculation formula is as shown in (16): Among them, {r1, r2... r N} are data features, while is the graph feature obtained through the graph neural network, represents the prediction result for the specified time t; (2) Use the mean squared error as the loss function to measure the predicted output and the observed data x (t) The difference between them is calculated by the formula (17): where T represents the total length of the entire time series, and w represents the size of the sliding window; (3) Abnormal score a i It is used to measure the degree of deviation of a data point from the normal mode, and the calculation formula is as shown in (18): wherein, is the predicted value, and x (t) is the measured value; (4) Anomaly scores a for multiple data variables i = {a1, a2, …, a N} are normalized, and the calculation formula is as shown in (19): where μ is the mean and σ is the standard deviation; (5) To further calculate the overall anomaly score AS(t), only consider the maximum quarter anomaly score at time t, and the calculation formula is as shown in (20): where N represents the number of multi-dimensional time series data variables, and sort{a i ′} represents the anomaly score sequence at time t after descending order sorting, and the top one-fourth of the anomaly scores are taken; (6) To reduce the influence of spike noise, use the exponential moving average method to smooth the anomaly score, and the calculation formula is as shown in (21): where α is the smoothing coefficient, and the smoothed anomaly score AS′(t) is used for the final anomaly detection result.

7. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and association difference according to claim 1, characterized in that In step 3, establish the local and global connections of multi-dimensional time series data, including the following steps: (1) By introducing the global self-attention mechanism, calculate the correlation between each data point in the multi-dimensional time series data and all other data points within the time window, and the calculation formulas are as shown in (22) and (23): Q = XW q , K = XW k (22) where Q and K represent the query and key respectively, and W q and W k are the corresponding parameter matrices, and d madei is the model input dimension; (2) Through the local self-attention mechanism, calculate the correlation between the data point at the current time step and its surrounding data points. Define a sliding window of length ω, which includes the current time point and its adjacent ω - 1 data points, and the calculation formula is as shown in (23); Identify the intermediate process anomalies in the long time series anomalies through the reconstruction structure; use the output parameters of the global self-attention to construct the reconstruction structure, and realize the reconstruction of the data through the Transformer structure; the reconstruction structure introduces the parameters and outputs of the correlation difference module.

8. The construction method of the multi-dimensional time series data anomaly detection system based on graph neural network and correlation difference as claimed in claim 1, characterized in that, In step 3, calculate the correlation difference loss and calculate the anomaly score, including the following steps: (1) Use the stacked KL divergence to calculate the distribution gap between the local correlation and the global correlation, and the calculation formula is as shown in (24): Loss ass (S,P) = [KL(S i ,P i ) + KL(P i ,S i )] i=t,…,w (24) where S represents the global correlation, P represents the local correlation, w is the window length, KL is the KL divergence, and the calculation formula of the KL divergence is as shown in (25): where P represents the probability distribution of the local correlation, Q represents the probability distribution of the global correlation, and log represents the natural logarithm; (2) Reconstruct the multi-dimensional time series data through the reconstruction layer, and the calculation formula is as shown in (26): where X is the original input data, is the reconstructed value; (3) Realize the reconstruction loss by fusing the reconstruction difference and the correlation difference; in the minimization stage, the goal is to reduce the reconstruction difference and the correlation difference, and the calculation formula is as shown in (27): In the maximization stage, the training goal is to reduce the reconstruction difference and increase the correlation difference, and the calculation formula is as shown in (28): where, ‖·‖ represents the norm, X is the original input data, is the reconstructed value, λ is the control coefficient, stopgrad represents stopping the gradient, S represents global correlation, and P represents local correlation; (4) Calculate the correlation difference score. The smaller the correlation difference, the more likely it is an anomaly. Therefore, take the negative of the difference and output it through the softmax function, and the calculation formulas are as shown in (29) and (30): L score = [-KL(S i , stopgrad(P i )) - KL(P i , stopgrad(S i ))] i=1,…,w (29) Score ass = softmax(L score ) (30) Among them, S i represents the global association distribution at the i-th position, and P i represents the local association distribution at the i-th position. stopgrad represents stopping the gradient, and w is the sequence length of the input multi-dimensional time series data; (5) Calculate the anomaly score at time t by fusing the reconstruction difference and the association difference, and adopt two calculation methods AS mul and AS add respectively to adapt to different scenarios. The calculation formulas are as shown in (31) and (32): Among them, AS mul is used in scenarios where it is necessary to mark the start and end of an anomaly. AS add is universal and applicable to detecting multiple types of anomalies. After normalizing the anomaly scores of both, weighted merging is used to obtain the final anomaly score, and the calculation formula is as shown in (33): Anomaly Score=uAS spatial +(1 - u)AS temporal (33) Among them, AS spatial represents the anomaly score obtained by the spatial dimension model, and AS temporal represents the anomaly score obtained by the time dimension model, while the parameter u is a weight parameter used to reflect the emphasis on the two models; The maximum anomaly score is used as the anomaly threshold, which specifically includes: traversing the validation set to calculate the anomaly score of each window, selecting the maximum score as the threshold, and determining the anomaly data through this threshold.

9. The method for constructing a multi-dimensional time series data anomaly detection system based on a graph neural network and association differences according to claim 1, characterized in that, In step 4, based on the idea of associated partitioning, spatial partitioning of multi-dimensional time series data is performed through an association graph and spectral clustering; based on the distributed computing framework implemented by Apache Flink, partitioning and detection of multi-dimensional time series data are completed through real-time stream processing; the real-time anomaly detection system includes a data storage module, a data partitioning module, an anomaly detection module, an anomaly reasoning module, and a Web front-end module; The data storage module is used for the differential storage of multi-dimensional time series data. The time series database TDengine is used to store real-time monitoring data and historical test results, the relational database MySQL is used to store the file addresses of the models, as well as the data partitions corresponding to the data measurement points, and Neo4J is used to store the association graph structure of multi-dimensional time series data; The data partitioning module is used to classify the original multi-dimensional time series data; the anomaly detection module is used to detect anomalies in real-time stream data, and at the same time solve the problems of data merging and missing value processing, and uses the DJL framework to implement the loading of a deep anomaly detection model by a Java program; the anomaly reasoning module is used to locate anomaly data and find the anomaly propagation chain; finally, the visualization of anomaly detection is realized through the Web front-end.

10. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, the steps of the method for constructing a multi-dimensional time series data anomaly detection system according to any one of claims 1-8 are implemented.

Citation Information

Cited By

  • Intelligent data monitoring method and monitoring platform

    CN120492213A

  • Mine ventilator anomaly detection system and method based on anti-deductive learning

    CN121093223A

  • Time sequence anomaly detection method and device for association difference mining, equipment and medium

    CN121435079A