Anomaly detection method and system based on GCN-LSTM and attention mechanism

By using GCN-LSTM and attention mechanism methods in the abnormal detection of cloud computing services, combining GCN network and LSTM network to extract multi-dimensional features, and using Copula function to accurately define the abnormal threshold, the problems of limited application scenarios and insufficient feature utilization in the existing technology are solved, and high-precision abnormality detection is achieved.

CN115168443BActive Publication Date: 2025-05-02GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210719478.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-05-02
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

In the abnormal detection of cloud computing services, the existing technology has problems such as limited application scenarios, unbalanced data classification, insufficient feature utilization, and insufficient correlation of data dimensions, and it is especially impossible to effectively utilize multidimensional features.

Method used

Using anomaly detection method based on GCN-LSTM and attention mechanism, the GCN-LSTM model combining GCN network and LSTM network is constructed, and the encoder-decoder model of the attention model is integrated into the attention model, the multi-dimensional features of the data sequence to be detected are extracted, and the abnormality threshold is accurately defined through the Copula function anomaly detection model.

Benefits of technology

It realizes accurate detection of abnormalities in a wider range of application scenarios, makes full use of multi-dimensional features, and improves the accuracy and interpretability of abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115168443B_ABST
    Figure CN115168443B_ABST
Patent Text Reader

Abstract

The present invention discloses an anomaly detection method and system based on GCN-LSTM and attention mechanism, which relates to the field of artificial intelligence detection technology, and includes the following steps: S1. constructing a GCN-LSTM model combining a GCN network and a LSTM network, and constructing a sequence reconstruction model based on the GCN-LSTM model; S2. training and testing the obtained sequence reconstruction model to obtain a tested sequence reconstruction model; S3. arranging the data sequence to be detected in time series, inputting the trained sequence reconstruction model, and generating a reconstructed sequence; S4. constructing an error sequence, and dividing the error sequence into a training data set and a test data set; S5. constructing an anomaly detection model based on a Copula function, and inputting the training data set into the anomaly detection model for training; S6. inputting the test data set into the trained anomaly detection model to obtain the test set abnormal sequence data detection result. The present invention solves the problem that the prior art has a small application scenario and cannot utilize multidimensional features, and has the characteristics of accurate results and clear steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence detection technology, and more specifically, to an anomaly detection method and system based on GCN-LSTM and attention mechanism. Background Art

[0002] In today's society, cloud computing services are being widely used. Cloud computing services refer to the increase, use and interaction of related services based on the Internet, providing scalable and virtualized resources through Internet technology. As the architecture of cloud computing services becomes increasingly large, the corresponding performance indicator data of the system will also increase, so a large number of operation and maintenance personnel are needed to deploy and maintain the system environment.

[0003] There is no unified definition of anomalies in cloud computing services, so anomalies in different application scenarios may have different definitions. The currently commonly defined outliers refer to observations that deviate from the overall sample. After years of development, different types of anomaly detection have been formed, such as rule-based anomaly detection methods, statistics-based anomaly detection methods, machine learning-based anomaly detection methods, and probability statistics-based anomaly detection methods.

[0004] Existing anomaly detection methods face the following difficulties and challenges:

[0005] First, the usage scenarios are limited. Many methods can only maintain excellent performance under specific requirements, such as having a complete rule base and expert knowledge base, and having data that meets the model requirements. If these requirements are not met, these methods often perform poorly.

[0006] Second, data classification is unbalanced and highly dependent on labels. Since the frequency of abnormalities in the actual operation of cloud servers is low and the amount of abnormal data collected is small, most of the current anomaly detection algorithms with good results are supervised algorithms, which require a large amount of labeled abnormal data; therefore, how to solve the problem of data shortage is an issue that the current anomaly detection task has to face and think about.

[0007] The third is insufficient feature utilization. The data recorded by the cloud server is first of all a time series data, and its complex dynamic changes and periodicity need to be considered; because cloud computing services cover a wide range of regions and fields, and are also a type of data with spatial dimension attributes, different spatial distributions make its internal operations different. Therefore, how to make full use of the changes between time series context variables and changes in spatial distribution is one of the issues that need to be considered.

[0008] Fourth, there is the problem of correlation between different dimensions of data. The dimensions of data recorded by the current cloud server are not independent of each other. There may be non-steady-state dependencies between data of different dimensions, but the existing anomaly detection algorithm cannot fully consider the correlation between data of different dimensions.

[0009] Fifth, the abnormal threshold definition standard used by the current detection method is vague. Data sets in different fields have different definitions of the concept of abnormality. In the actual cloud computing service engineering application, how to select the appropriate threshold to distinguish whether it is abnormal or not is also a major difficulty of the current algorithms and models.

[0010] The prior art has an unsupervised learning image anomaly detection method based on autoencoders, which is as follows: divide samples into training samples and test samples, preprocess the training samples and test samples respectively, and then input the preprocessed training samples / test samples into the autoencoder for reconstruction to obtain the reconstruction results, and calculate the reconstruction loss, the weighted feature consistency loss between the corresponding layers of the encoder and decoder during the reconstruction process, the feature discrimination loss and the adversarial loss respectively; then the weighted sum of the above losses is used as the total loss function; finally, the anomaly score of the test sample is calculated. Then, feature normalization is used to map the anomaly score of each sample to [0,1], and the area under the receiver operating characteristic curve is calculated as the evaluation indicator.

[0011] However, the existing technology has the problem of small application scenarios and inability to utilize multi-dimensional features. Therefore, how to invent an anomaly detection method with a large application scenario and capable of utilizing multi-dimensional features is an urgent problem to be solved in this technical field. Summary of the invention

[0012] In order to solve the problems that the prior art has small application scenarios and cannot utilize multi-dimensional features, the present invention provides an anomaly detection method and system based on GCN-LSTM and attention mechanism, which has the characteristics of accurate results and clear steps.

[0013] In order to achieve the above-mentioned purpose of the present invention, the technical scheme adopted is as follows:

[0014] An anomaly detection method based on GCN-LSTM and attention mechanism includes the following steps:

[0015] S1. Construct a GCN-LSTM model that combines a GCN network and an LSTM network, and construct a sequence reconstruction model based on the GCN-LSTM model; the sequence reconstruction model is an encoder-decoder model based on the GCN-LSTM model that incorporates an attention model;

[0016] S2. training and testing the obtained sequence reconstruction model to obtain a tested sequence reconstruction model;

[0017] S3. Arrange the data sequence to be detected in time sequence and input it into the trained sequence reconstruction model, extract the features of the data sequence to be detected through the encoder, and obtain the weighted vector of the features of the data sequence to be detected through the attention model, and finally generate the reconstructed sequence by combining the features and the weighted vector through the decoder;

[0018] S4. Subtract the reconstructed sequence from the data sequence to be tested to construct an error sequence, and divide the error sequence into a training data set and a test data set;

[0019] S5. Build an anomaly detection model based on Copula function, and input the training data set into the anomaly detection model for training;

[0020] S6. Input the test data set into the trained anomaly detection model for anomaly detection to obtain the test set anomaly sequence data detection result.

[0021] Preferably, in step S3, the data sequence to be detected is arranged in time sequence and input into the trained sequence reconstruction model, the features of the data sequence to be detected are extracted by the encoder, and the weighted vector of the features of the data sequence to be detected is obtained by the attention model, and finally the process of combining the features and the weighted vector to generate the reconstructed sequence by the decoder is specifically as follows;

[0022] S301. Obtain a data sequence to be detected, arrange the data sequence to be detected in time sequence, input the arranged data sequence to be detected into an encoder in a sequence reconstruction model, and extract features of the data sequence to be detected through the encoder;

[0023] S302. Input the feature sequence into the attention model, and calculate the weighted vector by assigning weights through the attention model;

[0024] S303. Input the features combined with the weighted vector into the decoder, and obtain the reconstructed sequence through decoding by the decoder.

[0025] Furthermore, in step S301, the arranged data sequence to be detected is input into the encoder in the encoder-decoder model in the sequence reconstruction model, and the process of extracting the features of the data sequence to be detected by the encoder is specifically as follows:

[0026] A1. Assume that the initial time t0 = t-s+1, where s is the length of the data sequence to be detected, and t is the time variable. The arranged data sequence to be detected is represented as the data sequence to be detected

[0027] A2. Combine the edge set E of the GCN network and extract the data in the data sequence to be detected through the GCN network unit of the first GCN-LSTM model The spatial features are input into the LSTM network of the first GCN-LSTM model to extract its temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model;

[0028] A3. Combined with the hidden layer state obtained from the previous GCN-LSTM model The edge set E of the GCN network is used to extract the data in the data sequence to be detected through the GCN network of the second GCN-LSTM model. The spatial features are input into the LSTM network of the second GCN-LSTM model to extract its temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model; in this way, until the hidden layer state of the entire data sequence to be detected is obtained, and the hidden layer state of the entire data sequence to be detected is organized into a sequence of features

[0029] Furthermore, in the step S302, the sequence of features is input into the attention model, and the process of calculating the weighted vector by the attention model in a weighted manner is as follows:

[0030] B1. Input the feature sequence into the attention model and calculate the time attention weight vector a corresponding to each historical moment at the current moment t t , specifically expressed as:

[0031]

[0032] Among them, W a is the trainable weight matrix; b a is the bias vector of the attention weight; tanh represents the activation function; finally:

[0033]

[0034] Among them, a i is the attention weight value at the i-th moment;

[0035] B2. Normalize the attention weight coefficients of each time through the softmax function to obtain the time attention weight have:

[0036]

[0037]

[0038] B3. Weight the sequence of temporal attention weights and their corresponding features to obtain the weighted vector c t , specifically expressed as:

[0039]

[0040] Furthermore, in step S6, the test data set is input into the trained anomaly detection model for anomaly detection, and the process of obtaining the test set anomaly sequence data detection result is specifically as follows:

[0041] S601. For each dimension of the test data set, use a non-parametric method to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and right tail of the probability distribution, and calculate the skewness coefficient;

[0042] S602. Calculate the empirical copula observation value of each time snapshot in the test data set according to the obtained empirical cumulative joint distribution of the abnormal samples of each dimension of the test data set in the left tail and the right tail of the probability distribution;

[0043] S603. Obtain the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set;

[0044] S604. Calculate the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set according to the obtained empirical copula observation value and the empirical copula observation value of the skewness coefficient of each time snapshot in the test data set;

[0045] S605. According to the probabilities of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, calculate the anomaly score and anomaly threshold corresponding to each time snapshot, and compare the anomaly score corresponding to each time snapshot with the anomaly threshold to obtain the test set anomaly sequence data detection result.

[0046] Furthermore, for each dimension of the test data set, a non-parametric method is used to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and right tail of the probability distribution, and the skewness coefficient is calculated, specifically:

[0047] C1. Use nonparametric methods to estimate the d-dimensional test data set E = (e 1,i ,e 2,i ,…,e d,i ) in the left tail of the empirical cumulative joint distribution and the empirical cumulative joint distribution in the right tail Specifically:

[0048]

[0049] Where n is the total number of data in the test data set;

[0050] C2. Calculate the skewness coefficient b of the test data set i , specifically:

[0051]

[0052] Furthermore, in step S603, the process of obtaining the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set is specifically as follows:

[0053] D1. Calculate the empirical copula observations in the left tail of each time snapshot and the empirical copula observations in the right tail

[0054]

[0055]

[0056] D2. According to the d-dimensional skewness coefficient b of the test data set d The value of the skewness coefficient determines the observed value of the empirical copula If b d ≥0, then on the contrary

[0057] Furthermore, the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set is calculated. Specifically, the left p of each time snapshot in the test data set is calculated. l , right side, and tail probabilities of skewness:

[0058]

[0059]

[0060]

[0061] Among them, p l is the left tail probability of the i-th time snapshot, p r is the right tail probability of the i-th time snapshot, p s is the tail probability of the skewness at the i-th time snapshot, is the left tail empirical copula observation at the j-th i-th time snapshot, is the right tail empirical copula observation at the j-th i-th time snapshot, is the empirical copula observation of the right tail skewness coefficient at the j-th i-th time snapshot.

[0062] Furthermore, according to the probabilities of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, the anomaly score and anomaly threshold corresponding to each time snapshot are calculated, and the anomaly score corresponding to each time snapshot is compared with the anomaly threshold to obtain the test set anomaly sequence data detection result. The specific steps are as follows:

[0063] E1. Get each time snapshot (e i ) corresponds to the anomaly score O(e i ):

[0064] O(e i )=max{p l ,p r ,p s}

[0065] Among them, the max function returns the largest data in a set of data;

[0066] E2. Calculate the abnormal threshold, specifically:

[0067] O Thershold (e i )=percentile(O(e i ),1-α)

[0068] Among them, O Thershold (e i ) is a time snapshot (e i ) is the abnormal threshold, α is the overall abnormal rate of the historical performance indicator sequence, and the percentile function calculates and analyzes the percentage value point;

[0069] E3. Compare the anomaly score corresponding to each time snapshot with the anomaly threshold to determine whether the data under each time snapshot has an anomaly:

[0070] If O(e i )>O Thershold (e i ), the current time snapshot is abnormal;

[0071] If O(e i )≤O Thershold (e i ), the current time snapshot is normal;

[0072] E4. Arrange the abnormal results of the data at each time snapshot to obtain the abnormal sequence data results of the test set.

[0073] The invention discloses an anomaly detection system based on GCN-LSTM and attention mechanism, comprising a reconstruction model construction module, a model training module, a sequence reconstruction model, a data processing module, a detection model construction module, a Copula training module and an anomaly detection model; the reconstruction model construction module is used to construct a GCN-LSTM model combining a GCN network and an LSTM network, and to construct a sequence reconstruction model based on the GCN-LSTM model; the model training module is used to train and test the sequence reconstruction model; the sequence reconstruction model is used to extract the features of a data sequence to be detected and obtain a weighted vector of the features of the data sequence to be detected, and to generate a reconstructed sequence by combining the features with the weighted vector through a decoder; the data processing module is used to make a difference between the reconstructed sequence and the data sequence to be detected, to construct an error sequence, and to divide the error sequence into two parts, a training data set and a test data set; the detection model construction module is used to construct an anomaly detection model based on a Copula function; the Copula training module is used to input the training data set into the anomaly detection model for training; the anomaly detection model is used to input the test data set into the trained anomaly detection model for anomaly detection to obtain a detection result.

[0074] The beneficial effects of the present invention are as follows:

[0075] The present invention constructs a GCN-LSTM model combining a GCN network and an LSTM network, constructs an encoder-decoder model based on the GCN-LSTM model incorporating an attention model, and incorporates a sequence reconstruction model of the attention model between the encoder and the decoder; by constructing an encoder with the GCN-LSTM model as a neural unit, the present invention can extract spatial and temporal multidimensional features of a data sequence to be detected, and by incorporating the attention model, the correlation problem of the data is considered; the present invention also constructs an error sequence, constructs a training data set and a test data set through the error sequence, and trains and detects the constructed anomaly detection model based on the Copula function through the training data set and the test data set, accurately defines the anomaly threshold, and improves the accuracy of anomaly detection; therefore, the present invention solves the problem of small application scenarios and inability to utilize multidimensional features in the prior art, and has the characteristics of accurate results and clear steps. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is a flowchart of the anomaly detection method based on GCN-LSTM and attention mechanism.

[0077] Figure 2It is a structural diagram of the sequence reconstruction model that introduces the attention mechanism.

[0078] Figure 3 It is a schematic diagram of GCN spatial feature extraction.

[0079] Figure 4 This is a schematic diagram of the internal module connections of the GCN-LSTM model.

[0080] Figure 5 This is a schematic diagram of the topological structure of the MBD data set in Example 2.

[0081] Figure 6 It is the ROC curve diagram of each method in Example 2.

[0082] Figure 7 Schematic diagram of moderate abnormality score in Example 2.

[0083] Figure 8 This is a schematic diagram of an anomaly detection system based on GCN-LSTM and attention mechanism in Example 3. DETAILED DESCRIPTION

[0084] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0085] Example 1

[0086] like Figure 1 As shown in FIG. 1 , an anomaly detection method based on GCN-LSTM and attention mechanism includes the following steps:

[0087] S1. Construct a GCN-LSTM model that combines a GCN network and an LSTM network, and construct a sequence reconstruction model based on the GCN-LSTM model; the sequence reconstruction model is an encoder-decoder model based on the GCN-LSTM model that incorporates an attention model;

[0088] S2. Obtain the index data training set and index data test set of the cloud computing server, train and test the obtained sequence reconstruction model, and obtain a tested sequence reconstruction model;

[0089] S3. Arrange the data sequence to be detected in time sequence and input it into the trained sequence reconstruction model, extract the features of the data sequence to be detected through the encoder, and obtain the weighted vector of the features of the data sequence to be detected through the attention model, and finally generate the reconstructed sequence by combining the features and the weighted vector through the decoder;

[0090] S4. Subtract the reconstructed sequence from the data sequence to be tested to construct an error sequence, and divide the error sequence into a training data set and a test data set;

[0091] S5. Build an anomaly detection model based on Copula function, and input the training data set into the anomaly detection model for training;

[0092] S6. Input the test data set into the trained anomaly detection model for anomaly detection to obtain the test set anomaly sequence data detection result.

[0093] In this embodiment, the data sequence to be detected is historical indicator data of the cloud computing server.

[0094] In this embodiment, Figure 2 As shown in the figure, in the encoder-decoder model, the encoder is used to read the input time series and encode it into a fixed-length vector, and the decoder is used to decode the fixed-length vector and output it as a result sequence. In anomaly detection in a cloud environment, we can understand it as a time series. When a time series is given, another time series can be obtained corresponding to it. In fact, the encoder-decoder structure uses lossy compression and has a noise reduction effect, so it has the ability to filter outliers or abnormal sequences, which is very effective in time series anomaly detection.

[0095] In this embodiment, the attention mechanism can selectively learn all hidden vectors and associate the decoder's input sequence with them at the output. Unlike the traditional encoder-decoder architecture that uses a fixed vector representation, when the encoder inputs a sequence, it selectively pays attention to the information generated by the sequence in the input. After introducing the attention mechanism, the decoder can integrate the input at each moment with the data trained by the attention model according to the different moments.

[0096] In this embodiment, the GCN-LSTM model includes a GCN network for extracting spatial features and an LSTM network for extracting temporal features; the GCN-LSTM model uses the GCN network to extract spatial information and attribute information, and deeply mines the characteristic rules in the graph module. Figure 3 As shown in the figure, the essence of the GCN network is to propagate the weighted average of the features of each node and the feature information of the nodes connected to it to the next layer, and as the number of layers increases, each node can aggregate node information farther away, thereby representing the structural characteristics of the entire graph module. The spatial data of the extracted sequence can achieve efficient and accurate sequence reconstruction. The sequence reconstruction model of the GCN network in this embodiment is integrated, so that the fused sequence reconstruction model has the spatial feature extraction of the data sequence to be detected, which better solves the problem of low accuracy of anomaly detection due to the lack of spatial features in the time series.

[0097] In this embodiment, the LSTM network can selectively obtain important information according to the characteristics of the sequence when processing the data sequence to be detected, and ignore unimportant information. Compared with the conventional recurrent neural network, it can have better effects in longer sequences. Therefore, it has been widely used in processing the data sequence to be detected, mainly in the fields of sentence generation, text classification and machine translation.

[0098] In a specific embodiment, in step S3, the data sequence to be detected is arranged in time sequence and input into the trained sequence reconstruction model, the features of the data sequence to be detected are extracted by the encoder, and the weighted vector of the features of the data sequence to be detected is obtained by the attention model, and finally the decoder combines the features and the weighted vector to generate the reconstructed sequence, which is specifically as follows;

[0099] S301. Obtain a data sequence to be detected, arrange the data sequence to be detected in time sequence, input the arranged data sequence to be detected into an encoder in a sequence reconstruction model, and extract features of the data sequence to be detected through the encoder;

[0100] S302. Input the feature sequence into the attention model, and calculate the weighted vector by assigning weights through the attention model;

[0101] S303. Input the features combined with the weighted vector into the decoder, and obtain the reconstructed sequence through decoding by the decoder.

[0102] like Figure 4 As shown, in a specific embodiment, in the step S301, the arranged data sequence to be detected is input into the encoder of the encoder-decoder model in the sequence reconstruction model, and the process of extracting the features of the data sequence to be detected by the encoder is specifically as follows:

[0103] A1. Assume that the initial time t0 = t-s+1, where s is the length of the data sequence to be detected, and t is the time variable. The arranged data sequence to be detected is represented as the data sequence to be detected

[0104] A2. Combine the edge set E of the GCN network and extract the data in the data sequence to be detected through the GCN network unit of the first GCN-LSTM model The spatial features are input into the LSTM network of the first GCN-LSTM model to extract its temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model;

[0105] A3. Combined with the hidden layer state obtained from the previous GCN-LSTM model The edge set E of the GCN network is used to extract the data in the data sequence to be detected through the GCN network of the second GCN-LSTM model. The spatial features are input into the LSTM network of the second GCN-LSTM model to extract its temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model; in this way, until the hidden layer state of the entire data sequence to be detected is obtained, and the hidden layer state of the entire data sequence to be detected is organized into a sequence of features

[0106] In this embodiment, the encoder calculates the hidden layer state h at time t t The formula can be expressed as:

[0107] h t =f([h t-1 ,x t ],E)

[0108] Among them, f represents the function of the simplified GCN-LSTM Encoder model.

[0109] In a specific embodiment, in step S302, the sequence of features is input into the attention model, and the process of calculating the weighted vector by the attention model in a weighted manner is as follows:

[0110] B1. Input the feature sequence into the attention model and calculate the time attention weight vector a corresponding to each historical moment at the current moment t t ={a t-s+1 ,a t-s+2 ,…,a i ,…,a t-1 ,a t}:

[0111]

[0112] Among them, a i is the attention weight value at the i-th moment, W a is the trainable weight matrix; b a is the bias vector of the attention weight; tanh represents the activation function;

[0113] B2. Normalize the attention weight coefficients of each time through the softmax function to obtain the time attention weight

[0114]

[0115] B3. Weight the sequence of temporal attention weights and their corresponding features to obtain the weighted vector c t :

[0116]

[0117] In this embodiment, the decoding formula of the decoder can be expressed as:

[0118] x′ i-1 =g([h′ i ,x′ i ,c t ],E)

[0119] Where g represents the function of the simplified decoder GCN-LSTM model, so that more information in the encoder is involved in the decoding stage, thereby optimizing the error between the reconstructed sequence and the input sequence, x′ i-1 represents the i-1th reconstructed data, x′ i represents the i-th reconstructed data, h′ i represents the hidden layer state output by the GCN-LSTM model of the i-th decoder.

[0120] In this embodiment, the reconstructed sequence and the data sequence to be detected are subtracted to construct an error sequence. Specifically, the error sequence is represented as E = [e1, ..., e n ], where e i =x i -x′ i , divided into training data set and test data set in a ratio of 2:1.

[0121] In this embodiment, since the GCN network can extract the internal information of the established graph data module, analyze its topological structure, and extract its spatial features. The features extracted by the GCN network at different times are input into the LSTM network in a time series manner to learn the time features, thereby improving the accuracy of the reconstructed data. This embodiment designs a reconstruction method that integrates GCN and LSTM, combines the ability of the GCN network to process spatial features, and uses the time series processing ability of LSTM to improve the data fitting effect, and uses it as a unit of the encoder-decoder, innovatively integrating the attention mechanism to achieve the reconstruction of cloud server data.

[0122] Example 2

[0123] like Figure 1 As shown in FIG. 1 , an anomaly detection method based on GCN-LSTM and attention mechanism includes the following steps:

[0124] S1. Construct a GCN-LSTM model that combines a GCN network and an LSTM network, and construct a sequence reconstruction model based on the GCN-LSTM model; the sequence reconstruction model is an encoder-decoder model based on the GCN-LSTM model that incorporates an attention model;

[0125] S2. training and testing the obtained sequence reconstruction model to obtain a tested sequence reconstruction model;

[0126] S3. Arrange the data sequence to be detected in time sequence and input it into the trained sequence reconstruction model, extract the features of the data sequence to be detected through the encoder, and obtain the weighted vector of the features of the data sequence to be detected through the attention model, and finally generate the reconstructed sequence by combining the features and the weighted vector through the decoder;

[0127] S4. Subtract the reconstructed sequence from the data sequence to be tested to construct an error sequence, and divide the error sequence into a training data set and a test data set;

[0128] S5. Build an anomaly detection model based on Copula function, and input the training data set into the anomaly detection model for training;

[0129] S6. Input the test data set into the trained anomaly detection model for anomaly detection to obtain the test set anomaly sequence data detection result.

[0130] In a specific embodiment, in step S3, the data sequence to be detected is arranged in time sequence and input into the trained sequence reconstruction model, the features of the data sequence to be detected are extracted by the encoder, and the weighted vector of the features of the data sequence to be detected is obtained by the attention model, and finally the decoder combines the features and the weighted vector to generate the reconstructed sequence, which is specifically as follows;

[0131] S301. Obtain a data sequence to be detected, arrange the data sequence to be detected in time sequence, input the arranged data sequence to be detected into an encoder in a sequence reconstruction model, and extract features of the data sequence to be detected through the encoder;

[0132] S302. Input the feature sequence into the attention model, and calculate the weighted vector by assigning weights through the attention model;

[0133] S303. Input the features combined with the weighted vector into the decoder, and obtain the reconstructed sequence through decoding by the decoder.

[0134] In a specific embodiment, in step S301, the arranged data sequence to be detected is input into the encoder of the encoder-decoder model in the sequence reconstruction model, and the process of extracting the features of the data sequence to be detected by the encoder is specifically as follows:

[0135] A1. Assume that the initial time t0 = t-s+1, where s is the length of the data sequence to be detected, and t is the time variable. The arranged data sequence to be detected is represented as the data sequence to be detected

[0136] A2. Combine the edge set E of the GCN network and extract the data in the data sequence to be detected through the GCN network unit of the first GCN-LSTM model The spatial features are input into the LSTM network of the first GCN-LSTM model to extract its temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model;

[0137] A3. Combined with the hidden layer state obtained from the previous GCN-LSTM model The edge set E of the GCN network is used to extract the data in the data sequence to be detected through the GCN network of the second GCN-LSTM model. The spatial features are input into the LSTM network of the GCN-LSTM model to extract their temporal features and obtain The hidden layer state The hidden layer state Input into the next GCN-LSTM model; in this way, until the hidden layer state of the entire data sequence to be detected is obtained, and the hidden layer state of the entire data sequence to be detected is organized into a sequence of features

[0138] In a specific embodiment, in step S302, the sequence of features is input into the attention model, and the process of calculating the weighted vector by the attention model in a weighted manner is as follows:

[0139] B1. Sequence of features Input the attention model and calculate the time attention weight vector a corresponding to each historical moment at the current moment t t ={a t-s+1 ,a t-s+2 ,…,a i ,…,a t-1 ,a t}:

[0140]

[0141] Among them, a i is the attention weight value at the i-th moment, W a is the trainable weight matrix; b ais the bias vector of the attention weight; tanh represents the activation function;

[0142] B2. Normalize the attention weight coefficients of each time through the softmax function to obtain the time attention weight

[0143]

[0144] B3. Weight the sequence of temporal attention weights and their corresponding features to obtain the weighted vector c t :

[0145]

[0146] In a specific embodiment, in step S6, the test data set is input into the trained anomaly detection model for anomaly detection, and the process of obtaining the test set anomaly sequence data detection result is specifically as follows:

[0147] S601. For each dimension of the test data set, use a non-parametric method to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and right tail of the probability distribution, and calculate the skewness coefficient;

[0148] S602. Calculate the empirical copula observation value of each time snapshot in the test data set according to the obtained empirical cumulative joint distribution of the abnormal samples of each dimension of the test data set in the left tail and the right tail of the probability distribution;

[0149] S603. Obtain the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set;

[0150] S604. Calculate the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set according to the obtained empirical copula observation value and the empirical copula observation value of the skewness coefficient of each time snapshot in the test data set;

[0151] S605. According to the probabilities of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, calculate the anomaly score and anomaly threshold corresponding to each time snapshot, and compare the anomaly score corresponding to each time snapshot with the anomaly threshold to obtain the test set anomaly sequence data detection result.

[0152] The anomaly detection method used in this embodiment is the COPOD method. COPOD is an anomaly detection method based on Copula. After years of development, Copula theory has become an effective means to solve the joint probability distribution problem of high-dimensional random variables. Since the copula function can summarize the correlation of each marginal distribution to a certain extent, COPOD can provide some explainability for which dimensions cause the anomaly. For example, operation and maintenance personnel can directly find the dimensions that cause the most anomalies and conduct in-depth analysis.

[0153] In a specific embodiment, for each dimension of the test data set, a non-parametric method is used to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and the right tail of the probability distribution, and the skewness coefficient is calculated, specifically:

[0154] C1. Use nonparametric methods to estimate the d-dimensional test data set E = (e 1,i ,e 2,i ,…,e d,i ) in the left tail of the empirical cumulative joint distribution and the empirical cumulative joint distribution in the right tail Specifically:

[0155]

[0156]

[0157] Where n is the total number of data in the test data set;

[0158] C2. Calculate the skewness coefficient b of the test data set i , specifically:

[0159]

[0160] In a specific embodiment, in step S603, the process of obtaining the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set is specifically as follows:

[0161] D1. Calculate the empirical copula observations in the left tail of each time snapshot and the empirical copula observations in the right tail Specifically:

[0162]

[0163]

[0164] D2. According to the d-dimensional skewness coefficient b of the test data setd The value of the skewness coefficient determines the observed value of the empirical copula If b d ≥0, then on the contrary

[0165] In a specific embodiment, the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set is calculated, specifically: the left side p of each time snapshot in the test data set is calculated l , right side, and tail probabilities of skewness:

[0166]

[0167]

[0168]

[0169] Among them, p l is the left tail probability of the i-th time snapshot, p r is the right tail probability of the i-th time snapshot, p s is the tail probability of the skewness at the i-th time snapshot, is the left tail empirical copula observation at the j-th i-th time snapshot, is the right tail empirical copula observation at the j-th i-th time snapshot, is the empirical copula observation of the right tail skewness coefficient at the j-th i-th time snapshot.

[0170] In a specific embodiment, according to the probabilities of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, the anomaly score and anomaly threshold corresponding to each time snapshot are calculated, and the anomaly score corresponding to each time snapshot is compared with the anomaly threshold to obtain the test set abnormal sequence data detection result. The specific steps are:

[0171] E1. Get each time snapshot (e i ) corresponds to the anomaly score O(e i ):

[0172] O(e i )=max{p l ,p r ,p s}

[0173] Among them, the max function returns the largest data in a set of data;

[0174] E2. Calculate the abnormal threshold, specifically:

[0175] O Thershold (e i )=percentile(O(e i ),1-α)

[0176] Among them, O Thershold (e i ) is a time snapshot (e i ) is the abnormal threshold, α is the overall abnormal rate of the historical performance indicator sequence, and the percentile function calculates and analyzes the percentage value point;

[0177] E3. Compare the anomaly score corresponding to each time snapshot with the anomaly threshold to determine whether the data under each time snapshot has an anomaly:

[0178] If O(e i )>O Thershold (e i ), the current time snapshot is abnormal;

[0179] If O(e i )≤O Thershold (e i ), the current time snapshot is normal;

[0180] E4. Arrange the abnormal results of the data at each time snapshot to obtain the abnormal sequence data results of the test set.

[0181] In this embodiment, in order to verify the effectiveness, rationality, feasibility and scientificity of the method proposed in this application. The MBD dataset is used to experiment with the anomaly detection method based on GCN-LSTM and attention mechanism. The dataset comes from an environment running a big data batch processing system, one of which contains 1 master node and 4 slave nodes. The monitored and collected data includes 26 indicators for each node, including CPU idle, CPU I / O wait, CPU software, CPU system, CPU system, CPU user, wait per second, disk I / O process, disk usage percentage, disk read speed, disk write speed, kernel entropy, load, etc. The graph module is as follows Figure 5 As shown. R740-3-1, R740-3-2, R740-3-3, R740-3-4 and R740-3-5 represent cloud servers located in different locations; the lines between the servers represent the interactions between them; the data indicators of the past operation stored by the servers can be used to construct the nodes, edges and attributes of the graph module respectively, so as to extract the spatial features within the graph module.

[0182] As shown in the following table, Table 1 shows the experimental results of accuracy, precision and recall of different algorithms. This experiment uses a variety of methods that have made achievements in the field of anomaly detection as comparative experiments, including distance-based KNN method, density-based LOF and COF methods, tree-based iForest method, probability-based ABOD method and deep learning-based AutoEncoder method. We also conducted experimental comparisons with the method of the benchmark module LSTM to better evaluate the module effect. At the same time, the model using this anomaly detection method based on GCN-LSTM and attention mechanism is named GL2GL-Att-Co.

[0183] Figure 6 The following table shows the ROC curves of different algorithms (selected top) and the AUC values ​​of different algorithms:

[0184] method KNN COF iForest COPOD AE LSTM_COPOD LOF ABOD GL2GL-Att-Co Accuracy 0.68 0.84 0.89 0.87 0.88 0.89 0.74 0.78 0.93 Accuracy 0.87 0.88 0.88 0.86 0.90 0.88 0.90 0.89 0.91 Recall 0.68 0.84 0.89 0.87 0.88 0.89 0.74 0.77 0.93 AUC 0.58 0.60 0.63 0.64 0.66 0.70 0.71 0.71 0.74

[0185] From the experimental results, it can be concluded that the anomaly detection method of GCN-LSTM and attention mechanism proposed in this invention is superior to other methods in terms of accuracy, precision, recall, AUC value, etc. on the MBD dataset. In addition to the improvement in accuracy, precision and recall to a certain extent, compared with other methods, the method proposed in this article can obtain a better AUC value, which is about 20% to 40% higher than other methods, which means that the effectiveness and detection value of this method in terms of anomalies have been improved. Figure 7 As shown in the figure, the ROC curve can well demonstrate the ability of various methods to detect anomalies, showing that the cloud service time series anomaly detection algorithm based on the GCN network proposed in this invention is superior to other types of methods and has high experimental accuracy. In addition, compared with the base module LSTM, GCN-LSTM has better applicability in acquiring spatial information, so that the module can learn more sufficient information during training, and can effectively reproduce spatial information and time information during sequence reconstruction, and the construction error at this time can achieve better results.

[0186] In terms of interpretability, the anomaly detection method used in the present invention can be regarded as a special case of a Copula function. When processing reconstructed error data, the method first estimates the distribution of the reconstructed error data by calculating the empirical cumulative distribution of each data dimension using a non-parametric method. For the data of each time node, the empirical cumulative distribution is used to estimate the tail probability of each dimension of the time node. Finally, the anomaly score of each time snapshot is calculated by aggregating the tail probabilities of each dimension. Finally, it is checked which dimension's anomaly score has the greatest contribution to the anomaly result, so that the method can be explained. Figure 7As shown in the figure, this is a schematic diagram of the anomaly scores of the last 20 dimensions when the MBD dataset is in an abnormal state. The dimension with number 14 has an anomaly score far exceeding that of other dimensions. The data represented by this dimension is the waiting time for establishing a TCP connection. Therefore, it means that the anomaly at this time is very likely caused by the establishment of TCP.

[0187] The present invention constructs a GCN-LSTM model combining a GCN network and an LSTM network, constructs an encoder-decoder model based on the GCN-LSTM model incorporating an attention model, and incorporates a sequence reconstruction model of the attention model between the encoder and the decoder; by constructing an encoder with the GCN-LSTM model as a neural unit, the present invention can extract spatial and temporal multidimensional features of a data sequence to be detected, and by incorporating the attention model, the correlation problem of the data is considered; the present invention also constructs an error sequence, constructs a training data set and a test data set through the error sequence, and trains and detects the constructed anomaly detection model based on the Copula function through the training data set and the test data set, accurately defines the anomaly threshold, and improves the accuracy of anomaly detection; therefore, the present invention solves the problem of small application scenarios and inability to utilize multidimensional features in the prior art, and has the characteristics of accurate results and clear steps.

[0188] Example 3

[0189] The invention discloses an anomaly detection system based on GCN-LSTM and attention mechanism, comprising a reconstruction model construction module, a model training module, a sequence reconstruction model, a data processing module, a detection model construction module, a Copula training module and an anomaly detection model; the reconstruction model construction module is used to construct a GCN-LSTM model combining a GCN network and an LSTM network, and to construct a sequence reconstruction model based on the GCN-LSTM model; the model training module is used to train and test the sequence reconstruction model; the sequence reconstruction model is used to extract the features of a data sequence to be detected and obtain a weighted vector of the features of the data sequence to be detected, and to generate a reconstructed sequence by combining the features with the weighted vector through a decoder; the data processing module is used to make a difference between the reconstructed sequence and the data sequence to be detected, to construct an error sequence, and to divide the error sequence into two parts, a training data set and a test data set; the detection model construction module is used to construct an anomaly detection model based on a Copula function; the Copula training module is used to input the training data set into the anomaly detection model for training; the anomaly detection model is used to input the test data set into the trained anomaly detection model for anomaly detection to obtain a detection result.

[0190] In this embodiment, Figure 8As shown in the figure, after the data sequence to be detected is input, the spatial and temporal features are first extracted through the trained sequence reconstruction model to generate a reconstructed sequence, and then the reconstructed sequence is input into the data sorting module to generate an error sequence and divided into a training data set and a test data set in a ratio of 2:1. Then, the training module of the anomaly detection model is trained with the training data set, and then the test data set is input into the trained anomaly detection model to obtain the final anomaly detection result.

[0191] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. An anomaly detection method based on GCN-LSTM and attention mechanism, characterized in that: The following steps are involved: S1. Construct a GCN-LSTM model that combines a GCN network and an LSTM network, and construct a sequence reconstruction model based on the GCN-LSTM model; the sequence reconstruction model is an encoder-decoder model based on the GCN-LSTM model that incorporates an attention model; S2. training and testing the obtained sequence reconstruction model to obtain a tested sequence reconstruction model; S3. Arrange the data sequence to be detected in time sequence and input it into the trained sequence reconstruction model. The encoder extracts the features of the data sequence to be detected, and the attention model obtains the weighted vector of the features of the data sequence to be detected. Finally, the decoder combines the features and the weighted vector to generate a reconstructed sequence. In step S3, the data sequence to be detected is arranged in time sequence and input into the trained sequence reconstruction model, the features of the data sequence to be detected are extracted by the encoder, and the weighted vector of the features of the data sequence to be detected is obtained by the attention model, and finally the decoder combines the features and the weighted vector to generate the reconstructed sequence. The process is as follows; S301. Obtain a data sequence to be detected, arrange the data sequence to be detected in time sequence, input the arranged data sequence to be detected into an encoder in an encoder-decoder model in a sequence reconstruction model, and extract features of the data sequence to be detected through the encoder; In step S301, the arranged data sequence to be detected is input into the encoder in the encoder-decoder model in the sequence reconstruction model, and the process of extracting the features of the data sequence to be detected by the encoder is specifically as follows: A1. Set the initial time , where s is the length of the data sequence to be detected, t is the time variable, and the arranged data sequence to be detected is represented as the data sequence to be detected ; A2. Combine the edge set E of the GCN network and extract the data in the data sequence to be detected through the GCN network unit of the first GCN-LSTM model The spatial features are input into the LSTM network of the first GCN-LSTM model to extract its temporal features and obtain The hidden state of ; The hidden layer state obtained Input into the next GCN-LSTM model; A3. Combined with the hidden layer state obtained from the previous GCN-LSTM model The edge set E of the GCN network is used to extract the data in the data sequence to be detected through the GCN network of the second GCN-LSTM model. The spatial features are input into the LSTM network of the second GCN-LSTM model to extract its temporal features and obtain The hidden layer state ; The hidden layer state obtained Input into the next GCN-LSTM model; in this way, until the hidden layer state of the entire data sequence to be detected is obtained, and the hidden layer state of the entire data sequence to be detected is organized into a sequence of features ; S302. Input the feature sequence into the attention model, and calculate the weighted vector by assigning weights through the attention model; In step S302, the sequence of features is input into the attention model, and the process of calculating the weighted vector by the attention model in a weighted manner is as follows: B1. Input the feature sequence into the attention model and calculate the current moment Time attention weight vector corresponding to each historical moment , specifically expressed as: in, is a trainable weight matrix; is the bias vector of attention weights; represents the activation function; finally: in, For the The attention weight value at the moment; B2. Normalize the attention weight coefficients of each time through the softmax function to obtain the time attention weight ,have: : = B3. Weight the sequence of temporal attention weights and their corresponding features to obtain a weighted vector , specifically expressed as: S303. Input the feature combined with the weighted vector into the decoder, and decode the decoder to obtain a reconstructed sequence; S4. Subtract the reconstructed sequence from the data sequence to be tested to construct an error sequence, and divide the error sequence into a training data set and a test data set; S5. Build an anomaly detection model based on Copula function, and input the training data set into the anomaly detection model for training; S6. Input the test data set into the trained anomaly detection model for anomaly detection, and obtain the test set anomaly sequence data detection result.

2. The anomaly detection method based on GCN-LSTM and attention mechanism as claimed in claim 1, characterized in that: In step S6, the test data set is input into the trained anomaly detection model for anomaly detection, and the process of obtaining the test set anomaly sequence data detection result is specifically as follows: S601. For each dimension of the test data set, use a non-parametric method to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and right tail of the probability distribution, and calculate the skewness coefficient; S602. Calculate the empirical copula observation value of each time snapshot in the test data set according to the obtained empirical cumulative joint distribution of the abnormal samples of each dimension of the test data set in the left tail and the right tail of the probability distribution; S603. Obtaining the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set; S604. Calculate the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set according to the obtained empirical copula observation value and the empirical copula observation value of the skewness coefficient of each time snapshot in the test data set; S605. According to the probabilities of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, calculate the anomaly score and anomaly threshold corresponding to each time snapshot, and compare the anomaly score corresponding to each time snapshot with the anomaly threshold to obtain the test set anomaly sequence data detection result.

3. The anomaly detection method based on GCN-LSTM and attention mechanism as claimed in claim 2, characterized in that: For each dimension of the test data set, a non-parametric method is used to estimate the empirical cumulative joint distribution of its abnormal samples in the left tail and right tail of the probability distribution, and the skewness coefficient is calculated, specifically: C1. Estimation of d-dimensional test data set using nonparametric methods Empirical cumulative joint distribution of the middle left tail and the empirical cumulative joint distribution in the right tail , specifically: Where n is the total number of data in the test data set; C2. Calculate the skewness coefficient of the test data set , specifically: 。 4. The anomaly detection method based on GCN-LSTM and attention mechanism as claimed in claim 3, characterized in that: In step S603, the process of obtaining the empirical copula observation value of the skewness coefficient according to the empirical copula observation value of each time snapshot in the test data set is specifically as follows: D1. Calculate the empirical copula observations in the left tail of each time snapshot and the empirical copula observations in the right tail : ; D2. Based on the d-dimensional skewness coefficient of the test data set The value of the skewness coefficient determines the observed value of the empirical copula ,like ,but on the contrary .

5. The anomaly detection method based on GCN-LSTM and attention mechanism as claimed in claim 4, characterized in that: Calculate the probability of the left tail, right tail, and skewness coefficient of each time snapshot in the test data set. Specifically, calculate the left side of each time snapshot in the test data set. , right side, and tail probabilities of skewness: in, is the left tail probability of the i-th time snapshot, is the right tail probability of the i-th time snapshot, is the tail probability of the skewness at the i-th time snapshot, is the left tail empirical copula observation at the j-th i-th time snapshot, is the right tail empirical copula observation at the j-th i-th time snapshot, is the empirical copula observation of the right tail skewness coefficient at the j-th i-th time snapshot.

6. The anomaly detection method based on GCN-LSTM and attention mechanism as claimed in claim 5, characterized in that: According to the probability of the left tail, right tail and skewness coefficient of each time snapshot in the test data set, the anomaly score and anomaly threshold corresponding to each time snapshot are calculated, and the anomaly score corresponding to each time snapshot is compared with the anomaly threshold to obtain the test set anomaly sequence data detection result. The specific steps are: E1. Get each time snapshot ( ) corresponding to the anomaly score : Among them, the max function returns the largest data in a set of data; E2. Calculate the abnormal threshold, specifically: in, is a time snapshot ( ), is the overall abnormality rate of the historical performance indicator sequence, Function calculation analysis percentage value points; E3. Compare the anomaly score corresponding to each time snapshot with the anomaly threshold to determine whether the data under each time snapshot has an anomaly: like When , the current time snapshot is abnormal; like When , the current time snapshot is normal; E4. Arrange the abnormal results of the data at each time snapshot to obtain the abnormal sequence data results of the test set.

7. An anomaly detection system based on GCN-LSTM and attention mechanism, used to execute the anomaly detection method according to any one of claims 1 to 6, characterized in that: The invention comprises a reconstruction model construction module, a model training module, a sequence reconstruction model, a data processing module, a detection model construction module, a Copula training module and an anomaly detection model; the reconstruction model construction module is used to construct a GCN-LSTM model combining a GCN network and an LSTM network, and to construct a sequence reconstruction model based on the GCN-LSTM model; the model training module is used to train and test the sequence reconstruction model; the sequence reconstruction model is used to extract the features of a data sequence to be detected and obtain a weighted vector of the features of the data sequence to be detected, and to generate a reconstructed sequence by combining the features with the weighted vector through a decoder; the data processing module is used to make a difference between the reconstructed sequence and the data sequence to be detected, to construct an error sequence, and to divide the error sequence into two parts, a training data set and a test data set; the detection model construction module is used to construct an anomaly detection model based on a Copula function; the Copula training module is used to input the training data set into the anomaly detection model for training; the anomaly detection model is used to input the test data set into the trained anomaly detection model for anomaly detection to obtain a detection result.

Citation Information

Patent Citations

  • Mobile cluster trajectory prediction method based on space attention network

    CN114297529A

  • Method and system for identification of cerebrovascular abnormalities

    EP3593722A1