KPIs anomaly detection method based on improved Transform
By improving the residual network connection of the Transformer model and introducing contrastive learning, the problems of low accuracy and poor noise robustness in KPIs anomaly detection are solved, achieving higher detection accuracy and stability.
Patent Information
- Application Number
- CN202510644105.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing Transformer-based KPIs anomaly detection methods suffer from low accuracy and poor noise robustness when processing complex KPIs data. In particular, traditional residual network connections fail to effectively distinguish the importance of each layer, and the complex control of gradient perturbations in adversarial training leads to unstable training.
The residual network connection with a gating mechanism is used to replace the traditional residual network connection. Combined with the idea of contrastive learning, the linear gating mechanism and contrastive loss are introduced to train the Transformer model to enhance the model's robustness to noise, and anomalies are identified through the SPOT algorithm.
It improves the accuracy and noise robustness of KPIs anomaly detection, significantly improves the detection effect in high-noise environments, and achieves higher F1 scores and more stable anomaly detection performance.
Smart Images

Figure CN120744735A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network system operation and maintenance technology, and specifically to a KPIs anomaly detection method based on an improved Transformer. Background Art
[0002] With the development of internet technology, enterprises are increasingly dependent on IT systems, while also facing challenges in system operations and maintenance. Traditional manual operations and maintenance methods are unable to cope with massive amounts of monitoring data and complex business scenarios, making artificial intelligence for IT operations (AIOps) an inevitable trend. As a core component of intelligent operations and maintenance systems, KPI anomaly detection, with its accurate detection results, plays a vital role in maintaining the stability and reliability of IT systems. In the KPI-based intelligent operations and maintenance process, the operations and maintenance team continuously collects key performance indicators of service components and infrastructure through a multi-dimensional monitoring mechanism, and uses intelligent analysis models to detect anomalies in KPI time series data, further providing early warning and troubleshooting of system failures. This operation and maintenance approach significantly improves the speed and accuracy of identifying system anomalies, enabling operations and maintenance personnel to more quickly and accurately locate the source of problems, reducing the risk of business interruptions caused by system failures.
[0003] Currently, deep learning-based KPI anomaly detection methods have become the mainstream approach for KPI anomaly detection due to their ability to maintain high accuracy and efficiency when dealing with large-scale, high-dimensional KPI data. Deep learning models such as LSTM (Long Short-Term Memory), VAE (Variational Autoencoder), and Transformer are widely used in this category. For example, researchers have proposed the C-LSTM method, which uses convolutional kernels of various sizes to extract local features from KPI time series and combines it with an LSTM for anomaly detection. However, LSTM is sensitive to noise and lacks robustness when processing noisy data, resulting in low anomaly detection accuracy. With the popularity of VAE models, researchers have begun applying them to KPI anomaly detection. For example, researchers have proposed the MST-VAE method, which extracts temporal features of KPIs by combining temporal convolution kernels of varying granularity. Although VAE models exhibit strong robustness when processing noisy KPI data, they do improve detection accuracy to a certain extent. However, using VAEs to process complex KPIs data can easily lead to bias, resulting in distortion in the reconstructed KPI data, which in turn affects anomaly detection results. Subsequently, due to the significant advantages of the Transformer model in natural language processing tasks, numerous researchers have conducted in-depth research on its application in KPI anomaly detection. For example, some researchers have proposed targeted improvements to the Transformer model, enhancing the model's ability to capture local temporal patterns in KPIs by introducing a local attention mechanism. Others have proposed the Informer model, which, by optimizing its attention mechanism, effectively addresses the efficiency and memory limitations faced by traditional Transformers when processing long KPI sequences. However, these methods often struggle to distinguish between normal data and noise when using the Transformer model for anomaly detection, resulting in poor robustness to noise. To enhance the model's robustness to noise, researchers have further combined the Transformer model with adversarial training. For example, some researchers have proposed the TranAD model, which combines adversarial learning with self-regulation mechanisms through a dynamic focus score mechanism. The self-adjustment mechanism utilizes dynamic focus scores to identify and focus on key features in the data, while adversarial training introduces adversarial samples, enabling the model to accurately capture abnormal patterns in KPIs despite noise interference. However, the complex control of gradient perturbations in adversarial training can lead to training instability. Furthermore, the generator in the adversarial network tends to generate highly similar samples during training, failing to fully capture the diversity of real data. These factors can lead to reduced detection accuracy. Summary of the Invention
[0004] Based on the above-mentioned shortcomings of the current mainstream KPIs anomaly detection method in terms of detection accuracy and noise robustness. The present invention proposes a KPIs anomaly detection method based on an improved Transformer. First, LSTM, FFT and GAT are used to extract the time, frequency domain and spatial features of KPIs, and the attention mechanism is used for feature fusion. Then, the improved Transformer model is used to reconstruct the time-frequency-spatial features of KPIs. The improved Transformer model adopts a residual network connection with a gating mechanism to improve the accuracy of anomaly detection, and at the same time introduces the idea of contrastive learning to enhance the robustness of the model to noise. Finally, the SPOT algorithm is used to identify abnormal KPIs, thereby improving the anomaly detection effect of KPIs.
[0005] The technical means adopted in the present invention are as follows:
[0006] A KPIs anomaly detection method based on an improved Transformer includes the following steps:
[0007] S1. Obtain a KPIs time series and preprocess the KPIs time series;
[0008] S2. Extract the time domain features of the KPIs time series through the long short-term memory network; process the time domain features of the KPIs time series through the fast Fourier transform to extract the frequency domain features of the KPIs time series;
[0009] S3. constructing a KPIs relationship graph based on the KPIs time series, calculating a graph adjacency matrix based on the relationships between the KPI nodes in the KPIs relationship graph, and performing a sparse processing on the graph adjacency matrix to obtain a sparse graph adjacency matrix;
[0010] S4. Processing the KPIs time series and the sparse graph adjacency matrix based on a graph attention network to extract spatial features of the KPIs time series;
[0011] S5. Based on the attention mechanism, the time domain features of the KPIs time series, the frequency domain features of the KPIs time series, and the spatial features of the KPIs time series are fused to obtain fused features;
[0012] S6. Reconstructing the fused features based on an improved Transformer model and calculating the reconstruction error of the KPIs time series, wherein the improved Transformer model is configured to adopt a residual network connection based on linear gating;
[0013] S7. Use the SPOT algorithm to process the reconstruction error of the KPIs time series, generate the anomaly score of the KPIs time series in each dimension, and output the detection results based on the anomaly score.
[0014] Furthermore, the KPIs time series is preprocessed, including:
[0015] The input data is normalized and preprocessed through translation and scaling operations.
[0016] Furthermore, the calculation formula of the graph adjacency matrix is:
[0017] A′=ReLU(tanh(CI))
[0018] Where A′ represents the graph adjacency matrix, C represents the node similarity matrix of each KPI in the KPIs relationship graph, ReLU and tanh both represent activation functions, and I represents an identity matrix.
[0019] Furthermore, the calculation formula of the gating weight of the linear gating is:
[0020] G=σ(W g ·[X]+b g )
[0021] Among them, G is the gate weight, X is the KPIs time series, W g is the weight matrix of the fully connected layer, b g is the bias term of the fully connected layer.
[0022] Furthermore, when training the improved Transformer model, the contrast loss and the reconstruction loss are jointly trained.
[0023] Furthermore, the improved Transformer model adopts two networks with the same structure, and the two networks share parameters.
[0024] Furthermore, the calculation formula of the abnormality score is:
[0025]
[0026] Among them, S is the abnormality score, μ e ,σ e are the mean and standard deviation of the KPIs time series reconstruction error, Loss recon Reconstruct error for KPIs time series.
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] This application improves the Transformer model from two aspects: improving the accuracy and noise robustness of KPIs anomaly detection. On the one hand, the residual network connection with a gating mechanism is used to replace the residual network connection of the original Transformer, so that the model can learn the weight information of each network layer according to the importance of different network layers, thereby improving the accuracy of anomaly detection. On the other hand, the contrastive learning idea is introduced into the Transformer model, and the contrastive loss and reconstruction error loss are used to jointly train the Transformer model, making the model more robust.
[0029] This application conducted experimental comparisons on different KPIs datasets. The comparison results with the current mainstream methods show that the KPIs anomaly detection effect of the method of the present invention has obvious advantages. Ablation experiments show that compared with the residual network connection of the original Transformer without a gating mechanism, the model of the present invention achieves the highest anomaly detection accuracy on both PSM and ASD datasets. Noise robustness experiments show that the average F1 score of the model of the present invention on the PSM contaminated training set with a noise ratio of up to 20% shows only a slight decrease. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0031] Figure 1 This is a flow chart of the KPIs anomaly detection method based on the improved Transformer of the present invention.
[0032] Figure 2 This is a structural diagram of the encoder with a gating mechanism of the present invention.
[0033] Figure 3 This is the structural diagram of the time-frequency feature comparison learning of the present invention.
[0034] Figure 4 This is a structural diagram of the KPIs reconstruction model based on the improved Transformer in the present invention. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0036] like Figure 1 As shown, the present invention provides 1. a KPIs anomaly detection method based on an improved Transformer, comprising the following steps:
[0037] S1. Obtain a KPIs time series and preprocess the KPIs time series.
[0038] In order to improve the effect of model training, it is necessary to preprocess the KPIs time series data to eliminate dimensional differences. Specifically, for the KPIs input sequence X = {x1, x2, ..., x t ,…,x T}, the input data is normalized and preprocessed by translation and scaling operations. First, calculate μ x The process is shown in formula (1).
[0039]
[0040] Then, calculate the variance The process is shown in formula (2).
[0041]
[0042] Finally, the process of normalizing the input data is shown in formula (3).
[0043]
[0044] Where ⊙ is the element-wise product, σ x is the standard deviation, μ x is the mean.
[0045] S2. Extract the time domain features of the KPIs time series through the long short-term memory network; process the time domain features of the KPIs time series through the fast Fourier transform to extract the frequency domain features of the KPIs time series.
[0046] This step mainly extracts KPIs time-frequency features based on LSTM and FFT.
[0047] LSTM can effectively capture long-term dependencies in KPIs time series through forget gates, input gates, and output gates. This paper uses LSTM to extract the temporal features H of KPIs to capture these dependencies. The process is shown in formula (4).
[0048] H=LSTM(X) (4)
[0049] LSTM focuses on modeling the dependencies of KPIs in the time domain, but periodic and seasonal patterns are more pronounced in the frequency domain, typically manifesting as concentrated energy at specific frequencies. Therefore, LSTM's ability to capture these patterns is limited. FFT (Fast Fourier Transform) is a classic method for converting KPIs from the time domain to the frequency domain. It can decompose KPIs into sine and cosine waves of different frequencies and has the advantages of strong global frequency information extraction capabilities, high computational efficiency, and ease of implementation and interpretation. This paper uses FFT to transform the time domain features H obtained by LSTM to obtain the frequency domain features F of KPIs to capture the seasonality and periodicity of KPIs. The process is shown in Formula (5).
[0050] F=FFT(H) (5)
[0051] S3. Construct a KPIs relationship graph based on the KPIs time series, calculate a graph adjacency matrix based on the relationship between each KPI node in the KPIs relationship graph, and perform sparsification processing on the graph adjacency matrix to obtain a sparse graph adjacency matrix.
[0052] KPIs are multidimensional time series. Changes in some indicators will cause changes in other indicators within the same time or time interval. To better capture the connections between different indicators, it is necessary to learn a graph adjacency matrix A to describe the associations between various KPIs. This paper uses graph structure learning methods to construct a sparse adjacency matrix for KPIs. The relevant steps are as follows:
[0053] 1) Randomly initialize a d-dimensional embedding vector v for each KPI node i , and its calculation process is shown in formula (6). The initialized vector is updated as the model is trained.
[0054] v i =R d ,i∈{1,2,…,N} (6)
[0055] Among them, R d is a d-dimensional real vector space, and N is the dimension of KPIs.
[0056] 2) Calculate the cosine similarity C between the embedding vectors of each KPI node ij, and its calculation process is shown in formula (7).
[0057]
[0058] in, Represents vector v i and vector v j The dot product of ‖v i ‖ and ‖v j ‖ respectively represent vector v i and v j The Euclidean norm of .
[0059] 3) Describe the relationship between each KPI node through calculation. The calculation process is shown in formula (8).
[0060] A′=ReLU(tanh(CI)) (8)
[0061] 4) A sparse graph adjacency matrix A is obtained by retaining the highest K related neighbors of each KPI node in the graph adjacency matrix A′. The calculation process is shown in formula (9).
[0062] A ij =1{j∈topK(A i: )} (9)
[0063] Among them, topK(A i: ) is a set of indices representing the largest K numbers in the i-th row of the graph adjacency matrix A. 1{·} is an indicator function. If the expression is true, the function value is equal to 1, otherwise it is equal to 0. It is used to determine whether j belongs to the index set of the largest K numbers in the i-th row. A ij It is the element located in the i-th row and j-th column in the graph adjacency matrix A, indicating the association relationship between KPI node i and KPI node j. If the value is 1, it means there is a relationship between them, otherwise there is no relationship.
[0064] S4. Process the KPIs time series and the sparse graph adjacency matrix based on the graph attention network to extract the spatial features of the KPIs time series.
[0065] GAT (Graph Attention Network) is a neural network that uses an attention mechanism to learn graph-structured data. It extracts the feature vectors of each node in the graph by analyzing the interactions between nodes. Based on the KPIs relationship graph A obtained by S3, this paper uses GAT to extract the spatial features G of the KPIs. The process is shown in Equation (10).
[0066] G=GAT(X,A) (10)
[0067] S5. Based on the attention mechanism, the time domain features of the KPIs time series, the frequency domain features of the KPIs time series, and the spatial features of the KPIs time series are fused to obtain fused features.
[0068] In order to make the subsequent anomaly detection model pay more attention to important feature information and highlight the importance of different features to the anomaly detection results, the present invention adopts the attention mechanism to fuse the time-frequency-space features of KPIs to obtain O. The process of calculating the attention weight factor is shown in formula (11).
[0069] a=tanh(G·W g +H·W h +F·W f +b) (11)
[0070] Where a is the weight factor and b is the bias term that adjusts the calculation of the attention score. g , W h and W f are the weight matrices for the KPIs' spatial, temporal, and frequency-domain feature vectors, respectively. These three are randomly initialized and updated during model training. Combining the weight factors with the KPIs' temporal, frequency, and spatial features, H, F, and G, yields the fused feature O = a·(G+H+F).
[0071] S6. Reconstruct the fused features based on the improved Transformer model and calculate the reconstruction error of the KPIs time series. The improved Transformer model is set to adopt a residual network connection based on linear gating and introduce contrastive learning ideas to improve the robustness of the model.
[0072] Due to the global modeling capability of the self-attention mechanism in the Transformer model, it can effectively capture the complex patterns and long-distance dependencies in the KPIs time series. In addition, the Transformer model avoids the loss of the temporal order information of the KPIs sequence during the processing process by introducing positional encoding technology. In view of these advantages of the Transformer model in processing KPIs time series, the present invention uses the Transformer model to reconstruct KPIs. However, this model still has certain limitations, which are mainly manifested in the following two aspects:
[0073] 1) The Transformer's residual network connections fail to adequately distinguish the importance of each layer, resulting in a lack of effective attention to key information when processing complex KPI time series. The Transformer model consists of a series of encoders and decoders, each of which incorporates a large number of residual network connections. The design of these connections is inspired by residual networks. The introduction of skip connections allows gradients to bypass certain layers during propagation, effectively overcoming the vanishing gradient problem that arises with increasing network layers and ensuring efficient parameter updates in deep networks. However, the Transformer model's residual connections simply sum the outputs of different layers through direct addition, implicitly assuming that each layer contributes equally to the final result. In the task of KPI anomaly detection, different network layers exhibit significant differences in how they process KPI features. The traditional Transformer's residual network connection method fails to account for the differences in the importance of information transfer between layers, potentially preventing the model from effectively identifying and utilizing important information provided by key layers. Therefore, in the task of KPIs anomaly detection, the traditional residual network is often not sophisticated enough in processing complex KPIs information, which easily leads to the loss of key information, resulting in poor anomaly detection effect.
[0074] 2) Transformer models lack robustness in handling KPI noise. Whether used as a prediction model or a reconstruction model, the core goal of the Transformer model is to learn underlying features and patterns from normal KPI data. The model calculates the loss by comparing the error between the reconstructed data and the original input data, and trains accordingly to improve the model's reconstruction capabilities. In the real world, KPI sequences are commonly subject to various factors, including data acquisition errors, internal system failures, and random fluctuations in the external environment. This leads to noise and perturbations in KPI data sequences. This noise is unavoidable during the data acquisition process and, in turn, affects the quality of KPI data. During training, the Transformer model cannot accurately distinguish this noise from true data features and may mistakenly learn from the noise as valuable information. As a result, when the model reconstructs or predicts data, it not only reproduces normal data patterns but also inappropriately reconstructs the noise, making the model overly sensitive to small changes in the input data and susceptible to noise. This sensitivity to noise reduces the robustness of the Transformer model, resulting in reduced accuracy in anomaly detection when the model encounters new, unseen noise or data that is slightly different from the training data distribution.
[0075] The present invention addresses the deficiencies of the Transformer model and proposes a corresponding improvement method.
[0076] 1) Residual network connection based on linear gating
[0077] The present invention adopts a linear gating mechanism to replace the traditional residual connection in the original Transformer model to obtain an encoder structure with a gating mechanism, such as Figure 2 Compared to other complex gating strategies, this mechanism significantly reduces the number of parameters required during training. More importantly, it dynamically controls the flow of information between different network layers. By assigning different weights to each layer, it can flexibly adjust their relative contributions to the overall model, further improving model performance.
[0078] The core of the linear gating mechanism is a gating weight, calculated using a simple fully connected layer network, that determines the strength and proportion of information transfer. To more efficiently weight the input information, this layer first performs a linear transformation on the input data, which better integrates and transfers the information. This linear gating mechanism not only optimizes the information fusion process but also improves the model's ability to handle complex time series data, thereby improving overall model performance and efficiency.
[0079] The input data X is transformed linearly and activated to generate the gate weight G. The calculation process is shown in formula (12).
[0080] G=σ(W g ·[X]+b g ) (12)
[0081] Among them, W g is the weight matrix of the fully connected layer, b g is the bias term of the fully connected layer.
[0082] By introducing a gating mechanism, we can achieve effective fusion of input data from different network layers. The following takes the fusion of feature information from two layers as an example, and the process is shown in formula (13).
[0083] y=G1⊙X1+G2⊙X2 (13)
[0084] Among them, G1 and G2 correspond to the gating weights of the input data X1 and X2 respectively, ⊙ represents element-by-element multiplication, and y is the final fusion result.
[0085] The linear gating mechanism is integrated into the residual network of Transformer to obtain the encoder structure with gating mechanism. The process is shown in formulas (14) and (15).
[0086] H attn =LayerNorm(G attn1 ⊙X+G attn2 ⊙Mulihead(X)) (14)
[0087] Among them, G attn1 and G attn2 are the linear gating weights of the self-attention layer, which respectively control the contribution of the input information X and the output information Mulihead(X) of the self-attention layer to the residual connection, H attn It is the output after layer normalization and residual connection, which will serve as the input of the next network layer.
[0088] X out =LayerNorm(G ffn1 ⊙H attn +G ffn2 ⊙FFN(H attn )) (15)
[0089] Among them, G ffn1 and G ffn2 are the linear gating weights of the feedforward network layer, which control the output H of the self-attention layer respectively. attn and the output of the feedforward network FFN(H attn ) contribution to the residual connection, X out It is the final output after layer normalization and residual connection, which will be passed to the next encoder or used as the output of the entire encoder.
[0090] 2) Introducing contrastive learning to improve the robustness of the model
[0091] The present invention introduces the idea of contrastive learning in the process of reconstructing the Transformer model. By jointly training contrastive loss and reconstruction loss, the model is guided to have a deeper understanding of the KPIs data features during training. Contrastive learning constructs positive and negative sample pairs. On the one hand, it reduces the distance between the two samples in the positive sample pair in the feature space, and on the other hand, it widens the feature difference between the two samples in the negative sample pair. This design aims to make the model pay more attention to the KPIs data features, suppress the interference of noise on the model performance, and thus significantly enhance the robustness and generalization ability. This dual loss function training strategy ensures that the model can maintain high reconstruction accuracy while maintaining stability in complex and changeable practical application scenarios.
[0092] The present invention inputs the time domain features and frequency domain features of KPIs into the Transformer model for comparative learning. Specifically, by performing positive and negative comparisons on the time domain features and frequency domain features of the same time step, the model is strengthened to learn the internal consistency of the data at the same moment. Since the KPIs training set only contains normal data, the present invention uses the time domain features and frequency domain features of different time steps to construct negative sample pairs. The KPIs time-frequency feature comparative learning structure is as follows: Figure 3 shown.
[0093] In order to accurately quantify the similarities and differences between positive and negative sample pairs, this paper adopts the normalized temperature scaled cross entropy function as the contrast loss function. The purpose of this loss function is to reduce the distance between the time-frequency features of KPIs at the same time step, while increasing the distance between the time-frequency features of KPIs at different time steps. Through this time-frequency feature contrast learning method, the Transformer model can focus more on the intrinsic characteristics of KPIs data during training, while ignoring the interference of potential noise and irrelevant variables. con The calculation process of is shown in formula (16).
[0094]
[0095] Among them, h i ′ Indicates the KPIs time feature h at timestamp i i The output result after model processing, f i ′ Represents the frequency domain features f of KPIs at timestamp i i The output result after model processing, f j ′ Represents the frequency domain features f of KPIs at timestamp j j The output result after model processing. sim(·) is the similarity function used to measure the time feature vector h ′ and frequency domain eigenvector f ′ τ is the temperature parameter, which is a hyperparameter used to adjust the impact of the similarity score.
[0096] The KPIs reconstruction model structure based on the improved Transformer is as follows Figure 4 shown.
[0097] In the training phase of the model, the input is the original sequence of KPIs X = {x1, x2, ..., x t ,…,x T}, the time feature H and frequency domain feature F obtained by S2, and the fused feature O obtained by S4. In order to optimize the training process of the model, reduce the number of parameters to reduce overfitting and improve training efficiency, the present invention adopts a Transformer architecture that only contains an encoder to encode the input KPIs data. In order to ensure that the reconstructed KPIs sequence maintains the same dimension as the original sequence, a linear projection layer is introduced. The function of this layer is to convert the high-dimensional features extracted by the encoder back to the dimensional space of the original KPIs sequence. Such a design not only simplifies the model structure, but also helps the model focus more on learning the key features of the sequence to achieve accurate reconstruction and subsequent anomaly detection tasks.
[0098] In order to enable the model to better understand and distinguish the difference between the two samples in the positive and negative sample pairs, the present invention adopts two networks with the same structure, one network is used to input time domain features, and the other network is used to input frequency domain features. The time domain features and frequency domain features of the same time step are used as positive sample pairs, and the time domain features and frequency domain features of different time steps are used as negative sample pairs. The difference between the positive and negative sample pairs is distinguished by increasing the similarity between the two samples in the positive sample pair and reducing the similarity between the two samples in the negative sample pair in the feature space. The two networks share a set of parameters, which means that their weights are updated synchronously, thereby ensuring the synchronization and consistency of the two networks during the learning process. The present invention takes the input of the fused KPIs time-frequency-space feature O into a Transformer network as an example to explain in detail its processing flow in the Transformer model. First, the fused time-frequency-space feature O is preliminarily vectorized by the embedding layer to convert the original features into a form that the network can process. Subsequently, these embedded features are encoded by position encoding to ensure that the time order information in the sequence is retained. After completing the position encoding, the features are input to the encoder of the Transformer model. The encoder uses its self-attention mechanism to learn the features in depth. Afterwards, a linear projection layer is used to map the high-dimensional features extracted by the encoder back to the dimensional space of the original KPIs sequence, thus completing the reconstruction of O and obtaining O. ′ 1. The time feature H and frequency domain feature F of KPIs also follow the same processing steps to obtain H ′ and F ′ The relevant formulas are shown in (17)-(20).
[0099] H ′ =Projection1(Encoder1(Embedding1(H)+PE)) (17)
[0100] O ′1=Projection1(Encoder1(Embedding1(O)+PE)) (18)
[0101] F ′ =Projection2(Encoder2(Embedding2(F)+PE)) (19)
[0102] O ′ 2=Projection2(Encoder2(Embedding2(O)+PE)) (20)
[0103] Next, by obtaining O ′ 1 and O ′ 2 and the initial input data X to reconstruct the error Loss recon The calculation of is shown in formula (21).
[0104]
[0105] According to the obtained H ′ and F ′ , calculate the contrast loss Loss by formula (16) con , and then a linear weighted method is used to combine the contrast loss and the reconstruction error loss to construct a total loss function. By incorporating the contrast loss into the reconstruction model, the model's noise resistance is enhanced, thereby training a reconstruction model with excellent robustness. Such a model will be able to more effectively reconstruct data in the test phase to cope with various noise and interference. The total loss function formula is shown in (22).
[0106] Loss = αLoss con +βLoss recon (twenty two)
[0107] Among them, Loss con is the contrast loss calculated according to formula (16), Loss recon is the reconstruction error loss calculated according to formula (21), α and β are weight coefficients used to control the weights of contrast loss and reconstruction error loss.
[0108] The KPIs reconstruction model training algorithm based on the improved Transformer is shown in Algorithm 1.
[0109]
[0110] S7. Use the SPOT algorithm to process the reconstruction error of the KPIs time series, generate the anomaly score of the KPIs time series in each dimension, and output the detection results based on the anomaly score.
[0111] SPOT is a streaming data anomaly detection algorithm based on extreme value theory. It can automatically adjust the threshold according to the change of data without manual setting. Therefore, the present invention uses the SPOT algorithm to detect anomalies in KPIs. After calculating the reconstruction error of the KPIs time series, these errors are standardized and converted into anomaly scores to facilitate more accurate evaluation and identification of potential anomalies. The process of calculating the anomaly score is shown in formula (23), through which the anomaly score of each KPIs dimension i at a specific timestamp can be obtained. By comparing the anomaly score of each dimension with the threshold calculated by SPOT, it is determined whether the KPI of each dimension is abnormal. The process is shown in formula (24).
[0112]
[0113] Among them, μ e ,σ e are the mean and standard deviation of KPIs data reconstruction error respectively.
[0114] y i =1(s i ≥SPOT(s i )) (twenty four)
[0115] Among them, s i is the abnormal score of KPI in the i-th dimension, y i Indicates whether the current timestamp is marked as abnormal in the i-th dimension. i ≥SPOT(s i ), then y i is 1, indicating that an abnormality is detected; otherwise, y i A value of 0 indicates that no abnormality is detected.
[0116] The scheme and effects of the present invention are further described below based on specific application examples.
[0117] This paper uses two public datasets: the Pooled Server Metrics (PSM) dataset, a public dataset from eBay with 25 KPI indicators. The other is the Application Server Dataset (ASD), a multidimensional KPI dataset collected by Tsinghua Netman Studio from a large internet company and publicly released on Github. It contains 19 KPI indicators. Both datasets include CPU-related and memory-related metrics, reflecting the overall performance of the server. Details of the datasets are shown in Table 1.
[0118] Table 1. Details of the dataset
[0119]
[0120] (2) Comparison method
[0121] To verify the effectiveness of the method of the present invention, a comparative analysis was conducted with some popular methods in recent years. These comparative methods are all based on currently commonly used deep learning architectures, such as VAE and Transformer, which are described in detail below.
[0122] Anomaly-Transformer: This method is an innovative anomaly detection model for time series data. It cleverly combines the powerful representation capabilities of the Transformer with a novel anomaly attention mechanism. By leveraging the concepts of global sequence correlation and local prior correlation, it can capture the complex dynamics of KPI time series and detect anomalies in KPIs using a minimax strategy.
[0123] GRELEN: This method is a VAE-based anomaly detection method. It uses VAE as the overall framework for feature extraction and system representation, fully considering the dependencies between different indicators in KPI anomaly detection. GRELEN improves anomaly detection accuracy by converting the dependencies between indicators into a graph and using a graph relationship learning network to refine the latent variables in the VAE architecture. This method performs well when processing complex, multi-dimensional time series data.
[0124] Dcdetector: This method is an advanced multi-scale dual-attention contrastive learning model designed for anomaly detection in time series data. It extracts differentiated features from the same anomaly point through two independent channels. It then compares the similarities between these extracted features to score and identify anomalies at each point in the time series.
[0125] MUTANT: This method considers the correlation between variables in each time period and uses a graph convolutional network to learn embeddings between all variables, which effectively captures the spatiotemporal phase characteristics between variables in multidimensional time series. Then, an attention-based reconstruction model is proposed to capture the normal patterns of KPIs.
[0126] (3) Evaluation indicators
[0127] KPIs anomaly detection is regarded as a classification problem, and the performance of the anomaly detection method is evaluated using precision, recall, and F1-score.
[0128] Precision refers to the proportion of time points predicted to be outliers that are actually outliers. A high precision means a low false positive rate, so it is also called the precision rate. Its calculation formula is shown in formula (25).
[0129]
[0130] TP stands for True Positive, which is the number of time points that the model correctly predicts as anomalies. FP stands for False Positive, which is the number of time points that the model incorrectly predicts as anomalies (i.e., false positives).
[0131] Recall rate refers to the proportion of all actual anomalies that are correctly predicted by the model as anomalies. A high recall rate means a low missed detection rate, so it is also called the recall rate. Its calculation formula is shown in formula (26).
[0132]
[0133] Where TP is the True Positive, which is the number of time points that the model correctly predicts as abnormal. FN is the False Negative, which is the number of time points that the model incorrectly predicts as normal (i.e., missed reports).
[0134] F1-score is the weighted harmonic mean of Precision and Recall, which is suitable for evaluating the comprehensive performance of the algorithm. Its calculation formula is shown in formula (27).
[0135]
[0136] (4) Experimental comparative analysis
[0137] The experimental environment of the present invention uses PyCharm 2020, Python version 3.8.16, CUDA version 11.1, and Pytorch version 1.5.1 in terms of software. In terms of hardware configuration, an Intel Core i5-8300H CPU is used. The model is optimized using the Adam optimizer, using the cross-entropy loss function, setting the initial learning rate to 0.001, and setting the number of training iterations to 30.
[0138] This paper conducts comparative experiments in three aspects. The first part is a conventional experimental comparative analysis, comparing the proposed method with current mainstream methods. The second part is an ablation experiment to verify the effectiveness of each component of the proposed model. The third part is a robustness experiment to further verify the effectiveness of the contrastive learning proposed in this paper.
[0139] 1) Comparative experiment
[0140] In this paper, the proposed KPIs sequence anomaly detection method (Ours) is experimentally compared with four comparison methods to verify the effectiveness of the proposed method. Tables 2 and 3 show the comparison results of Ours and other baseline methods on the PSM and ASD datasets, respectively.
[0141] Table 2 Comparative experimental results of KPIs anomaly detection on the PSM dataset
[0142]
[0143] Table 3 Comparative experimental results of KPIs anomaly detection on the ASD dataset
[0144]
[0145] Comparative experimental results show that the F1 scores of Ours method in the present invention reach 0.9798 and 0.9715 on the PSM dataset and ASD dataset, respectively.
[0146] Anomaly-Transformer performs better on the PSM dataset than on the ASD dataset. This is because the temporal patterns contained in the PSM dataset are more regular and consistent, which is consistent with the global sequence correlation characteristics of the Anomaly-Transformer. This method can more accurately identify anomalies on the PSM dataset and reduce the missed detection rate. Therefore, the recall rate on the PSM dataset is slightly higher than that of the method of the present invention. The ASD dataset may contain more noise and non-stationarity. These factors may interfere with the Anomaly-Transformer model's recognition of anomalies and may affect the accuracy of anomaly detection. This method uses Transformer for anomaly detection. Although it uses an improved attention mechanism, it ignores the problem of its internal residual network connection. Therefore, the final anomaly detection result is lower than the method proposed in the present invention.
[0147] The GRELEN method, by adopting a variational autoencoder structure, deeply considers the interactions between different indicators and optimizes the model's performance by fine-tuning the VAE's latent variables. This adjustment not only enhances the model's ability to represent data features, but also improves the model's accuracy in identifying anomalous instances. The VAE's reconstruction error is used as the basis for anomaly detection, and in theory, it can effectively identify anomalies that are inconsistent with the normal data distribution. Although the GRELEN method has improved accuracy, the long-term dependencies in KPIs time series data often involve complex patterns and associations spanning long time intervals. The VAE structure may have difficulty capturing these long-distance dependencies. In addition, the use of the VAE model will lead to a mismatch between the posterior distribution and the prior distribution of the latent variables when reconstructing the KPIs, which in turn affects the final anomaly detection results.
[0148] Dcdetector outperforms GRELEN on both datasets due to its innovative dual-branch attention mechanism, which is based on the principle of contrastive learning and can effectively enhance the discrimination between normal and abnormal samples. Since the PSM dataset contains more complex time series patterns, Dcdetector's dual-branch structure can focus on both global and local features, which adapts to this complexity and reduces the missed detection rate, thus performing slightly better than the method proposed in this invention in terms of recall rate. However, the extraction of time dependencies in KPIs data is not sufficient, and the dynamic changes in the extracted KPIs time series data are not sensitive enough. Therefore, the final anomaly detection result is lower than the method proposed in this invention.
[0149] MUTANT's performance on the dataset was relatively stable. This is primarily because it effectively captures the spatiotemporal correlations in multidimensional time series by considering the correlations between variables at different time points and using a graph convolutional network to learn embedding representations between variables. However, reconstruction-based models can overfit noisy KPI data, resulting in poor robustness. Consequently, the final anomaly detection performance was suboptimal.
[0150] The experimental results and analysis above demonstrate that the proposed method is not only more stable but also achieves superior detection results. This is primarily due to the adoption of the Transformer model for anomaly detection and improvements to the model's residual network structure and robustness, effectively overcoming potential model challenges. In summary, the proposed KPI anomaly detection method is stable and highly effective.
[0151] 2) Ablation experiment
[0152] To verify the effectiveness of each component of the model, we conducted further ablation experiments. We replaced the residual network connections using the gating mechanism in our invention with the residual connections of the original Transformer, naming this variant model No-Gate. We then removed the contrastive loss from the loss function and used only the reconstruction error loss, naming this variant No-Contrastive Learning. Table 4 shows the ablation results on the PSM dataset, and Table 5 shows the ablation results on the ASD dataset.
[0153] Table 4 Results of Ours and variant models on the PSM dataset
[0154]
[0155] Table 5 Results of Ours and variant models on the ASD dataset
[0156]
[0157]
[0158] In the "No-Gate" approach, the present invention uses the original residual connections in the Transformer model without a gating mechanism. This connection method is unable to distinguish the importance of information at different levels, which may result in the model being unable to effectively filter out the information layers that are more critical for anomaly detection when processing complex KPIs data. As a result, this Transformer model using original residual connections will ignore certain key information when processing KPIs data, thereby reducing the accuracy of anomaly detection.
[0159] In the "No-Contrastive Learning" experiment, we observed that when a model trained solely on reconstruction error performs anomaly detection during the test phase, the final anomaly detection results show a certain degree of degradation. This approach, which uses only reconstruction error training, indiscriminately reconstructs noise information when reconstructing KPI data, weakening the model's ability to resist noise interference. Consequently, its anomaly detection performance is reduced when detecting KPIs containing noise. These experimental results further emphasize the importance of contrastive learning in improving a model's resistance to noise interference.
[0160] 3) Robustness experiment
[0161] In order to further explore the key role of contrastive learning in improving the model's ability to resist noise, when conducting an in-depth analysis of the ablation experiment in the previous section, the changes in F1-score were observed by removing the contrast loss. In addition, the present invention selected the PSM dataset as the research object, and by mixing a specific proportion of noise data into it, it highly simulated the complex state presented by the data in the real environment, as an effective means to test the model's ability to resist noise. With this method, the performance of the model in the face of noise interference can be more accurately evaluated, and the intrinsic relationship between contrastive learning and the model's ability to resist noise can be further explored.
[0162] Table 6 details the experimental results of the proposed method using a PSM-contaminated training set. A careful analysis of the data clearly shows that performance does decline somewhat as the noise content in the training set increases. Even with a noise content as high as 20% in the training set, the average F1 score of the proposed method only drops by 1.67%. This small decrease fully demonstrates the method's superior performance in dealing with noise interference.
[0163] Table 6 KPIs anomaly detection results of robustness experiments on PSM dataset
[0164]
[0165]
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A KPIs anomaly detection method based on an improved Transformer, characterized by: The following steps are involved: S1. Obtain a KPIs time series and preprocess the KPIs time series; S2. Extract the time domain features of the KPIs time series through the long short-term memory network; process the time domain features of the KPIs time series through the fast Fourier transform to extract the frequency domain features of the KPIs time series; S3. constructing a KPIs relationship graph based on the KPIs time series, calculating a graph adjacency matrix based on the relationships between the KPI nodes in the KPIs relationship graph, and performing a sparse processing on the graph adjacency matrix to obtain a sparse graph adjacency matrix; S4. Processing the KPIs time series and the sparse graph adjacency matrix based on a graph attention network to extract spatial features of the KPIs time series; S5. Based on the attention mechanism, the time domain features of the KPIs time series, the frequency domain features of the KPIs time series, and the spatial features of the KPIs time series are fused to obtain fused features; S6. Reconstructing the fused features based on an improved Transformer model and calculating the reconstruction error of the KPIs time series, wherein the improved Transformer model is configured to adopt a residual network connection based on linear gating; S7. Use the SPOT algorithm to process the reconstruction error of the KPIs time series, generate the anomaly score of the KPIs time series in each dimension, and output the detection results based on the anomaly score.
2. The KPIs anomaly detection method based on the improved Transformer according to claim 1 is characterized in that: Preprocessing the KPIs time series includes: The input data is normalized and preprocessed through translation and scaling operations.
3. The KPIs anomaly detection method based on improved Transformer according to claim 1 is characterized in that: The calculation formula of the graph adjacency matrix is: A′=ReLU(tanh(CI)) Where A′ represents the graph adjacency matrix, C represents the node similarity matrix of each KPI in the KPIs relationship graph, ReLU and tanh both represent activation functions, and I represents an identity matrix.
4. The KPIs anomaly detection method based on improved Transformer according to claim 1 is characterized in that: The calculation formula of the gating weight of the linear gating is: G=σ(W g ·[X]+b g ) Among them, G is the gate weight, X is the KPIs time series, W g is the weight matrix of the fully connected layer, b g is the bias term of the fully connected layer.
5. The KPIs anomaly detection method based on improved Transformer according to claim 1 is characterized in that: When training the improved Transformer model, the contrast loss and the reconstruction loss are jointly trained.
6. The KPIs anomaly detection method based on improved Transformer according to claim 1 is characterized in that: The improved Transformer model adopts two networks with the same structure, and the two networks share parameters.
7. The KPIs anomaly detection method based on improved Transformer according to claim 1 is characterized in that: The calculation formula of the abnormality score is: Among them, S is the abnormality score, μ e ,σ e are the mean and standard deviation of the KPIs time series reconstruction error, Loss recon Reconstruct error for KPIs time series.