Log anomaly detection method based on multi-head GRU

Through the multi-head GRU network, the problem of insufficient learning of sequence patterns in large-scale log data is solved, efficient and flexible log anomaly detection is achieved, and detection accuracy and computing efficiency are improved.

CN120256179APending Publication Date: 2025-07-04ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510337262.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing log anomaly detection methods process large-scale unstructured log data, it is difficult to effectively learn sequence patterns, resulting in inaccuracy and inefficiency of detection, and traditional models lack computing resources and adaptability.

Method used

The multi-head GRU network is adopted to learn the local mode of log sequences in parallel through multiple GRU models, and integrate the basic model under the multi-head mechanism, and flexibly adjust the detection configuration to weigh accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of log exception detection, reduces the computing resource requirements, adapts to different log data scenarios, and meets real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256179A_ABST
    Figure CN120256179A_ABST
Patent Text Reader

Abstract

The invention provides a multi-head GRU-based log anomaly detection method, which comprises the following steps: S1, an offline training stage: on the basis of a multi-head mechanism, constructing a multi-head GRU as a core model, and training the multi-head GRU by using a generated log vector; s2, an online detection stage: inputting a vector sequence into the trained multi-head GRU model, and calculating conditional probability distribution; and setting a threshold value z of the candidate event to judge whether the log sequence is abnormal or not. According to the method, the learning ability of the model for the log sequence mode is effectively enhanced, so that log abnormity is identified more accurately. The multi-head mechanism can flexibly integrate a plurality of basic GRU models and provide a plurality of model options with excellent detection performance. Due to the flexibility, the detection method can balance the detection accuracy and the detection efficiency according to actual requirements. The method is superior to a traditional baseline method in terms of evaluation of mass log data, and the accuracy of log anomaly detection is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting log anomalies of large-scale multi-log events, and more particularly to a method for detecting log anomalies based on multi-head GRU. Background Art

[0002] An anomaly refers to a small number of instances whose patterns are inconsistent with most instances. These anomalies exist in various types of data. Therefore, researchers have proposed corresponding anomaly detection methods in different fields, such as video surveillance, health monitoring, and software defect prediction. In software systems, anomaly detection aims to identify abnormal behaviors or states that occur during system operation. As the scale and functional diversity of systems continue to grow, their complexity has increased sharply, resulting in an increased probability of anomalies. This may affect critical applications and a large number of users in software systems, and thus cause significant losses to service providers. Therefore, timely and accurate anomaly detection is crucial for ensuring the stable operation of the system and reducing potential losses.

[0003] Logs are common monitoring data in software systems, used to record system states and important events at different times. They are one of the most valuable data sources in the field of anomaly detection and are easily obtainable in various systems. By extracting valuable information from logs, developers and maintainers can identify and locate runtime faults of the system. With the explosion of the log volume in large systems, manual anomaly detection has become infeasible. Therefore, researchers have proposed many automated anomaly detection methods. Among existing methods, data-driven deep learning methods have shown greater potential compared to traditional machine learning methods. These deep learning-based methods can obtain better performance when trained on a large amount of training data. They use neural networks as the basic model and train the detection model to learn the sequence patterns in log data.

[0004] When a software system fails, the logs generated by it are important data sources for analyzing the failure. However, the log volume generated by modern systems has grown explosively, and logs have characteristics such as unstructured and imbalanced, which pose great challenges to the anomaly detection task. Manual retrieval-based anomaly detection is not only costly but also difficult to quickly and accurately discover anomalies. Existing detection methods are usually built based on a single (type) of sequence model and have limited ability to learn sequence patterns and cannot accurately identify anomalies. In addition, the quality of log data is crucial for anomaly detection. Therefore, how to effectively extract accurate structured information from logs and use this information to accurately identify anomalies in the system is an urgent problem to be solved.

[0005] The invention patent with the application number 202410028353.6 proposes a log anomaly detection method and device based on multi-feature fusion. First, the original logs are divided into several log sequences, and different preprocessing schemes are selected according to whether each log message content in the sequence contains parameters; then each log is used as the processing object; features are extracted from the internal parameters and templates of each log respectively, and the parameter and template features are aggregated to form corresponding complete log features; finally, taking a sequence as a unit, the relevance of different positions inside the log sequence is captured through multi-head attention, the output of multi-head attention is modeled by a GRU network for sequence information, the result output by GRU is concatenated with the output of the original multi-head attention, and the concatenated features are integrated through a standard feed-forward layer to output a sequence representation, and the sequence representation is passed through a Softmax layer to complete classification. This invention detects abnormal logs from multiple features, improving the accuracy and robustness of log anomaly detection.

[0006] The essence of the multi-head mechanism is to use multiple groups of weight parameters (denoted as Q, K, V in Transformer) to perform calculations on the same input to obtain different mapping spaces, thereby expanding the model's ability to focus on different positions. The aforementioned invention patent adopts the multi-head attention mechanism in Transformer.

[0007] The multi-head self-attention mechanism originates from Transformer and consists of two parts: self-attention and multi-head mechanism. The self-attention mechanism is an important component in Transformer, which allows the model to pay attention to information in other positions in the sequence when processing sequence data, so as to better capture context relationships. Specifically, (1) the input vector generates query (Q), key (K), and value (V) respectively through linear transformation, and these vectors are used to calculate attention weights; (2) by calculating the similarity between the query and all keys (usually using dot product), attention scores are obtained. These scores are converted into a probability distribution through the softmax function, indicating the importance of each position for the currently processed position; (3) finally, all value vectors are weighted and summed according to the above probability distribution to form the final output representation.

[0008] Moreover, in the aforementioned invention patent, the model for calculating probabilities is two modules: the multi-head attention mechanism and GRU, and the structure of the model is in series and fixed, and basic models cannot be flexibly added or deleted. The flexibility of the model is insufficient. Summary of the Invention

[0009] Object of the Invention: The present invention aims to automatically and accurately detect log anomalies, and proposes a log anomaly detection method based on multi-head GRU, aiming to automatically and accurately detect log anomalies.

[0010] The prior art adopts the multi - head attention mechanism in Transformer. The present invention only adopts the multi - head mechanism, which is used to integrate GRU instead of the weight matrix in Transformer, saving computing power.

[0011] The log anomaly detection method of the present invention effectively enhances the model's learning ability for log sequence patterns, thus enabling more accurate identification of log anomalies. The log feature extraction method of the present invention is more flexible. Under the multi - head mechanism, the base model can be flexibly integrated. This feature endows the detection method with more flexibility, enabling it to provide multiple excellent detection performance model options under a unified framework and allowing trade - offs between the requirements of detection accuracy and detection efficiency.

[0012] Technical solution of the invention:

[0013] Inspired by the multi - head mechanism in Transformer, the present invention proposes a log anomaly detection method based on a multi - head GRU network. The process of detecting log anomalies is roughly as follows:

[0014] (1) Construct a multi - head GRU as the core model and train it with the generated log vectors;

[0015] (2) Input a vector sequence X = {v1,..., v T} into the trained multi - head GRU model and calculate the conditional probability distribution

[0016] (3) Set the threshold z of the candidate event to determine whether the log sequence is abnormal. The log anomaly detection is divided into two stages, including the offline training stage and the online detection stage.

[0017]

[0018] In the offline training stage, use the existing log parsing method to convert the original log data into structured data and extract log templates from it (the log data includes log parameter words and log template words. Log parsing deletes the log parameter words in the log data through a certain algorithm and retains the log template words. The log template consists of log template words). These log templates are used to represent the original logs for vectorization to train the detection model. Each unique log template represents a log event. Assuming that e log events are extracted from the log dataset, then define C = {c1, c2,..., c e} as the log event set. Anomaly detection is essentially a multi - class classification model, and each log event represents a class. Given an input as a vector sequence, the output is the conditional probability distribution corresponding to each log event, as shown in formula (2).

[0019] After training, the model has learned a high-dimensional non-linear relationship mapping that can map a vector sequence X (with length T) to the corresponding conditional probability distribution And the parameter θ endows the model with this non-linear mapping.

[0020] Denote the event predicted by the detection model under the condition of the given log vector sequence X and parameter θ Belonging to class c i (log event i) probability.

[0021]

[0022] Due to the diversity of log anomalies and the scarcity of abnormal data required for training, the present invention focuses on learning sequence patterns from normal log datasets (dividing sequence patterns into multiple subsequences during the learning process). This method enables the model to effectively identify abnormal log events by only learning the potential sequence patterns of normal events without a large number of abnormal samples.

[0023] Specifically, after receiving a vector sequence, predict the log events that may occur after the sequence. For a vector sequence L = {v1, v2,..., v m}, use a sliding window to extract subsequences X with length T i ={v i , v (i+1) ,..., v (i+T-1)}. The training objective is to find a model weight such that the predicted event is as close as possible to the actual label event y (i+T) , that is, θ in formula (3) max . During this process, the Adam optimizer and batch gradient descent (BGD) are used to update the model weights. During weight update, the cross-entropy function is used to calculate the loss of the model, as shown in formula (4). I(·) is an indicator function that outputs 0 when the condition is satisfied (y k =c), and 1 otherwise. At the end of the training phase, the model weights with the minimum (validation) loss are retained.

[0024] During the online detection phase, the input log sequence is parsed and vectorized according to the same procedure as in the training phase. The trained model is reloaded to calculate the probabilities of subsequent events for each log sequence. For example, for the log sequence {E5→…→E 11}, the predicted log event probability distribution is {p(E1), p(E2),..., p(E e )}, and by matching the top z high-probability events with the label events, it is determined whether the log sequence is abnormal.

[0025]

[0026] The present invention does not directly compare the labeled event with the predicted event because the log event that occurs after a sequence of log events is not unique. Therefore, several candidate events are selected to match the labeled event. As shown in formula (5), Top z is a function designed to select z predicted events with higher probabilities, where z is called the candidate event quantity threshold. Determining log anomalies is equivalent to checking whether the labeled event y is among the top z high-probability predicted events. If the decision condition A(X, y) is satisfied, the log sequence X is considered normal; otherwise, the sequence is regarded as abnormal. Description of the Drawings

[0027] Figure 1 is the overall design diagram of the present invention;

[0028] Figure 2 is the schematic diagram of the word log vectorization process of the present invention;

[0029] Figure 3 is the model design diagram of the present invention;

[0030] Figure 4 is the schematic diagram of the forward propagation of the data of the present invention;

[0031] Figure 5 is the schematic diagram of the log anomaly determination of the present invention.

[0032] Advantages of the Invention:

[0033] 1. The log anomaly detection method based on multi-head GRU of the present invention can enhance the learning ability of log sequence patterns, thereby improving the detection accuracy of log anomalies. The present invention uses the multi-head mechanism to integrate multiple GRU networks, achieving excellent anomaly detection accuracy. The multi-head mechanism is used to integrate multiple GRU networks to learn the sequence patterns hidden in the log data. Each GRU network is only responsible for learning a local sequence pattern. Subsequently, these local sequence patterns are summarized, and the detection of log anomalies is achieved through global analysis, thereby enhancing the sequence pattern learning ability of the model as a whole and more accurately identifying log anomalies.

[0034] 2. The log anomaly detection method based on multi-head GRU of the present invention allows the basic (sequence) model to be flexibly integrated under the multi-head mechanism. This flexibility allows the model configuration to be adjusted according to efficiency or accuracy requirements. Under the multi-head mechanism, the basic model can be flexibly integrated. This feature endows the detection method with more flexibility, enabling multiple model options with excellent detection performance to be provided externally under a unified framework, allowing a trade-off between the requirements of detection accuracy and detection efficiency.

[0035] 3. The log anomaly detection method based on multi-head GRU of the present invention adopts a relatively simple and direct log feature extraction method, without the need for large-scale pre-trained models and complex calculations like BERT. Through a step-by-step feature transformation and aggregation method, while ensuring the extraction of effective features, it greatly reduces the computational amount and the demand for computing resources. When processing log data, it can complete the feature extraction process faster, improving the real-time performance and efficiency of anomaly detection, and is more suitable for processing large-scale and time-sensitive log data. The evaluation results on the public dataset show that the present invention demonstrates its effectiveness when applying the multi-head mechanism to achieve log anomaly detection. The experimental results verify the performance and practicality of this anomaly detection method.

[0036] 4. The log anomaly detection method based on multi-head GRU of the present invention has the following advantages compared with the prior art:

[0037] (1) Differences between multi-head GRU and multi-head self-attention mechanism:

[0038] The goal of multi-head GRU is to learn sequence patterns on the same log sequence through multiple GRUs (dividing the sequence patterns into multiple subsequences during the learning process). In other words, a set of parameters in this patent is the GRU weight parameters, while for the multi-head attention mechanism used in the prior art patent literature, a set of its parameters is the Q, K, V weight matrices.

[0039] From the perspective of the model structure, multi-head GRU (multi-head + GRU) and multi-head attention (multi-head + self-attention) in the prior art literature are models at the same level. The prior art connects a GRU above the multi-head attention, while the present invention directly calculates the probability after the multi-head GRU. There are essential differences between the two in the model structure.

[0040] (2) Differences in computing resources and efficiency

[0041] From the perspective of the overall model, the prior art literature uses BERT to generate embeddings, multi-head self-attention and GRU to predict probabilities. The BERT model has a large number of parameters and a complex structure. When processing large-scale log data, both in the training and inference stages, it requires a large amount of computing resources. Its computational amount increases significantly with the increase in data volume, resulting in a long processing time and poor real-time performance in practical applications. The present invention uses Skip-gram to generate embeddings and multiple GRUs to predict probabilities. Compared with the prior art, the model in the present invention has fewer parameters, and thus requires less computing resources, and can obtain higher computing efficiency.

[0042] (3) Advantages of the present invention (multi-head GRU): Flexibility of the model

[0043] The multi-head GRU in the present invention is a parallel and flexible model structure that can assemble an appropriate number of basic models (GRU) to meet specific detection accuracy and detection efficiency requirements. For example, assembling 1 GRU can achieve an efficient detection model; assembling 4 GRUs can achieve the best detection accuracy, while 2 GRUs are a balanced model between detection accuracy and detection efficiency. The flexibility of this model is not available in the prior art.

[0044] Model flexibility and adaptability issues

[0045] The BERT model is relatively fixed, and its pre-trained model structure and parameters are difficult to quickly adjust to adapt to the characteristics of different log data. Due to the wide range of log data sources, diverse formats, and complex business logic, the log data features in different scenarios vary greatly. When faced with these differences, it is difficult for the BERT model to quickly optimize the feature extraction method for specific log data, and the construction and parameter adjustment costs are high, resulting in poor adaptability in different log data scenarios. In addition, the amount of log data is large, and the process of BERT's table lookup to construct features will bring additional computing overhead and time costs, affecting the model processing efficiency.

[0046] Computing resources and efficiency issues

[0047] The BERT model has many parameters and a complex structure. When processing large-scale log data, both the training and inference stages require a lot of computing resources. The amount of computing increases significantly with the amount of data, resulting in long processing time and poor real-time performance in practical applications. It is unable to quickly extract features and detect anomalies from new log data, and is difficult to meet some application scenarios with high time requirements, especially system anomaly detection with high real-time requirements.

[0048] The log anomaly detection method based on multi-feature fusion of the present invention has a more flexible log feature extraction method. The log parsing method published in the top journal of SCI Zone 1 is used to remove the template parameter words and construct a log template. Then, it divides the log template into independent log words through the natural language toolkit, and then converts the log words into vector sequences as log features through a series of steps. This method can flexibly adjust the processing method of each step according to the characteristics of different log data. For example, in terms of selecting log encoders, language models, and parameter settings of the fully connected layer, they can all be optimized according to actual data, without the need to face the complex table lookup feature construction process and heavy parameter construction and adjustment problems like BERT. It can process different log data more efficiently and quickly adapt to diverse log data and anomaly detection needs. DETAILED DESCRIPTION

[0049] To make the technical concept and advantages of the present invention for achieving its invention purpose clearer and more understandable, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the following embodiments are only for explaining and illustrating the preferred implementation modes of the present invention, and should not be regarded as and do not constitute a limitation to the scope of patent protection required by the present invention.

[0050] Embodiment 1

[0051] See Figure 1 , the detection process of the log anomaly of the present invention is specifically as follows:

[0052] (1) Construct a multi-head GRU as the core model and train it using the generated log vectors;

[0053] (2) Input a vector sequence X = {v1,..., v T} into the trained multi-head GRU model, and calculate the conditional probability distribution

[0054] (3) Set the threshold z of the candidate event to determine whether the log sequence is abnormal. The log anomaly detection is divided into two stages, including the offline training stage and the online detection stage.

[0055]

[0056] In the offline training stage, use the existing log parsing method to convert the original log data into structured data, and extract the log templates from it. These log templates are used to represent the original logs for vectorization to train the detection model. Each unique log template represents a log event. Assuming that e log events are extracted from the log dataset, then define C = {c1, c2,..., c e} as the log event set. Anomaly detection is essentially a multi-class classification model, and each log event represents a class. Given the input as a vector sequence, the output is the conditional probability distribution corresponding to each log event, as shown in formula (2).

[0057] After training, the model learns a high-dimensional non-linear relationship mapping, which can map the vector sequence X (with length T) to the corresponding conditional probability distribution And the parameter θ endows the model with this non-linear mapping.

[0058] Represents the probability that the event predicted by the detection model belongs to class c i (log event i) under the condition of the given log vector sequence X and parameter θ.

[0059]

[0060] Due to the diversity of log anomalies and the scarcity of abnormal data required for training, in the learning process of the present invention, the sequence pattern is divided into multiple subsequences, and the sequence pattern is learned from the normal log dataset (the learning results of multiple subsequences are aggregated). This method enables the model to effectively identify abnormal log events by only learning the potential sequence patterns of normal events without a large number of abnormal samples. Specifically, after receiving a vector sequence, predict the log events that may occur after this sequence. For a vector sequence L = {v1, v2,..., v m}, use a sliding window to extract a subsequence X i = {v i , v (i+1) , …, v (i+T-1)} of length T. The training objective is to find a model weight such that the predicted event is as close as possible to the actual labeled event y (i+T) , that is, θ max in formula (3). During this process, the Adam optimizer and batch gradient descent (BGD) are used to update the model weights. During weight update, the cross-entropy function is used to calculate the loss of the model, as shown in formula (4). I(·) is an indicator function that outputs 0 when the condition is satisfied (y k = c), and 1 otherwise. At the end of the training phase, the model weight with the minimum (validation) loss is retained.

[0061] In the online detection phase, the input log sequence is parsed and vectorized according to the same procedure as in the training phase. The trained model is reloaded to calculate the probability of subsequent events for each log sequence. For example, for the log sequence {E5 → … → E 11}, the predicted log event probability distribution is {p(E1), p(E2),..., p(E e )}. By matching the top z high-probability events with the labeled event, it is determined whether the log sequence is abnormal. See Figure 5 .

[0062]

[0063] The present invention does not directly compare the labeled event with the predicted event because the log event that occurs after a log event sequence is not unique. Therefore, several candidate events are selected to match with the labeled event. As shown in formula (5), Top z is a function designed to select z predicted events with higher probabilities, where z is called the candidate event number threshold. Determining log anomalies is equivalent to checking whether the labeled event y is among the top z high-probability predicted events. If the decision condition A(X, y) is satisfied, the log sequence X is considered normal; otherwise, the sequence is considered abnormal.

[0064] The log anomaly detection method based on multi - head GRU of the present invention has two characteristics:

[0065] (1) The multi - head mechanism is used to integrate multiple GRU networks to learn the execution sequence patterns hidden in the log data. Each GRU network is only responsible for learning a local sequence pattern. Subsequently, these local sequence patterns are summarized, and the detection of log anomalies is realized through global analysis, thereby enhancing the model's sequence pattern learning ability as a whole and more accurately identifying log anomalies.

[0066] The execution sequence pattern is a complete log sequence, and the local sequence pattern is understood as a part of this sequence. A complete execution sequence pattern is composed of multiple local sequence patterns. That is to say, the features of the execution sequence pattern include the features of the local sequence pattern.

[0067] (2) Under the multi - head mechanism, the basic model can be flexibly integrated. This feature endows the detection method with more flexibility, enabling it to provide multiple excellent detection performance model options under a unified framework and allowing it to balance between the requirements of detection accuracy and detection efficiency.

[0068] Example 2

[0069] The present invention aims to automatically and accurately detect log anomalies. Inspired by the multi - head mechanism in Transformer, a log anomaly detection method based on multi - head GRU network is proposed. The purpose of the multi - head mechanism is to enhance the learning ability of sequence patterns because it can prompt multiple GRU networks to learn log sequence patterns at different positions.

[0070] Once the log is parsed, the features applicable to the multi - head GRU network should be extracted next. The present invention uses natural language processing (NLP) technology to extract log features because log content can be regarded as text data. Specifically, semantic vectors generated by a language model are used as training data. These semantic vectors are obtained in the feature extraction stage and are used to train the model in the offline training stage. The trained model will have the ability to detect log anomalies.

[0071] Let \(L=(l_1, l_2, \ldots, l m )\) be a log sequence, where \(l i \) is the parsed structured log data (log template), where \(1\leq i\leq m\), representing the index of the log template in this log sequence. Each log template is a string composed of multiple log words. Use an existing natural language toolkit (e.g., NLTK) to split \(l i \) into multiple independent log words. This process is defined as a transformation function where s is the sequence of log words generated after splitting a log. Split the result of l i is represented as where is a vocabulary containing all log words. t j is a log template word, which is the smallest and indivisible unit in l. After processing, the log sequence is represented as L = {s1, s2,..., s m}.

[0072] Log feature extraction (mining potential features in the log template, such as the position feature of the template word, the length feature of the log template, and the association feature of the log template words):

[0073] First, use a log encoder to convert log words into vectors of a fixed dimension. The generated vectors are called (log word) embedding vectors.

[0074] Second, vectorize the log template based on the log word embedding vectors. This vectorization is defined as whose input is the sequence of log words of a log template. φ uses a language model to transform these log words, generating |s| embedding vectors, and d represents the dimension of the embedding vectors.

[0075] Finally, use a fully connected layer to aggregate all the log word embedding vectors in the template. This process is defined as The fully connected layer aims to convert the vector dimension of the log template from d×|s| to a. This process is called template2vector, where a is a variable used to adapt to the change of the input dimension under the multi-head GRU. The log template can be converted into an embedding vector through template2vector. Finally, the log sequence can be represented as a sequence of vectors, that is, L = {v1, v2,..., v m}, and this sequence of vectors is the log feature directly available to the model. The foregoing description is shown in Figure 2 .

[0076] Log feature extraction uses a language model instead of one-hot to construct log word vectors. The advantage of this method is to avoid the sparsity of log word vectors under one-hot encoding. At the same time, the vectors generated by the language model contain some semantic information, which helps to improve the performance of the detection model.

[0077] Learning sequence patterns. LSTM and GRU are two popular variants of RNNs and are often used as the basic models for log anomaly detection. By introducing memory cells to store the information accumulated at previous time steps, these two models overcome the long-term dependence problem in traditional RNNs and demonstrate excellent performance in sequence prediction. As a simplified variant of LSTM, GRU uses two gating units (i.e., the reset gate and the update gate) to achieve performance comparable to that of LSTM. Due to the significant sequential characteristics of log sequences, GRU is selected as the base model to construct a multi-head GRU detection model. The experimental results show that GRU performs slightly better than LSTM in the log anomaly detection task.

[0078]

[0079] Equation (1) shows the calculation process of each component in the GRU memory cell, where t represents the current time step. Each memory cell contains a reset gate and an update gate, denoted by Z t and R t respectively. The weight matrix W is used for the connections between the gates, the input, and the hidden state. Different subscript combinations represent their respective independent weights. For example, W xr represents the computational weight of the input X t on the reset gate, and W hr represents the computational weight of the hidden state on the reset gate. In addition, σ represents the sigmoid function, and tanh is the hyperbolic tangent function.

[0080] The output of GRU at each time step is the current hidden state H t , and the hidden state H t-i from the previous time step participates in the calculation of H t . At the same time, H t also participates in the calculation of the output at the next time step t + 1. In this recursive process, the reset gate controls how to combine the hidden state H t-1 to generate the candidate hidden state at the current time step From a functional perspective, the reset gate determines how much information from the previous hidden state should be forgotten or retained. The update gate is responsible for determining how to fuse the candidate hidden state and the previous hidden state H t-1 to update the hidden state H t at the current time step. In this way, the update gate balances the contributions of new information and old information, ensuring that the model can effectively capture the long-term dependencies in the sequence.

[0081] The multi - head mechanism in the Transformer enhances the model's ability to focus on different position information by providing multiple representation sub - spaces, thus improving the performance of the attention layer. The essence of the multi - head mechanism is to project the input into different representation sub - spaces using multiple sets of weights, enabling the model to capture different features at multiple positions in the sequence simultaneously. Inspired by this, the multi - head GRU - based anomaly detection method draws on the advantages of the multi - head mechanism and constructs a multi - head GRU model using a two - layer GRU network. Each GRU unit is responsible for learning the local sequence patterns at its corresponding position (the sequence patterns are divided into multiple sub - sequences during the learning process), thereby enhancing the model's ability to capture complex sequence patterns. This method not only enables the RNN to focus on sequence patterns at different positions but also enhances the overall model's sequence pattern learning ability, thus improving the accuracy of log anomaly detection.

[0082] Specifically, a multi - head GRU network is constructed using N GRUs. As Figure 3 shown, consistent with the number of GRUs, in the log vector sequence, each vector is also evenly divided into N parts. Consider a vector X = {v1,..., v T}, where v t (1 ≤ t ≤ T) represents a log embedding vector. Divide v t evenly into N parts. The divided vectors are represented as: where is a vector segment extracted from v t . Vector segments extracted from the same position in different vectors form a vector segment sequence, which is defined as a local vector sequence. For example, X 1 is a local vector sequence composed of It is composed of vector segments extracted from all vectors in X at position 1. The uniform partitioning of the log embedding vectors is to achieve the effect of GRU under the multi-head mechanism. The multi-head mechanism is sourced from the Transformer network, and the queries, keys, and values (i.e., the main weights Q, K, and V of the model) are obtained through multiple linear projections. In other words, the purpose of the multi-head mechanism is to create multiple sets of query / key / value weight matrices. In the present invention, multiple sets of weights of the base model (e.g., GRU) are constructed by random initialization and direct parameter copying. The multi-head mechanism designed in the present invention aims to expand the input data rather than the weight parameters. Compared with the matrix multiplication in Transformer, GRU has a stronger learning ability for sequential data. When the same data is input into each GRU, it may cause their weight parameters to approach similarity. To prevent this, the log embedding vectors are evenly segmented and input into each GRU separately. Such a processing method ensures that all GRUs receive different data sources, enabling each GRU to learn local sequence patterns from the corresponding positions, thereby enhancing the model's learning ability for log sequence patterns as a whole.

[0083] Assume that the input dimension of the base model (such as GRU or LSTM) and the dimension of the log embedding vector are both 256. The present invention needs to adapt the log embedding vector to each base model, which includes two steps: (1) Dimension expansion: The initial log embedding is converted into a higher-dimensional embedding through a linear layer. The converted dimension is 256×N, where N represents the number of "heads". (2) Data segmentation: The expanded data is evenly divided into N "heads", each with a dimension of 256, and they are input into each base model separately. From the perspective of the model, a single GRU is responsible for learning the corresponding local sequence patterns from the local vector sequence (e.g., X n ). From the perspective of the input vector, a log embedding vector is divided into N parts, which are input into N GRU blocks in different GRU networks. Taking a vector v T as an example, it is divided into N segments and input into N GRU networks at T time steps, that is where, represents the GRU block at the T-th time step in the first GRU network. Each vector segment outputs the corresponding hidden state at the T-th time step In practice, the hidden state output at the final time step is selected to represent the local sequence pattern, is the hidden state output by the n-th GRU at the final time step (T). Combining these local sequence patterns can thus learn the sequence pattern from a global perspective.

[0084] For a log sequence, each local sequence pattern has a specific impact on the prediction result. Therefore, a weight is assigned to each hidden state These weights reflect the importance of each hidden state in the global sequence pattern and are automatically learned during the training phase. Aggregate all the weighted hidden states and use a fully connected layer (W to output the global sequence pattern. Subsequently, add a softmax layer to calculate the prediction scores fc where X represents the vector sequence, and represents the predicted next log event. The foregoing description is shown in Figure 4 .

[0085] The log anomaly detection method based on the multi-head GRU network of the present invention effectively enhances the model's learning ability of log sequence patterns, thereby more accurately identifying log anomalies. With the support of the multi-head mechanism, this method can flexibly integrate multiple basic GRU models, providing multiple model options with excellent detection performance. This flexibility enables the detection method to balance between detection accuracy and detection efficiency according to actual needs. Experimental results show that the present invention outperforms traditional baseline methods in the evaluation of massive log data, significantly improving the accuracy of log anomaly detection.​

Claims

1. A log anomaly detection method based on multi-head GRU, characterized in that: The process of realizing log anomaly detection includes: Step S1, offline training phase: Based on the multi-head mechanism, construct a multi-head GRU as the core model and train it using the generated log vectors. Step S2, online detection phase: Input a vector sequence X = {v1, …, v T} into the trained multi-head GRU model to calculate the conditional probability distribution Set a threshold z for candidate events to determine whether the log sequence is abnormal.

2. The log anomaly detection method based on multi-head GRU according to claim 1, characterized in that: Step S1, offline training phase: Adopt the GRU neural network, use the multi-head mechanism to integrate multiple GRU network models to learn the execution sequence patterns hidden in the log data; each GRU network model is responsible for learning a local sequence pattern in the log data, and then summarize these local sequence patterns, and realize the detection of log anomalies through global analysis, so as to enhance the execution sequence pattern learning ability of the multi-head GRU model as a whole and more accurately identify log anomalies.

3. The method for log anomaly detection based on multi-head GRU according to claim 1 or 2, characterized in that: In the offline training phase, use the log parsing method to convert the original log data into structured data and extract log templates from it. These log templates are used to represent the original logs for vectorization to train the detection model. Each unique log template represents a log event. Suppose e log events are extracted from the log dataset, then define C = {c1, c2, …, c e} as the log event set; Anomaly detection is essentially a multi-class classification model, and each log event represents a class; Given the input as a vector sequence, the output is the conditional probability distribution corresponding to each log event, as shown in formula (2). After training, the model learns a high-dimensional non-linear relationship mapping that maps the vector sequence X (with length T) to the corresponding conditional probability distribution The parameter θ endows the model with this non-linear mapping; Denotes the probability that, given a sequence of log vectors X and parameters θ, the detection model predicts that the event belongs to class c i (of log event i); 4. The method for detecting log anomalies based on multi-head GRU according to claim 3, characterized in that: Learn sequence patterns from the normal log dataset, and effectively identify abnormal log events by only learning the potential sequence patterns of normal events: after receiving a vector sequence, predict the log events that may occur after the sequence. For a vector sequence \(L = \{v_1, v_2, \ldots, v\) m \}, a subsequence \(X\) i of length \(T\) is extracted using a sliding window i \(=\{v\) (i+1) , \(v\) (i+τ-1) \}; The training objective is to find a set of model weights such that the predicted event is as close as possible to the actual labeled event y (i+T) , i.e., θ in Equation (3) max ; the Adam optimizer and batch gradient descent (BGD) are used to update the model weights; during weight update, the cross-entropy function is used to calculate the loss of the model, as shown in Equation (4): where I(·) is the indicator function that outputs 0 when the condition (y k = c) is satisfied and 1 otherwise; at the end of the training phase, the model weights with the smallest (validation) loss are retained.

5. The method for detecting log anomalies based on multi-head GRU according to claim 3, characterized in that: In the online detection phase, the input log sequence is parsed and vectorized according to the same procedure as in the training phase; the trained multi-head GRU model is reloaded to calculate the probability of subsequent events for each log sequence. For the log sequence {E5 → … → E 11}, the predicted probability distribution of log events is {p(E1), p(E2),..., p(E e )}. By matching the first z high-probability events with the labeled events, it is determined whether the log sequence is abnormal: In formula (5), Top z is a function designed to select z predicted events with higher probabilities, where z is called the candidate event quantity threshold; determining a log anomaly is equivalent to checking whether the labeled event y is among the top z predicted high-probability events; if the decision condition A(X, y) is satisfied, the log sequence X is considered normal; otherwise, the sequence is regarded as abnormal.

6. The method for detecting log anomalies based on multi-head GRU according to claim 4 or 5, characterized in that: Adopt natural language processing technology to extract log features, use the semantic vectors generated by the language model as training data, and extract features suitable for the multi-head GRU network: these semantic vectors are obtained in the feature extraction phase and used to train the model in the offline training phase; the trained multi-head GRU model will have the ability to detect log anomalies. Let \(L=(l_1, l_2, \ldots, l\) m ) be a log sequence, where \(l\) i is the parsed structured log data / log template, where \(1\leq i\leq m\), representing the index of the log template in this log sequence; each log template is a string composed of multiple log words; use the existing natural language toolkit to split \(l\) i into multiple independent log words, and this process is defined as a conversion function where \(s\) is the sequence of log words generated after splitting a log; represent the splitting result of \(l\) i as wherein is a vocabulary containing all log words; t j is a log template word, which is the smallest and indivisible unit in l; After The log sequence after processing is represented as L = {s1, s2,..., s m}.

7. The method for detecting log anomalies based on multi-head GRU according to claim 6, wherein: The process of log feature extraction is as follows: First, use the log encoder to convert log words into vectors of a fixed dimension, and the generated vectors are called (log word) embedding vectors. Secondly, the log template is vectorized based on the log word embedding vectors; this vectorization is defined as Its input is the log word sequence of a log template; φ uses a language model to transform these log words and generate |s| embedding vectors, where d represents the dimension of the embedding vectors; Finally, a fully connected layer is used to aggregate all the log word embedding vectors in the template, and this process is defined as The fully connected layer aims to convert the vector dimension of the log template from d×|s| to a; this process is named template2vector, where a is a variable used to adapt to the change of the input dimension under the multi-head GRU; the log template can be converted into an embedding vector through template2vector; Finally, the log sequence can be represented as a sequence of vectors, i.e., L = {v1, v2, …, v m}, and this sequence of vectors is the log feature directly available to the model.

8. The method for detecting log anomalies based on multi-head GRU according to claim 3 or 7, characterized in that: The process of learning sequence patterns is as follows: LSTM and GRU are two popular RNN variants and are often used as the basic models for log anomaly detection; select GRU as the basic model to construct a multi-head GRU detection model. Equation (1) shows the calculation process of each component in the GRU memory unit, where t represents the current time step; each memory unit contains a reset gate and an update gate, denoted by Z t and R t respectively; the weight matrix W is used for the connection between the gates, the input, and the hidden state; different subscript combinations represent their respective independent weights, W xr represents the calculation weight of the input X t on the reset gate, W hr represents the calculation weight of the hidden state on the reset gate, σ represents the sigmoid function, and tanh is the hyperbolic tangent function. The output of the GRU at each time step is the current hidden state H t , the hidden state H of the previous time step t-i participates in the calculation of H t ; at the same time, H t also participates in the calculation of the output at the next time step t+1; During this recursive process, the reset gate controls how the hidden state H t-1 is combined to generate the candidate hidden state at the current time step 9. The method for detecting log anomalies based on multi-head GRU according to claim 8, wherein: Construct a multi - head GRU model using a double - layer GRU network. Each GRU unit is responsible for learning the local sequence pattern at its corresponding position, and an N - head GRU network is constructed using N GRUs; consistent with the number of GRUs, in the log vector sequence, each vector is also evenly divided into N parts; Consider a vector X = {v1,..., v T}, where v t (1 ≤ t ≤ T) represents a log embedding vector; Divide v t evenly into N parts, and the divided vector is represented as: Among them is a vector segment extracted from v t . A sequence of vector segments is formed by extracting vector segments from the same positions in different vectors, and it is defined as a local vector sequence. X 1 is a local vector sequence, which is composed of and is formed by the vector segments extracted from all vectors in X at position 1; Divide the log embedding vectors evenly, and obtain queries, keys, and values (i.e., the main weights Q, K, and V of the model) through multiple linear projections, and construct the weights of multiple groups of basic models by random initialization and direct copying of parameters.

10. The method for log anomaly detection based on multi-head GRU according to claim 9, characterized in that: Assume that the input dimension of the basic model such as GRU or LSTM and the dimension of the log embedding vector are both 256. Adapting the log embedding vector to each basic model includes two steps: 1) Dimension expansion: Convert the initial log embedding into a higher-dimensional embedding through a linear layer; the converted dimension is 256×N, where N represents the number of "heads". 2) Data segmentation: The extended data is evenly divided into N "headers", each with a dimension of 256, and they are respectively input into each basic model; From the perspective of the model, a single GRU is responsible for learning the corresponding local sequence patterns from the local vector sequence; from the perspective of the input vector, a log embedding vector is divided into N parts, which are input into N GRU blocks in different GRU networks; a vector v T is divided into N segments and input into N GRU networks at T time steps, that is where represents the GRU block at the T-th time step in the first GRU network, and each vector segment outputs the corresponding hidden state at the T-th time step Select the hidden state output at the final time step to represent the local sequence pattern is the hidden state output by the n-th GRU at the final time step (T); by combining these local sequence patterns, the sequence pattern can be learned from a global perspective; For each hidden state Assign a weight These weights reflect the importance of each hidden state in the global sequence pattern and are automatically learned during the training phase; aggregate all weighted hidden states and use a fully connected layer (W fc ) to output the global sequence pattern; then add a softmax layer to calculate the predicted score where X represents the vector sequence, and represents the predicted next log event.

Citation Information

Patent Citations

  • Log anomaly detection method and device based on multi-feature fusion

    CN117971594A