Log Sequence Anomaly Detection Method Based on Time Interval-Aware Self-Attention Mechanism

By introducing a time interval-aware self-attention mechanism in log sequence abnormality detection, combined with BERT and Transformer encoder, the problem of time interval information in log sequence is solved, and the accuracy and robustness of detection are improved.

CN115617614BActive Publication Date: 2025-07-22DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211339210.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-07-22
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

The existing log sequence abnormality detection methods fail to make full use of the time interval information in the log, resulting in limited abnormality detection performance and inability to adapt to word changes in log statements, affecting the accuracy and robustness of detection.

Method used

A time interval perceptual self-attention mechanism is introduced, the word vectors of the log template are extracted through the BERT model, and a one-dimensional convolutional neural network and a Transformer encoder are used to detect log sequence anomalies in combination with the time interval information.

Benefits of technology

It improves the accuracy and robustness of log sequence abnormality detection, can adapt to word changes in log statements, and enhances the model's adaptability during system upgrade.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115617614B_ABST
    Figure CN115617614B_ABST
Patent Text Reader

Abstract

The present invention provides a log sequence anomaly detection method based on a time interval-aware self-attention mechanism. In the process of using a Transformer encoder to extract log sequence features, the present invention introduces a time interval-aware self-attention mechanism, and utilizes the semantic information and time interval information of log templates in the process of calculating attention scores, improving the ability to obtain the correlation information between logs, enabling the model to learn the impact of the time interval between logs in the sequence on anomaly detection, and improving the accuracy of anomaly detection. At the same time, on the basis of using a BERT language model to extract word vectors, a CNN is used to aggregate word vectors to generate a vector representation of the log template, learning semantic information of words themselves and their contexts at different scales, enabling the model to adapt to the word changes in log statements during the software update process, and improving the robustness of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a method for detecting anomalies in log sequences based on a time interval-aware self-attention mechanism. Background Art

[0002] Log data is an important and valuable data source in online service systems, which records detailed information about the system running status and user behavior. By analyzing whether the log sequence generated by the system deviates from the normal working mode, errors occurring during system operation can be effectively detected, improving the reliability of the software.

[0003] Currently, mainstream methods for detecting anomalies in log sequences can be divided into machine learning-based methods and deep learning-based methods.

[0004] Most machine learning-based methods for detecting anomalies in log sequences use the statistical features of log data combined with machine learning algorithms to detect anomalies in log sequences.

[0005] For example, the unsupervised PCA algorithm is used. By analyzing the source code, the event status and occurrence times expressed by the logs are statistically analyzed, and the state proportion vector and event count vector are used as the input of the model for anomaly detection. Although this method extracts log features from different aspects, it does not extract the semantic information of log statements, resulting in a lack of robustness.

[0006] Another example is using the SVM algorithm. By quantifying the log sequence based on the quantity and distribution of various log levels in a sliding window and inputting it into the model, anomaly detection is performed through supervised training. Although it has certain effects, it does not consider the sequential relationship of the logs and cannot detect the sequential anomalies of the logs, resulting in a low detection accuracy.

[0007] Most deep learning-based methods for detecting anomalies in log sequences focus on using the sequential relationship between logs and use methods such as recurrent neural networks and attention mechanisms to detect anomalies in log sequences. Some scholars apply technical methods in the field of natural language processing to log sequence anomaly detection.

[0008] For example, the DeepLog log anomaly detection model based on LSTM. Although it learns the sequential relationship between logs, the encoding method using log template indexing cannot fully extract the semantic information of log templates, lacking robustness.

[0009] For another example, LogRobust converts the words in the log template into word vectors through natural language processing methods, aggregates the word vectors based on TF-IDF (term frequency-inverse document frequency) to generate log template vectors, and uses a Bi-LSTM model based on the attention mechanism to detect anomalies. Although it considers the semantic information contained in the log statements, the method of aggregating word vectors using statistical methods cannot cope with the changes in words in the log statements and cannot learn the context features of words, resulting in low robustness. At the same time, the LSTM-based model does not consider the location information of log events, is troubled by the long-term dependence problem, and there is a risk of losing past information, resulting in low accuracy.

[0010] Another example is NeuralLog, which extracts word vectors through BERT. The log template vector is calculated as the average of its corresponding word vectors. The log template vector sequence is input into a Transformer-based model for anomaly detection, improving the accuracy of detection. Although these methods have achieved certain effects in the accuracy of anomaly detection, using statistical methods to aggregate word vectors to generate log template vectors cannot distinguish the semantic information of different words, cannot learn the local features of the log template, and will reduce the accuracy of anomaly detection when the words in the log template change. When the system is running normally, the response times of different tasks will be stable within the normal range. However, when anomalies such as hardware problems, network communication congestion, and component performance defects occur, the response times of tasks will fluctuate greatly, and the time intervals between the log statements generated by the tasks will be too long.

[0011] The above methods only focus on the sequential information of the logs and fail to fully utilize the time information useful for the anomaly detection task in the logs, resulting in limited performance of anomaly detection. In summary, the current anomaly detection methods for log sequences mainly have the following disadvantages:

[0012] (1) Most methods use techniques in the field of natural language processing to obtain word vectors and use statistical methods to aggregate word vectors to generate log template vectors. Although they can obtain the semantic information of the words in the logs, they cannot well distinguish the impact of the semantics of different words on the anomaly detection task. At the same time, they cannot obtain the context features of the words, capture the dependencies between the words, and cannot adapt to the word changes caused by the irregular updates of log statements during the upgrade of the system or service, affecting the accuracy of anomaly detection.

[0013] (2) Most methods mainly focus on the impact of the sequential information of logs on anomaly detection. Although most system failures will cause the logs to deviate from the normal log sequence, and system anomalies can be effectively detected through the sequential information of logs, the time interval between logs is also an important information source for judging system anomalies. Ignoring the time interval between logs while only considering the sequential information of logs limits the performance of anomaly detection. Summary of the Invention

[0014] The present invention provides a log sequence anomaly detection method based on a time interval-aware self-attention mechanism. The present invention introduces a time interval-aware self-attention mechanism, inputs the time interval between logs and semantic information into a multi-head self-attention mechanism to obtain the correlation information between logs, optimizes the expression ability of the extracted features, and improves the accuracy of anomaly detection.

[0015] The technical means adopted by the present invention are as follows:

[0016] A log sequence anomaly detection method based on a time interval-aware self-attention mechanism, comprising:

[0017] Obtain a log set for training, perform data parsing on the log set for training to generate a log template sequence and a timestamp sequence; calculate the relative time interval between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix;

[0018] Input the log template sequence and the time interval matrix into an anomaly detection model based on a time interval-aware self-attention mechanism to perform supervised training on the entire model.

[0019] Obtain a log set to be detected, and input the log set to be detected into the trained log sequence anomaly detection model based on a time interval-aware self-attention mechanism for anomaly detection.

[0020] Further, the training steps of the log sequence anomaly detection model based on a time interval-aware self-attention mechanism include:

[0021] Input the log template sequence into a BERT model for word vector processing to obtain the word vectors of the log template sequence;

[0022] Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors; determine whether there are still unprocessed log templates in the log template sequence, and if so, perform word vector extraction and template vector extraction on the next log template until all log templates are processed;

[0023] Input the log template vector and the time interval matrix into the Transformer encoder to obtain the log sequence vector;

[0024] Input the log sequence vector into a classifier based on a fully connected neural network to output the detection result, calculate the loss based on the detection result, and update the model parameters.

[0025] Further, calculate the relative time interval between each log event in the sequence according to the timestamp sequence, so as to obtain the time interval matrix, including:

[0026] For a timestamp sequence T = {t1, t2, … t n}, calculate the time difference t ij and take the absolute value to represent the relative time interval between the log at the i-th position and the log at the j-th position, and then standardize it where μ is the mean of all time interval data in the training set, and σ is the standard deviation, to obtain a new time interval matrix

[0027] Further, input the log template sequence into the BERT model to extract word vectors, so as to obtain the word vectors of the log template sequence, including:

[0028] Regard the log template as a sentence in natural language. For a log template sequence X = {x1, x2, … x n}, n is the sequence length, and x i represents the i-th log template, and split it into a word sequence expressed as x i = {w1, w2, … w m}, where m represents the sentence length of the log template;

[0029] Encode the word sequence through the BERT language model, map each word to a d-dimensional vector, then this log template is represented as a word vector sequence z i = {v1, v2, … v m}, where

[0030] Further, extract the template vector from the word vectors of the log template sequence through a one-dimensional convolutional neural network, so as to obtain the log template vector, including:

[0031] For the log template word vector sequence z i , assuming that v k:k+j represents all word vectors from v k to v k+j , input the word vectors into the convolutional layer, and there are multiple convolutional kernels in the convolutional layer. The operation process of each convolutional kernel is shown in formula (1).

[0032] c i = f(W·v k:k+h-1 + b)(1)

[0033] where h is the height of the convolutional kernel, d is the width of the convolutional kernel, is the bias term, and f is the non-linear activation function ReLU;

[0034] Each convolutional kernel can obtain a new feature map C = {c1, c2, … c m-h+1} by operating on all word vectors in the log template, and the maximum value is obtained through the max pooling layer as the eigenvalue obtained by the operation of this convolutional kernel;

[0035] The eigenvalues generated by all convolutional kernels are concatenated to obtain the log template vector where y is the number of convolutional kernels.

[0036] Furthermore, the log template vector and the time interval matrix are input into the Transformer encoder based on the time interval-aware self-attention mechanism to obtain the log sequence vector, including:

[0037] Perform positional encoding on the log template vector, and use sine and cosine functions to generate vectors for each log event in the log sequence as shown in formula (2):

[0038]

[0039] where t = 1, 2, … y, representing different dimensions of the vector, and PE i is added to the log template vector e i at position i, so that the model can learn the relative position information of each log event;

[0040] The log template vector sequence with added position information is input into the time interval-aware self-attention layer, and the time interval information of the log event is added in the attention calculation. The time interval matrix is extended by one dimension to become and multiplied by two parameter matrices to obtain the key-value matrix of the time interval and the value matrix The attention calculation method for the vector e i in the sequence is as shown in formula (3):

[0041]

[0042] where h is the h-th attention head, D q = Dk = D v = y / H, where H is the number of attention heads, and each attention head uses different learnable parameter matrices during calculation. The operation results of each head are concatenated and multiplied by the parameter matrix to obtain a new log template vector representation, as shown in formula (4):

[0043] z i = Concat(A1, A2…, A H )W O (4);

[0044] The log template vector obtained through attention calculation is input into the feed-forward fully connected layer for two linear transformations, as shown in formula (5), to obtain the final log sequence vector:

[0045] r i = max(0, z i ·W1 + b1)W2 + b2 (5)

[0046] where The final log sequence vector is represented as R = {r1, r2,... r n}, r i ∈ R y .

[0047] Furthermore, the log sequence vector is input into a classifier based on a fully connected neural network to obtain the classification result of the log sequence and the entire model is trained in a supervised manner. Based on the detection result, the loss is calculated and the model parameters are updated, including:

[0048] The log sequence vector is input into a classifier based on a fully connected neural network to output the detection result, calculate the loss, and update the model parameters, as shown in formula (6):

[0049] pre = Softmax(R·W s + b s ) (6)

[0050] where W s , b s are the weights and biases of the fully connected neural network, and pre is the probability of the model output being normal and abnormal;

[0051] The cross-entropy loss function is used to calculate the loss between the model detection result and the label provided in the dataset, as shown in formula (7):

[0052]

[0053] Update the parameters through backpropagation using the Adam optimizer. The parameters to be updated include those in the CNN, Transformer encoder, and fully connected neural network.

[0054] Further, the step of inputting the log set to be detected into the trained log sequence anomaly detection model based on the time interval-aware self-attention mechanism for anomaly detection includes:

[0055] Obtain the log set to be detected, perform data parsing on the log set to generate a log template sequence and a timestamp sequence; calculate the relative time intervals between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix;

[0056] Input the log template sequence into the BERT model to extract word vectors, thereby obtaining the word vectors of the log template sequence;

[0057] Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors; determine whether there are still unprocessed log templates in the log template sequence. If so, perform word vector extraction and template vector extraction on the next log template until all log templates are processed;

[0058] Input the log template vectors and the time interval matrix into the Transformer encoder based on the time interval-aware self-attention mechanism to obtain log sequence vectors;

[0059] Input the log sequence vectors into a classifier based on a fully connected neural network to output the detection result.

[0060] Compared with the prior art, the present invention has the following advantages:

[0061] The present invention relates to a log sequence anomaly detection method based on the time interval-aware self-attention mechanism. In order to improve the accuracy of anomaly detection, the present invention introduces the time interval-aware self-attention mechanism in the process of using the Transformer encoder to extract log sequence features, utilizes the time interval information between logs, improves the ability to obtain the correlation information between logs, and enables the model to learn the influence of the time interval between each log in the sequence on anomaly detection. In addition, on the basis of using the pre-trained BERT language model to extract the semantic representation of words in the log, the present invention uses CNN to aggregate word vectors to generate the vector representation of the log template, learns the semantic information of words themselves and their contexts at different scales, enables the model to adapt to the word changes in the log statements during the software update process, and improves the robustness of anomaly detection. The experimental results on two log datasets show that this method is superior to most existing log sequence-based anomaly detection methods. Brief Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0063] Figure 1 It is a flowchart of a method for detecting abnormal log sequences based on a time interval-aware self-attention mechanism of the present invention.

[0064] Figure 2 It is a schematic diagram of the LogT framework in the embodiment.

[0065] Figure 3 It is a schematic diagram of the LogT training process in the embodiment.

[0066] Figure 4 It is a schematic diagram of the detection process of LogT in the embodiment.

[0067] Figure 5 It is the influence of different numbers of attention heads on the detection performance in the HDFS dataset in the embodiment.

[0068] Figure 6 It is the influence of different numbers of attention heads on the detection performance in the BGL dataset in the embodiment.

[0069] Figure 7 It is the performance of different methods on the HDFS dataset in the embodiment.

[0070] Figure 8 It is the performance of different methods on the BGL dataset in the embodiment.

[0071] Figure 9 It is a schematic diagram of log update in the embodiment.

[0072] Figure 10 It is a comparison chart of the robustness of different methods in the embodiment. Detailed Embodiments

[0073] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] Such asFigure 1-4 As shown in Figure 1-4 , the present invention provides a log sequence anomaly detection method (LogT) based on a time interval-aware self-attention mechanism, which mainly includes:

[0075] S1. Obtain a log set for training, parse the data of the log set to be detected, so as to generate a log template sequence and a timestamp sequence; calculate the relative time interval between each log event in the sequence according to the timestamp sequence, so as to obtain a time interval matrix.

[0076] The goal of log parsing is to convert unstructured log text into structured data

[11] . In this method, the Drain algorithm with better performance is selected for log parsing. Its core idea is to construct a parsing tree with a fixed depth based on log data, and perform log parsing according to the template extraction rules contained in the tree, and extract log templates and timestamps from unstructured log text.

[0077] According to the characteristics of the log data set, the present invention preferably divides the log sequence according to its session id or sliding window, and obtains the log template sequence and timestamp sequence therein.

[0078] Further, calculating the relative time interval between each log event in the sequence according to the timestamp sequence, so as to obtain a time interval matrix, includes:

[0079] For a timestamp sequence T = {t1, t2,... t n}, calculate the time difference t ij and take the absolute value to represent the relative time interval between the log at the i-th position and the log at the j-th position, and then standardize it where μ is the mean of all time interval data in the training set, and σ is the standard deviation, to obtain a new time interval matrix

[0080] S2. Input the log template sequence and the time interval matrix into an anomaly detection model based on a time interval-aware self-attention mechanism for supervised training.

[0081] Specifically, as a preferred implementation manner of the present invention, in the step S2, the process of training the anomaly detection model based on a time interval-aware self-attention mechanism includes:

[0082] S201. Input the log template sequence into the BERT model to extract word vectors, so as to obtain the word vectors of the log template sequence.

[0083] In the present invention, in order to capture the semantic information of log events, the log template is regarded as a sentence in natural language. For a log template sequence X = {x1, x2,... x n}, n is the sequence length, xi Denote the i-th log template, and split it into a word sequence denoted as x i ={w1, w2, …, w m}, where m represents the sentence length of the log template; encode the word sequence through the BERT language model, map each word to a d-dimensional vector, then this log template is represented as a word vector sequence z i ={v1, v2, …, v m}, where

[0084] S202. Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors; determine whether there are still unprocessed log templates in the log template sequence. If so, perform word vector extraction and template vector extraction on the next log template until all log templates are processed.

[0085] Specifically, for the log template word vector sequence z i , assume that v k:k+j represents all word vectors from v k to v k+j . Input the word vectors into the convolutional layer. The convolutional layer contains multiple convolutional kernels, and the operation process of each convolutional kernel is shown in formula (1).

[0086] c i = f(W · v k:k+h-1 + b) (1)

[0087] where h is the height of the convolutional kernel, d is the width of the convolutional kernel, is the bias term, and f is the non-linear activation function ReLU;

[0088] Each convolutional kernel can obtain a new feature map C = {c1, c2, …, c m-h+1} by operating on all word vectors in the log template. Obtain the maximum value through the max pooling layer as the eigenvalue obtained by the operation of this convolutional kernel;

[0089] Connect the eigenvalues generated by all convolutional kernels to obtain the log template vector where y is the number of convolutional kernels.

[0090] Furthermore, determine whether there are still unprocessed log templates in the log template sequence. If so, execute S201; otherwise, execute S204.

[0091] S203. Input the log template vector and the time interval matrix into the Transformer encoder based on the time interval-aware self-attention mechanism to obtain the log sequence vector.

[0092] First, perform positional encoding on the log template vectors, and use sine and cosine functions to generate vectors for log events at each position in the log sequence As shown in formula (2):

[0093]

[0094] where t = 1, 2, … y, representing different dimensions of the vector, and PE i is added to the log template vector e at position i i , enabling the model to learn the relative position information of each log event.

[0095] Secondly, input the log template vector sequence with added position information into the time interval-aware self-attention layer. The time interval information of log events is added in the attention calculation. The time interval matrix is extended by one dimension to become and multiplied by two parameter matrices to obtain the key-value matrix of the time interval and the value matrix The attention calculation method for the vector e i in the sequence is as shown in formula (3):

[0096]

[0097] where h is the h-th attention head, D q = D k = D v = y / H, H is the number of attention heads. Each attention head uses different learnable parameter matrices during calculation. The operation results of each head are concatenated and multiplied by the parameter matrix to obtain the new log template vector representation, as shown in formula (4):

[0098] z i = Concat(A1, A2…, A H )W O (4).

[0099] Finally, input the log template vector after attention calculation into the feed-forward fully connected layer, perform two linear transformations, as shown in formula (5), to obtain the final log sequence vector:

[0100] r i = max(0, z i ·W1 + b1)W2 + b2 (5)

[0101] where The final log sequence vector is represented as R = {r1, r2,... r n}, r i ∈R y .

[0102] S204. Input the log sequence vector into a classifier based on a fully connected neural network to output a classification result and perform supervised training on the entire model.

[0103] Specifically, input the log sequence vector into a fully connected neural network with a Softmax function to output a detection result, calculate the loss, and update the model parameters, as shown in formula (6):

[0104] pre = Softmax(R · W s + b s ) (6)

[0105] where W s , b s are the weights and bias terms of the fully connected neural network, and pre is the probability of the model outputting normal and abnormal;

[0106] Use the cross-entropy loss function to calculate the loss between the model detection result and the label provided in the dataset, as shown in formula (7):

[0107]

[0108] Update the parameters by backpropagation through the Adam optimizer. The parameters to be updated include those in the CNN, Transformer encoder, and fully connected neural network.

[0109] S205. Determine whether there are still untrained sequences in the log test set. If so, go to S2; otherwise, end the training.

[0110] S3. Obtain the log combination to be detected, and input the log set to be detected into the trained anomaly detection model based on the time interval-aware self-attention mechanism for anomaly detection.

[0111] As a preferred embodiment of the present invention, the above detection steps include:

[0112] S301. Obtain the training log set, perform data parsing on the training log set to generate a log template sequence and a timestamp sequence; calculate the relative time interval between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix. This step is the same as the execution method of S1 and will not be elaborated here.

[0113] S302. Input the log template sequence into the BERT model to extract word vectors, thereby obtaining the word vectors of the log template sequence. The execution method of this step is the same as that of S201 and will not be elaborated here.

[0114] S303. Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network, thereby obtaining log template vectors; determine whether there are still unprocessed log templates in the log template sequence. If so, perform word vector extraction and template vector extraction on the next log template until all log templates are processed. The execution method of this step is the same as that of S202 and will not be elaborated here.

[0115] S304. Input the log template vectors and the time interval matrix into the Transformer encoder to obtain log sequence vectors. The execution method of this step is the same as that of S203 and will not be elaborated here.

[0116] S305. Input the log sequence vectors into a classifier based on a fully connected neural network to output the detection results.

[0117] S306. Determine whether there are still untrained sequences in the log test set. If so, go to S302; otherwise, end the training.

[0118] The following further illustrates the solution and effect of the present invention through specific application examples.

[0119] This part compares the algorithm of this patent with various current latest algorithms respectively from two aspects of the accuracy and robustness of log sequence anomaly detection, and describes the beneficial effects of this algorithm.

[0120] (1) Experimental environment

[0121] The experiments of the present invention are all carried out on a server with NVIDIA TESLA V100 32G GPU. The model is constructed based on Pytorch using the Python3.6 environment. The LogT model is trained using the Adam optimizer, and the cross-entropy function is used as the loss function during training. The training process is terminated after 10 iterations.

[0122] (2) Dataset

[0123] In 2019, Shilin He et al. released a large number of logbooks from 16 different systems on Loghub. For this invention, two representative log datasets, HDFS and BGL, were selected. The details of the datasets are shown in Table 1. The HDFS dataset was generated by running Hadoop on more than 200 Amazon EC2 (ec2) nodes. It is a commonly used benchmark data for anomaly detection based on logs. It contains a total of 11,175,629 original log messages, and corresponding labels were assigned to 575,061 sessions to indicate their normal and abnormal states. The BGL dataset was collected from the BlueGene / L supercomputer system of Lawrence Livermore National Laboratory (LLNL) in Livermore, California. It contains a total of 4,747,963 original log messages, and each log was marked as normal or abnormal. In the experiment, the top 80% were taken as training data according to the timestamp information of the logs from top to bottom, and the remaining 20% were used as test data. Among them, the sequences of the HDFS dataset were divided by Block_id, and the sequences of the BGL dataset were framed by a sliding window of size 20. Since the labels of the HDFS and BGL datasets were manually marked, these marks were used as the factual basis for evaluation.

[0124] Table 1 Details of the datasets

[0125]

[0126] (3) Baseline methods

[0127] This invention selects PCA, SVM, DeepLog, LogRobust, and NeuralLog as the baseline methods for comparative experiments.

[0128] PCA statistically analyzes the event states and occurrence times expressed by the logs by analyzing the source code, and uses the state proportion vector and event count vector as the input of the model for unsupervised anomaly detection; SVM vectorizes the log sequence through the quantity and distribution of various log levels in the sliding window and performs anomaly detection through supervised training; DeepLog adopts an encoding method of log template indexing and uses LSTM to learn the sequential relationship between normal logs to detect anomalies; LogRobust extracts word vectors through FastText, generates log template vectors by aggregating word vectors through TF-IDF, and inputs them into an attention-based Bi-LSTM model to detect anomalies; NeuralLog extracts word vectors through Bert, obtains template vectors by calculating the average value of word vectors, and inputs them into a Transformer-based model for anomaly detection.

[0129] (4) Evaluation metrics

[0130] Anomaly detection is a binary classification problem. The present invention uses widely used metrics, namely accuracy, recall, and F1-score, to evaluate the accuracy of LogT and each baseline method in anomaly detection.

[0131] Accuracy is the percentage of truly anomalous log sequences among all log sequences determined to be anomalous by the model.

[0132] Recall represents the percentage of anomalous log sequences correctly identified as anomalous by the model among all anomalous log sequences.

[0133] F1-score represents the harmonic mean of accuracy and recall.

[0134] Among them, TP is the number of anomalous log sequences correctly detected by the model. FP is the number of normal log sequences misidentified as anomalous by the model. FN is the number of anomalous log sequences not detected by the model.

[0135] (5) Experimental parameter settings

[0136] For the one-dimensional convolutional layer, according to the experience of applying convolutional neural networks in the field of natural language processing [14 - 15], the present invention uses three different sizes of convolutional kernels, namely 1, 2, and 3, and the number of convolutional kernels of each size is set to 100 to ensure that the convolutional neural network can obtain information at different scales of the word itself and its context; for the different numbers of attention heads in the self-attention mechanism, according to the results of multiple experiments (as Figure 5 , 6 shown), considering both the model complexity and accuracy, the number of attention heads for both the HDFS dataset and the BGL dataset is set to 4.

[0137] (6) Comparative analysis of experimental result sets

[0138] The present invention conducts experimental comparisons between LogT and six baseline methods to verify the advantages of the present invention in terms of the accuracy and robustness of log sequence anomaly detection.

[0139] 1) Accuracy

[0140] Figure 7 and Figure 8 respectively show the comparison results of the present method and six baseline methods in terms of accuracy on the HDFS and BGL datasets.

[0141] From Figure 7 and Figure 8It can be seen that the experimental results on the HDFS and BGL datasets show that LogT has the highest accuracy among the six methods, and the F1-score on both the HDGS dataset and the BGL dataset is 0.98. The PCA method obtains the statistical features of log data through analyzing the source code for anomaly detection, and cannot be used as a general anomaly detection method, resulting in poor performance on both datasets; although the SVM has a relatively high accuracy of 0.98 on the BGL dataset, its low recall rate affects the detection performance and is not ideal on the HDFS dataset, because it only considers the statistical features of log data and does not take into account the order information of the logs; DeepLog shows a relatively high accuracy of 0.96 on the HDFS dataset, but its recall rate is low and its performance is poor on the BGL dataset, because it uses the method of log template indexing and cannot obtain the semantic information of the logs; LogRobust achieves a relatively high recall rate on both datasets, but its high recall rate comes at the cost of low accuracy, and the low accuracy means that it cannot detect more anomalies. NeuralLog uses the models of BERT and Transformer for anomaly detection, and compared with the above methods, it has improved accuracy and recall rate on both datasets; LogT uses the time interval-aware self-attention mechanism to utilize the time interval information between logs during the system operation, obtaining higher accuracy and recall rate on the HDFS dataset and higher recall rate on the BGL dataset, enabling LogT to detect more anomalies and avoid false alarms, and improving the accuracy of anomaly detection.

[0142] 2) Robustness

[0143] With the upgrade of the system or service, developers often insert or delete some words in the log statements during the process of updating the system source code. The update of log statements will affect the accuracy of anomaly detection. Therefore, it is particularly important to improve the robustness of the model so that the model can cope with the update of log templates. To compare the robustness of LogT with other baseline methods, the present invention makes certain modifications to the original HDFS dataset according to the log update rules proposed by Zhang et al. [6]. The modified log statements will not significantly change the semantics of the original log statements, so the corresponding anomaly label status is not affected. Specific examples of log updates are as Figure 9 shown. Anomaly detection is performed again on the HDFS dataset updated at different ratios, and the comparison results of the F1 scores of different methods are as Figure 10 shown.

[0144] From Figure 9It can be seen that when the update ratio of the log reaches 5%, the F1 score of DeepLog begins to decrease significantly. When the update ratio of the log reaches 15%, the F1 scores of SVM, PCA, and NeuralLog also decrease significantly. When the update ratio of the log reaches 25%, LogRobust also shows a relatively obvious decrease. However, the F1 score of logT is little affected. Even when the update ratio of the log reaches 30%, it can still maintain a high F1 score. It can be concluded that the LogT method proposed in the present invention uses Bert and CNN to obtain log template vectors, enabling the model to better learn the semantic information and context features of different words, being able to adapt to the situation where words in the log are updated, and having better robustness compared with other baseline methods.

[0145] (7) Conclusion

[0146] Software logs are widely used in various reliability assurance tasks. Most methods perform anomaly detection by analyzing the correlation information between the logs contained in the log sequence. The present invention proposes a log sequence anomaly detection method (logT) based on a time interval-aware self-attention mechanism, constructs a hierarchical network structure composed of CNN and Transformer to obtain log features at different levels, uses CNN to obtain multi-scale features of words, and more carefully considers the correlation information between words in the log statement, enabling the model to adapt to the situation of log statement updates and improving the robustness of anomaly detection. At the same time, in the process of extracting the features of the log sequence, not only the order information of the logs but also the time information contained in the logs is considered. Using the time interval-aware self-attention mechanism, the time interval between the logs and the semantic information are input into the self-attention mechanism together to obtain the correlation information between the logs, optimizing the expression ability of the extracted features and improving the effect of anomaly detection. Experimental evaluations on large-scale system log datasets show that LogT achieves better anomaly detection results compared with current mainstream methods. In future work, the present invention will further explore feature extraction methods for log sequence anomaly detection, consider integrating more valuable information in the logs into the anomaly detection model, be able to more effectively distinguish the differences between normal log sequences and abnormal log sequences, and enhance the performance of anomaly detection. And test the efficiency of this method on more log datasets to verify the universality of this method.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A log sequence anomaly detection method based on a time interval-aware self-attention mechanism, characterized in that Including: Obtain a log set for training, and perform data parsing on the log set for training to generate a log template sequence and a timestamp sequence; Calculate the relative time intervals between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix; Input the log template sequence and the time interval matrix into an anomaly detection model based on the time interval-aware self-attention mechanism for supervised training of the entire model, including: Input the log template sequence into a BERT model for word vector processing to obtain the word vectors of the log template sequence; Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors, including: Input the log template vector and the time interval matrix into the Transformer encoder to obtain the log sequence vector, perform positional encoding on the log template vector, and use sine and cosine functions to generate vectors for log events at each position in the log sequence As shown in formula (2): where \(t = 1, 2, \ldots, y\), representing different dimensions of the vector, and PE i is added to the log template vector \(e\) at position \(i\) i so that the model can learn the relative position information of each log event Input the log template vector sequence with added location information into the time interval-aware self-attention layer. Add the time interval information of log events in the attention calculation, and expand the time interval matrix by one dimension to become and multiply it with two parameter matrices to obtain the key-value matrix of the time interval and the value matrix The attention calculation method for the vector e i in the sequence is shown in formula (3): Among them h is the h-th attention head, D q = D k = D v = y / H, where H is the number of attention heads, and each attention head uses different learnable parameter matrices during calculation. The operation results of each head are concatenated and multiplied by the parameter matrix to obtain a new log template vector representation, as shown in formula (4): z i = Concat(A1,A2...,A H )W O (4); Input the log template vectors after attention calculation into a feed-forward fully connected layer for two linear transformations, as shown in formula (5), to obtain the final log sequence vectors: r i = max(0, z i ·W1 + b1)W2 + b2 (5) Among them The final log sequence vector is represented as R = {r1, r2,... r n}, r i ∈R y ; Input the log sequence vectors into a classifier based on a fully connected neural network to output a detection result, calculate the loss based on the detection result, and update the model parameters; Obtain a log set to be detected, and input the log set to be detected into the trained log sequence anomaly detection model based on the time interval-aware self-attention mechanism for anomaly detection.

2. The log sequence anomaly detection method based on time interval-aware self-attention mechanism according to claim 1, characterized in that Calculate the relative time intervals between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix, including: For a timestamp sequence \(T = \{t_1,t_2,...t\}\), calculate the time difference \(t\) n and take the absolute value to represent the relative time interval between the log at the \(i\)-th position and the log at the \(j\)-th position. Then standardize it ij where \(\mu\) is the mean of all time interval data in the training set and \(\sigma\) is the standard deviation, to obtain a new time interval matrix ​ 3. The method for detecting anomalies in log sequences based on a time interval-aware self-attention mechanism according to claim 1, wherein, Input the log template sequence into a BERT model to extract word vectors to obtain the word vectors of the log template sequence, including: Regarding the log template as a sentence in natural language, for a sequence of log templates X = {x1, x2,... x n}, where n is the length of the sequence, and x i represents the i-th log template, which is split into a sequence of words and represented as x i = {w1, w2,... w m}, where m represents the sentence length of the log template; Encode the word sequence through the BERT language model, map each word to a d-dimensional vector, and the log template is represented as a sequence of word vectors z i ={v1, v2,... v m}, where 4. The method for detecting abnormal log sequences based on a time interval-aware self-attention mechanism according to claim 3, wherein, Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors, including: For the log template word vector sequence z i , assume that v k:k+j represents all word vectors from v k to v k+j . The word vectors are input into the convolutional layer, and the convolutional layer contains multiple convolutional kernels. The operation process of each convolutional kernel is shown in formula (1): c i = f(W·v k:k+h-1 + b)(1) where h is the height of the convolution kernel, and d is the width of the convolution kernel, is the bias term, and f is the non-linear activation function ReLU; Each convolutional kernel can obtain a new feature map C = {c1, c2,... c m-h+1} by performing operations on all word vectors in the log template, and the maximum value is obtained through the max pooling layer as the eigenvalue obtained by the operation of this convolutional kernel; Connect the eigenvalues generated by all convolution kernels to obtain a log template vector where y is the number of convolution kernels.

5. The anomaly detection method for log sequences based on the time interval-aware self-attention mechanism according to claim 1, wherein Input the log sequence vectors into a classifier based on a fully connected neural network to obtain the classification result of the log sequence and perform supervised training on the entire model, calculate the loss based on the detection result, and update the model parameters, including: Input the log sequence vectors into a classifier based on a fully connected neural network to output a detection result, calculate the loss, and update the model parameters, as shown in formula (6): pre = Softmax(R·W s + b s ) (6) Among which W s , b s are the weights and bias terms of the fully connected neural network, and pre is the probability of normal and abnormal outputs of the model; Use a cross-entropy loss function to calculate the loss between the model detection result and the label provided in the dataset, as shown in formula (7): Update the parameters by backpropagation through an Adam optimizer. The parameters to be updated include those in the CNN, Transformer encoder, and fully connected neural network.

6. The log sequence anomaly detection method based on the time interval-aware self-attention mechanism according to claim 1, wherein The step of inputting the log set to be detected into the trained log sequence anomaly detection model based on the time interval-aware self-attention mechanism for anomaly detection includes: Obtain a log set to be detected, perform data parsing on the log set to generate a log template sequence and a timestamp sequence; calculate the relative time intervals between each log event in the sequence according to the timestamp sequence to obtain a time interval matrix; Input the log template sequence into a BERT model to extract word vectors to obtain the word vectors of the log template sequence; Extract template vectors from the word vectors of the log template sequence through a one-dimensional convolutional neural network to obtain log template vectors; determine whether there are unprocessed log templates in the log template sequence, and if so, perform word vector extraction and template vector extraction on the next log template until all log templates are processed; Input the log template vectors and the time interval matrix into a Transformer encoder based on a time interval-aware self-attention mechanism to obtain log sequence vectors; Input the log sequence vectors into a classifier based on a fully connected neural network to output the detection results.

Citation Information

Patent Citations

  • Information pushing method and device, electronic equipment and storage medium

    CN113505292A

  • Log sequence anomaly detection method based on out-of-stream regularization

    CN114416479A