A log anomaly detection method based on parsing optimization and temporal convolutional network
Through the log anomaly detection method based on BERT and TCN, combined with log parsing tree and window technology, the semantic, sequence and quantity characteristics of the log are extracted, which solves the problem of low detection efficiency in the existing technology and realizes efficient and accurate log anomaly detection.
Patent Information
- Application Number
- CN202211711499.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing log anomaly detection methods find it difficult to effectively extract the semantic, sequence, and quantity features of logs, and RNN-based models cannot achieve parallel processing, resulting in low detection efficiency and insufficient accuracy.
A BERT-based pre-training model is used to extract the semantic features of logs, and window technology is used to obtain the sequence features of logs. Combined with the temporal convolutional network (TCN) and self-attention mechanism, semantic features and quantitative features are integrated. The log parsing process is optimized through the log parsing tree structure to improve detection efficiency and accuracy.
Through comprehensive feature extraction and parallel processing, the efficiency and accuracy of log anomaly detection are improved, the matching efficiency problem in the log parsing process is solved, and more efficient anomaly detection is achieved.
Smart Images

Figure CN115828180B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network anomaly detection, and in particular to a log anomaly detection method based on parsing optimization and temporal convolutional networks. Background Art
[0002] With the rapid development of internet technology and the increasing trend toward service-oriented network applications and cloud-based services, the scale of data is growing exponentially. We are now in the era of big data. "Big data" generally refers to data that cannot be acquired, managed, and processed by traditional tools within acceptable computing time. Current large-scale online service systems, such as those at Microsoft and Google, typically consist of hundreds of distributed components and support millions of users. The amount of data generated by these services is also growing exponentially. Simultaneously, these services also generate logs that record system status and application service execution information at key points. In many scenarios, the number and types of system operation anomalies caused by illegal user input or changes in the external environment are also increasing. Therefore, system logs play a vital role in identifying these anomalies and are one of the most valuable data sources for monitoring system status and detecting anomalies.
[0003] Traditional log anomaly detection relies on manual searches for suspicious logs that may indicate system issues. For example, software developers detect anomalies through simple rule matching or keyword searches, such as "error" and "exception." However, with the growth of data volumes and the increasing complexity of application services, relying solely on manual regular expression matching or string searches is unrealistic. Therefore, log anomaly detection methods based on machine learning and deep learning are becoming increasingly popular due to their high detection efficiency and accuracy.
[0004] Early log anomaly detection algorithms typically mined multidimensional data features, such as time series features, from logs and relied on machine learning algorithms to automatically detect anomalies. However, these methods had limited ability to capture features, making it difficult to analyze and extract deep, hidden relationships. In recent years, the development of deep learning has brought breakthroughs in anomaly detection, and an increasing number of deep learning methods are being used in the field of log analysis to achieve automation and improve accuracy. Many models based on recursive neural networks (RNNs), such as long short-term memory networks (LSTMs), have been widely used in anomaly detection due to their excellent performance in expressing sequential features.
[0005] Although most anomaly detection methods are effective in experiments, they remain insufficiently robust in practice. This is because real logs are volatile, meaning new yet similar log sequences frequently appear, posing a challenge for log parsing. Raw log messages are typically unstructured data, consisting of constants (unchanging parameters) and variables (variable parameters). The variable parameters often hinder automated log analysis. Therefore, following the practice of log-based anomaly detection, we need to quickly and efficiently parse logs into structured log events, or log templates, through log parsing. Searching and comparing are core and time-consuming processes, especially as log data grows in size. Therefore, it's important to reduce the number of message searches and comparisons. For example, the popular log parsing method Spell maintains a global LCSMap. When a specific log message arrives, the search process requires looping through it, significantly impacting the efficiency of the overall log parsing process.
[0006] At the same time, raw log messages contain semantic, quantitative, and sequence features. Existing methods only extract one or two of these aspects for feature representation, making it difficult to fully and comprehensively extract all of this information. Furthermore, existing log anomaly detection technologies, such as the classic DeepLog, LogRobust, and LogAnomaly, mostly rely on RNN models for final anomaly detection. However, RNN models require contextual dependencies and cannot implement parallel processing. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this paper proposes a log anomaly detection method based on parsing optimization and a temporal convolutional network. This method extracts semantic features from logs using a BERT pre-trained model, acquires sequential features using windowing techniques, and captures quantitative features based on the term frequency-inverse document frequency (TF-IDF) method. This method provides a more comprehensive perspective for log feature extraction, improving the robustness and efficiency of log parsing. Furthermore, the detection model incorporates a temporal convolutional network (TCN) and a self-attention mechanism, enhancing the accuracy of anomaly detection.
[0008] In order to achieve the above object, the present invention provides the following technical solutions:
[0009] A log anomaly detection method based on parsing optimization and temporal convolutional network includes the following steps:
[0010] S1. Log parsing optimization: Use the log parsing tree structure to parse the original log data and obtain the log template;
[0011] S2. Feature extraction: Use windowing technology to group logs and obtain the sequence features of log templates. Use the BERT pre-trained model to extract semantic features from log templates and use the TF-IDF representation layer to capture the quantitative features of log templates.
[0012] S3. Anomaly detection: Use a temporal convolutional network to capture the time series of sequence features, fuse semantic features and quantitative features, then use the self-attention mechanism to assign different weights to these log features, and finally use a fully connected layer to combine all features to obtain the final prediction output.
[0013] Furthermore, the specific process of step S1 is:
[0014] S11. Data preprocessing: Define regular expressions in advance based on the characteristics of the log data to replace variables and special symbols in the original log data. Use wildcard symbols to replace variable-type data, and merge multiple consecutive wildcard symbols to obtain preprocessed logs.
[0015] S12. Log parsing tree construction: The log parsing tree starts as an empty tree with only one root node without any data. Other nodes of the tree are continuously generated by continuously inserting the input preprocessed log;
[0016] S13. Log parsing: perform similarity matching on the string prefix and each character at the same position to obtain a log template.
[0017] Furthermore, the insertion rule of step S12 is: divide the log parsing tree into four layers, the root node is located in the first layer of the log parsing tree, and does not store any log-related information; the second layer of the log parsing tree is a length-based division layer, which is used to distinguish the length of the pre-processed log message; the third layer of the log parsing tree is a string prefix-based division layer, which is used to store the string sequence prefix of each log message; the fourth layer of the log parsing tree is a leaf node, which is used to store the log groups divided by the filtering conditions of the previous layers.
[0018] Furthermore, step S13 calculates the similarity of the string prefixes according to the edit distance. The edit distance represents the similarity between two strings by calculating the minimum number of operations required to convert one string into another.
[0019] Furthermore, the formula for similarity matching of each character at the same position in step S13 is:
[0020]
[0021] Where seq1(j) represents the jth character of the string seq1, and the eq function is defined as follows:
[0022]
[0023] w1, w2 are the characters at the same position in string seq1 and string seq2;
[0024] Calculate the similarity values between the log message and all templates in the leaf node, and then select the log template with the largest similarity that exceeds the predefined threshold as the log template; if no such template is found, the current leaf node will add this log message as a new log template.
[0025] Furthermore, step S2 uses three window technologies: fixed window, sliding window, and session window to group logs. Fixed window and sliding window are based on timestamps. Fixed window has a fixed time span or duration, and all logs in this time window are classified as log sequences; sliding window consists of two attributes: window size and step size; session window is based on identifier. In each sequence, all logs share a common identifier, indicating that they come from the same task operation.
[0026] Furthermore, step S2 uses a TF-IDF-based feature representation method to vectorize the obtained log template sequence to obtain the quantitative features of the log template.
[0027] Furthermore, step S2 uses word embedding to vectorize the obtained log template sequence, and obtains the semantic features of the log template by calculating the relationship between words and mining the connections between words.
[0028] Furthermore, dilated convolution is used in each unit block of the temporal convolutional network in step S3, and the dilated convolution output of a specific s element is defined as:
[0029]
[0030] in is the input sequence, is a filter of size k, d is the dilation rate, x s-d×i Represents the sequence element corresponding to the multiplication of the convolution kernel.
[0031] Furthermore, in step S3, each unit block of the temporal convolutional network adds a batch normalization layer, an activation layer, and a Dropout layer on the basis of the dilated convolutional layer, and the activation layer uses the ReLU function.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] (1) The log anomaly detection method based on parsing optimization and temporal convolutional network proposed in the present invention is a log parsing method based on a fixed-depth tree structure. It optimizes the log parsing process and solves the efficiency problem of matching according to logs in the log parsing stage to cope with the problem of variable parameters in the logs and the unstable log output. It optimizes the search process and similarity calculation algorithm to improve the efficiency of log matching.
[0034] (2) The log anomaly detection method based on parsing optimization and temporal convolutional network proposed in the present invention extracts semantic features from log templates based on the BERT pre-training model, obtains the sequence features of the log using window technology, captures the quantitative feature information of the log template based on TF-IDF, and fully extracts the features in the log parsing, providing a more comprehensive perspective for the log feature extraction link, and can effectively extract the semantic features, sequence features and quantitative features therein, enriching the expression ability of the model, making the final anomaly detection more efficient and comprehensive.
[0035] (3) The present invention processes sequence features based on the temporal convolutional network (TCN), which can achieve parallelization and thus improve the operating efficiency of the model.
[0036] (4) The present invention uses the self-attention mechanism to complete the learning of three types of features: semantic features, sequence features, and quantitative features. Considering the intrinsic connections between each other, different weights are assigned to different features, thereby improving the accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0038] Figure 1 Flowchart of a log anomaly detection method based on parsing optimization and temporal convolutional network provided by an embodiment of the present invention.
[0039] Figure 2 A flowchart for constructing a parse tree provided by an embodiment of the present invention.
[0040] Figure 3 Schematic diagram of three sliding window technologies provided by embodiments of the present invention.
[0041] Figure 4 This is a diagram showing the overall framework of the anomaly detection model based on a temporal convolutional network provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The log anomaly detection method based on parsing optimization and time series convolutional network proposed in this paper has made some improvements in the log parsing stage based on the existing technology to achieve robustness and accuracy of the log parsing stage. In addition, the semantic information of the log sequence is modeled using the BERT pre-training layer, and the sequence information of the log is extracted using the window technology. In addition, the present invention is based on
[0043] TF-IDF captures the quantitative feature information of the log template and introduces the temporal convolutional network (TCN) with self-attention mechanism for prediction in the anomaly detection stage to capture the patterns of normal sequences.
[0044] In order to better understand the present technical solution, the method of the present invention is described in detail below with reference to the accompanying drawings.
[0045] The log anomaly detection method based on parsing optimization and temporal convolutional network proposed in this paper includes optimized log parsing, feature representation and comprehensive anomaly detection. The overall process is as follows: Figure 1 shown.
[0046] ①(1) Optimize the log parsing stage
[0047] As log formats increase, some logs contain variable parameters. If unprocessed, the results of subsequent segmentation based on log length will not match expectations. Therefore, during data preprocessing, we handle variable parameters by merging consecutive wildcards to enhance the robustness of the log parsing process.
[0048] Specifically, during the data preprocessing phase, raw log data often contains common variable information such as IP, numbers, and time. We need to define regular expressions in advance based on the characteristics of the log data to replace some variables and special symbols in the raw log data. Variable type data will be replaced with the wildcard "*", which can effectively improve the accuracy of the mined log template. However, with the increase in log formats, variable parameters may sometimes appear, such as "Receive From Node:node1 node2" and "Receive From Node:node1"
[0049] Due to the different lengths, they will be parsed into two different log templates, but in fact both belong to the same template "Receive From node*". The present invention can effectively solve this problem of variable parameters by merging multiple consecutive wildcard symbols during the string replacement process of the preprocessing flow.
[0050] After log preprocessing, the construction of the log parsing tree is the key to this log parsing method. It is necessary to ensure the effectiveness and robustness of log parsing as much as possible. The log parsing tree is initially an empty tree with only one root node without any data. Other nodes of the tree are continuously generated by inserting the input preprocessed log in a certain rule. The specific construction process is as follows: Figure 2 shown.
[0051] like Figure 2 As shown, the depth of the leaf node is a fixed length of 4, which is used to limit the number of nodes that need to be traversed during the log parsing process. In this log parsing tree, the root node is at the first level of the log parsing tree, but does not store any log-related information. The second-level nodes of the log parsing tree are length-based partitioning layers, which are used to distinguish the lengths of pre-processed log messages. First, the logs are filtered by length, and log messages of different lengths flow to different third-level nodes, which can greatly reduce the number of subsequent matches between log templates. In order to further reduce the number of matches between pre-processed logs and log templates, the third level of the log parsing tree is a string prefix-based partitioning layer, which is used to store the string sequence prefix of each log message. The fourth level of the log parsing tree is a leaf node, which is used to store log groups divided by the filtering conditions of the previous levels. Therefore, each node in this layer usually contains multiple log template information, including the log template and log id belonging to the template.
[0052] In the log template matching and search process, search and comparison are the core and time-consuming processes. This invention optimizes the search process and similarity calculation algorithm. As can be seen from the previous log parsing tree construction steps, the log parsing process involves two similarity-based matches: one for matching the prefix string, and the other for matching every character at the same position in the log template. Based on the characteristics of these two stages, we designed different similarity matching algorithms.
[0053] First, in the previous layer, we search by the prefix of the log string. This uses a matching algorithm based on edit distance, which significantly reduces the number of matching characters. The prefix string consists of the first letter of each word in the log message, so it can represent the specific structure of the log and is very short. Therefore, we calculate the similarity between the two based on the edit distance. The edit distance indicates the similarity between two strings by calculating the minimum number of operations (additions, deletions, and substitutions) required to transform one string into the other.
[0054] Then, at the next level, a second similarity calculation is performed based on whether each position in the string is identical to complete the final search. When the preprocessed log message falls into a leaf node of the log parsing tree, it needs to match all log templates in the leaf node until it finds the log template with the highest similarity that exceeds the set threshold. Therefore, to improve efficiency, our similarity function is defined as follows:
[0055]
[0056] Where seq1(j) represents the jth character of the string seq1, and the eq function is defined as follows:
[0057]
[0058] w1, w2 are the characters at the same position in string seq1 and string seq2;
[0059] Calculate the similarity between the log message and all templates in the leaf node, and then select the log template with the largest similarity that exceeds the predefined threshold as the log template. If no such template is found, the current leaf node will add this log message as a new log template.
[0060] (2) Feature representation stage
[0061] For the log template after log parsing, we need to extract features from it and convert it into a vector form that can be recognized by the anomaly detection model. In order to further extract sequence features, we first need to group the logs and divide the original logs into different log sequences, where each group represents a log sequence. For this, we use window technology, which includes fixed window, sliding window and session window. The three window technologies are as follows: Figure 3 shown.
[0062] Fixed windows and sliding windows are based on timestamps. Fixed windows have a fixed time span or duration Δt. The window size is a constant value, such as one second or one minute. All logs in this time window are classified as log sequences. Unlike fixed windows, sliding windows consist of two properties: the window size and the step size, such as an hourly window that slides every five minutes. In addition, session windows are based on identifiers (a or b) rather than timestamps. In each sequence, all logs share a common identifier, indicating that they come from the same task operation.
[0063] After using windowing techniques to group logs, we next need to obtain a vectorized representation of the log template sequence. There are two commonly used log vectorization methods: one is TF-IDF-based feature representation, and the other is word embedding. This paper uses TF-IDF to calculate the quantitative relationship of log templates in the log sequence and then determine the importance of each template. The quantitative relationship characteristics of the templates are incorporated into the feature representation of the final model, providing a more comprehensive perspective on the overall model's expressive power. On the other hand, by calculating the relationship between words and mining the connections between words, word embedding can express the semantic information of the original text. In the field of word embedding, BERT is a method for pre-training language representation. This paper uses BERT, a general language understanding model trained on the large text corpus of Wikipedia, to directly obtain the semantic features of the text, and then uses this feature in our subsequent anomaly detection stage. BERT is the first unsupervised, deeply bidirectional pre-trained natural language processing system. Its emergence has enabled the widespread application of text representation methods based on pre-trained models, providing a more efficient way to obtain semantic representations of log templates.
[0064] After log parsing and grouping, logs are divided into a series of log templates, which contain log template sequence information, log template semantic information, and log template quantity information. Existing methods only extract one or two of these aspects for feature representation, failing to fully utilize all three aspects of information. Based on the log sequences converted using windowing technology, this paper uses the BERT pre-trained model to extract semantic information from the log templates, utilizes windowing technology to obtain log sequence information, and captures log template quantity feature information based on TF-IDF, providing a more comprehensive perspective for log feature extraction.
[0065] (3) Comprehensive anomaly detection stage
[0066] Most existing log anomaly detection technologies rely on RNN models for final anomaly detection, such as the classic DeepLog, LogRobust, and LogAnomaly. RNN models require contextual dependencies and cannot be processed in parallel. We propose a temporal convolutional network to process sequential information, which enables parallelization and improves model efficiency.
[0067] Specifically, the hybrid log anomaly detection module of the present invention includes three input parts, including sequence features, semantic features and quantitative features, which provides a more comprehensive perspective for the final anomaly detection model. At the same time, it uses the advanced implementation of the TCN layer to capture time series, realize parallelization and thus improve the operating efficiency of the model. In addition, it contains a pre-trained module to extract semantic information, and a TF-IDF-based layer to process the quantitative information of the log template. It also uses the self-attention mechanism to complete the learning of features and assign different weights to different features. The entire anomaly detection process can achieve high accuracy and high operating efficiency. The overall framework is as follows Figure 4 shown.
[0068] TCN uses dilated convolution to expand the model's receptive field. By stacking convolutional structures, TCN can even capture longer-term dependencies. Compared to the general convolutional neural network method of expanding the receptive field by stacking pooling layers, this method has a lower network depth and is therefore less prone to information attenuation. In each TCN unit block, we use dilated convolution. The dilated convolution output of a specific s element can be defined as:
[0069]
[0070] in is the input sequence, is a filter of size k, d is the dilation rate, x s-d×i Represents the sequence element corresponding to the multiplication of the convolution kernel.
[0071] Furthermore, the present invention adds batch normalization, activation, and dropout layers to the convolutional layers. The batch normalization layer effectively speeds up computation, while the dropout layer randomly inactivates neurons to prevent overfitting. Furthermore, the activation layer uses the ReLU function, introducing nonlinear factors to enhance the model's expressiveness.
[0072] This paper does not directly use the output of the TCN layer as the final prediction output. Instead, it fuses the TCN module output with semantic information and template quantity features to introduce more log features. Specifically, this paper uses the input of the TCN layer as a high-level representation of the time series features in the log. On top of the time series features, it combines the semantic representation features based on BERT and the template quantity features based on TF-IDF, thus providing a more comprehensive perspective on the log feature representation.
[0073] Since different log features have different effects on the final classification results, this paper introduces a self-attention mechanism to assign different weights to these log features. The self-attention mechanism can consider the intrinsic connections between each other and calculate the weights of different features. The specific weight calculation formula is as follows:
[0074]
[0075] Among them, W is the fused input matrix, is a scaling factor used to prevent the inner product from being too large.
[0076] Finally, the present invention uses a fully connected layer to combine all features to obtain the final prediction output. The output result is a binary classification, which is used to determine whether the input log sequence is abnormal.
[0077] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A log anomaly detection method based on parsing optimization and temporal convolutional network, characterized in that: The following steps are involved: S1. Log parsing optimization: Use the log parsing tree structure to parse the original log data and obtain the log template. The specific process of step S1 is as follows: S11. Data preprocessing: Define regular expressions in advance based on the characteristics of the log data to replace variables and special symbols in the original log data. Use wildcard symbols to replace variable-type data, and merge multiple consecutive wildcard symbols to obtain preprocessed logs. S12. Log Parsing Tree Construction: The log parsing tree starts as an empty tree with only one root node that does not contain any data. Other nodes of the tree are continuously generated by continuously inserting the input pre-processed log. The insertion rule is: the log parsing tree is divided into four layers. The root node is located in the first layer of the log parsing tree and does not store any log-related information. The second layer of the log parsing tree is a length-based partitioning layer used to distinguish the length of the pre-processed log messages. The third layer of the log parsing tree is a partitioning layer based on string prefixes, which is used to store the string sequence prefix of each log message. The fourth layer of the log parsing tree is a leaf node, which is used to store log groups divided by the filtering conditions of the previous layers. S13, log parsing: perform similarity matching on the string prefix and each character at the same position to obtain a log template; the formula for similarity matching on each character at the same position in step S13 is: Where seq1(j) represents the jth character of the string seq1, and the eq function is defined as follows: w1, w2 are the characters at the same position in string seq1 and string seq2; Calculate the similarity between the log message and all templates in the leaf node, and then select the log template with the largest similarity that exceeds the predefined threshold as the log template; if no such template is found, the current leaf node will add this log message as a new log template; S2. Feature Extraction: We group logs using fixed, sliding, and session window techniques to obtain sequence features of log templates. We extract semantic features from log templates using the BERT pre-trained model and capture quantitative features of log templates using the TF-IDF representation layer. S3. Anomaly detection: Use a temporal convolutional network to capture the time series of sequence features, fuse semantic features and quantitative features, then use the self-attention mechanism to assign different weights to these log features, and finally use a fully connected layer to combine all features to obtain the final prediction output.
2. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 1 is characterized in that: Step S13 calculates the similarity of the string prefixes according to the edit distance. The edit distance represents the similarity between two strings by calculating the minimum number of operations required to convert one string into another.
3. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 1 is characterized in that: Step S2 uses three window technologies: fixed window, sliding window, and session window to group logs. Fixed window and sliding window are based on timestamps. Fixed window has a fixed time span or duration. All logs in this time window are classified as log sequences. Sliding window consists of two attributes: window size and step size. Session window is based on identifier. In each sequence, all logs share a common identifier, indicating that they come from the same task operation.
4. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 1 is characterized in that: In step S2, the acquired log template sequence is vectorized using a TF-IDF-based feature representation method to obtain the quantitative features of the log template.
5. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 1 is characterized in that: Step S2 uses word embedding to vectorize the obtained log template sequence, and obtains the semantic features of the log template by calculating the relationship between words and mining the connections between words.
6. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 5 is characterized in that: In step S3, dilated convolution is used in each unit block of the temporal convolutional network. The dilated convolution output of a specific s element is defined as: in is the input sequence, is a filter of size k, d is the dilation rate, x s-d×i Represents the sequence element corresponding to the multiplication of the convolution kernel.
7. The log anomaly detection method based on analytical optimization and temporal convolutional network according to claim 6 is characterized in that: In step S3, each unit block of the temporal convolutional network adds a batch normalization layer, an activation layer, and a Dropout layer on the basis of the dilated convolution layer, and the activation layer uses the ReLU function.