System Log Anomaly Detection Method Based on Semantic Flow Graph Mining
By preprocessing log statements and clustering the semantic flow graph, and using graph convolutional neural networks for feature extraction, the anomaly detection challenges brought about by noise and system changes in unstructured logs are solved, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202310873970.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-07-17
AI Technical Summary
The prior art is difficult to effectively remove noise and handle log statement changes caused by system changes in massive unstructured logs, affecting the accuracy and efficiency of abnormal detection.
By preprocessing log statements, word vectors and sentence vectors are generated using Word2Vec and a bidirectional GRU network based on attention mechanism, semantic flow graphs are constructed using K-means clustering, and feature extraction and training are used for graph convolutional neural networks to achieve abnormal detection.
It improves the efficiency of abnormal detection, reduces the impact of log noise, solves the problem of log statement instability caused by system updates, and uses spatial structure information to improve the accuracy of detection.
Smart Images

Figure CN116910013B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and further relates to an anomaly detection method, specifically a system anomaly detection method based on log semantic information extraction, which can be used for anomaly detection and identification of mainstream computer systems. Background Art
[0002] System logs record the status information and running conditions of the system, including the anomaly information of the system. They are usually composed of static texts and variables and are valuable resources for understanding the system status. By analyzing the information contained in the system logs, the anomalies of the system can be analyzed, the system fault points can be located, and thus the security and reliability of the system can be improved. Log files are not the same as documents written in natural language. First, similar messages in the logs are constantly repeated because programs usually execute in a loop, resulting in repeated events, and most logs are generated by a limited number of log printing statements, that is, predefined functions in the code write formatted strings to the output and generate log messages. Second, some messages in the logs are highly correlated because the execution of system programs follows certain control flows and the components that generate logs are linked to each other. At present, the logs generated by the vast majority of systems are semi-structured or unstructured, and the log formats and types vary between different systems. Therefore, even though the logs contain information about important events of the system, it has become a difficult point to efficiently parse the logs and extract the event information in the logs for anomaly detection. In addition to the complex semi-structured log format that is difficult to parse and extract effective information, anomaly detection for system logs also faces the problems of huge log data volume and the influence of garbage data and noise in the logs.
[0003] Currently, the mainstream methods for anomaly detection of system logs mainly include the following steps: 1) Parse the system logs and extract the templates in the logs; 2) Construct the log templates into log sequences and extract the feature vectors in the log template sequences; 3) Use machine learning or deep learning methods for anomaly detection. However, in the face of a large amount of unstructured and semi-structured system logs, parsing the logs to extract log templates poses a challenge to the log parser. In addition, log parsing has a large time and space consumption in the anomaly detection system, resulting in low efficiency of the anomaly detection task. More importantly, it is difficult to remove the noise data existing in the system logs, which will directly affect the accuracy of anomaly detection. Therefore, how to overcome the low efficiency brought by log parsing and the influence of log noise on the accuracy of anomaly detection when performing anomaly detection on system logs has become an urgent problem to be solved. Summary of the Invention
[0004] The object of the present invention is to propose a system log anomaly detection method based on a semantic flow graph in view of the deficiencies of the above-mentioned existing technologies, and solve the problems such as the difficulty in removing log noise and the change of log statements caused by system changes in the anomaly detection task for massive unstructured logs. During the process of performing the log-based anomaly detection task, the generation of log noise mainly has the following reasons: 1) During the collection and transmission of logs, the log data is chaotic and missing due to transmission delay or data loss; 2) The system or application program repeatedly records event information; 3) During the log parsing process, the incorrect recognition of the log template is caused by the instability of the parser. In addition, the upgrade of the computer system during the delivery process causes the log printing statements in the code to change accordingly, and the resulting update of the log template will also increase the false alarm rate of anomaly detection. The present invention can extract the semantic vector of the log statement as the node feature of the semantic flow graph, and train the semantic flow graph through a graph convolutional neural network model, which can effectively solve the influence of log noise on anomaly detection and further improve the accuracy of anomaly detection by using the spatial structure information between log statements.
[0005] The idea of implementing the solution of the present invention is as follows: First, perform simple preprocessing on the original log statements, remove meaningless symbols and perform word segmentation; secondly, use Word2Vec to calculate in combination with the importance of words in the log statements to obtain the word vectors of the log statements, and then use a bidirectional gated recurrent unit (GRU) network based on the attention mechanism to calculate to obtain the sentence vector representation of the log statements; then use the K-means clustering method to cluster the log sentence vectors, and divide the log statements with higher similarity into the same class, which is regarded as the same node of the semantic flow graph. After converting the log sequence into a non-duplicate node sequence, construct a directed acyclic graph according to the sequence order of the nodes, which is called the semantic flow graph in the present invention. Finally, perform feature extraction and training on the semantic flow graph through a graph convolutional neural network to realize system anomaly detection based on the semantic flow graph.
[0006] The specific steps for the present invention to achieve the above object are as follows:
[0007] (1) Segment the original system logs into log statements, remove meaningless symbols, and retain combined words with special meanings to obtain the initially preprocessed system logs;
[0008] (2) Divide the initially preprocessed system logs into log sequences according to the session or window mechanism, and use the Word2Vec model to convert the words or phrases in the log sequence into word vectors, and let the word vector of the nth word in the mth log be v n m ;
[0009] (3) Calculate the term frequency-inverse document frequency TF-IDF of the words in the log statement. Among them, the TF-IDF of the nth word in the mth log is expressed as T mn ;
[0010] (4) Combine with T mn to obtain the vector representation W of the final log statement word according to the following formula mn :
[0011]
[0012] Among them, α represents the weight factor;
[0013] (5) Use W mn as the input of the bidirectional GRU model based on the attention mechanism to obtain a set of sentence vectors S = {s1, s2,..., s m ,..., s L} containing L log sequences. Among them, s m represents the mth log statement vector;
[0014] (6) Use the K-means clustering method to cluster the log statement vectors in the set S, match the log sequences according to the clustering results to obtain a node sequence, and construct a directed acyclic graph of nodes in the order of this node sequence. Finally, obtain the semantic flow graph G = (V, E) of the log sequence, where V represents the set of nodes and E represents the set of edges;
[0015] (7) Use a graph convolutional neural network to extract features and train the semantic flow graph G. By propagating and aggregating the node features in the semantic flow graph, map them to classification labels to obtain the system log anomaly detection result.
[0016] The present invention has the following advantages compared with the prior art:
[0017] First, since the present invention extracts the semantics of the original log statement and converts it into the form of a semantic flow graph for anomaly detection, only simple preprocessing of the original log is required, and there is no need to perform difficult and inefficient log parsing work on the original log statements, thus greatly improving the efficiency of anomaly detection.
[0018] Second, the present invention uses the K-means clustering method to perform clustering analysis on the log statement vectors, greatly reducing the impact of log noise on anomaly detection, and at the same time solving the problem of log statement instability caused by continuous system updates and iterations during system delivery.
[0019] Thirdly, the present invention uses the form of a semantic flow graph for anomaly detection. The graph structure can contain spatial structure information that sequences do not have. Therefore, the graph convolutional neural network model can extract the node and edge features and spatial information features in the semantic flow graph, and can also detect anomalies hidden in the spatial structure information. Description of the Drawings
[0020] Figure 1 is the implementation flowchart of the present invention;
[0021] Figure 2 is the schematic diagram of the semantic flow graph construction process in the present invention;
[0022] Figure 3 is the simplified example diagram of the semantic flow graph constructed by the present invention; Detailed Embodiments
[0023] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present invention will be further clearly and completely described below in conjunction with specific embodiments.
[0024] Embodiment 1: Referring to Figure 1 , a system anomaly detection method based on a log semantic flow graph proposed by the present invention specifically includes the following steps:
[0025] Step 1. Split the original system logs into log statements, and remove meaningless symbols, such as punctuation marks like semicolons and commas, as well as special symbols like #, @, and *; then tokenize the original log statements. For some specific combined words, such as PacketResponder, the present invention retains these combined words with specific meanings without tokenization, and obtains the initially preprocessed system logs.
[0026] Step 2. Divide the initially preprocessed system logs into log sequences according to the session or window mechanism to obtain a log sequence L with a scale of N. Use the Word2Vec model to convert the words or phrases in the log sequence into word vectors. Let the word vector of the nth word in the mth log be . In this embodiment, the skip-gram variant of the Word2Vec model is used to extract semantic vectors from the input log sequence. The semantic vector of the ith log in the log sequence L is where represents the semantic vector of the jth word in this log.
[0027] Step 3. Calculate the term frequency-inverse document frequency TF-IDF of the words in the log statement. Among them, the TF-IDF of the nth word in the mth log is represented as T mnSince the present invention only performs simple preprocessing on the original logs without log parsing and retains most of the words, in order to consider the importance of different words in the log statements, the present invention calculates TF-IDF for the words in the log statements.
[0028] The steps to calculate the term frequency-inverse document frequency TF-IDF of the words in the log statements are as follows:
[0029] (3.1) Form a document set D from the preprocessed log statements. For each log statement and a given word in the set, calculate the term frequency TF(w, d) of the word w in the log statement d;
[0030] (3.2) Calculate the inverse document frequency IDF(w, D) of the word w:
[0031]
[0032] where |D| represents the total number of documents in the document set D, and |{d ∈ D: w ∈ d}| represents the number of documents containing the word w;
[0033] (3.3) Calculate the TF-IDF value T of the word w in the log statement mn :
[0034] T mn = TF-IDF(w, d, D) = TF(w, d) × IDF(w, D).
[0035] Step 4. Combine with T mn to obtain the vector representation W of the final log statement words according to the following formula mn :
[0036]
[0037] where α represents the weight factor;
[0038] Step 5. Use W mn as the input of the bidirectional GRU model based on the attention mechanism to obtain a set of sentence vectors S = {s1, s2,..., s m ,..., s L} containing L log sequences, where s m represents the m-th log statement vector. The bidirectional GRU model consists of a forward GRU and a backward GRU. The forward GRU processes the sequence from front to back, and the backward GRU processes the sequence from back to front.
[0039] Step 6. After obtaining the statement vector representation of the log sequence, in order to construct a semantic flow graph, it is necessary to match the log statements with the nodes of the semantic flow graph. Therefore, the K-means clustering method is used to cluster the log statement vectors in the set S, and the log sequence is matched according to the clustering result to obtain a node sequence. A directed acyclic graph of nodes is constructed in the order of this node sequence, and finally the semantic flow graph G=(V, E) of the log sequence is obtained, where V represents the set of nodes and E represents the set of edges.
[0040] The implementation of clustering the log statement vectors in the set S using the K-means clustering method is as follows:
[0041] (6.1) Initialize the clustering center points. Let S a and S b represent the sentence vectors of two different log statements, which are respectively denoted as the first vector and the second vector. Calculate the Euclidean distance Distance(S a and S b ) between them according to the following formula: a ,S b ):
[0042]
[0043] where n is the vector dimension, and S z a , S z b respectively represent the element values of the first vector S a and the second vector S b on the z-th dimension;
[0044] (6.2) Iteratively update the clustering centers. In each iteration, assign each vector to the cluster to which the nearest clustering center belongs, and then update the clustering center of each cluster to the mean vector of all vectors in the cluster. Assume that C k represents the k-th cluster, N k represents the number of vectors in the k-th cluster, and u k represents the clustering center of the k-th cluster. Then the update formula for the clustering center of the k-th cluster is as follows:
[0045]
[0046] where x e represents the e-th vector, which is used as a sample point. When the clustering center no longer changes or reaches the maximum number of iterations, it is considered that the algorithm converges and the final clustering result is obtained.
[0047] In this step, the semantic flow graph G=(V, E) of the log sequence is obtained. Specifically, log statements with the same log template in the clustering cluster are regarded as the same node in the semantic flow graph, and the node type is matched with the log entry in the original log sequence. Different nodes are connected according to the structure of the log sequence to construct a directed acyclic semantic flow graph, and the sentence vector of the log statement is embedded into the semantic flow graph as the node feature; where, represents the node set of G, represents the p-th node v in G p pointing to the set of directed edges of the q-th v q ; is a positive integer.
[0048] By calculating the distance and updating the clustering center, the K-means algorithm can iteratively optimize the clustering result. When the clustering center no longer changes or reaches the maximum number of iterations, it is considered that the algorithm converges, and the final clustering result is obtained. Each log statement vector will belong to a clustering cluster. Log statements in each clustering cluster have the same log template. We regard log statements with the same log template in these clustering clusters as the same node in the semantic flow graph, and match the node type with the log entry in the original log sequence. Different nodes are connected according to the structure of the log sequence to construct a directed acyclic semantic flow graph, where the sentence vector of the log statement obtained in step 5 is embedded into the semantic flow graph as the node feature.
[0049] Step 7. After obtaining the semantic flow graph, the present invention uses a graph convolutional neural network model to extract and aggregate features of the semantic flow graph, combines the structure of the graph and the features of the nodes to learn the high-level representation vector of the nodes in the graph or the representation vector of the entire graph, and realizes the classification of the semantic flow graph, that is, predicts the label of the graph.
[0050] Use a graph convolutional neural network to extract features and train the semantic flow graph G. By propagating and aggregating the node features in the semantic flow graph, map them to the classification label to obtain the system log anomaly detection result. The implementation is as follows:
[0051] (7.1) Represent the semantic flow graph as a set of nodes and edges. For graph G, which contains k n nodes and k e edges, use the adjacency matrix A to represent the connection relationship of the graph. Among them, A is a matrix of size k n ×k n , and A[p][q] represents that there is an edge between node p and node q; use the feature matrix X to represent the feature of each node, where X is a matrix of size k n ×k f , X[p] represents the feature vector of node p, and k f represents the feature dimension of each node;
[0052] (7.2) The graph convolutional neural network is a deep learning model for graph data. The characteristic of the graph convolutional neural network is to learn the representation of nodes by iteratively aggregating the neighbor information of nodes. The graph convolutional neural network aggregates the features of neighbor nodes through the following formula to obtain the feature representation of nodes:
[0053]
[0054] where, H (l) represents the node feature matrix of the l-th layer, is the normalized adjacency matrix, where I is the identity matrix, is the diagonal degree matrix of, σ is the activation function, W (l) is the weight matrix of the l-th layer;
[0055] (7.3) We regard the anomaly detection task as a graph classification task, and divide the semantic flow graph into two categories: normal graphs and abnormal graphs. In order to implement the graph classification task, that is, to obtain the predicted value of the graph label, the present invention adds a global pooling layer to the last layer of the graph convolutional neural network to aggregate the node-level representation into a graph-level representation. The formula for the specific pooling operation is as follows:
[0056]
[0057] where, H (L) represents the node feature representation matrix of the last layer, h G is the final representation vector of graph G. After obtaining the representation vector of graph G, it is used as the input of the graph classification task, mapped through a fully connected layer, and the normal and abnormal probabilities of the given log sequence are calculated using the softmax function. The formula is as follows:
[0058]
[0059] where, represents the probability vector, W represents the weight matrix of the fully connected layer, and b represents the bias vector;
[0060] (7.4) Use the cross-entropy loss function to measure the difference between the output result of the graph convolutional neural network model and the true label, and use the backpropagation algorithm and the gradient descent algorithm to minimize the loss function Loss and update the network parameters; the formula for the loss function is as follows:
[0061]
[0062] where, y G represents the true label of graph G, represents the label of graph G predicted by the model.
[0063]
[0064] Example 2: The overall implementation steps of this example are the same as those of Example 1. Now, the process of generating log statement vectors based on the attention mechanism-based bidirectional GRU model will be further described:
[0065] The GRU model is a variant of the recurrent neural network (RNN) that can model sequences and has a gating mechanism to control the flow of information. The GRU contains two internal structures, namely the reset gate and the update gate. The reset gate determines the influence degree of the previous hidden state on the current moment and can be used to reduce the information considered irrelevant in the previous unit. The update gate determines the update degree of the new information at the current moment in the current hidden state and can be used to determine how much information from the previous unit needs to be passed to the next unit.
[0066] The bidirectional GRU model consists of two GRUs in opposite directions. One processes the sequence from front to back, and the other processes the sequence from back to front. In the forward direction, the update formula of the GRU is as follows:
[0067] z t = σ(W z · [h t-1 , x t )
[0068] r t = σ(W r · [h t-1 , x t )
[0069]
[0070]
[0071] where h t represents the hidden state at the t-th time step, where t ranges from 1 to T, and T represents the length of the input sequence; z t and r t represent the update gate and the reset gate respectively, represents the temporary hidden state, and W z , W r and W h represent the first, second, and third learnable parameters respectively, and σ is the sigmoid function, representing element-wise multiplication;
[0072] In the backward direction, the formula is the same as that in the forward direction. The backward GRU processes the sequence starting from the end of the sequence, that is, from x T to x1.
[0073] The present invention adds an attention mechanism to the bidirectional GRU model, enabling the model to focus on the important parts of the input sequence. In the bidirectional GRU model, the attention mechanism is used to combine the hidden states in the forward and backward directions to obtain the attention weight coefficient α at the t-th time step. t :
[0074]
[0075] Among them, W α represents the learnable attention weight, and h context represents the context vector used to calculate the attention weight; the output context of the bidirectional GRU model based on the attention mechanism is obtained according to the following formula t :
[0076]
[0077] The parts not detailed in the present invention belong to the common general knowledge of those skilled in the art.
[0078] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A system log anomaly detection method based on semantic flow graph mining, characterized in that The steps are as follows: (1) Split the log statements in the original system log, remove meaningless symbols, and retain combined words with special meanings to obtain the initially preprocessed system log; (2) Divide the initially preprocessed system logs into log sequences according to the session or window mechanism, and use the Word2Vec model to convert the words or phrases in the log sequence into word vectors. Let the word vector of the nth word in the mth log be (3) Calculate the term frequency-inverse document frequency TF-IDF of the words in the log statement, where the TF-IDF of the nth word in the mth log is represented as T mn ; (4) Combine with T mn to obtain the vector representation W of the final log statement word according to the following formula mn : where α represents the weight factor; (5) Take W mn as the input of the attention-based bidirectional GRU model, and obtain a set of sentence vectors S = {s1, s2,..., s m ,..., s L} containing L log sequences, where s m represents the m-th log sentence vector; (6) Use the K-means clustering method to cluster the log statement vectors in the set S, match the log sequences according to the clustering results to obtain the node sequences, construct a directed acyclic graph of the nodes in the order of the node sequences, and finally obtain the semantic flow graph G=(V, E) of the log sequences, where V represents the set of nodes and E represents the set of edges; the semantic flow graph G=(V, E) of the log sequences is specifically to consider the log statements with the same log template in the clustering cluster as the same node in the semantic flow graph, match the node type with the log entries in the original log sequence, connect different nodes according to the structure of the log sequence, construct a directed acyclic semantic flow graph, and embed the sentence vectors of the log statements as node features into the semantic flow graph; among them, represents the set of nodes of G, represents the p-th node v in G p pointing to the set of directed edges of the q-th v q ; is a positive integer; (7) Use a graph convolutional neural network to perform feature extraction and training on the semantic flow graph G. By propagating and aggregating the node features in the semantic flow graph, map them to classification labels to obtain the system log anomaly detection result.
2. The method according to claim 1, wherein: In step (3), calculate the term frequency-inverse document frequency TF-IDF of the words in the log statement. The implementation steps are as follows: (3.1) Form a document set D with the preprocessed log statements. For each log statement and a given word in the set, calculate the term frequency TF(w, d) of word w in log statement d; (3.2) Calculate the inverse document frequency IDF(w, D) of word w: where |D| represents the total number of documents in document set D, and |{d ∈ D: w ∈ d}| represents the number of documents containing word w; (3.3) Calculate the TF-IDF value T of word w in the log statement mn : T mn = TF-IDF(w, d, D) = TF(w, d) × IDF(w, D).
3. The method according to claim 1, characterized in that: In step (5), the bidirectional GRU model based on the attention mechanism. The bidirectional GRU model is composed of a forward GRU and a backward GRU. The forward GRU processes the sequence from front to back, and the backward GRU processes the sequence from back to front; In both directions, the update formula of the GRU is as follows: z t = σ(W z · [h t-1 , x t ) r t = σ(W r · [h t-1 , x t ) where h t represents the hidden state at the t-th time step, where t ranges from 1 to T, and T represents the length of the input sequence; z t and r t represent the update gate and the reset gate respectively, represents the temporary hidden state, W z , W r and W h represent the first, second, and third learnable parameters respectively, σ is the sigmoid function, and ⊙ represents element-wise multiplication; In the bidirectional GRU model, the attention mechanism is used to combine the hidden states in the forward and backward directions to obtain the attention weight coefficient α at the t-th time step t : Among them, W α represents learnable attention weights, and h context represents the context vector used to calculate the attention weights; the output context of the bidirectional GRU model based on the attention mechanism is obtained according to the following formula t :
4. The method according to claim 1, wherein: In step (6), use the K-means clustering method to cluster the log statement vectors in set S. The implementation is as follows: (6.1) Initialize the cluster center points, let S a and S b represent the sentence vectors of two different log statements, denoted as the first vector and the second vector respectively. Calculate the Euclidean distance Distance(S a and S b using the following formula: Distance(S a , S b ): Among them, n is the vector dimension, and S z a , S z b respectively represent the element values of the first vector S a and the second vector S b on the z-th dimension; (6.2) Iteratively update the cluster centers. In each iteration, assign each vector to the cluster to which the nearest cluster center belongs, and then update the cluster center of each cluster to the mean vector of all vectors in that cluster; assume C k represents the k-th cluster, and N k represents the number of vectors in the k-th cluster, and u k represents the cluster center of the k-th cluster. Then the update formula for the cluster center of the k-th cluster is expressed as follows: Among them, x e represents the e-th vector, which is used as a sample point. When the cluster centers no longer change or reach the maximum number of iterations, the algorithm is considered to converge and the final clustering result is obtained.
5. The method according to claim 1, wherein: In step (7), use a graph convolutional neural network to perform feature extraction and training on the semantic flow graph G. By propagating and aggregating the node features in the semantic flow graph, map them to classification labels to obtain the system log anomaly detection result. The implementation is as follows: (7.1) Represent the semantic flow graph as a set of nodes and edges. For graph G, which contains k n nodes and k e edges, use the adjacency matrix A to represent the connection relationship of the graph. Among them, A is a matrix of size k n ×k n . A[p][q] indicates that there is an edge between node p and node q; use the feature matrix X to represent the features of each node. Among them, X is a matrix of size k n ×k f . X[p] represents the feature vector of node p, and k f represents the feature dimension of each node. (7.2) The graph convolutional neural network aggregates the features of neighbor nodes through the following formula to obtain the feature representation of the node: Among them, H (l) represents the node feature matrix of the l-th layer, is the normalized adjacency matrix, where I is the identity matrix, is the diagonal degree matrix of, σ is the activation function, W (l) is the weight matrix of the l-th layer; (7.3) Divide the semantic flow graph into two categories: normal graph and abnormal graph, obtain the predicted value of the graph label, add a global pooling layer to the last layer of the graph convolutional neural network, and aggregate the node-level representation into a graph-level representation. The formula for the specific pooling operation is as follows: Among them, H (L) represents the node feature representation matrix of the last layer, and h G is the final representation vector of graph G; after obtaining the representation vector of graph G, it is used as the input of the graph classification task, mapped through a fully connected layer, and the softmax function is used to calculate the normal and abnormal probabilities of the given log sequence. The formula is as follows: Among them, represents a probability vector, W represents the weight matrix of the fully connected layer, and b represents the bias vector; (7.4) Use the cross-entropy loss function to calculate the difference between the output result of the graph convolutional neural network model and the true label. Use the backpropagation algorithm and the gradient descent algorithm to minimize the loss function Loss and update the network parameters. The loss function formula is as follows: Among them, y G represents the true label of graph G, and represents the label of graph G predicted by the model.
Citation Information
Patent Citations
Mass alarm data processing method and system, medium, computer equipment and application
CN112312443A
Semi-supervised log anomaly detection method based on probability label estimation
CN113312447A