Log anomaly detection method and system based on improved TCN
By using ReMish activation function and dynamic threshold weight pruning regularization technology in the TCN-based log anomaly detection model, the shortcomings of the existing models in activation function and regularization technology are solved, and the learning ability and output stability of the model are improved.
Patent Information
- Application Number
- CN202311572697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing TCN-based log anomaly detection model has shortcomings in activation functions and regularization techniques, resulting in difficulty in network training, insufficient learning ability and unstable output.
ReMish activation function is used instead of ReLU function to improve the learning ability and convergence speed of the network, and through the weight pruning regularization technology based on dynamic thresholds, the network complexity is reduced and overfitting is prevented.
It effectively solves the problem of gradient vanishing, improves the network's learning ability and convergence speed, reduces training time and improves the stability of the output.
Smart Images

Figure FDA0004565997840000011 
Figure FDA0004565997840000012 
Figure FDA0004565997840000013
Abstract
Description
Technical Field
[0001] The present invention relates to a log anomaly detection method and system, and in particular to a log anomaly detection method and system of an improved TCN, belonging to the field of neural networks. Background Art
[0002] With the rapid development and widespread application of Internet technology, log data plays an important role in the operation and maintenance of complex network monitoring systems, troubleshooting, and security audits. Analyzing log data to discover potential abnormal behaviors is an important means of rapid troubleshooting and safe operation and maintenance. Temporal Convolutional Network (TCN), as a model based on convolutional neural networks, has the ability to capture long-term dependencies and translation invariance. The anomaly detection model based on TCN can automatically and efficiently learn the time series features in log data and detect abnormal behaviors based on these features.
[0003] Convolutional neural networks use activation functions to introduce nonlinear properties to help the network better fit data and extract features. The derivative of the activation function ReLU in the original TCN is close to 0, which makes network training difficult. Although PReLU has improved this problem, it is still a linear activation function and cannot introduce nonlinear relationships, which limits the fitting ability of the neural network. In 2019, Sergey Ioffe proposed a novel smooth nonlinear activation function Mish. Compared with ReLU and PReLU functions, the Mish function maps the input value to a nonlinear differentiable output value, helping the neural network achieve better performance and effects. However, the Mish function has poor learning ability for shallow neurons, and the output range is [-1,∞]. The non-standardized output will cause the value range of the network output to change, affecting the network training effect. In addition, the use of Dropout regularization in the TCN-based structure will randomly discard some neurons, increase training time and may cause network output instability, limiting the network's learning ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 This is the log sequence anomaly detection model of the present invention.
[0005] Figure 2 Comparison chart of ReMish, Mish, ReLU and PReLU function curves. Summary of the invention
[0006] In view of the above deficiencies in the prior art, the present invention proposes a log anomaly detection method and system based on improved TCN. The system comprises three modules: a word embedding module, an improved TCN module, and a fully connected module.
[0007] The word embedding module represents the words in the original log text as vectors, maps each word to a low-dimensional space, and encodes the semantic information in the text as a continuous real number vector. The output of the word embedding module serves as the input of the improved TCN module.
[0008] The improved TCN module includes input layer, convolution layer, stacking layer, residual connection, spatial dimension pooling and output layer. The output of the word embedding module is used as the input of the improved TCN module, and the improved TCN maps the feature representation to the corresponding output category for output. The convolution layer in the improved TCN module consists of a one-dimensional hole convolution layer, ReMish function and pruning regularization, where ReMish(x) = e x-1 tanh(ln(1+e x-1 )). Setting the weights of some connections or neurons with lower scores in the weight importance scoring network to 0 can reduce the complexity of the network and prevent overfitting. i =|E(w i x i )·H i |, where i, j represent the i-th and j-th neurons respectively, and w i represents the weight of the connection between neuron i and neuron j, x i Represents the input value. The pruning rate of the network is p, and some parameters with lower activation-entropy are selected for pruning. The pruning threshold is Let the mash vector Furthermore, the pruning threshold for the qth round of training is After training is completed, the threshold is calculated first, and then the weight is updated, and the mash vector is combined with Multiply to get the pruned weight Then the forward propagation calculates the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight Where η represents the learning rate, w i The gradient of is the partial derivative of the pruned weight. Therefore, the pruned neuron weights are updated and may exceed the threshold in the next iteration and reconnect to the network.
[0009] The fully connected module will fully connect all nodes of the improved TCN module with all nodes of the current fully connected module to achieve information transmission and feature extraction between nodes.
[0010] Based on the above system, the present invention proposes a log anomaly detection method based on improved TCN, which includes the following steps:
[0011] (1) Represent the words in the original log text as vectors, map each word to a low-dimensional space, and encode the semantic information in the text as a continuous real number vector.
[0012] (2) The above real vector is mapped to the corresponding output category through the improved TCN module for output. The convolution layer in the improved TCN module consists of a one-dimensional hole convolution layer, a ReMish function, and pruning regularization, where ReMish(x) = e x- 1 tanh(ln(1+e x-1 )). Setting the weights of some connections or neurons with lower scores in the weight importance scoring network to 0 can reduce the complexity of the network and prevent overfitting. i =|E(w i x i )·H i |, where i, j represent the i-th and j-th neurons respectively, and w i represents the weight of the connection between neuron i and neuron j, x i Represents the input value. The pruning rate of the network is p, and some parameters with lower activation-entropy are selected for pruning. The pruning threshold is Let the mash vector Furthermore, the pruning threshold for the qth round of training is After training is completed, the threshold is calculated first, and then the weight is updated, and the mash vector is combined with Multiply to get the pruned weight Then the forward propagation calculates the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight Where η represents the learning rate, w i The gradient of is the partial derivative of the pruned weight. Therefore, the pruned neuron weights are updated and may exceed the threshold in the next iteration and reconnect to the network.
[0013] (3) All nodes of the improved TCN module are fully connected with all nodes of the current fully connected module to achieve information transmission and feature extraction between nodes.
[0014] The present invention adopts the ReMish activation function to replace the ReLU function, improves the learning ability and convergence speed of TCN, and uses regularization based on dynamic threshold weight pruning to discard some low-weight neurons, combines the dynamic threshold with the loss function, adjusts the pruning weights of the next iteration, and reconnects the weights exceeding the threshold back to the network. The ReMish function retains a non-zero slope at zero, and at the same time is smaller than the slope of Mish, effectively solving the problem of gradient vanishing, ensuring that the error can be transmitted to the shallower neurons, and improving the learning ability and convergence speed of the network. DETAILED DESCRIPTION
[0015] In view of the shortcomings of the prior art, the present invention proposes a log anomaly detection method based on an improved TCN network. The ReMish activation function is used to replace the ReLU function to improve the learning ability and convergence speed of the TCN network. The regularization based on dynamic threshold weight pruning is used to discard some low-weight neurons. The dynamic threshold is combined with the loss function to adjust the pruning weight of the next iteration and reconnect the weights exceeding the threshold back to the network.
[0016] Log anomaly detection model based on improved TCN network Figure 1 As shown in the figure, it contains 3 layers: word embedding layer, TCN layer, and fully connected layer.
[0017] The word embedding layer represents the words of the original log text in vector form, and encodes the semantic information in the text into a continuous real number vector by mapping each word to a vector in a low-dimensional space. The output of the word embedding layer will serve as the input of the improved TCN module.
[0018] TCN module: The temporal convolutional network is based on the convolutional neural network and can process time series data, including input layer, convolution layer, stacking layer, residual connection, spatial dimension pooling and output layer. After the output of the word embedding layer is used as the input of the TCN module, the feature representation is mapped to the corresponding output category for output through the TCN network. The general TCN convolution layer consists of a one-dimensional hole convolution layer, a ReLU activation function and Dropout. The one-dimensional hole convolution refers to the introduction of a hole (or expansion) rate parameter in the convolution operation, which can increase the receptive field without increasing the number of convolution layer parameters and capture longer-range time relationships. The ReLU activation function is usually used to enhance the nonlinear characteristics of the network and increase the expressive power of the model. Dropout is a regularization technique that reduces the dependency between neurons and prevents overfitting by randomly setting the output of some neurons to zero during training.
[0019] Fully connected module: fully connects all nodes of the previous module with all nodes of the current module to achieve information transmission and feature extraction between nodes.
[0020] The present invention replaces the ReLU activation function with the proposed new nonlinear activation function ReMish(x)=e x-1 tanh(ln(1+e x-1 )). The ReMish function retains a non-zero slope at the zero point, and at the same time has a smaller slope than the Mish function. It can effectively solve the problem of gradient vanishing while ensuring that the error can be transmitted to the shallower neurons, thereby improving the learning ability and convergence speed of the network.
[0021] Secondly, the present invention uses the proposed pruning regularization to replace the Dropout regularization. The pruning strategy sets the weights of some connections or neurons with lower scores in the weight importance scoring network to 0 to reduce the complexity of the network and prevent overfitting. i x i ) or information entropy H i The higher the value, the greater the weight of the neuron in the model, and the importance score of the weight is i =|E(w i x i )·H i |, where i, j represent the i-th and j-th neurons respectively, and w i represents the weight of the connection between neuron i and neuron j, x i Represents the input value. The pruning rate of the network is p, and some parameters with lower activation-entropy are selected for pruning. The pruning threshold is Let the mash vector
[0022] Furthermore, the pruning threshold for the qth round of training is After training is completed, the threshold is calculated first, and then the weight is updated, and the mash vector is combined with Multiply to get the pruned weight Then the forward propagation calculates the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight Where η represents the learning rate, w i The gradient of is the partial derivative of the pruned weight. Therefore, the pruned neuron weights are updated and may exceed the threshold in the next iteration and reconnect to the network.
[0023] After the system log template sequence is initialized by the word embedding layer, the log template sequence feature representation is generated through the following steps:
[0024] 1. Input a one-dimensional hole convolution layer and normalize it;
[0025] 2. According to the input of the first convolutional layer and the size of padding, the convolutional tensor is sliced to implement causal convolution. Causal convolution is used to process the log template sequence before the last log template in the log sequence of the current input model to predict the next log template. The output of this convolutional layer is only related to the log sequence of the previous or current time node, and will not depend on the input after the next log template.
[0026] 3. Add ReMish activation function;
[0027] 4. The first convolution block is completed based on regularization of dynamic threshold pruning. Pruning regularization is used to discard the output of some neurons to reduce overfitting.
[0028] 5. Stacking a second convolutional block with the same structure constitutes a residual block.
[0029] Specifically, ReMish(x)=e x-1 tanh(ln(1+e x-1 )), its output range is restored to [0,∞]. Compared with the Mish function, it reduces the difficulty of network training and produces better optimization results. It also increases the nonlinear characteristics of the activation function, allowing the neural network to learn more complex patterns and representations. The comparison of the ReMish function curve with the Mish and ReLU function curves is shown in the figure below. Figure 2 shown.
[0030] Specifically, the pruning regularization described in step 4 determines the importance score of the neuron based on the activation value and the information entropy value, and resets the weights of neurons with lower scores to 0 according to a certain pruning ratio. After each round of training, the pruned neuron weights are updated, and the loss function is calculated using forward propagation, and the error is back-propagated to obtain the gradient of the weight.
[0031] Let i be the number of the neuron in the previous layer, j be the number of the neuron in the current layer, and there are n neurons in the previous layer. Specifically, the importance score is used and the weight value is used to combine the sparse regularization term as follows:
[0032] L w =∑ (x,y) loss(f(x,W),y)+λ∑‖w‖, where, L w represents the loss value of the entire network, loss represents the network classification loss function, ‖·‖ represents L 1 Regularization function, λ is a hyperparameter used to balance the two terms, w is the weight value. x, y are the network input and true output respectively, W is the trainable parameter in the network, and f(x, W) is the predicted output calculated by the network. The closer y is to f(x, W), the better the network training effect is.
[0033] According to the forward propagation formula of the neural network, when When , the activation value of neuron j is Its expected value is Where W j Represents the weight matrix connecting neuron j with the neurons in the previous layer. Let p(x i ) is x i The probability of occurrence, then the information entropy value of the discrete variable X is Similarly, randomly select K samples from the training set and set w i x i The value of is divided into D intervals, using p i,z Indicates w i x i The probability of the value of is distributed in z can be calculated to get w i x i Information entropy
[0034] The expected absolute value of the neuron activation value |E(w i x i )| is larger, the weight w i The higher the importance, the higher the information entropy H i The smaller the weight w i The lower the importance of , the activation value and information entropy are positively correlated with the importance of weight. Let the importance score of weight be i =|E(w i x i )·H i |. According to the set pruning rate p, select some parameters with lower activation-entropy for pruning, and the pruning threshold is Vector via Mash
[0035] The pruning threshold for the qth round of training is After training is completed, the threshold is calculated first and then the weight is updated. Multiply to get the pruned weight M i (q) , and then forward propagation is used to calculate the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight: Therefore, the pruned neuron weights can be updated and may exceed the threshold in the next iteration and reconnect back to the network.
[0036] The output feature vector of the TCN module will be used as the input of the fully connected module for linear transformation and nonlinear activation to achieve the integration and extraction of input features, thereby outputting more representative feature representations and providing more meaningful feature representations for subsequent tasks.
[0037] In the training phase, the normal log template sequence generates the input sequence and the target data is input into the anomaly detection model for training, and the adaptive gradient descent method is used to optimize the loss function. In the detection phase, for the log sequence of length h in the log stream, the trained model is loaded to obtain a prediction result of the current sequence, and the predicted output is compared with the actual log template. The actual log templates that rank in the top g range of the predicted output probability are all normal sequences.
Claims
1. A log anomaly detection system based on improved TCN, the system includes a word embedding module, an improved TCN module, and a fully connected module. Features: The word embedding module represents the words in the original log text as vectors, maps each word to a low-dimensional space, and encodes the semantic information in the text into a continuous real number vector as the input of the improved TCN module; The improved TCN module includes input layer, convolution layer, stacking layer, residual connection, spatial dimension pooling and output layer. The real vector is mapped to the corresponding output category by the improved TCN. The convolution layer in the improved TCN module consists of a one-dimensional hole convolution layer, ReMish function and pruning regularization, where ReMish(x) = e x-1 tanh(ln(1+e x-1 )). The pruning strategy is to set the weights of some connections or neurons with lower scores in the weight importance scoring network to 0, so that the weight importance score i =|E(w i x i )·H i |, where i, j represent the i-th and j-th neurons respectively, and w i represents the weight of the connection between neuron i and neuron j, x i Represents the input value. The pruning rate of the network is p, and some parameters with lower activation-entropy are selected for pruning. The pruning threshold is Let the mash vector The pruning threshold for the qth round of training is After training is completed, the threshold is calculated first, and then the weight is updated, and the mash vector is combined with Multiply to get the pruned weight Then the forward propagation calculates the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight Where η represents the learning rate, w i The gradient of is the partial derivative of the pruned weights, thereby updating the pruned neuron weights, which may exceed the threshold in the next iteration and reconnect to the network. The fully connected module will fully connect all nodes of the improved TCN module with all nodes of the current fully connected module to achieve information transmission and feature extraction between nodes.
2. A log anomaly detection method based on improved TCN, Features: The method comprises the following steps: (1) Represent the words in the original log text as vectors, map each word to a low-dimensional space, and encode the semantic information in the text as a continuous real number vector. (2) The above real vector is mapped to the corresponding output category through the improved TCN module for output. The convolution layer in the improved TCN module consists of a one-dimensional hole convolution layer, a ReMish function, and pruning regularization, where ReMish(x) = e x-1 tanh(ln(1+e x-1 )). The pruning strategy is to set the weights of some connections or neurons with lower scores in the weight importance scoring network to 0, which can reduce the complexity of the network and prevent overfitting. i =|E(w i x i )·H i |, where i, j represent the i-th and j-th neurons respectively, and w i represents the weight of the connection between neuron i and neuron j, x i Represents the input value. The pruning rate of the network is p, and some parameters with lower activation-entropy are selected for pruning. The pruning threshold is Let the mash vector The pruning threshold for the qth round of training is After training is completed, the threshold is calculated first, and then the weight is updated, and the mash vector is combined with Multiply to get the pruned weight Then the forward propagation calculates the loss function L w , and then back-propagate the error to get the gradient of the weight and update the weight Where η represents the learning rate, w i The gradient of is the partial derivative of the pruned weight. The pruned neuron weights are updated and may exceed the threshold in the next iteration and reconnect to the network. (3) All nodes of the improved TCN module are fully connected with all nodes of the current fully connected module to achieve information transmission and feature extraction between nodes.