Real-time log anomaly detection method and device, computer equipment and storage medium

Through the stream processing framework and online deep learning method, combined with bi-LSTM and Transformer models, the model parameters are updated online using the FTRL algorithm, which solves the accuracy and real-time problems of abnormal detection in large-scale log data, and achieves efficient and accurate abnormal detection.

CN120257973APending Publication Date: 2025-07-04HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283400.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

It is difficult for the prior art to detect abnormal information quickly and efficiently in large-scale log data, and traditional methods are difficult to take into account both accuracy and real-time.

Method used

The stream processing framework is used to collect log data, and deep learning training is performed through the prediction model combined with bi-LSTM and Transformer, and online update of model parameters is combined with the FTRL algorithm to enhance the time series modeling ability and model stability.

Benefits of technology

It improves the efficiency and accuracy of abnormal detection, can dynamically adapt to log flow characteristics, support distributed deployment, reduce maintenance costs, and improve the intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257973A_ABST
    Figure CN120257973A_ABST
Patent Text Reader

Abstract

The invention relates to a real-time log anomaly detection method and device, computer equipment and a storage medium, and the method comprises the following steps: collecting log data by adopting a stream processing framework, and analyzing the log data to generate structured data; carrying out deep learning training on the off-line model based on the bi-LSTM based on the structured data to obtain a prediction model of combining Transform with LSTM; performing anomaly detection on a real-time log obtained through a stream processing framework by using the prediction model to obtain a prediction result of the real-time log; and on the basis of the prediction result of the real-time log, performing model parameter online updating on the prediction model by using an FTRL algorithm. According to the method, the LSTM is introduced into a Transform encoder structure, the time sequence modeling capability is enhanced, the neglect of local dependence when the Transform processes a long sequence is avoided, the FTRL algorithm is used to optimize an online model in combination with a prediction result, the long-term stability of the model is kept, and the anomaly detection efficiency and accuracy are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of anomaly detection, and particularly to a real-time log anomaly detection method, apparatus, computer device, and storage medium. Background Art

[0002] With the rapid development of information technology, more and more enterprises and organizations rely on computer systems for business management and production operations, and these systems generate a large amount of log data every day. Logs are important records of the system operation status, which contain rich information, such as system operations, user behaviors, and error messages. Log analysis is not only the basis for system maintenance and troubleshooting, but also plays an important role in fields such as network security, anomaly behavior detection, and business optimization. However, in the face of the increasing scale and complexity of log data, how to quickly and efficiently detect abnormal information from massive logs has become an urgent challenge to be solved. Traditional log anomaly detection methods, including rule-based detection methods, traditional machine learning methods, and deep learning methods, often have difficulty in balancing accuracy and real-time performance when dealing with large-scale log streams, and have the disadvantage of low anomaly detection efficiency. Summary of the Invention

[0003] Based on this, it is necessary to provide a real-time log anomaly detection method, apparatus, computer device, and storage medium that can improve the anomaly detection efficiency for the above technical problems.

[0004] In a first aspect, the present application provides a real-time log anomaly detection method. The method includes:

[0005] Collect log data using a stream processing framework, and parse the log data to generate structured data;

[0006] Perform deep learning training on a bi-LSTM-based offline model based on the structured data to obtain a prediction model that combines Transformer and LSTM;

[0007] Use the prediction model to perform anomaly detection on real-time logs obtained through the stream processing framework to obtain prediction results of the real-time logs;

[0008] Based on the prediction results of the real-time logs, use the FTRL algorithm to perform online update of the model parameters of the prediction model.

[0009] In one embodiment, performing deep learning training on a bi-LSTM-based offline model based on the structured data includes:

[0010] Perform semantic mining based on the structured data, convert it into an initial embedding vector using a word embedding matrix, and then add a position vector using positional encoding to obtain a semantic vector sequence of the log;

[0011] According to the semantic vector sequence, use a Transformer encoding layer for feature extraction, and jointly focus on different matrices based on the multi-head self-attention mechanism layer to obtain subspaces with different representations;

[0012] Take the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the temporal dependencies in the sequence;

[0013] Based on the output of the LSTM layer, map the feature vector of the log from a low dimension to a high-dimensional space through a feed-forward neural network layer, and then compress the feature vector back to the original dimension in the high-dimensional space;

[0014] Based on the output of the feed-forward neural network layer, use a softmax layer for normalization processing, and output the predicted labels of the semantic vector sequence, which are used to determine whether the real-time log is abnormal.

[0015] In one embodiment, the position vector includes:

[0016]

[0017] Among them, Pos (i,j) represents the position vector, i represents the position in the corresponding semantic vector sequence, d represents the dimension of the position vector, and k represents the dimension encoding; add the initial embedding vector and the position vector to obtain the semantic vector sequence of the log;

[0018] Take the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the temporal dependencies in the sequence, including:

[0019] h t = LSTM(Z t , h t-1 )

[0020] Among them, Z t is the feature vector obtained from the multi-head attention mechanism layer, h t-1 is the hidden state of the previous moment, and h t is the output of the LSTM layer.

[0021] In one embodiment, the feed-forward neural network layer includes a fully connected layer that maps the feature vector from a low dimension to a high-dimensional space, and another fully connected layer that compresses the feature vector back to the original dimension in the high-dimensional space, and a ReLU activation function is used between the two fully connected layers; the calculation formula of the feed-forward neural network layer is:

[0022] output = ReLU(LN(h t ))

[0023] Among them, output represents the output of the feedforward neural network layer, LN represents the normalization operation, and ReLU represents the activation function.

[0024] In one embodiment, based on the prediction result of the real-time log, the model parameters of the prediction model are updated online using the FTRL algorithm, including:

[0025] Calculate the cross-entropy loss function according to the difference between the actual label and the prediction result of the real-time log, and calculate the gradient vector of the cross-entropy loss function with respect to the model parameters;

[0026] Based on the calculated gradient vector, use the FTRL algorithm to maintain the cumulative gradient and the cumulative squared gradient, and adjust the model parameters according to the maintained cumulative gradient and cumulative squared gradient.

[0027] In one embodiment, the cross-entropy loss function includes:

[0028]

[0029] Among them, L(θ) represents the cross-entropy loss function, y represents the actual label, represents the prediction result of the model, and θ represents all the parameters of the model;

[0030] The update formula of the cumulative squared gradient includes:

[0031]

[0032] Among them, n i is the cumulative squared gradient corresponding to the i-th model parameter, and g t,i is the i-th element of the gradient vector gt;

[0033] The update formula of the cumulative gradient includes:

[0034]

[0035] Among them, Z i is the cumulative gradient corresponding to the i-th model parameter, α is the learning rate, n i and n i,t-1 respectively represent the current cumulative squared gradient and the cumulative squared gradient of the previous iteration of the i-th model parameter, and ω i is the current value of the i-th model parameter.

[0036] In one embodiment, the adjusting the model parameters according to the maintained cumulative gradient and cumulative squared gradient includes:

[0037]

[0038] Among them, ω t+1,i represents the i-th model parameter after update, λ1 and λ2 are regularization coefficients, and β is a constant.

[0039] In a second aspect, the present application also provides a real-time log anomaly detection device. The device includes:

[0040] A data processing module, configured to collect log data using a stream processing framework, and parse the log data to generate structured data;

[0041] A model training module, configured to perform deep learning training on an offline model based on bi-LSTM based on the structured data to obtain a prediction model that combines Transformer and LSTM;

[0042] An anomaly detection module, configured to perform anomaly detection on real-time logs obtained through the stream processing framework using the prediction model to obtain a prediction result of the real-time logs;

[0043] A model update module, configured to perform online update of the model parameters of the prediction model based on the prediction result of the real-time logs using the FTRL algorithm.

[0044] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0045] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0046] The above real-time log anomaly detection method, device, computer device and storage medium collect log data using a stream processing framework, parse the log data to generate structured data, perform deep learning training on an offline model based on bi-LSTM using the structured data to obtain a prediction model combining Transformer and LSTM, use the prediction model to perform anomaly detection on the real-time log obtained through the stream processing framework to obtain the prediction result of the real-time log, and online update the model parameters of the prediction model using the FTRL algorithm based on the prediction result of the real-time log. By introducing LSTM into the Transformer encoder structure, the time series modeling ability is enhanced, the neglect of local dependencies by Transformer when processing long sequences is avoided, and the FTRL algorithm is used to optimize the online model in combination with the prediction result to maintain the long-term stability of the model, effectively improving the efficiency and accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic flowchart of the real-time log anomaly detection method in an embodiment;

[0048] Figure 2 It is a schematic flowchart of performing deep learning training on an offline model based on bi-LSTM using structured data in an embodiment;

[0049] Figure 3 It is an offline model diagram in an embodiment;

[0050] Figure 4 It is a flowchart of online model update in an embodiment;

[0051] Figure 5 It is a structural block diagram of the real-time log anomaly detection device in an embodiment;

[0052] Figure 6 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] In practical applications, log anomaly detection technology first needs to address the dynamic nature of logs and the real-time detection requirements: Log data is generated in a streaming manner, and its data distribution may change over time. This "concept drift" makes it difficult for traditional static models to cope. In network security or critical business scenarios, abnormal logs may indicate attack behaviors or system failures, and they must be detected and responded to within an extremely short time. At the same time, log information is massive. Modern systems may generate hundreds of gigabytes or even terabytes of log data every day, which poses a huge challenge to storage and processing capabilities. To address these challenges, researchers have proposed various log anomaly detection methods, including rule-based detection methods, traditional machine learning methods, and deep learning methods. However, when dealing with large-scale log streams, these methods often struggle to balance accuracy and real-time performance. Therefore, there is an urgent need for an efficient anomaly detection method that can dynamically adapt to the characteristics of log streams and support distributed deployment.

[0055] Based on this, this application provides a real-time log anomaly detection method that combines a stream processing framework and online deep learning. By building a data pipeline based on the stream processing framework, high-throughput and low-latency data processing can be achieved, enabling log data to be parsed, analyzed, and anomaly-detected while flowing into the system. The Transformer model automatically learns log patterns through the self-attention mechanism, which can accurately capture abnormal patterns in log sequences, reduce the dependence on manual feature engineering, and improve the accuracy of detection. In addition, by combining the incremental learning mechanism, the model can continuously adapt to changes in log patterns and improve the detection ability for new types of anomalies. Compared with traditional methods, this method not only has real-time detection capabilities but also can enhance the intelligence level of the system and reduce maintenance costs.

[0056] In one embodiment, as Figure 1 shown, a real-time log anomaly detection method is provided, including the following steps:

[0057] Step S110: Use a stream processing framework to collect log data and parse the log data to generate structured data. Among them, the stream processing framework can be used to collect log data in real time from multi-source log systems. The log data includes system logs, security logs, access logs, etc. The log stream is divided based on time windowing to avoid the lag of batch processing methods and ensure the real-time performance of anomaly detection.

[0058] Specifically, a distributed log stream processing framework can be built by combining Kafka and Flink, which supports high concurrency, load balancing, and low-latency processing to ensure the efficient distribution and streaming calculation of log data. After obtaining the log data, the log can be parsed based on the parse tree to parse the original information into structured data that can be recognized by machines. Among them, parsing the log based on the parse tree includes the following steps:

[0059] S11: For the log text data, use a custom regular expression to clean specific information (such as common variables like IP addresses). Through the log overview, formulate basic regular matching rules to determine the regular expression.

[0060] S12: Split the log statements based on heuristic rules. During the splitting process, increase the recognition and processing of symbols such as '=', ':', '(', ')', ',' etc.

[0061] S13: According to the consistency of the log types, count the token number of the current log. In the first layer of the parse tree, if a child node of the corresponding length is found, update to this node and continue to S14; otherwise, go to S17.

[0062] S14: Search for the leaf nodes under the current length node. If a leaf node containing the current token content is found, go to S15; otherwise, go to S17.

[0063] S15: Calculate the similarity and perform threshold judgment. If the similarity between the current token and the leaf node path information exceeds the set threshold, return the matching log template (the leaf node path content) and go to S16; otherwise, go to S17.

[0064] S16: Compare the log content. If the current result is consistent with the token at the same position in the log message, it is recognized as a constant; otherwise, it is regarded as a variable and blurred with a wildcard <*>.

[0065] S17: Update the parse tree structure: If there is no similar leaf node, create a new node; if there is a node of the current length, update or create a child node; if the node depth exceeds the set maximum value, define this node as a leaf node and stop creating.

[0066] Step S120: Perform deep learning training on the offline model based on bi-LSTM using the structured data to obtain a prediction model that combines Transformer and LSTM. Use the structured data parsed from historical log data to preliminarily train the deep learning model to identify normal log patterns and potential abnormal logs.

[0067] A Transformer is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks. An offline model based on bi-LSTM is designed. After obtaining structured data, the deep learning model is initially trained to identify normal log patterns and potential abnormal logs. The self-attention mechanism is used to calculate the global dependencies of the logs, enhancing the understanding ability of logs with complex structures. LSTM (Long Short-Term Memory) is introduced into the Transformer encoder structure to enhance the time series modeling ability and avoid the neglect of local dependencies when the Transformer processes long sequences. In addition, positional encoding and time window features can be combined to enhance the modeling ability of time information and improve the temporal sensitivity of anomaly detection.

[0068] In one embodiment, as Figure 2 shown, the deep learning training of the offline model based on bi-LSTM using structured data in step S120 includes:

[0069] Step S121: Conduct semantic mining based on structured data, convert it into an initial embedding vector using a word embedding matrix, and then add a positional vector using positional encoding to obtain a semantic vector sequence of the logs.

[0070] According to the parsed structured data, conduct semantic mining of the logs and log sequences, convert the log text into an initial embedding vector using a word embedding matrix, and then add a positional vector using positional encoding. Add the initial embedding vector and the positional vector to obtain a semantic vector sequence X of the logs. Among them, the positional vector includes:

[0071]

[0072] Among them, Pos (i,j) represents the positional vector, i represents the position in the corresponding semantic vector sequence, d represents the dimension of the positional vector, and k represents the dimension encoding.

[0073] Step S122: According to the semantic vector sequence, use the Transformer encoding layer for feature extraction, and jointly focus on different matrices based on the multi-head self-attention mechanism layer to obtain subspaces with different representations. Use the semantic vector sequence X as the input of the Transformer encoding layer, process the semantic vector sequence X using the encoder, and obtain the context information of each word through the attention mechanism to obtain the semantic embedding vector S of the logs. Finally, also use the Transformer encoding layer to merge the semantic embedding vector sequences of the logs to generate the semantic representation of the log sequence.

[0074] Specifically, the Transformer encoding layer is used for feature extraction, and the multi-head self-attention mechanism layer jointly focuses on different Q, K, and V matrices to obtain subspaces with different representations. According to the semantic vector sequence X of the log, vectors Q (query), K (key), and V (value) are obtained by multiplying with matrices Wq, Wk, and Wv. According to the self-attention mechanism, the calculation formula is as follows:

[0075]

[0076] where T represents transpose, and d k represents the dimension of the position vector encoded as K.

[0077] Using the multi-head attention mechanism allows the model to jointly focus on subspace information from different representations:

[0078] MultiHead(Q,K,V)=Concat(head1,head2,...,headn)

[0079] Step S123: Take the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the temporal dependencies in the sequence. Specifically, take the output of the multi-head attention mechanism layer as the input of the LSTM layer, and the LSTM layer can effectively capture the temporal dependencies in the sequence, including:

[0080] h t =LSTM(Z t ,h t-1 )

[0081] where Z t is the feature vector obtained from the multi-head attention mechanism layer, h t-1 is the hidden state at the previous moment, and h t is the output of the LSTM layer.

[0082] In addition, as Figure 3 shown, after the semantic vector sequence X of the log is processed by the multi-head attention mechanism layer and the LSTM layer, the output of the LSTM layer can also be subjected to residual connection and normalization processing with the semantic vector sequence X, and the processing result is used as the input of the feed-forward neural network layer.

[0083] Step S124: Based on the output of the LSTM layer, the feed-forward neural network layer maps the feature vectors of the log from a low dimension to a high-dimensional space, and then compresses the feature vectors back to the original dimension in the high-dimensional space.

[0084] where, as Figure 3As shown in the figure, the feedforward neural network layer includes a fully connected layer that maps the feature vector from a low dimension to a high-dimensional space, and another fully connected layer that compresses the feature vector back to the original dimension in the high-dimensional space. The ReLU activation function is used between the two fully connected layers, which can help the model learn more non-linear representations and effectively increase the sparsity of the output, thereby improving the model's expressive ability. The calculation formula of the feedforward neural network layer is:

[0085] output = ReLU(LN(h t ))

[0086] Among them, output represents the output of the feedforward neural network layer, LN represents the normalization operation, and ReLU represents the activation function.

[0087] Step S125: Based on the output of the feedforward neural network layer, use the softmax layer for normalization processing, and output the predicted labels of the semantic vector sequence, which are used to determine whether the real-time log is abnormal. Specifically, linearly process and perform softmax normalization on the output of the feedforward neural network layer, output the predicted labels of the log sequence, map the output to the interval (0, 1), and determine whether the input log data is abnormal.

[0088] It can be understood that after completing the above model training process, the accuracy of the model prediction can also be determined by comparing the predicted labels and the actual labels of the historical log data. Optimize the model parameters through iterative training until the model prediction accuracy meets the set requirements (such as being greater than the set threshold), and obtain a prediction model that meets the requirements, which can be used for anomaly detection of real-time logs.

[0089] Step S130: Use the prediction model to perform anomaly detection on the real-time log obtained through the stream processing framework, and obtain the prediction result of the real-time log.

[0090] After completing the model training, the stream processing framework can also be used to obtain the real-time generated log data, divide the log output by time window, and establish an efficient real-time data processing pipeline through the stream processing framework, which can quickly receive, parse, and preprocess the incoming log data. For each real-time log obtained in a time window, convert it into an initial embedding vector representation through the trained word embedding matrix, and use the prediction model to add position vectors to generate the semantic vector sequence of the log. After passing through each layer of the model for processing, finally output the predicted labels to complete the anomaly detection and obtain the prediction result of the current log.

[0091] Step S140: Based on the prediction result of the real-time log, use the FTRL algorithm to perform online update of the model parameters of the prediction model. Refer to Figure 4, after obtaining the real-time log, the operator can add corresponding actual tags according to whether the real-time log belongs to an abnormal log, calculate the cross-entropy loss function based on the difference between the actual tag and the predicted tag, and update the model parameters using the FTRL algorithm. By adding an online learning mechanism for real-time logs, online updates are performed using the FTRL (Follow-The-Regularized-Leader) algorithm to optimize the online model parameters. Utilize its sparsity control and dynamic learning rate to enhance the adaptability to new log patterns. In this embodiment, the online update of model parameters can adopt a sliding window incremental training mechanism to only update the model parameters within the recent window, reducing the computational overhead while maintaining the long-term stability of the model.

[0092] In one embodiment, step S140 includes: calculating the cross-entropy loss function according to the difference between the actual tag of the real-time log and the prediction result, and calculating the gradient vector of the cross-entropy loss function with respect to the model parameters; based on the calculated gradient vector, use the FTRL algorithm to maintain the cumulative gradient and the cumulative squared gradient, and adjust the model parameters according to the maintained cumulative gradient and cumulative squared gradient. In this embodiment, the cross-entropy loss function includes:

[0093]

[0094] where L(θ) represents the cross-entropy loss function, y represents the actual tag, represents the prediction result of the model, and θ represents all the parameters of the model. The cross-entropy loss function characterizes the gap between the model output and the actual result. After determining the cross-entropy loss function, further calculate the gradient vector gt of the loss function with respect to the model parameters, that is, the direction in which the cross-entropy loss function decreases fastest. The gradient vector gt can be calculated by taking partial derivatives. Each element in the gradient vector gt represents the change rate of the corresponding element of the loss function with respect to the weight.

[0095] Specifically, first initialize the parameters required by the FTRL algorithm, use the weights of the last fully connected layer of the offline model as the initial weight vector ω, initialize the corresponding cumulative squared sum of gradients n to zero, and initialize the learning rate α and the regularization coefficients λ1 and λ2. Based on the calculated gradient vector, the FTRL algorithm can be used to maintain the cumulative gradient and the cumulative squared gradient, and adjust the model parameters according to the maintained cumulative gradient and cumulative squared gradient. In this embodiment, the update formula for the cumulative squared gradient includes:

[0096]

[0097] where n i is the cumulative squared gradient corresponding to the i-th model parameter, and g t,iis the i-th element of the gradient vector gt.

[0098] Furthermore, the update formula for the cumulative gradient includes:

[0099]

[0100] where Z i is the cumulative gradient corresponding to the i-th model parameter, α is the learning rate, n i and n i,t-1 respectively represent the current cumulative squared gradient and the cumulative squared gradient of the previous iteration of the i-th model parameter, ω i is the current value of the i-th model parameter.

[0101] Adjusting the model parameters according to the maintained cumulative gradient and cumulative squared gradient includes:

[0102]

[0103] where ω t+1,i represents the updated i-th model parameter, λ1 and λ2 are regularization coefficients, which respectively control the weights of L1 regularization and L2 regularization. L1 regularization and L2 regularization respectively control the feature selection of the model and prevent overfitting, and β is a small constant.

[0104] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0105] Based on the same inventive concept, an embodiment of the present application also provides a real-time log anomaly detection device for implementing the real-time log anomaly detection method involved above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the real-time log anomaly detection device provided below can refer to the limitations on the real-time log anomaly detection method in the above text, and will not be repeated here.

[0106] In one embodiment, as Figure 5As shown in the figure, a real-time log anomaly detection device is also provided, including a data processing module 110, a model training module 120, an anomaly detection module 130, and a model update module 140, where:

[0107] The data processing module 110 is used to collect log data by using a stream processing framework and parse the log data to generate structured data.

[0108] The model training module 120 is used to perform deep learning training on an offline model based on bi-LSTM based on the structured data to obtain a prediction model that combines Transformer and LSTM.

[0109] The anomaly detection module 130 is used to perform anomaly detection on the real-time log obtained through the stream processing framework by using the prediction model to obtain the prediction result of the real-time log.

[0110] The model update module 140 is used to perform online update of the model parameters of the prediction model by using the FTRL algorithm based on the prediction result of the real-time log.

[0111] In one embodiment, the model training module 120 is used to perform semantic mining based on the structured data, convert it into an initial embedding vector by using a word embedding matrix, and then add a position vector by using position encoding to obtain a semantic vector sequence of the log; according to the semantic vector sequence, use the Transformer encoding layer for feature extraction, and jointly focus on different matrices based on the multi-head self-attention mechanism layer to obtain subspaces with different representations; use the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the time-dependent relationship in the sequence; based on the output of the LSTM layer, map the feature vector of the log from a low dimension to a high-dimensional space through a feed-forward neural network layer, and then compress the feature vector back to the original dimension in the high-dimensional space; based on the output of the feed-forward neural network layer, use the softmax layer for normalization processing, and output the prediction label of the semantic vector sequence, which is used to judge whether the real-time log is abnormal.

[0112] In one embodiment, the model update module 140 is used to calculate the cross-entropy loss function according to the difference between the actual label and the prediction result of the real-time log, and calculate the gradient vector of the cross-entropy loss function with respect to the model parameters; based on the calculated gradient vector, use the FTRL algorithm to maintain the cumulative gradient and the cumulative squared gradient, and adjust the model parameters according to the maintained cumulative gradient and cumulative squared gradient.

[0113] Each module in the above real-time log anomaly detection device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0114] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a real-time log anomaly detection method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0115] Those skilled in the art can understand that Figure 6 the structure shown in

[0116] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in each of the above method embodiments.

[0118] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0120] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0122] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A real-time log anomaly detection method, characterized in that, Including: Using a stream processing framework to collect log data and parse the log data to generate structured data; Based on the structured data, performing deep learning training on an offline model based on bi-LSTM to obtain a prediction model that combines Transformer and LSTM; Using the prediction model to perform anomaly detection on the real-time log obtained through the stream processing framework to obtain the prediction result of the real-time log; Based on the prediction result of the real-time log, using the FTRL algorithm to perform online update of the model parameters of the prediction model.

2. The method according to claim 1, characterized in that, Based on the structured data, performing deep learning training on an offline model based on bi-LSTM, including: Performing semantic mining based on the structured data, converting it into an initial embedding vector using a word embedding matrix, and then adding a position vector using position encoding to obtain a semantic vector sequence of the log; According to the semantic vector sequence, using a Transformer encoding layer for feature extraction, and jointly focusing on different matrices based on the multi-head self-attention mechanism layer to obtain subspaces with different representations; Using the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the temporal dependence relationship in the sequence; Based on the output of the LSTM layer, mapping the feature vector of the log from a low dimension to a high-dimensional space through a feed-forward neural network layer, and then compressing the feature vector back to the original dimension in the high-dimensional space; Based on the output of the feed-forward neural network layer, using a softmax layer for normalization processing, and outputting the prediction label of the semantic vector sequence, which is used to determine whether the real-time log is abnormal.

3. The method according to claim 2, wherein The position vector includes: Among them, Pos (i,j) represents the position vector, i represents the position in the corresponding semantic vector sequence, d represents the dimension of the position vector, and k represents the dimension encoding; adding the initial embedding vector and the position vector to obtain the semantic vector sequence of the log; Using the output of the multi-head attention mechanism layer as the input of the LSTM layer to capture the temporal dependence relationship in the sequence, including: h t = LSTM(Z t , h t-1 ) Among them, Z t is the feature vector obtained from the multi-head attention mechanism layer, h t-1 is the hidden state at the previous moment, h t is the output of the LSTM layer.

4. The method according to claim 3, wherein The feed-forward neural network layer includes a fully connected layer that maps the feature vector from a low dimension to a high-dimensional space, and another fully connected layer that compresses the feature vector back to the original dimension in the high-dimensional space. The ReLU activation function is used between the two fully connected layers; the calculation formula of the feed-forward neural network layer is: output = ReLU(LN(h t )) Where output represents the output of the feed-forward neural network layer, LN represents the normalization operation, and ReLU represents the activation function.

5. The method according to any one of claims 1 to 4, characterized in that, Based on the prediction result of the real-time log, using the FTRL algorithm to perform online update of the model parameters of the prediction model, including: According to the difference between the actual label and the prediction result of the real-time log, calculating the cross-entropy loss function and calculating the gradient vector of the cross-entropy loss function with respect to the model parameters; Based on the calculated gradient vector, using the FTRL algorithm to maintain the cumulative gradient and the cumulative squared gradient, and adjusting the model parameters according to the maintained cumulative gradient and cumulative squared gradient.

6. The method according to claim 5, characterized in that, The cross-entropy loss function includes: Among them, \(L(\theta)\) represents the cross-entropy loss function, \(y\) represents the actual label, represents the prediction result of the model, and \(\theta\) represents all the parameters of the model; The update formula of the cumulative squared gradient includes: where n i is the cumulative squared gradient corresponding to the i-th model parameter, and g t,i is the i-th element of the gradient vector gt; The update formula of the cumulative gradient includes: Among them, Z i is the cumulative gradient corresponding to the i-th model parameter, α is the learning rate, n i and n i,t-1 respectively represent the current cumulative squared gradient and the cumulative squared gradient of the previous iteration of the i-th model parameter, ω i is the current value of the i-th model parameter.

7. The method according to claim 6, characterized in that, The adjusting the model parameters according to the maintained cumulative gradient and cumulative squared gradient includes: where ω t+1,i represents the updated i-th model parameter, λ1 and λ2 are regularization coefficients, and β is a constant.

8. A real-time log anomaly detection device, characterized in that, Including: A data processing module for using a stream processing framework to collect log data and parse the log data to generate structured data; A model training module, which is used to perform deep learning training on an offline model based on bi-LSTM using the structured data to obtain a prediction model that combines Transformer and LSTM; An anomaly detection module, which is used to perform anomaly detection on the real-time log obtained through the stream processing framework using the prediction model to obtain the prediction result of the real-time log; A model update module, which is used to perform online update of the model parameters of the prediction model using the FTRL algorithm based on the prediction result of the real-time log.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.