User behavior log anomaly detection model training method and user behavior log anomaly detection method
By integrating the time and frequency characteristics of multi-source log data, combining LSTM and BERT models, log analysis and model training are optimized, the problem of difficulty in comprehensively capturing user behavior abnormality detection characteristics in the existing technology is solved, and more efficient abnormality detection effects are achieved.
Patent Information
- Application Number
- CN202510263138.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-08
AI Technical Summary
现有的用户行为异常检测方法在处理多源日志数据时,难以全面捕捉时序、频率和内容特征,导致异常检测准确率受限。
By integrating the time and frequency characteristics of multi-source log data, combining LSTM and BERT models for feature extraction and fusion, optimizing log analysis and model training, using Drain tool for log analysis, using LSTM model for extraction, and self-supervised learning of the BERT model, designing Masked Log Key Prediction and Temporal-Frequency Joint Anomaly Detection tasks.
It significantly improves the accuracy and robustness of abnormal detection, can more accurately identify complex abnormal patterns, and improves the performance of user behavior abnormal detection.
Smart Images

Figure CN120277384A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of network security and anomaly detection, and particularly relates to a method for training a user behavior log anomaly detection model and a user behavior log anomaly detection method. Background Art
[0002] With the rapid development of the information age, the Internet has been deeply integrated into people's daily lives and various business activities. However, the accompanying cyberattacks and security threats have become increasingly severe, posing great risks to the security of individuals, enterprises, and even countries. To effectively maintain the security of the cyberspace, ensure the availability of various network resources, and prevent the occurrence of various malicious attack behaviors, user behavior anomaly detection technology, as an active defense means, has become a research hotspot in the current network security field.
[0003] User Behavior Anomaly Detection aims to identify abnormal or malicious behaviors that deviate from normal behaviors by monitoring and analyzing the behavior patterns of users. Such technologies have important application values in preventing internal threats, account hijacking, data leakage, etc. By timely detecting abnormal behaviors, potential security incidents can be effectively prevented, economic losses can be reduced, and the overall security of the system can be maintained. Early user behavior anomaly detection mainly relied on traditional machine learning algorithms, such as Isolation Forest, Support Vector Machine (SVM), etc. These methods modeled the user behavior characteristics and identified abnormal behaviors significantly different from normal behaviors through sequence matching and similarity calculation. However, these methods rely highly on feature engineering, and with the increase in massive high-dimensional data in the network and the improvement of network bandwidth, the complexity of data and the diversity of features are also constantly increasing, making it difficult for shallow learning to achieve the purpose of analysis and prediction.
[0004] In recent years, deep neural network technology has achieved great success in the fields of image recognition, natural language processing, speech recognition, etc. Deep neural network is a method for representing learning of data, which can learn the internal laws of data and adapt to the requirements of high-dimensional learning and prediction through a non-linear network structure composed of multiple hidden layers. With the development of deep learning technology, more and more research has begun to explore user behavior anomaly detection methods based on deep neural networks, such as Long Short-Term Memory (LSTM) and BERT (Bidirectional Encoder Representations from Transformers) models, for anomaly detection by capturing the time-dependent relationships of log sequences. Although these deep learning models can effectively process sequence data and have achieved certain success, these methods still have some problems.
[0005] For example, during the feature extraction process, many studies directly extract features from the original log fields for model training, often relying on single features or simple feature fusion. It is difficult to comprehensively capture the temporal, frequency, and content features in multi-source log data, and thus it is impossible to fully mine the deep associations in the log data, resulting in limited accuracy in anomaly detection. Summary of the Invention
[0006] The present invention aims to achieve efficient and accurate anomaly detection of user behavior by integrating the time features and frequency features of multi-source log data and combining advanced deep learning models. This method significantly improves the accuracy and robustness of anomaly detection by optimizing steps such as log parsing, feature extraction and fusion, and deep model training.
[0007] The present invention proposes a method for training an anomaly detection model for user behavior logs, including:
[0008] Obtain a log dataset;
[0009] According to the corresponding features of each log in the log dataset, parse the log using the corresponding splitting method to obtain log keys;
[0010] Perform a serialization operation on all log keys to obtain a log key sequence;
[0011] Extract time features and frequency features from the log key sequence;
[0012] Concatenate the log key sequence, the time features, and the frequency features and input them into a neural network model for training to obtain a trained anomaly detection model for user behavior logs.
[0013] Furthermore, according to the number and length of tokens in each log, a set tree structure is used as the corresponding splitting method.
[0014] Furthermore, use the Drain tool to parse the log.
[0015] Furthermore, use an LSTM model to extract the time features.
[0016] Furthermore, use a BERT model as the neural network model; the BERT model uses self-supervised learning tasks.
[0017] Furthermore, the last layer of the BERT model outputs a scalar value for scoring the log.
[0018] The present invention also proposes an anomaly detection method for user behavior logs, including:
[0019] Use the trained anomaly detection model for user behavior logs obtained by the above method to perform anomaly detection on user behavior logs.
[0020] The present invention also provides a training system for a user behavior log anomaly detection model, including:
[0021] A data acquisition module for acquiring a log data set;
[0022] A log parsing module for parsing the log according to the corresponding features of each log in the log data set by using a corresponding splitting method to obtain log keys;
[0023] A time series sorting module for performing a serialization operation on all log keys to obtain a log key sequence;
[0024] A feature extraction module for extracting time features and frequency features from the log key sequence;
[0025] A model training module for concatenating the log key sequence, the time features, and the frequency features and inputting them into a neural network model for training to obtain a trained user behavior log anomaly detection model.
[0026] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.
[0027] The present invention also provides a storage medium storing a computer program, and when the computer program is executed by a computer, the above method is implemented.
[0028] The beneficial effects of the present invention are as follows:
[0029] The reason why the present invention selects the LSTM model for extracting time series features is that it has the following advantages: (1) The special gating mechanism of the LSTM model can effectively capture long-term dependence relationships, which is crucial for analyzing time patterns in user behavior sequences. (2) Compared with traditional RNNs, LSTM can better solve the problem of gradient disappearance and maintain long-term memory ability.
[0030] The neural network model of the present invention selects the BERT model instead of the traditional unidirectional model, mainly based on the following considerations: (1) The bidirectional encoding ability of BERT enables it to consider the context information of the log sequence before and after at the same time, which is crucial for understanding the context of user behavior. (2) The pre-training - fine-tuning paradigm enables the model to better adapt to different scenarios. (3) The multi-head self-attention mechanism can capture various dependence relationships in the sequence.
[0031] Through the multi - feature fusion strategy, the comprehensiveness of feature expression is improved, and the user behavior pattern can be characterized more accurately. By adopting the bidirectional encoding ability of the BERT model and combining the enhancement of the LSTM for time - series features, the recognition ability of complex abnormal patterns is significantly improved. Description of the Drawings
[0032] Figure 1 is the overall flowchart of the present invention.
[0033] Figure 2 is the overall process of the Masked Log Key Prediction (MLKP) self - supervised learning task designed in the model training stage of the present invention.
[0034] Figure 3 is the overall process of the Temporal - Frequency Joint Anomaly Detection (TFJAD) self - supervised learning task designed in the model training stage of the present invention.
[0035] Figure 4 is the experimental data comparison chart of the present invention on the HDFS dataset and the BGL dataset. Detailed Implementation Modes
[0036] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and make the objectives, features, and advantages of the present invention more obvious and understandable, the following further details the technical core of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] The embodiments of the present invention provide a method for training a user behavior log anomaly detection model. The general idea of this method is to parse and pre - process system log data, use the LSTM model to extract time and frequency features, and combine the BERT model for feature fusion and anomaly detection.
[0038] The overall flowchart of the present invention is as Figure 1 shown.
[0039] S100 Obtain the log dataset
[0040] The data sets used in this invention are the HDFS data set and the BGL data set. Specifically, the HDFS data set (Hadoop Distributed File System Log) is a widely used log data set, which mainly records the log information when the Hadoop distributed file system runs on the Amazon EC2 cluster. This data set contains tens of millions of log records. The log content includes the start, execution, completion and other status information of various tasks, as well as potential system errors and abnormal behaviors. With its large scale and characteristics of structured logs, the HDFS data set has become an important benchmark in anomaly detection research. It contains 22 million logs, and after parsing, 575,061 log template sequences are obtained, among which the proportion of abnormal logs is about 2.90%. The BGL data set (BlueGene / L Log) comes from the BlueGene / L supercomputer system, which is one of the most powerful computers in the world. The BGL data set records millions of log messages generated during system operation, including system startup, task scheduling, changes in the status of computing nodes and system failures, etc. A significant feature of this data set is the high complexity and high dimension of log entries. In particular, the alarm information contained in it can reflect the potential faults and abnormal behaviors of the system. The BGL data set contains 4.74 million logs, and after parsing, 348,460 log template sequences are obtained, and the proportion of abnormal logs is about 4.00%. It is widely used in anomaly detection research, especially suitable for verifying the ability of algorithms in processing complex system logs.
[0041] S200 Log Parsing
[0042] Log parsing is the basic work of anomaly detection. Logs generated by different systems have different formats. Directly using these raw logs for analysis will lead to problems of data diversity and noise. This is because system logs are usually unstructured or semi-structured data, and there are great challenges in directly using them for analysis. Through log parsing, logs can be converted into structured log keys, which not only extracts the key information in the logs, but also filters out redundant content, laying a foundation for subsequent analysis and processing. This step of log parsing greatly simplifies the complexity of data processing and improves the efficiency and accuracy of subsequent anomaly detection.
[0043] This step has made improvements to the adaptability of log parsing based on the Drain log parsing tool. Drain is an effective log parsing method that can convert unstructured logs into structured log keys by constructing a parsing tree with a fixed depth. When dealing with high-dimensional log data, there are often two situations in the fixed-depth parsing tree of Drain: one is that the depth of the tree is not enough to capture all the important information in the logs; the other is that the tree structure may be overly split, resulting in too much redundant information and reducing the parsing efficiency. Therefore, in this step, the tree structure is optimized, and multi-dimensional splitting based on the characteristics of log content is introduced at different levels of the tree. When performing log parsing, instead of relying on a single fixed depth, different splitting methods are selected at different levels according to the different characteristics of the logs. For some logs with large amounts of information and short lengths, the present invention uses a shallower tree structure for processing, while for those logs with large amounts of information and diverse changes, the present invention will further subdivide and construct a deeper parsing tree. Further, for tokens with numbers and lengths less than 4 in the information fields of the logs, a shallow tree structure is used, and a deep tree structure is used for the others. According to the characteristics of different types of logs, the depth of the parsing tree is set. This improvement makes log parsing more flexible, can effectively process high-dimensional data, and avoids the problem of over-splitting, thus improving the accuracy and efficiency of log parsing.
[0044] In this step, the improved Drain log parsing tool is used to convert logs from different sources into structured log keys. Through a hierarchical tree structure, irregular log messages are matched to fixed log templates, the key information in the logs is extracted, and it is mapped to a unique log key.
[0045] The specific changes to the Drain log parsing tool are as follows: a search algorithm is adopted to traverse each token of the log, and the depth of the tree structure is selected according to the token type. A shallow tree structure is used to process simple tokens, and a deep tree structure is used to process complex tokens, and finally the leaf node is located; at the same time, the prefix tree construction function follows the same rule, inserts the logs into the tree structure in layers according to the token complexity, and stores the log clusters at the leaf nodes.
[0046] S300 Log Serialization:
[0047] After log parsing, the log keys are often discrete and disordered. Therefore, in this step, through log serialization, the parsed log keys are organized into an ordered sequence of log keys. Log serialization preserves the time and order relationship of log events, which is crucial for capturing the temporal correlation between log events. Through this step, the model can more effectively analyze the temporal patterns of log events and provide richer context information for subsequent anomaly detection.
[0048] S400 Extract Temporal Features and Frequency Features
[0049] Log events often exhibit temporal dependencies, and certain abnormal behaviors may occur at specific time intervals or frequency patterns. Therefore, by extracting temporal features and frequency features, the system can better capture these potential abnormal patterns. This step further extracts temporal features and frequency features based on the log key sequence, aiming to capture the temporal behavior patterns of log events. Temporal features reflect the time intervals at which log events occur, while frequency features represent the occurrence frequencies of log events within a specific time window. These features can reveal potential abnormal behaviors. For example, frequent operations within a short period may indicate abnormal activities in the system. By extracting temporal and frequency features, the model can identify more complex log patterns, thereby improving the accuracy of anomaly detection.
[0050] S410 Extract Temporal Features
[0051] Using the LSTM model, based on the timestamps of the log key sequence, extract the time interval features between each log key and its preceding and succeeding log keys. The LSTM network can effectively process time series data, learn the temporal dependencies of log events, and thus perform well in capturing the time patterns of abnormal events.
[0052] S420 Extract Frequency Features
[0053] Extract frequency features by analyzing the occurrence frequencies of log keys within a specific time window. Frequency features can help the model identify abnormal frequent or abnormal reduction situations of log events, which are usually closely related to potential abnormal behaviors.
[0054] S500 Model Training
[0055] To overcome the limitations of the LSTM model in capturing bidirectional dependencies in the log key sequence, the present invention further selects the BERT model. The BERT model is based on the Transformer architecture and can capture more comprehensive context information in the log sequence, especially performing outstandingly when dealing with log sequences with complex dependencies. By using the BERT model and through multi-feature fusion, the performance of user behavior anomaly detection can be improved.
[0056] After feature extraction is completed, the vectors of the log key sequence, temporal features, and frequency features are concatenated to form a complete input feature vector. The purpose of feature fusion is to combine the features of multiple information sources so that the model can consider both the content of log events and their temporal frequency features simultaneously.
[0057] In the model training part, two self-supervised learning tasks are adopted for the training of BERT, namely the "Masked LogKey Prediction (MLKP)" task and the "Temporal-Frequency Joint Anomaly Detection (TFJAD)" task.
[0058] The purpose of the "Masked Log Key Prediction (MLKP)" task is to effectively capture the context dependencies in the log key sequence. The process of the MLKP task is as Figure 2 shown.
[0059] By randomly masking some log keys during training and requiring the model to predict these masked log keys, MLKP enables the model to learn the common patterns in normal log sequences. The design of this task helps the model to identify abnormal log sequences that deviate from these normal patterns in actual applications, thereby improving the accuracy of anomaly detection.
[0060] In the MLKP task, first, a part of the log keys are randomly replaced with special "[MASK]" tokens. BERT receives the log sequence containing these "[MASK]" tokens as input and generates the context embedding representation of each log key through the Transformer encoder. For the i-th masked log key, its corresponding context embedding representation is The model inputs this representation into the Softmax function and outputs a probability distribution over the set of log keys K:
[0061]
[0062] where, W C and b C are trainable parameters.
[0063] Next, the present invention adopts the cross-entropy loss function as the objective function of the MLKP task:
[0064]
[0065] where, represents the true value of the i-th masked log key in the j-th sequence, M is the total number of masked log keys, and N is the number of log sequences in the training set.
[0066] In this way, BERT can effectively learn the global context patterns in normal log sequences. After training, the model can predict the log keys randomly masked in normal log sequences, thereby capturing the normal patterns of log sequences. Utilizing this ability, BERT can identify abnormal log sequences that deviate from these normal patterns.
[0067] In addition to the MLKP task, the present invention also designs a self-supervised learning task called "Temporal-Frequency Joint Anomaly Detection (TFJAD)".
[0068] The TFJAD task combines the temporal features and frequency features of log keys, enabling the model to not only capture the context information of log keys but also identify log patterns that are abnormal in terms of time and frequency. This design helps the model identify more complex and potential attack behaviors, enhancing the robustness and effectiveness of anomaly detection. The overall process of the TFJAD task is as Figure 3 shown.
[0069] In the TFJAD task, the present invention not only focuses on the context information of the log keys themselves but also considers the temporal gap and frequency features of each log occurrence. Specifically, the present invention jointly encodes each log key in the log sequence with its corresponding temporal gap and frequency features, enabling the model to learn the temporal and frequency patterns between log keys in this way.
[0070] The present invention defines the temporal gap as the time difference between two consecutive log keys, denoted as Δt i , while the frequency feature represents the frequency of occurrence of a certain log key within a specific time window, denoted as f i . In the TFJAD task, the input to the model includes not only the original log key sequence k1, k2, …, k T , but also the corresponding temporal gap sequence Δt1, Δt2, …, Δt T-1 and the frequency feature sequence f1, f2, …, f T .
[0071] For each log key k i , the present invention fuses its context embedding representation with the temporal and frequency features:
[0072]
[0073] where represents the context embedding vector of the log key k i , and W t and W f are the trainable weight matrices for the temporal and frequency features respectively.
[0074] To make full use of the self-attention mechanism of the BERT model, the TFJAD task does not directly predict the log keys themselves, but focuses on capturing the abnormal patterns of time and frequency features in the log sequence. Specifically, the present invention fuses the time-frequency vectors as the input, and captures the mutual relationships between the log keys in the sequence and the dependencies in time and frequency through the self-attention mechanism of BERT. Through the self-attention mechanism, the model can learn the potential associations between time and frequency features across multiple time steps, so as to better identify abnormal patterns.
[0075] TFJAD's focus on context information mainly focuses on capturing abnormal behaviors related to time and frequency, while MLKP can improve the model's understanding of the overall log key sequence by learning context dependencies. MLKP and TFJAD complement each other to jointly enhance the BERT anomaly detection ability.
[0076] Finally, the ReLU activation function is selected, and the mean squared error is used to calculate the loss value. The present invention designs a scoring mechanism to evaluate the anomaly degree of each log sequence. Specifically, the present invention inputs the last-layer representation H BERT of the BERT model into the non-linear activation function ReLU and a linear transformation layer to generate a scalar value s i to represent the anomaly score of the log sequence:
[0077] s i = ReLU(W s H BERT + b s )
[0078] where W s is the trainable linear transformation weight matrix, and b s is the bias term. The linear transformation layer is used to map the high-dimensional BERT output to a scalar space, and the non-linear activation function introduces the non-linear expression ability of the model, so that the model can better capture complex abnormal patterns. The higher the anomaly score s i , the more likely the time and frequency features contained in the log sequence are to show anomalies.
[0079] The present invention uses the mean squared error (MSE) loss function to measure the difference between the predicted anomaly score and the true label:
[0080]
[0081] where y i is the corresponding true label (0 represents normal, 1 represents abnormal), and N is the number of log sequences in the training set.
[0082] After training the model, the trained model is used to classify the test data, and evaluation metrics are combined to evaluate the detection results.
[0083] This embodiment also provides comparative experimental data of the method of the present invention and the prior art on several real log data sets, and uses precision, recall rate and F1 value to evaluate the performance of the model. In order to verify the effectiveness of the proposed method (OurMethod), the present invention conducts comparative experiments on various user behavior anomaly detection methods. Four existing user behavior anomaly detection methods are used for comparison, namely: iForest Isolation Forest, LogAnomaly, DeepLog, and LogBERT.
[0084] The performance comparison results of this experiment on the HDFS data set and the BGL data set are as Figure 4 shown, where the left figure is the experimental result of user behavior anomaly detection on the HDFS data set, and the right figure is the experimental result of user behavior anomaly detection on the BGL data set. From Figure 4 it can be seen that for the LogBERT method that also uses the BERT model, the anomaly detection ability of the method of the present invention is more excellent, because this method extracts more-dimensional features. From Figure 4 it can also be seen that the F1 values of the present invention on the HDFS data set and the BGL data set are 84.5 and 94.2 respectively, which are better than other existing user behavior anomaly detection methods, proving the effectiveness of the method of the present invention.
[0085] The performance comparison results with existing user behavior anomaly detection methods show that the user behavior anomaly detection model proposed by the present invention has better prediction accuracy for the detection of abnormal logs and also has the potential for practical applications.
[0086] The above-described embodiments only express the implementation manners of the present invention, and the description thereof is relatively specific, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for training an abnormal detection model of user behavior logs, comprising: Obtaining a log data set; Parsing the logs according to the corresponding features of each log in the log data set by using a corresponding splitting method to obtain log keys; Performing a serialization operation on all the log keys to obtain a log key sequence; Extracting time features and frequency features from the log key sequence; Inputting the concatenated log key sequence, the time features, and the frequency features into a neural network model for training to obtain a trained abnormal detection model of user behavior logs.
2. The method according to claim 1, wherein Using a set tree structure as the corresponding splitting method according to the number and length of tokens in each log.
3. The method according to claim 2, wherein Parsing the logs by using the Drain tool.
4. The method according to claim 1, wherein Extracting the time features by using an LSTM model.
5. The method according to claim 1, characterized in that, Using a BERT model as the neural network model; the BERT model adopts a self-supervised learning task.
6. The method according to claim 5, wherein The last layer of the BERT model outputs a scalar value for scoring the logs.
7. An abnormal detection method for user behavior logs, comprising: Using the trained abnormal detection model of user behavior logs according to any one of claims 1-6 to perform abnormal detection on user behavior logs.
8. A training system for an abnormal detection model of user behavior logs, comprising: A data acquisition module for acquiring a log data set; A log parsing module for parsing the logs according to the corresponding features of each log in the log data set by using a corresponding splitting method to obtain log keys; A time series sorting module for performing a serialization operation on all the log keys to obtain a log key sequence; A feature extraction module for extracting time features and frequency features from the log key sequence; A model training module for inputting the concatenated log key sequence, the time features, and the frequency features into a neural network model for training to obtain a trained abnormal detection model of user behavior logs.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for executing the method according to any one of claims 1 to 7.
10. A storage medium storing a computer program, which when executed by a computer, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
DCS system long-term health evaluation method and system based on spatial-temporal feature extraction
CN121030706A