Log anomaly detection model updating method based on experience playback

By adopting a model update method based on empirical playback in log exception detection, the problem that log exception detection model in the prior art is difficult to adapt to data flow updates, and the continuous learning and efficient detection of the model are realized, reducing the occurrence of 'catastrophic forgetting' phenomenon.

CN120216306APending Publication Date: 2025-06-27NANJING NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133019.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing log anomaly detection methods are difficult to effectively adapt to data flow updates, which leads to the model that it may forget the knowledge it has learned before when learning new data, especially when the distribution of the training data changes, it is prone to 'catastrophic forgetting'.

Method used

The log anomaly detection model update method based on empirical playback is adopted, and the paradigm samples are extracted from the original samples through Kmeans clustering, and the full empirical playback is performed in the replay buffer. The model is updated to adapt to data changes in combination with the dark empirical playback strategy and distillation loss.

Benefits of technology

It realizes learning new data while retaining existing knowledge, enhances the model's continuous learning ability, reduces the occurrence of "catastrophic forgetting", improves the detection model's adaptability to new models, and the detection accuracy rate reaches more than 98%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216306A_ABST
    Figure CN120216306A_ABST
Patent Text Reader

Abstract

The invention discloses a log anomaly detection model updating method based on experience playback, and the method comprises the steps: extracting a paradigm sample from an original sample through employing a Kmeans clustering method, and putting the paradigm sample into a playback buffer area; retaining an original category score zT (S) corresponding to the example sample and an intermediate layer feature; in a model updating process, based on a dark experience playback strategy, jointly training the model by using an example sample in a playback buffer area and a new sample; and performing complete experience playback on the example sample in the playback buffer area on the CNN feature fusion layer of the MLog, and reserving an original category score zT (S) and middle layer features corresponding to the example sample until incremental updating of the log anomaly detection model is completed. The log anomaly detection method can be helped to adapt to data changes, new data are learned under the condition that existing knowledge is reserved, and continuous and effective detection is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security, relates to log anomaly detection technology, and particularly relates to a method for updating a log anomaly detection model based on experience replay. Background Art

[0002] The scale of modern software systems is continuously increasing, and the structure is becoming more and more complex, which greatly increases the difficulty of software system maintenance and poses huge challenges to system management and security. Logs have become an important information source for maintaining large-scale software systems. The industry has begun to focus on log-based anomaly detection research to discover and identify abnormal events in software systems, so as to identify weak links and avoid the occurrence of system risks. Log anomaly detection plays an important role in maintaining system security, improving system reliability and stability, and preventing failures. At present, log anomaly detection has been applied in various information systems such as supercomputers, industrial cloud systems, and Internet of Things application systems.

[0003] In recent years, the development of machine learning and deep learning technologies has provided more effective solutions to the problem of log anomaly detection. Many studies are dedicated to solving the problem of unstable log data caused by log evolution and noise. However, the distribution or pattern of log data changes over time, which is also one of the reasons for the instability of logs, thus affecting the effect of log anomaly detection. Current researchers have noticed this problem and carried out relevant research. After investigation, it is found that although existing solutions attempt to adapt to data changes by using misjudgment information incremental update or periodic full-scale update, they still lack the ability of continuous learning and cannot fully utilize new data for incremental update.

[0004] Introducing incremental learning technology in log anomaly detection research will face the problem of balancing old and new knowledge. In an actual log detection system, the log detection model will face continuously incoming log data. When the detection model learns new data, it may forget the knowledge learned before. Especially when the distribution of training data changes, the phenomenon of "catastrophic forgetting" is likely to occur.

[0005] Therefore, a new technical solution is needed to solve this problem. Summary of the Invention

[0006] Object of the Invention: To solve the problem that the log anomaly detection method cannot well adapt to the update of data streams, a method for updating a log anomaly detection model based on experience replay is provided, which can help the log anomaly detection method adapt to data changes, learn new data while retaining existing knowledge, and achieve continuous and effective detection.

[0007] Technical Solution: To achieve the above object, the present invention provides a method for updating a log anomaly detection model based on experience replay, including the following steps:

[0008] S1: Use Kmeans clustering method to extract example samples from original samples and put them into replay buffer; retain the original category score z corresponding to the example samples T (S) and intermediate layer characteristics;

[0009] S2: During the model update process, based on the dark experience replay strategy, the model is trained using the example samples in the replay buffer and the new samples;

[0010] The example samples in the replay buffer are fully replayed on the CNN feature fusion layer of MLog, retaining the original category score z corresponding to the example samples T (S) and intermediate layer features until the incremental update of the log anomaly detection model is completed.

[0011] Furthermore, the replay buffer in step S1 is represented as

[0012]

[0013] in, are the old model parameters, and are the original sample scores and intermediate layer features of the example samples, respectively.

[0014] Furthermore, when the complete experience playback is performed in step S2, the complete loss function of the example sample for each round of update is expressed as:

[0015]

[0016] in, is the loss of the intermediate features of the convolutional layer; is the loss of the original category score of the example sample; Represents the label loss of the example samples in the buffer;

[0017] Furthermore, in the complete loss function, cross entropy loss is used to measure the 0 / 1 label loss. The expression is:

[0018]

[0019] Furthermore, in the complete loss function, the current model parameters θ and the old model parameters are used Calculate the response to the example sample The expression is:

[0020]

[0021] Furthermore, in the complete loss function, The calculation formula is:

[0022]

[0023] Furthermore, in step S2, the incremental update updates the model based on the trained model; given the model representation obtained from the current training and the old data set L T , where T represents the historical time interval; assume that the new data set collected after the new time interval ΔT is L ΔT , and the incremental learning task is to update the new classifier based on the existing model where L' ∈ L T ∪L ΔT , Y' is the corresponding label set, and the updated detection model maintains a good detection effect for both new and old data instances.

[0024] Furthermore, the dark experience replay strategy in step S2 includes:

[0025] The incremental update is regarded as multiple sequential classification tasks, and the classifier f with parameter θ is optimized in chronological order; use z θ (S) to represent the original class score output by the model, to represent the probability distribution of the class; the learning objective is to correctly classify the log sequence samples of t ∈ {1, …, T + ΔT} when the current updated time interval ΔT is given:

[0026]

[0027] Furthermore, in step S2, the model mimics its original response to the example samples and minimizes the following objective:

[0028]

[0029] Under mild assumptions, the optimization of the KL divergence is equivalent to minimizing the Euclidean distance between the corresponding type of original score and the predicted class score; the optimization objective during model update is finally transformed into:

[0030]

[0031] where, is the old model parameter, and θ is the model parameter that needs to be updated and trained.

[0032] Furthermore, in step S2, DER++ introduces an additional coefficient ζ to balance the last term in the objective function, and the optimization objective is:

[0033]

[0034] To improve the incremental learning ability of the detection model, the present invention introduces the Experience Replay (ER) technology of Incremental Learning in the research of log anomaly detection. Experience Replay uses the additionally stored historical data samples or the learned features to help the model "recall" knowledge, reintroducing past experiences into the model training process, thereby achieving the continuous learning of the model and effectively coping with the adverse effects caused by data distribution and pattern changes.

[0035] Incremental learning methods need to adopt appropriate strategies to retain the key knowledge learned and constrain model training during the update stage, balancing the weights of new and old knowledge to maintain the overall performance of the model.

[0036] The main protection point of the present invention is an update method for a log anomaly detection model based on Experience Replay. Existing methods mainly focus on the evolution of log statements caused by system updates or parsing errors, dealing with a small number of newly added sequences caused by noise or new execution paths, and there is relatively limited research on the distribution and pattern changes of log data. Although existing solutions attempt to adapt to data changes using misjudgment information incremental updates or periodic full updates, they still lack the ability of continuous learning and cannot fully utilize new data for incremental updates. To address the above problems, the present invention designs and implements an update method for a log anomaly detection model based on Experience Replay. The basic idea is to improve on an advanced log anomaly detection method to enhance the continuous learning ability of the model. When a small batch of new data is generated, the model is updated using the new data and the exemplar samples extracted from the old data. The update method is based on the Dark Experience Replay (DER) strategy and retains important intermediate features for Full Experience Replay (FER). The difference from training is that during update, the exemplar samples use the distillation loss, storing the original class scores and intermediate features of the exemplar samples and replaying them during model update, thereby effectively retaining the previously learned knowledge.

[0037] Beneficial effects: Compared with the prior art, the present invention can continuously and effectively detect new data, improving the adaptability of the anomaly detection model to newly emerging patterns. At the same time, the present invention extracts the key features of the exemplar samples to achieve full experience replay, improving the overall learning ability and performance of the model, effectively coping with the catastrophic forgetting phenomenon, and achieving a detection accuracy of over 98% for both new and old samples under the premise of consuming less time resources. Brief Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the model incremental update method provided by the present invention;

[0039] Figure 2It is a graph showing the change of accuracy with model updates;

[0040] Figure 3 It is a graph showing the change of recall with model updates;

[0041] Figure 4 It is a graph showing the change of F1-score with model updates;

[0042] Figure 5 It is a graph showing the loss change of full update;

[0043] Figure 6 It is a comparison graph of different update strategies. Detailed implementation manners

[0044] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0045] As Figure 1 shown, the present invention provides a method for updating a log anomaly detection model based on experience replay, including the following steps:

[0046] S1: Use the Kmeans clustering method to extract exemplar samples from the original samples and put them into the replay buffer; retain the original class scores z T (S) and the intermediate layer features;

[0047] The replay buffer is represented as

[0048]

[0049] Wherein, are the old model parameters, and are respectively the original sample scores and intermediate layer features of the exemplar samples.

[0050] After each update is completed, the replay buffer is used to store the exemplar samples and their corresponding key features and original class scores. When updating next time, the exemplar samples in the buffer are read to train the detection model, which helps the model consolidate the knowledge it has learned and reduce forgetting.

[0051] The overall framework for anomaly detection model update is mainly divided into two modules: the log anomaly detection module and the incremental update module. In the log anomaly detection part, the present invention adopts a new log anomaly detection method MLog based on log template semantic information and hybrid neural network, which includes three steps: log parsing, log template vectorization, and anomaly detection model training and prediction. The incremental update module implements an incremental update method based on the Dark Experience Replay (DER) strategy. As Figure 1 shown, the model parameters are updated using the new data and representative old data (i.e., exemplar samples) collected over a period of time. Specifically, on the basis of having fully trained the model using the original dataset, after collecting new samples with a sufficient time span, representative samples are extracted from the old data to generate an exemplar sample set, and the exemplar samples and a small batch of new samples are jointly used to update the model, where the distillation loss and cross-entropy loss are respectively applied to the exemplar samples and new samples. To reduce the impact of samples with large training biases being sampled and replayed, the incremental update module also uses true labels for some exemplar samples to enhance the classification consistency between the old and new models. The method of the present invention further uses Full Experience Replay (FER) to strengthen the supervision of the feature fusion layer for the MLog model structure, so as to retain more learned key features.

[0052] Cross-entropy loss:

[0053] Distillation loss:

[0054] S2: During the model update process, based on the Dark Experience Replay strategy, the exemplar samples and new samples in the replay buffer are jointly used to train the model;

[0055] The exemplar samples in the replay buffer perform full experience replay on the CNN feature fusion layer of MLog, retaining the original class scores z T (S) and intermediate layer features until the incremental update of the log anomaly detection model is completed;

[0056] Incremental update is to update the model on the basis of the trained model; given the current model representation obtained by training and the old data set L T , T represents the historical time interval; assuming that the new data set collected after the new time interval ΔT is L ΔT , the incremental learning task is to update a new classifier on the basis of the existing model where L' ∈ L T ∪L ΔT , Y' is the corresponding label set, and the updated detection model maintains good detection effects for both old and new data instances.

[0057] The dark experience replay strategy includes:

[0058] The incremental update is regarded as multiple ordered classification tasks, and the classifier f with parameter θ is optimized in chronological order; using z θ (S) represents the original class scores output by the model, representing the probability distribution of the classes; the learning objective is to be able to correctly classify the log sequence samples of t ∈ {1, …, T + ΔT} given the current updated time interval ΔT:

[0059]

[0060] In the training set of example samples the original class scores z T (S) corresponding to the samples are retained, and these scores provide a richer information description of the data points. To retain as much knowledge as the model has already learned, the model mimics its original response to the example samples and minimizes the following objective:

[0061]

[0062] Under mild assumptions, the optimization of the KL divergence is equivalent to minimizing the Euclidean distance between the corresponding type of original scores and the predicted class scores, and the optimization objective during model update is finally transformed into:

[0063]

[0064] where, are the old model parameters, and θ is the model parameters to be updated and trained.

[0065] DER++ introduces an additional coefficient ζ to balance the last term in the objective function, and the optimization objective is:

[0066]

[0067] This coefficient is a hyperparameter that is predefined before the start of model training, and the setting of the hyperparameter varies depending on the dataset and buffer size. Adding the last term of the objective function is to reduce the adverse effects caused by the replay of mispredicted samples.

[0068] DER++ promotes the model to reach a flatter minimum by optimizing the objective function. This flatter minimum helps the model explore neighboring regions in the parameter space, improves the model's tolerance to local perturbations, and reduces the forgetting of knowledge.

[0069] Full experience replay:

[0070] The present invention performs Full Experience Replay (FER) on the CNN feature fusion layer of MLog. The CNN feature fusion layer can capture both the global and local dependencies in the log sequence, enhancing the attention to important features that contain more key knowledge and enabling the model to better retain the memory of past tasks. Full Experience Replay (FER) preserves the complete information of past samples (including inputs, original class scores, and intermediate layer features), keeping the knowledge stable across different layers of the feature fusion layer and further effectively reducing knowledge forgetting.

[0071] When performing full experience replay, the complete loss function for the exemplar samples used in each round of update is expressed as:

[0072]

[0073] Where, is the loss of the intermediate features of the convolutional layer; is the loss of the original class scores of the exemplar samples; represents the label loss of the exemplar samples in the buffer;

[0074] In the complete loss function, the cross-entropy loss is used to measure the 0 / 1 label loss The expression is:

[0075]

[0076] In the complete loss function, the responses of the exemplar samples using the current model parameters θ and the old model parameters are used to calculate The expression is:

[0077]

[0078] In the complete loss function, The calculation formula is:

[0079]

[0080] In the above way, FER provides an effective mechanism to enhance the learning ability of the model. Especially when facing multi-round update tasks, it can significantly improve the performance and stability of the model.

[0081] The method of the present invention reduces the loss of information in the probability space by introducing an incremental learning method, replaying the original class scores and intermediate layer features of the exemplar samples, improving the learning efficiency of the model, enhancing the adaptability of the model, enabling it to maintain high detection performance in practical applications, and providing support for long-term monitoring and detection.

[0082] To verify the effectiveness and effect of the incremental update method provided by the present invention, this embodiment conducts a comparative verification through experiments, which is specifically as follows:

[0083] This embodiment uses a real dataset to evaluate the update method of the incremental update log anomaly detection model to verify the effectiveness of the research method. First, the experimental settings are described in detail, including the dataset, baseline methods, experimental preparations, and evaluation metrics. Secondly, the experiment verifies the impact of each round of update on the detection effect and designs experiments to compare the effects of different update strategies and strategy combinations. This embodiment also compares the incremental update log anomaly detection method with advanced baseline methods and analyzes the advantages and disadvantages brought by incremental update. Specifically, it includes:

[0084] I. Experimental Settings

[0085] The experiment uses the open-source real-world dataset BGL provided by LogHub, which is an open log dataset collected from the BlueGene / L supercomputer system of Lawrence Livermore National Laboratory in Livermore, California, and contains 4,747,963 log entries. BGL sets an alarm label for each log entry. In the first column of the log, "-" indicates a non-alarm message, while other symbols indicate alarm messages. Compared with another commonly used log dataset HDFS, the BGL dataset has a longer time span. After statistical analysis of the BGL dataset, it is found that the distribution of log sequence data changes over time, which is consistent with the motivation of the present invention. The sequence generation adopts a sliding window mechanism, and a total of 16,432 valid sequences are generated. For the generated sequence set, it is divided into a training set, a validation set, and a test set, and then the training set is further divided into an original training set and an update set according to the time span.

[0086] Two baseline methods are adopted in the experiment: PCA and DeepLog. PCA generates normal and abnormal subspaces by finding patterns in the event count vector, and then determines whether a log sequence is abnormal by calculating the projection length of the template count vector on the abnormal subspace. If the projection length corresponding to the log sequence is greater than a predefined threshold, it will be regarded as abnormal. DeepLog uses a long short-term memory network (LSTM) to learn the log sequence pattern, predicts the next possible log event for the current given sequence, and reports it as abnormal when the log pattern deviates from the model prediction.

[0087] The research on log-based anomaly detection generally adopts the common evaluation metrics in classification tasks, namely precision, recall, and F1-score. The calculation formulas are as follows:

[0088]

[0089] II. Experimental Results and Analysis

[0090] To verify the effectiveness of the incremental learning method, this embodiment designs experiments from two aspects: detection performance and time efficiency, and compares the method of the present invention with other advanced detection methods. First, the change in the detection performance of the basic method MLog using the incremental update method for learning is shown. Specifically, 65% of the BGL dataset is extracted as the training set, 5% as the validation set, and 30% as the test set. Then, the BGL training dataset and the validation set are divided according to time. The time span of the BGL dataset is six months. The data in the first two months is used as the initial training set, and the data generated every 15 days thereafter is used as the updated training set. Specifically, first, the initial model is trained using the initial training set from July to August. Then, starting from September 16th, the model is updated every half month. The detection effects after learning using the incremental update method based on experience replay and the full update method are compared, as Figures 2 to 4 respectively showing the changes in the accuracy, recall rate, and F1-score of the two update methods with the number of update rounds;

[0091] From Figures 2 to 4 the experimental results, it can be seen that the change trends of the incremental update and the full update methods are generally the same, and the index gap is small, indicating that the incremental update method using only a small amount of data can learn the information contained in the new samples while retaining the old knowledge, achieving an effect comparable to that of the full update. Generally speaking, the incremental update method is slightly inferior to the full update method in terms of detection effect because the full update can ensure data integrity and consistency, and the experimental results are in line with expectations. However, the full update requires more time resources and computing resources. When performing full training, an early stopping strategy is adopted, and the training termination condition is set according to the change in loss. After multiple experiments, the changes in the training loss and the validation loss are plotted as Figure 5 shown.

[0092] The time consumed by the full update method and the incremental method in each round of update is compared in Table 1. The training duration of the full update mainly shows an increasing trend as the training samples increase. The incremental update only uses the newly generated samples and the exemplar samples in this time period for training, and the time used is about 20% - 30% of the full training, and the improvement effect of the time index is significant.

[0093] Table 1 Statistical Situation of the Time Used by the Update Method

[0094]

[0095] In this experiment, the final detection effect on the BGL dataset after the update was compared with the basic method MLog and other non-incremental baseline methods. The method of the present invention showed slightly lower performance than the full-scale baseline MLog after multiple rounds of incremental updates, but higher than PCA and DeepLog trained with the full-scale data, as shown in Table 2:

[0096] Table 2 Comparison of the effects of the method of the present invention and the baseline methods

[0097]

[0098] This embodiment also designed an ablation experiment to compare the detection effects of different update strategies on the BGL dataset. Figure 6 The F1-scores of multiple-round update experiments using three incremental update methods, DER, DER++, and FER, were shown. The experimental results showed that FER with intermediate layer feature constraints performed better than DER and DER++ that only used the category score constraint model for update, proving the effectiveness of using complete empirical constraints in the feature fusion layer of the method of the present invention.

Claims

1. A log anomaly detection model updating method based on experience playback, characterized in that: The steps include: S1: Use Kmeans clustering method to extract example samples from original samples and put them into replay buffer; retain the original category score z corresponding to the example samples T (S) and intermediate layer characteristics; S2: During the model update process, based on the dark experience replay strategy, the model is trained using the example samples in the replay buffer and the new samples; The example samples in the replay buffer are fully replayed on the CNN feature fusion layer of MLog, retaining the original category score z corresponding to the example samples T (S) and intermediate layer features until the incremental update of the log anomaly detection model is completed.

2. According to claim 1, a log anomaly detection model updating method based on experience playback is characterized in that: The replay buffer in step S1 is represented as in, are the old model parameters, and are the original sample scores and intermediate layer features of the example samples, respectively.

3. According to claim 1, a log anomaly detection model updating method based on experience playback is characterized in that: When the complete experience replay is performed in step S2, the complete loss function of the example sample for each round of update is expressed as: in, is the loss of the intermediate features of the convolutional layer; is the loss of the original category score of the example sample; Represents the label loss of the exemplar samples in the buffer.

4. According to the method for updating the log anomaly detection model based on experience playback according to claim 3, it is characterized in that: In the complete loss function, cross entropy loss is used to measure 0 / 1 label loss. The expression is:

5. According to claim 3, a log anomaly detection model updating method based on experience playback is characterized in that: In the complete loss function, the current model parameters θ and the old model parameters are used Calculate the response to the example sample The expression is:

6. According to the method for updating the log anomaly detection model based on experience playback in claim 3, it is characterized in that: In the complete loss function, The calculation formula is:

7. The log anomaly detection model updating method based on experience playback according to claim 1 is characterized in that: The incremental update in step S2 is to update the model based on the trained model; given the model representation obtained by the current training and the old data set L T , T represents the historical time interval; assuming that the new data set collected after the new time interval ΔT is L ΔT The incremental learning task is to update the new classifier based on the existing model. where L'∈L T ∪L ΔT , Y' is the corresponding label set.

8. The log anomaly detection model updating method based on experience playback according to claim 7 is characterized in that: The dark experience playback strategy in step S2 includes: The incremental update is viewed as multiple sequential classification tasks, and the classifier f with θ as parameter is optimized in time order; θ (S) represents the original category score output by the model, Represents the probability distribution of the category; the learning goal is to correctly classify the log sequence samples of t∈{1,…,T+ΔT} given the current updated time interval ΔT:

9. The log anomaly detection model updating method based on experience playback according to claim 7 is characterized in that: In step S2, the model imitates its original response to the example sample and minimizes the following objectives: Under mild assumptions, the optimization of KL divergence is equivalent to minimizing the Euclidean distance between the corresponding type raw score and the predicted category score; the optimization objective during model update is ultimately transformed into: in, is the old model parameter, and θ is the model parameter that needs to be updated and trained.

10. The log anomaly detection model updating method based on experience playback according to claim 9 is characterized in that: In step S2, DER++ introduces an additional coefficient ζ to balance the last term in the objective function. The optimization objective is: