Self-updating method and device of log anomaly detection model and electronic equipment

By building an experience pool and online detection mechanism for log exception detection models, the problem of low accuracy of log exception detection in the existing technology is solved, and the self-update and real-time detection capabilities of the model are improved.

CN120066896APending Publication Date: 2025-05-30CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152379.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing log anomaly detection model has low detection accuracy due to feature offsets and untimely data updates, making it difficult to adapt to the rapidly changing online environment.

Method used

Build a model experience pool by obtaining the initial log dataset and build an initial exception detection model in the offline state of the log exception detection system. In the online state, abnormal detection is performed on real-time log stream data, model suitability is evaluated, and model self-updation is performed based on the suitability.

Benefits of technology

It effectively improves the real-time detection capability of the model and the flexibility of model updates, makes full use of the value of real-time data, meets the real-time detection needs of massive log data, improves the accuracy of abnormal detection results, and reduces the cost of system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066896A_ABST
    Figure CN120066896A_ABST
Patent Text Reader

Abstract

The invention provides a self-updating method and device of a log anomaly detection model and electronic equipment. The method comprises the steps that an initial log data set is obtained, a model experience pool is constructed based on the initial log data set, an initial anomaly detection model is constructed when a log anomaly detection system is in an offline state, the model experience pool is adopted to train the initial anomaly detection model, and the log anomaly detection model is obtained, the initial log data set comprises abnormal log data of multiple log modes and log abnormal types corresponding to the abnormal log data; under the condition that the log anomaly detection system is in an online state, carrying out anomaly detection on the real-time log stream data by adopting a log anomaly detection model to obtain a detection result; and evaluating the adaptability of the log anomaly detection model according to the detection result, and updating the log anomaly detection model according to the adaptability. The problem that log anomaly detection in the prior art is low in detection accuracy is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of self-update of log anomaly detection models, and more specifically, to a method, device, computer-readable storage medium, and electronic device for self-update of a log anomaly detection model. Background Art

[0002] Enterprises and institutions involve many business processes during operation, and a series of log data will be generated during the execution of the processes. By using these log data, enterprises and institutions can analyze, monitor, and predict the processes, discover anomalies and risks in the business processes, and accordingly adjust, improve, and enhance the business processes. With the continuous upgrade of systems and businesses, traditional offline anomaly detection for process execution historical logs cannot meet the growing needs of real-time analysis, prediction, and early warning due to serious information lag, and online anomaly detection for real-time log streams has become a new development trend.

[0003] In practical applications, the characteristics of log stream data change dynamically over time, that is, the "feature drift" phenomenon occurs, resulting in a gradual decline in the detection performance of existing models. This situation poses a severe challenge to real-time anomaly detection in particular. Existing anomaly real-time detection algorithms based on automata usually rely on static data sets for training, and the anomaly determination threshold is fixed, making it difficult to adapt to the rapidly changing online environment. Therefore, there are significant limitations in terms of anomaly detection accuracy and flexibility.

[0004] In addition, as new log data continuously pours in, the log anomaly detection model based on a deep neural network needs to be continuously updated to learn new features in the new data. To ensure that the model maintains its memory of old knowledge while learning new knowledge, a simple strategy is to jointly train the old and new data. However, this method requires storing a large amount of historical data, resulting in high consumption of storage and computing resources and low efficiency in practical applications. If only new data is used for model fine-tuning, it will lead to knowledge forgetting and a decline in the overall detection performance of the model. Summary of the Invention

[0005] The main objective of the present application is to provide a method, device, computer-readable storage medium, and electronic device for self-update of a log anomaly detection model, so as to at least solve the problem of low detection accuracy in the log anomaly detection of the prior art.

[0006] To achieve the above object, according to one aspect of the present application, there is provided a self-update method for a log anomaly detection model, including: obtaining an initial log data set, constructing a model experience pool based on the initial log data set, and constructing an initial anomaly detection model when the log anomaly detection system is in an offline state, and training the initial anomaly detection model with the model experience pool to obtain a log anomaly detection model, wherein the initial log data set includes abnormal log data of various log patterns and the corresponding log anomaly types; when the log anomaly detection system is in an online state, using the log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result; evaluating the applicability of the log anomaly detection model according to the detection result, and updating the log anomaly detection model according to the applicability.

[0007] Optionally, evaluating the applicability of the log anomaly detection model according to the detection result includes: according to the first formula: determining the applicability of the log anomaly detection model, where S is the applicability, λ is a weight parameter, σmax is the maximum value of the standard deviation, e t is the error rate in the t-th time window, σ is the standard deviation of the error rates in multiple time windows, T is the total number of time windows, and μ is the average value of the error rates in all time windows.

[0008] Optionally, updating the log anomaly detection model according to the applicability includes: determining the self-update threshold and replacement threshold of the log anomaly detection model; when it is determined that the applicability is greater than the self-update threshold and less than the replacement threshold, retraining the log anomaly detection model to update the log anomaly detection model; when it is determined that the applicability is greater than or equal to the replacement threshold, performing iterative processing on the log anomaly detection model to update the log anomaly detection model.

[0009] Optionally, determining the self-update threshold and replacement threshold of the log anomaly detection model includes: according to the second formula: determining the self-update threshold and the replacement threshold of the log anomaly detection model, where θ u is the self-update threshold, θ r is the replacement threshold, is the average error rate in multiple time windows, is the average standard deviation, max(e t ) is the maximum error rate in multiple time windows, max(σ) is the maximum standard deviation, and α, β, γ, δ are weight coefficients.

[0010] Optionally, during the process of performing anomaly detection on real-time log stream data using the log anomaly detection model, the method further includes: determining the degree of feature deviation of data features according to the real-time log stream data, where the data features include data association features, data time series features, and data statistical features; determining the feature deviation log data in the real-time log stream data according to the degree of feature deviation, and adding the feature deviation log data to the model update pool of the log anomaly detection model.

[0011] Optionally, updating the log anomaly detection model according to the applicability includes: when it is determined that the applicability is greater than the self-update threshold, performing an update operation on the log anomaly detection model using the feature deviation log data in the model update pool and the initial log data set in the model experience pool.

[0012] Optionally, determining the degree of feature deviation of data features according to the real-time log stream data includes: using the chi-square test algorithm to determine the data association features of the real-time log stream data, using the dynamic time warping algorithm to determine the data time series features of the real-time log stream data, using the KL divergence algorithm to determine the data statistical features of the real-time log stream data; determining the degree of feature deviation according to the data association features, the data time series features, and the data statistical features.

[0013] According to another aspect of the present application, there is provided a self-update device for a log anomaly detection model, including: a training unit, configured to obtain an initial log data set, construct a model experience pool based on the initial log data set, and construct an initial anomaly detection model in an offline state of the log anomaly detection system, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model, where the initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types; an anomaly detection unit, configured to perform anomaly detection on real-time log stream data using the log anomaly detection model in an online state of the log anomaly detection system to obtain a detection result; an update unit, configured to evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

[0014] According to still another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the self-update methods of the log anomaly detection model.

[0015] According to another aspect of the present application, an electronic device is provided, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for self-updating any one of the log anomaly detection models.

[0016] Applying the technical solution of the present application, an initial log data set is obtained, a model experience pool is constructed based on the initial log data set, an initial anomaly detection model is constructed when the log anomaly detection system is in an offline state, and the initial anomaly detection model is trained using the model experience pool to obtain a log anomaly detection model, wherein the initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types; when the log anomaly detection system is in an online state, the log anomaly detection model is used to perform anomaly detection on real-time log stream data to obtain a detection result; the applicability of the log anomaly detection model is evaluated according to the detection result, and the log anomaly detection model is updated according to the applicability. Through model applicability evaluation and model adaptive update mechanism, the real-time detection ability of the model and the flexibility of model update are effectively improved, the value of real-time data is fully utilized, the real-time detection requirements of a large amount of log data in actual application scenarios are met, the accuracy of anomaly detection results is improved, and the system operation cost is reduced. It solves the problem of low detection accuracy in the existing log anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The specification drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0018] Figure 1 It shows a hardware structure block diagram of a mobile terminal for executing a self-update method of a log anomaly detection model provided in an embodiment of the present application;

[0019] Figure 2 It shows a schematic flowchart of a self-update method of a log anomaly detection model provided in an embodiment of the present application;

[0020] Figure 3 It shows a schematic flowchart of a specific self-update method of a log anomaly detection model provided in an embodiment of the present application;

[0021] Figure 4 It shows a schematic diagram of segmented processing of log data provided in an embodiment of the present application;

[0022] Figure 5Shows a schematic flow diagram of feature offset analysis based on real-time log streams and empirical pool log data provided according to an embodiment of the present application;

[0023] Figure 6 Shows a structural block diagram of a self-updating device for a log anomaly detection model provided according to an embodiment of the present application.

[0024] Wherein, the above-mentioned drawings include the following reference numerals:

[0025] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed implementation manners

[0026] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0027] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so as to describe the embodiments of the present application herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:

[0030] Sliding window: A fixed-length window that moves on sequential data, used to handle problems such as time series analysis and pattern recognition. The window size and moving step can be adjusted according to the application scenario to optimize performance.

[0031] Feature offset: Refers to the phenomenon of the distribution difference between training data and actual application data, resulting in a decline in model performance.

[0032] Chi-square test: The chi-square test is a non-parametric test method in statistics used to test the significance of the difference between categorical data and the theoretical distribution.

[0033] Dynamic Time Warping (DTW) algorithm: It is a method for measuring the similarity between two sequences, especially suitable for pattern recognition on non-linear time axes, such as time alignment in speech recognition.

[0034] KL divergence: The KL divergence (Kullback-Leibler divergence) is an asymmetric measure of the difference between two probability distributions P and Q, commonly used in information theory and statistics.

[0035] Softmax function: The softmax function converts a real-valued vector into a probability distribution vector, where each element value ranges between (0,1) and the sum is 1, and is commonly used in the output layer of multi-classification problems.

[0036] CrossEntropyLoss: A commonly used loss function for measuring the difference between the probability distribution predicted by the model and the true label distribution, widely used in classification tasks.

[0037] MSE: Mean Squared Error, which calculates the average of the squares of the differences between the predicted values and the true values, and is used to evaluate the prediction accuracy of the model.

[0038] LSTM (Long Short-Term Memory): This model is a special type of Recurrent Neural Network (RNN) designed to address the vanishing or exploding gradient problems faced by traditional RNNs when dealing with long sequence data.

[0039] As introduced in the background art, the existing technology for log anomaly detection has a relatively low detection accuracy. To solve the problem of low detection accuracy in the existing technology for log anomaly detection, the embodiments of the present application provide a self-update method, device, computer-readable storage medium, and electronic device for a log anomaly detection model.

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0041] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a self-update method of a log anomaly detection model in an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1Only one processor 102 is shown (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA), and a memory 104 for storing data. Among them, the mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0042] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the self-update method of the log anomaly detection model in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0043] In this embodiment, a self-update method of a log anomaly detection model running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] Figure 2 is a flowchart of the self-update method of the log anomaly detection model according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0045] Step S201: Obtain an initial log data set, construct a model experience pool based on the above initial log data set, construct an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the above initial anomaly detection model using the above model experience pool to obtain a log anomaly detection model. Among them, the above initial log data set includes abnormal log data of multiple log patterns and the corresponding log anomaly types of the above abnormal log data;

[0046] Specifically, segment the initial log data set based on special fields in the log (such as IP, session ID, etc.) and the sliding window technique, and divide the log data into multiple sequence segments arranged according to timestamps; based on the log sequences, calculate the similarity between the sequences. The calculation of similarity uses an improved formula based on the length of the matching segment, and the similarity is obtained by comparing the continuous matching degree of the same segments in two log sequences. For log sequences with a similarity higher than the preset threshold ∈, classify them. Use the generated unique identifier to deduplicate the log sequences. Remove duplicate log sequences with the same identifier, and store the deduplicated log data in the experience pool. This can reduce storage redundancy, and at the same time ensure that the log data in the experience pool is highly representative, providing reliable data support for model update and feature analysis.

[0047] Step S202: When the above log anomaly detection system is in an online state, use the above log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result;

[0048] Step S203: Evaluate the applicability of the above log anomaly detection model according to the above detection result, and update the above log anomaly detection model according to the above applicability.

[0049] Specifically, evaluate the applicability of the model by quantitatively processing the indicators of the model log anomaly detection effect in the online environment, and finally implement the self-update and iterative process of the model according to the dynamic double-threshold decision mechanism. This method can automatically adjust the update strategy according to the performance of the model in actual applications, reduce the need for manual intervention, and improve the automation level and maintenance efficiency of the system.

[0050] Through this embodiment, an initial log data set is obtained, a model experience pool is constructed based on the initial log data set, an initial anomaly detection model is constructed when the log anomaly detection system is in an offline state, and the initial anomaly detection model is trained using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types. When the log anomaly detection system is in an online state, the log anomaly detection model is used to detect anomalies in real-time log stream data to obtain detection results. The applicability of the log anomaly detection model is evaluated based on the detection results, and the log anomaly detection model is updated according to the applicability. Through model applicability evaluation and model adaptive update mechanism, the real-time detection ability of the model and the flexibility of model update are effectively improved, the value of real-time data is fully utilized, the real-time detection requirements of massive log data in actual application scenarios are met, the accuracy of anomaly detection results is improved, and the cost of system operation is reduced. It solves the problem of low detection accuracy in log anomaly detection of the prior art.

[0051] In the specific implementation process, evaluating the applicability of the log anomaly detection model according to the above detection results includes: According to the first formula: Determine the applicability of the log anomaly detection model, where S is the applicability, λ is the weight parameter, σmax is the maximum value of the standard deviation, e t is the error rate in the t-th time window, σ is the standard deviation of the error rates in multiple time windows, T is the total number of time windows, and μ is the average value of the error rates in all time windows.

[0052] Specifically, updating the log anomaly detection model according to the above applicability includes: determining the self-update threshold and replacement threshold of the log anomaly detection model; when it is determined that the applicability is greater than the self-update threshold and less than the replacement threshold, retraining the log anomaly detection model to update the log anomaly detection model; when it is determined that the applicability is greater than or equal to the replacement threshold, performing iterative processing on the log anomaly detection model to update the log anomaly detection model.

[0053] More specifically, determining the self-update threshold and replacement threshold of the log anomaly detection model includes: According to the second formula: Determine the self-update threshold and the replacement threshold of the log anomaly detection model, where θ u is the self-update threshold, θ r is the replacement threshold, is the average error rate in multiple time windows, is the average standard deviation, max(e t) is the maximum error rate in multiple time windows, max(σ) is the maximum standard deviation, and α, β, γ, and δ are weight coefficients.

[0054] To evaluate the performance and stability of the model at different time periods, this method introduces a comprehensive evaluation index S to comprehensively evaluate the applicability of the model, and adopts a dynamic double-threshold decision mechanism to achieve the self-update and iterative process of the model. In the dynamic double-threshold decision mechanism, two thresholds are set: the self-update threshold (θu) and the replacement threshold (θr), which are used to trigger the self-update of the model and the model iteration operation respectively. When the S value exceeds θu, the system will automatically trigger the self-update operation of the model and retrain the model. If the S value exceeds the more stringent θr, it indicates that the performance of the current model has deviated seriously from the expectation and needs to be replaced by a new model version. This mechanism helps to maintain the effectiveness and adaptability of the model in a changing data environment.

[0055] Furthermore, in the process of using the above log anomaly detection model to detect anomalies in real-time log stream data, the above method further includes: determining the feature deviation degree of data features according to the above real-time log stream data, where the above data features include data association features, data time series features, and data statistical features; determining the feature deviation log data in the above real-time log stream data according to the above feature deviation degree, and adding the above feature deviation log data to the model update pool of the above log anomaly detection model.

[0056] This method conducts a comparative analysis on three features (association feature deviation degree ΔR, time series feature deviation degree ΔT, and statistical feature deviation degree ΔP), calculates the deviation degree on each feature dimension, and comprehensively obtains the total feature deviation degree. To ensure that the deviation magnitudes of different features are consistent, it is necessary to standardize the deviation degree of each feature. The standardized deviation degrees are weighted and synthesized to calculate the total feature deviation degree, and finally, the feature deviation log data in the above real-time log stream data is determined according to the total feature deviation degree. Calculate the feature deviation degree between the empirical pool and the real-time log stream from three aspects: association features, time series features, and statistical features, conduct a comparative analysis for each of these three features respectively, calculate the deviation degree on each feature dimension, and comprehensively obtain the total feature deviation degree, ensuring that the model can promptly perceive and respond to changes in data distribution, and improving the sensitivity and rapid response ability of the system to the feature of newly emerging log data sequences.

[0057] Even further, updating the above log anomaly detection model according to the above applicability includes: in the case of determining that the above applicability is greater than the self-update threshold, using the above feature deviation log data in the above model update pool and the above initial log data set in the above model empirical pool to perform an update operation on the above log anomaly detection model.

[0058] When the applicability is determined to be greater than the self-update threshold, the method will automatically trigger the self-update operation of the model and retrain the model. This mechanism helps to maintain the effectiveness and adaptability of the model in a constantly changing data environment. The training dataset is jointly constructed based on the experience pool and the update pool, and a new model is initialized for training. Through the designed joint loss function, the weights of each loss term are adaptively adjusted, aiming to minimize the loss function, and the self-update process of the model is realized, so that the model can achieve the effect of not forgetting historical knowledge and learning new knowledge during the update process, further enhancing the learning ability of the model.

[0059] Specifically, determining the feature deviation degree of data features according to the above real-time log stream data includes: using the chi-square test algorithm to determine the above data association features of the above real-time log stream data, using the dynamic time warping algorithm to determine the above data time series features of the above real-time log stream data, and using the KL divergence algorithm to determine the above data statistical features of the above real-time log stream data; determining the above feature deviation degree according to the above data association features, the above data time series features and the above data statistical features.

[0060] The association features of this method reflect the relationships between log segments, such as loop, sequential, and parallel relationships. First, extract the association features of log segments from the experience pool and the real-time log stream, and construct a loop relationship set, a sequential relationship set, and a parallel relationship set respectively. Each set contains multiple specific associated log segments, and calculate the frequency distributions of these relationship types in the experience pool and the real-time log stream; the time series features reflect the law of the log sequence arranged in time. By calculating the offset degree between the experience pool and the real-time log stream in the time series feature dimension, the matching degree between the two can be evaluated. First, arrange the log entries in the experience pool and the real-time log stream in chronological order to form two independent time series respectively. Then, use the dynamic time warping (DTW) algorithm to measure the time series similarity between the two time series, that is, align the log entries in the experience pool with the log entries in the real-time log stream, and calculate the minimum distance between the two to quantify the similarity between the two. The statistical features reflect the occurrence frequency of each log entry in the log sequence. During the analysis process, it is necessary to count the occurrence frequencies of each log entry in the experience pool and the real-time log stream, and calculate the frequency distributions of the log entries in the experience pool and the frequency distributions of the log entries in the real-time log stream respectively. Use the Kullback-Leibler (KL) divergence to quantify the difference between the two frequency distributions, so as to evaluate the deviation degree of the statistical features.

[0061] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the self-update method of the log anomaly detection model of the present application will be described in detail below in combination with specific embodiments.

[0062] This embodiment relates to a self - updating method for a specific log anomaly detection model. By detecting feature drift in real - time log stream data, continuously evaluating the applicability of the online model, and corresponding model self - updating methods, it meets the real - time detection requirements of massive log data in actual application scenarios, improves the accuracy of anomaly detection results, and reduces the system operation cost.

[0063] First, select an initial log data set and set relevant model parameters. Train an initial log anomaly detection model in an offline state, and construct an experience pool based on the initial log data set.

[0064] Secondly, the present invention proposes a method for evaluating the degree of decline in model detection performance based on applicability and feature drift analysis. At the data level, the present invention monitors and detects feature drift in real - time log stream data, calculates the degree of feature drift from three aspects: associated features, temporal features, and statistical features, and adds the log data with feature drift to the update pool; at the model evaluation level, combined with the manual labeling situation, evaluate the model applicability, consider the evaluation algorithm from two aspects: false positives and false negatives, and according to the dynamic double - threshold decision mechanism, realize the self - updating and iterative process of the model.

[0065] Finally, during the model self - updating training process of this embodiment, a model update method for adaptively adjusting the weights of the joint loss function is proposed. Based on the experience pool and the update pool, jointly construct a training data set, initialize a new model for training, design a joint loss function, and adaptively adjust the weights of each loss term to minimize the loss function as the goal to complete the model self - updating process, so as to ensure that the model neither forgets historical log data nor fails to learn the features of the drifted log data. The specific technical solution is as Figure 3 shown, including the following steps:

[0066] Step S1: Select an initial log data set and set relevant model parameters, construct an initial log anomaly detection model in an offline state, and construct an experience pool based on the initial log data set;

[0067] Specifically:

[0068] Based on special fields in the log (such as IP, session ID, etc.) and the sliding window technique, segment the initial log data set. Arrange the log data according to the timestamp and divide it into multiple sequence segments. The log data segmentation process is as Figure 4 shown;

[0069] Based on the log sequences, calculate the similarity between the sequences. The calculation of similarity uses an improved formula based on the matching segment length. By comparing the continuous matching degree of the same segments in two log sequences, the similarity is obtained. The formula is as follows:

[0070]

[0071] Among them, L is the number of matching segments, mi is the length of the i-th matching segment, and |seq1| and |seq2| are the lengths of sequences seq1 and seq2;

[0072] For log sequences with a similarity higher than the preset threshold ∈, they are classified. For the classified log sequences, a hash function H(seq) is used to generate a unique identifier ID(seq) to ensure the uniqueness of each log sequence in subsequent processing;

[0073] The generated unique identifier is used to deduplicate the log sequences. Duplicate log sequences with the same identifier are removed, and the deduplicated log data is stored in the experience pool.

[0074] This step can reduce storage redundancy, and at the same time ensure that the log data in the experience pool is highly representative, providing reliable data support for model update and feature analysis.

[0075] Step S2: Deploy the initial model to the online environment to receive and monitor anomalies in the log stream data in real time;

[0076] Step S3: Monitor and detect feature drift in the real-time log stream data, and add the log data with feature drift to the update pool;

[0077] This embodiment proposes a solution for monitoring and detecting feature drift, calculating the feature drift degree between the experience pool and the real-time log stream from three aspects: associated features, temporal features, and statistical features, respectively conducting comparative analysis for these three features, calculating the drift degree on each feature dimension, and comprehensively obtaining the total feature drift degree, as Figure 5 shown, specifically including the following content:

[0078] (1) Associated feature calculation:

[0079] Associated features reflect the relationships between log segments, such as loop, sequential, and parallel relationships. First, extract the associated features of log segments from the experience pool and the real-time log stream, and construct a loop relationship set, a sequential relationship set, and a parallel relationship set respectively, where each set contains multiple specific associated log segments. Calculate the frequency distributions of these relationship types in the experience pool and the real-time log stream, denoted as Fexp(Ri) and Freal(Ri), where Ri represents the i-th specific relationship type. To quantify the associated feature drift degree between the two, a chi-square test is used for calculation, and the chi-square statistic formula is as follows:

[0080]

[0081] Among them, M is the number of specific relationship types in each set. Through the chi-square test, the deviation degree of the associated features is calculated, and the deviation degrees of the cyclic relationship set, sequential relationship set, and parallel relationship set are obtained respectively, denoted as Δc, Δs, and Δp. Further, these deviation degrees are weighted and synthesized to obtain the total deviation degree of the associated features ΔR:

[0082] ΔR = α c ·Δc + β s ·Δs + γ p ·Δp;

[0083] Among them, αc, βs, and γp are the weight parameters of the cyclic relationship, sequential relationship, and parallel relationship respectively.

[0084] Calculation of temporal features:

[0085] The temporal features reflect the law of the log sequence arranged in time. By calculating the deviation degree between the empirical pool and the real-time log stream in the dimension of temporal features, the matching degree between the two can be evaluated. First, the log entries in the empirical pool and the real-time log stream are arranged in chronological order to form two independent time series respectively. Then, the dynamic time warping (DTW) algorithm is used to measure the temporal similarity between the two time series, that is, the log entries in the empirical pool are aligned with the log entries in the real-time log stream, and the minimum distance between them is calculated to quantify the similarity between the two. The calculation formula is as follows:

[0086]

[0087] Among them, T exp (i) represents the i-th log entry in the empirical pool, T real (j) represents the j-th log entry in the empirical pool, and dist represents the distance between log entries. By dynamically adjusting the matching of log entries, the DTW algorithm finds the minimum alignment distance, and then obtains the deviation degree of the temporal features, denoted as ΔT.

[0088] Calculation of statistical features:

[0089] The statistical features reflect the occurrence frequency of each log entry in the log sequence. In the analysis process, it is necessary to count the occurrence frequencies of each log entry in the empirical pool and the real-time log stream, and calculate the frequency distribution Pexp(x) of the log entries in the empirical pool and the frequency distribution Preal(x) of the log entries in the real-time log stream respectively. The Kullback-Leibler (KL) divergence is used to quantify the difference between the two frequency distributions to evaluate the deviation degree of the statistical features. The formula is as follows:

[0090]

[0091] By calculating the DKL value, the deviation degree of the statistical features is obtained, denoted as ΔP.

[0092] Calculation of the overall feature deviation degree:

[0093] To comprehensively evaluate the feature differences between the real-time log stream and the experience pool, the present invention proposes to conduct a comparative analysis on the aforementioned three features (the association feature deviation degree ΔR, the temporal feature deviation degree ΔT, and the statistical feature deviation degree ΔP), calculate the deviation degree in each feature dimension, and comprehensively obtain the overall feature deviation degree. To ensure that the deviation magnitudes of different features are consistent, it is necessary to standardize the deviation degree of each feature. The standardized deviation degrees are weighted and synthesized to calculate the overall feature deviation degree:

[0094] Δ total = λ R ·ΔR + λ T ·ΔT + λ P ·ΔP;

[0095] Among them, λ R , λ T and λ P are the weight parameters of the association feature, the temporal feature, and the statistical feature respectively. If the overall feature deviation degree Δ total exceeds the preset threshold θ, it is determined that a significant feature deviation has occurred between the real-time log stream and the experience pool. Once a significant feature deviation is detected, the log data with feature deviation is added to the update pool, and this process occurs within the time period between the completion of the last model self-update and the next trigger of the model self-update, so as to adjust the model in time to adapt to the new log features.

[0096] Step S4, quantitatively process the index of the model log anomaly detection effect in the online environment, further evaluate the applicability of the model, and finally, according to the dynamic double-threshold decision mechanism, realize the self-update and iterative process of the model.

[0097] Evaluation of model applicability:

[0098] In the online environment, to more accurately evaluate the performance of the model, it is necessary to continuously track the error rate in each time window and the standard deviation reflecting the volatility of the error rate.

[0099] The error rate reflects the accuracy of the model in identifying log anomalies, and the calculation formula of the error rate is as follows:

[0100]

[0101] Among them, e t is the error rate in the t-th time window, FP t represents the number of false alarm anomaly logs detected in this time window, FNt Indicates the number of abnormal logs with missed reports, N t Is the total number of logs detected within this time window.

[0102] To evaluate the degree of change of the error rate over time (i.e., volatility), the standard deviation of the error rate within multiple time windows can be calculated. The formula is:

[0103]

[0104] Among them, σ is the standard deviation, T is the total number of time windows, and e t Is the error rate of the t-th time window, and μ is the average error rate within all time windows.

[0105] (2) Model self-update mechanism:

[0106] To evaluate the performance and stability of the model at different time periods, the present invention introduces a comprehensive evaluation index S. Through this index, the applicability of the model is comprehensively evaluated, and a dynamic double-threshold determination mechanism is adopted to realize the self-update and iteration process of the model. In the dynamic double-threshold determination mechanism, two thresholds are set: the self-update threshold (θu) and the replacement threshold (θr), which are used to trigger the self-update of the model and the model iteration operation respectively. When the S value exceeds θu, the system will automatically trigger the self-update operation of the model and perform retraining of the model. If the S value exceeds the more stringent θr, it indicates that the performance of the current model has deviated seriously from the expectation and needs to be replaced by a new model version. This mechanism helps to maintain the effectiveness and adaptability of the model in a changing data environment.

[0107] Set the applicability evaluation index S, which comprehensively considers the error rate e within the current time window t And the standard deviation σ of the error rate within multiple time windows to reflect the performance and stability of the model at different time periods. Its calculation formula is:

[0108]

[0109] Among them, λ is the weight parameter, and σ max Is the maximum value of the standard deviation for normalization processing.

[0110] The calculation process of the dynamic double-threshold and the formula of the determination mechanism are as follows:

[0111]

[0112] Among them, Is the average error rate within multiple time windows, Is the average standard deviation, and max(e t) is the maximum error rate in multiple past time windows, max(σ) is the maximum standard deviation, and the remaining parameters are all weight coefficients.

[0113] Step S5, during the self-update process, a training data set is jointly constructed based on the experience pool and the update pool, and a new model is initialized for training. The present invention designs a joint loss function, and by adaptively adjusting the weights of each loss term, with the goal of minimizing the loss function, the self-update process of the model is completed;

[0114] The joint loss function consists of a knowledge retention loss term Lexp, a new knowledge learning loss term Lnew, and an online model retention loss term Lonline. Among them, the knowledge retention loss term Lexp is used to ensure that the model maintains good recognition ability for historical data in the experience pool; the new knowledge learning loss term Lnew aims to enhance the model's adaptability to the data in the update pool; the online model retention loss term Lonline is used to minimize the difference between the new model and the current online model, and avoid large fluctuations in the prediction results of the new model for real-time data streams. The formulas for each loss term are as follows:

[0115]

[0116] Among them, Y e ′ xp is the predicted output of the model on the data in the experience pool, Y exp is the true label of the data in the experience pool, Y′ new is the predicted output of the model on the data in the update pool, Y new is the true label of the data in the update pool, Y′ online is the predicted output of the new model on the online data, Y current is the predicted output of the current online model on the same data.

[0117] When calculating the joint loss function, it is necessary to adaptively adjust the weights of each loss term. The initial weights can be set according to the importance of each loss term in model training, and then adjusted in real time by calculating the relative change amount of each loss term. Calculate the change rate for each loss term, and record them as Δexp(t), Δnew(t), and Δonline(t) respectively. In the t-th round of training, the weight parameters αt, βt, and γt are calculated through the softmax function to ensure that the sum of the three weights is 1. The calculation formulas are as follows:

[0118] α t , β t , γ t = softmax(Δ exp (t), Δ new (t), Δ online (t))

[0119] Among them, the weight α t , β t and γ t are respectively the relative proportions of each loss term in the total loss during the t-th round of training.

[0120] Finally, based on the dynamic adjustment of the above loss terms and weight parameters, with the goal of minimizing the joint loss function until the convergence condition is reached, the self-update process of the model is completed.

[0121] Step S6, after the self-update process of the model ends, synchronize the log data in the update pool to the experience pool, and clear the log data in the update pool.

[0122] When this embodiment monitors and detects the feature drift of real-time log stream data, it calculates the feature drift degree between the experience pool and the real-time log stream from three aspects: associated features, temporal features, and statistical features, conducts comparative analysis for each of these three features respectively, calculates the drift degree on each feature dimension, and comprehensively obtains the total feature drift degree, ensuring that the model can timely perceive and respond to changes in data distribution, improving the sensitivity and rapid response ability of the system to the characteristics of newly emerging log data sequences; and by quantitatively processing the abnormal detection effect of the model logs in the online environment, evaluating the applicability of the model, and finally according to the dynamic double-threshold decision mechanism, to realize the self-update and iterative process of the model. This method can automatically adjust the update strategy according to the performance of the model in actual applications, reduce the need for manual intervention, and improve the automation level and maintenance efficiency of the system; in addition, based on the joint construction of the training data set by the experience pool and the update pool, and initialize a new model for training. Through the joint loss function designed by the present invention, adaptively adjust the weights of each loss term, with the goal of minimizing the loss function, to realize the self-update process of the model, so that the model achieves the effect of not forgetting historical knowledge while learning new knowledge during the update process, and further enhances the learning ability of the model.

[0123] This application embodiment also includes:

[0124] (1). Select an initial log data set and set model-related parameters, and construct an experience pool based on the initial log data set. The initial data set should cover a variety of common log patterns and typical abnormal types to ensure that the model has a wide range of recognition capabilities. Based on this data set, offline train the initial log anomaly detection model. The selection of the model can be based on the actual scenario requirements, and it is preferred to select a temporal deep learning model, such as a long short-term memory network (LSTM) or a bidirectional LSTM, etc., to make full use of the time series characteristics of the log data. After training, the model can effectively detect and identify common log abnormal types, laying a foundation for subsequent real-time anomaly detection.

[0125] (2) Deploy the initial log anomaly detection model to the log anomaly detection system, so as to store the offset feature log data in the update pool through feature offset analysis, and complete the self-update and self-replacement of the model through fitness evaluation.

[0126] The embodiment of the present application also provides a self-update device for a log anomaly detection model. It should be noted that the self-update device for the log anomaly detection model in the embodiment of the present application can be used to execute the self-update method for the log anomaly detection model provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0127] The following introduces the self-update device for the log anomaly detection model provided in the embodiment of the present application.

[0128] Figure 6 is a schematic diagram of the self-update device for the log anomaly detection model according to the embodiment of the present application. As Figure 3 shown, the device includes:

[0129] A training unit 61, configured to obtain an initial log data set, construct a model experience pool based on the initial log data set, construct an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model with the model experience pool to obtain a log anomaly detection model, where the initial log data set includes abnormal log data of multiple log patterns and the corresponding log anomaly types;

[0130] An anomaly detection unit 62, configured to perform anomaly detection on real-time log stream data using the log anomaly detection model when the log anomaly detection system is in an online state to obtain a detection result;

[0131] An update unit 63, configured to evaluate the fitness of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the fitness.

[0132] In this embodiment, the training unit is used to obtain an initial log data set, construct a model experience pool based on the initial log data set, build an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types; the anomaly detection unit is used to perform anomaly detection on real-time log stream data using the log anomaly detection model when the log anomaly detection system is in an online state to obtain a detection result; the update unit is used to evaluate the applicability of the log anomaly detection model according to the detection result and update the log anomaly detection model according to the applicability. Through the model applicability evaluation and the model adaptive update mechanism, the real-time detection ability of the model and the flexibility of model update are effectively improved, the value of real-time data is fully utilized, the real-time detection requirements of a large amount of log data in actual application scenarios are met, the accuracy of anomaly detection results is improved, and the system operation cost is reduced. It solves the problem of low detection accuracy in the log anomaly detection of the prior art.

[0133] As an optional solution, the update unit includes a first determination module, which is used to determine the applicability of the above-mentioned log anomaly detection model according to the first formula: where S is the above-mentioned applicability, λ is a weight parameter, σ is the maximum value of the standard deviation, e max is the error rate in the t-th time window, σ is the standard deviation of the error rates in multiple time windows, T is the total number of time windows, and μ is the average error rate in all time windows. t

[0134] As an optional solution, the update unit further includes a second determination module, a retraining module, and an iterative processing module; the second determination module is used to determine the self-update threshold and the replacement threshold of the above-mentioned log anomaly detection model; the retraining module is used to retrain the above-mentioned log anomaly detection model when it is determined that the above-mentioned applicability is greater than the self-update threshold and less than the replacement threshold to update the above-mentioned log anomaly detection model; the iterative processing module is used to perform iterative processing on the above-mentioned log anomaly detection model when it is determined that the above-mentioned applicability is greater than or equal to the replacement threshold to update the above-mentioned log anomaly detection model.

[0135] As an optional solution, the second determination module includes a determination sub-module, which is used to determine the above-mentioned self-update threshold and the above-mentioned replacement threshold of the above-mentioned log anomaly detection model according to the second formula: where θ u is the above-mentioned self-update threshold, θ r is the above-mentioned replacement threshold, is the average error rate in multiple time windows, is the average standard deviation, max(e t ) is the maximum error rate in multiple time windows, max(σ) is the maximum standard deviation, and α, β, γ, and δ are weight coefficients.

[0136] An optional solution is that the device further includes a first determination unit and an addition unit; the first determination unit is configured to determine the degree of feature deviation of data features according to the real-time log stream data during the process of performing anomaly detection on the real-time log stream data by using the above-mentioned log anomaly detection model, where the above-mentioned data features include data association features, data timing features, and data statistical features; the addition unit is configured to determine the feature deviation log data in the above-mentioned real-time log stream data according to the above-mentioned degree of feature deviation, and add the above-mentioned feature deviation log data to the model update pool of the above-mentioned log anomaly detection model.

[0137] An optional solution is that the update unit further includes an update module, which is configured to, when determining that the above-mentioned fitness degree is greater than the self-update threshold, perform an update operation on the above-mentioned log anomaly detection model by using the above-mentioned feature deviation log data in the above-mentioned model update pool and the above-mentioned initial log data set in the above-mentioned model experience pool.

[0138] An optional solution is that the first determination unit includes a third determination module and a fourth determination module; the third determination module is configured to determine the above-mentioned data association features of the above-mentioned real-time log stream data by using the chi-square test algorithm, determine the above-mentioned data timing features of the above-mentioned real-time log stream data by using the dynamic time warping algorithm, and determine the above-mentioned data statistical features of the above-mentioned real-time log stream data by using the KL divergence algorithm; the fourth determination module is configured to determine the above-mentioned degree of feature deviation according to the above-mentioned data association features, the above-mentioned data timing features, and the above-mentioned data statistical features.

[0139] The self-update device of the above-mentioned log anomaly detection model includes a processor and a memory. The above-mentioned training unit, anomaly detection unit, update unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is located in different processors in any combination form.

[0140] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the problem of low detection accuracy in the existing log anomaly detection is solved by adjusting the kernel parameters.

[0141] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0142] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the self-update method of the log anomaly detection model.

[0143] Specifically, the self-update method of the log anomaly detection model includes:

[0144] Step S201: Obtain an initial log data set, construct a model experience pool based on the initial log data set, construct an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types;

[0145] Step S202: When the log anomaly detection system is in an online state, use the log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result;

[0146] Step S203: Evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

[0147] An embodiment of the present invention provides a processor. The processor is used to run a program. When the program runs, it executes the self-update method of the log anomaly detection model.

[0148] Specifically, the self-update method of the log anomaly detection model includes:

[0149] Step S201: Obtain an initial log data set, construct a model experience pool based on the initial log data set, construct an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types;

[0150] Step S202: When the log anomaly detection system is in an online state, use the log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result;

[0151] Step S203: Evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

[0152] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:

[0153] Step S201: Obtain an initial log data set, build a model experience pool based on the initial log data set, build an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types of the anomaly log data;

[0154] Step S202: When the log anomaly detection system is in an online state, use the log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result;

[0155] Step S203: Evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

[0156] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0157] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:

[0158] Step S201: Obtain an initial log data set, build a model experience pool based on the initial log data set, build an initial anomaly detection model when the log anomaly detection system is in an offline state, and train the initial anomaly detection model using the model experience pool to obtain a log anomaly detection model. The initial log data set includes anomaly log data of multiple log patterns and the corresponding log anomaly types of the anomaly log data;

[0159] Step S202: When the log anomaly detection system is in an online state, use the log anomaly detection model to perform anomaly detection on real-time log stream data to obtain a detection result;

[0160] Step S203: Evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

[0161] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0162] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0163] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0164] These computer program instructions can also be stored in a computer-readable memory capable of guiding the computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0165] These computer program instructions can also be loaded onto the computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one or more blocks

[0166] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0167] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0168] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0169] It should also be noted that the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity, or device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, commodity, or device that comprises the element.

[0170] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A self-updating method for a log anomaly detection model, characterized in that: include: Acquire an initial log data set, build a model experience pool based on the initial log data set, build an initial anomaly detection model when the log anomaly detection system is offline, and use the model experience pool to train the initial anomaly detection model to obtain a log anomaly detection model, wherein the initial log data set includes abnormal log data of multiple log modes and log anomaly types corresponding to the abnormal log data; When the log anomaly detection system is in an online state, the log anomaly detection model is used to perform anomaly detection on the real-time log stream data to obtain a detection result; The applicability of the log anomaly detection model is evaluated according to the detection result, and the log anomaly detection model is updated according to the applicability.

2. The method according to claim 1, characterized in that Evaluating the applicability of the log anomaly detection model according to the detection result includes: According to the first formula: Determining the fitness of the log anomaly detection model, wherein: S is the fitness, λ is the weight parameter, σ max is the maximum value of the standard deviation, e t is the error rate in the tth time window, σ is the standard deviation of the error rate in multiple time windows, T is the total number of time windows, and μ is the average error rate in all time windows.

3. The method according to claim 1, characterized in that Updating the log anomaly detection model according to the applicability includes: Determining a self-update threshold and a replacement threshold of the log anomaly detection model; When it is determined that the fitness is greater than the self-update threshold and less than the replacement threshold, retraining the log anomaly detection model to update the log anomaly detection model; When it is determined that the applicability is greater than or equal to the replacement threshold, the log anomaly detection model is iteratively processed to update the log anomaly detection model.

4. The method according to claim 3, characterized in that Determining a self-update threshold and a replacement threshold of the log anomaly detection model includes: According to the second formula: Determine the self-update threshold and the replacement threshold of the log anomaly detection model, where θ u is the self-renewal threshold, θ r is the replacement threshold, is the average error rate in multiple time windows, is the mean standard deviation, max(e t ) is the maximum error rate in multiple time windows, max(σ) is the maximum standard deviation, and α, β, γ, ∈ are weight coefficients.

5. The method according to claim 1, characterized in that: In the process of using the log anomaly detection model to perform anomaly detection on real-time log stream data, the method further includes: Determine a feature deviation degree of data features according to the real-time log stream data, wherein the data features include data association features, data time series features, and data statistical features; Determine feature shift log data in the real-time log stream data according to the feature shift degree, and add the feature shift log data to a model update pool of the log anomaly detection model.

6. The method according to claim 5, characterized in that Updating the log anomaly detection model according to the applicability includes: When it is determined that the fitness is greater than the self-update threshold, the log anomaly detection model is updated using the feature offset log data in the model update pool and the initial log data set in the model experience pool.

7. The method according to claim 5, characterized in that Determining a feature deviation degree of a data feature according to the real-time log stream data includes: A chi-square test algorithm is used to determine the data association characteristics of the real-time log stream data, a dynamic time warping algorithm is used to determine the data timing characteristics of the real-time log stream data, and a KL divergence algorithm is used to determine the data statistical characteristics of the real-time log stream data; The feature shift degree is determined according to the data association feature, the data timing feature and the data statistical feature.

8. A self-updating device for a log anomaly detection model, characterized in that: include: A training unit, used to obtain an initial log data set, and build a model experience pool based on the initial log data set, and build an initial anomaly detection model when the log anomaly detection system is in an offline state, and use the model experience pool to train the initial anomaly detection model to obtain a log anomaly detection model, wherein the initial log data set includes abnormal log data of multiple log modes and log anomaly types corresponding to the abnormal log data; an anomaly detection unit, configured to, when the log anomaly detection system is in an online state, use the log anomaly detection model to perform anomaly detection on the real-time log stream data to obtain a detection result; An updating unit is used to evaluate the applicability of the log anomaly detection model according to the detection result, and update the log anomaly detection model according to the applicability.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the self-updating method of the log anomaly detection model according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a self-update method for executing the log anomaly detection model described in any one of claims 1 to 7.