Log anomaly detection method and device, equipment and medium
By using periodic log data sequence and abnormal trend analysis in log exception detection, the candidates and target abnormal time points and their weights are determined, and the accuracy of abnormal detection in massive log data is solved, and the effect of log abnormal detection is improved.
Patent Information
- Application Number
- CN202510159615.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively detect and locate abnormal situations in massive log data, resulting in the neglect of problems or delays.
By obtaining the periodic log data sequence in two adjacent acquisition cycles, determining the relative difference value sequence of logs, and determining whether there is a log abnormal trend. If it exists, the candidate abnormal time point and the corresponding candidate abnormal data will be determined, and the target abnormal time point and its weight will be further determined, and the data will be input into the trained log abnormality detection model to obtain the current log detection result.
The accuracy of log exception detection is improved, and the ability to identify and locate log exceptions is enhanced by focusing on the log data of target exception time points corresponding to the target exception weight.
Smart Images

Figure CN119988145A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a log anomaly detection method, device, equipment and medium. Background Art
[0002] Log anomaly detection is a key technology, usually used in application monitoring, fault analysis and other fields. With the emergence of a large number of applications, in order to locate the problems in application operation faster and more accurately, a large amount of log data is generated. The large scale and complexity of log data make manual analysis and monitoring very difficult, which easily leads to problems being ignored or delayed. Therefore, it is crucial to improve the accuracy of log anomaly detection. Summary of the invention
[0003] The present invention provides a log anomaly detection method, device, equipment and medium to improve the accuracy of log anomaly detection.
[0004] According to one aspect of the present invention, a log anomaly detection method is provided, comprising:
[0005] Acquire a periodic log data sequence within two adjacent collection periods, and determine a log relative difference sequence based on the periodic log data sequence; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence;
[0006] Determine whether the log relative difference sequence has a log abnormality trend, and if so, determine candidate abnormal time points and corresponding candidate abnormal data within the current collection period according to the current log data sequence and the previous log data sequence;
[0007] Determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence;
[0008] The current log data sequence, the preset normal point weights and the target abnormal point weights are input into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
[0009] According to another aspect of the present invention, a log anomaly detection device is provided, comprising:
[0010] A relative difference sequence determination module is used to obtain a periodic log data sequence within two adjacent collection periods, and determine a log relative difference sequence based on the periodic log data sequence; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence;
[0011] A candidate abnormal data determination module is used to determine whether the log relative difference sequence has a log abnormal trend, and if so, determine the candidate abnormal time point and the corresponding candidate abnormal data in the current collection period according to the current log data sequence and the previous log data sequence;
[0012] A target abnormal point weight determination module is used to determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence;
[0013] The log detection result determination module is used to input the current log data sequence, the preset normal point weight and the target abnormal point weight into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
[0014] According to another aspect of the present invention, there is provided an electronic device, comprising:
[0015] one or more processors;
[0016] A memory for storing one or more programs;
[0017] When one or more programs are executed by one or more processors, the one or more processors can execute any one of the log anomaly detection methods provided by the embodiments of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions, and the computer instructions are used to enable a processor to implement any log anomaly detection method provided by an embodiment of the present invention when executed.
[0019] An embodiment of the present invention provides a log anomaly detection solution, which obtains periodic log data sequences within two adjacent collection cycles, and determines a log relative difference sequence according to the periodic log data sequences; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence; determines whether the log relative difference sequence has a log anomaly trend, and if so, determines candidate anomaly time points and corresponding candidate anomaly data within the current collection cycle according to the current log data sequence and the previous log data sequence; determines a target anomaly time point and a target anomaly point weight corresponding to the target anomaly time point according to the candidate anomaly time point, the candidate anomaly data and the log relative difference sequence; inputs the current log data sequence, the preset normal point weight and the target anomaly point weight into a trained log anomaly detection model, and obtains a current log detection result corresponding to the current log data sequence. The above scheme determines the log relative difference sequence according to the current log data sequence and the previous log data sequence to determine whether there is a log abnormal trend. If there is a log abnormal trend, the candidate abnormal time points and the corresponding candidate abnormal data in the current acquisition cycle are determined, and then the target abnormal time point and the corresponding target abnormal point weight are determined from the candidate abnormal time points. The current log data sequence, the preset normal point weight and the target abnormal point weight are input into the log anomaly detection model to obtain the current log detection result, and the target abnormal time point is determined from each time point associated with the current log data sequence, and the corresponding target abnormal point weight is determined for the log data at the target abnormal time point. At the same time, the current log data sequence, the normal point weight or the target abnormal point weight corresponding to each time point is input into the log anomaly detection model, and the current log detection result is output, so that the log anomaly detection model can focus on the log data at the target abnormal time point corresponding to the target abnormal point weight, thereby improving the accuracy of the current log detection result, that is, improving the accuracy of the log anomaly detection.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a flow chart of a log anomaly detection method provided by Embodiment 1 of the present invention;
[0023] Figure 2 is a flow chart of a log anomaly detection method provided by Embodiment 2 of the present invention;
[0024] Figure 3 is a flow chart of a method for training a log anomaly detection model provided in Embodiment 3 of the present invention;
[0025] Figure 4 It is a structural schematic diagram of a log anomaly detection device provided by Embodiment 4 of the present invention;
[0026] Figure 5 It is a structural schematic diagram of an electronic device for implementing a log anomaly detection method provided in Embodiment 5 of the present invention. DETAILED DESCRIPTION
[0027] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0028] Embodiment 1
[0029] Figure 1 This is a flow chart of a log anomaly detection method provided in Example 1 of the present invention. This embodiment is applicable to the situation where anomaly detection is performed on a log data sequence. The method can be executed by a log anomaly detection device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the log anomaly detection function.
[0030] See also Figure 1 The log anomaly detection method shown includes:
[0031] S110, acquiring periodic log data sequences within two adjacent collection periods, and determining a log relative difference sequence according to the periodic log data sequences; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence.
[0032] Among them, the collection cycle refers to the pre-set period for collecting log data. The periodic log data sequence refers to the log data sequence within a collection cycle. The current log data sequence refers to the sequence of log data within the current collection cycle. The previous log data sequence refers to the sequence of log data in the previous collection cycle based on the current collection cycle. The log relative difference sequence refers to the difference sequence between the current log data sequence and the previous log data sequence.
[0033] Exemplarily, a current log data sequence in a current collection cycle and a previous log data sequence in a previous collection cycle are obtained; and a log relative difference sequence is obtained by calculation based on the log data at corresponding time points in the current log data sequence and the previous log data sequence.
[0034] S120, determining whether there is a log abnormality trend in the log relative difference sequence, and if so, determining candidate abnormal time points and corresponding candidate abnormal data in the current collection cycle according to the current log data sequence and the previous log data sequence.
[0035] The current collection cycle refers to the collection cycle of the log data for which log anomaly detection is currently performed. The previous collection cycle refers to the collection cycle of the log data for which log anomaly detection was performed at the previous moment.
[0036] The candidate abnormal time point refers to the time point in the current acquisition cycle where abnormal log data may exist. The candidate abnormal data refers to the absolute value of the difference between the fitting value corresponding to the candidate abnormal time point in the current data waveform and the fitting value corresponding to the corresponding candidate abnormal time point in the previous data waveform, which is the abnormal absolute value.
[0037] The current data waveform refers to the waveform obtained by fitting according to the current log data sequence, and the previous data waveform refers to the waveform obtained by fitting according to the previous log data sequence.
[0038] Specifically, according to the preset difference sequence algorithm, it is determined whether there is a log abnormality trend in the log relative difference sequence; if there is a log abnormality trend, the candidate abnormal time points and the corresponding candidate abnormal data in the current acquisition cycle are determined according to the current log data sequence and the previous log data sequence. The embodiment of the present invention does not impose any limitation on the preset difference sequence algorithm, which can be set by the technician according to experience or needs. Exemplarily, the preset difference sequence algorithm can be a non-parametric statistical method for detecting trend changes in time series.
[0039] The log abnormal trend can be determined by calculating the log relative difference sequence according to a preset difference sequence algorithm. For example, if the algorithm result is 1, it indicates that there is a log abnormal trend; if the algorithm result is 0, it indicates that there is no log abnormal trend.
[0040] In an optional embodiment, candidate abnormal time points and corresponding candidate abnormal data within a current acquisition cycle are determined based on a current log data sequence and a previous log data sequence, including: determining a current data waveform corresponding to the current log data sequence and a previous data waveform corresponding to the previous log data sequence; determining candidate abnormal time points and corresponding candidate abnormal data within the current acquisition cycle based on the current data waveform and the previous data waveform.
[0041] Exemplarily, based on a preset waveform fitting algorithm, determine the current data waveform corresponding to the current log data sequence, and the previous data waveform corresponding to the previous log data sequence; for any time point, determine the fitting value at the time point in the current data waveform, and the fitting difference between the fitting values at the time point in the previous data waveform; if the aforementioned fitting difference is greater than a preset data fitting threshold, the time point corresponding to the fitting difference is used as a candidate abnormal time point, and the fitting difference is used as a candidate abnormal data. The embodiment of the present invention does not impose any limitation on the preset waveform fitting algorithm, which can be set by a technician based on experience or needs. Exemplarily, the preset waveform fitting algorithm can be the least squares method. The embodiment of the present invention does not impose any limitation on the size of the preset data fitting threshold, which can be set by a technician based on experience or needs, or determined repeatedly through a large number of experiments.
[0042] It should be noted that at least part of the current log data with log abnormality trends can be intercepted from the current log data sequence, and the corresponding part of the previous log data can be intercepted from the previous log data sequence, and then the current data waveform can be determined based on the intercepted current log data, and the previous data waveform can be determined based on the intercepted previous log data.
[0043] It can be understood that by determining the candidate abnormal time points and the corresponding candidate abnormal data within the current acquisition cycle based on the current data waveform corresponding to the current log data sequence and the previous data waveform corresponding to the previous log data sequence, the accuracy of the determined candidate abnormal time points and candidate abnormal data is improved.
[0044] S130. Determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence.
[0045] The target anomaly time point refers to the candidate time point at which abnormal log data exists in the current collection cycle, that is, the target anomaly time point can be understood as the candidate time point that the log anomaly detection model needs to focus on. The target anomaly point weight refers to the weight corresponding to the log data at the target anomaly time point.
[0046] Specifically, according to the candidate abnormal time points, the candidate abnormal data and the log relative difference sequence, the target abnormal time point is determined from the candidate abnormal time points, and the target abnormal point weight of the log data at the target abnormal time point is determined.
[0047] S140: Input the current log data sequence, the preset normal point weights and the target abnormal point weights into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
[0048] The normal point weight refers to the weight of the log data at a preset time point without abnormality. The embodiment of the present invention does not limit the size of the normal point weight, which can be set by technicians based on experience or needs, or can be repeatedly determined through a large number of experiments.
[0049] Among them, the current log detection result refers to the result obtained by the log anomaly detection model for anomaly detection of the current log data sequence. The log anomaly detection model can be used to perform anomaly detection on log data and output the log detection result. The embodiment of the present invention does not impose any restrictions on the network structure of the log anomaly detection model, which can be set by technicians based on experience or needs. Exemplarily, the log anomaly detection model can be a model for processing and analyzing time series data with long-term dependencies.
[0050] Specifically, the log data at the target abnormal time point in the current log data sequence is used as abnormal log data, and the log data corresponding to other time points in the current log data sequence except the target abnormal time point is used as normal log data; the normal log data corresponds to the normal point weight, and the abnormal log data corresponds to the target abnormal point weight determined above; the current log data sequence, the preset normal point weight and the target abnormal point weight are input into the trained log anomaly detection model, and the current log detection result corresponding to the current log data sequence is output.
[0051] An embodiment of the present invention provides a log anomaly detection solution, which obtains periodic log data sequences within two adjacent collection cycles, and determines a log relative difference sequence according to the periodic log data sequences; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence; determines whether the log relative difference sequence has a log anomaly trend, and if so, determines candidate anomaly time points and corresponding candidate anomaly data within the current collection cycle according to the current log data sequence and the previous log data sequence; determines a target anomaly time point and a target anomaly point weight corresponding to the target anomaly time point according to the candidate anomaly time point, the candidate anomaly data and the log relative difference sequence; inputs the current log data sequence, the preset normal point weight and the target anomaly point weight into a trained log anomaly detection model, and obtains a current log detection result corresponding to the current log data sequence. The above scheme determines the log relative difference sequence according to the current log data sequence and the previous log data sequence to determine whether there is a log abnormal trend. If there is a log abnormal trend, the candidate abnormal time points and the corresponding candidate abnormal data in the current acquisition cycle are determined, and then the target abnormal time point and the corresponding target abnormal point weight are determined from the candidate abnormal time points. The current log data sequence, the preset normal point weight and the target abnormal point weight are input into the log anomaly detection model to obtain the current log detection result, and the target abnormal time point is determined from each time point associated with the current log data sequence, and the corresponding target abnormal point weight is determined for the log data at the target abnormal time point. At the same time, the current log data sequence, the normal point weight or the target abnormal point weight corresponding to each time point is input into the log anomaly detection model, and the current log detection result is output, so that the log anomaly detection model can focus on the log data at the target abnormal time point corresponding to the target abnormal point weight, thereby improving the accuracy of the current log detection result, that is, improving the accuracy of the log anomaly detection.
[0052] Embodiment 2
[0053] Figure 2It is a flow chart of a log anomaly detection method provided in the second embodiment of the present invention. Based on the above embodiments, this embodiment further refines the operation of "determining the target anomaly time point and the target anomaly point weight corresponding to the target anomaly time point according to the candidate anomaly time point, the candidate anomaly data and the log relative difference sequence" into "determining the basic anomaly data from the log relative difference sequence according to the candidate anomaly time point, and determining the data anomaly rate of the candidate anomaly time point according to the basic anomaly data and the candidate anomaly data; determining the target anomaly time point and the anomaly distance value of the target anomaly time point according to the data anomaly rate and the preset anomaly rate threshold; determining the target anomaly point weight corresponding to the corresponding target anomaly time point according to the anomaly distance value" to improve the determination mechanism of the target anomaly point weight. It should be noted that for the parts not described in detail in the embodiments of the present invention, reference may be made to the descriptions of other embodiments.
[0054] See also Figure 2 The log anomaly detection method shown includes:
[0055] S210, acquiring periodic log data sequences within two adjacent collection periods, and determining a log relative difference sequence according to the periodic log data sequences; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence.
[0056] S220, determining whether there is a log abnormality trend in the log relative difference sequence, and if so, determining candidate abnormal time points and corresponding candidate abnormal data in the current collection cycle according to the current log data sequence and the previous log data sequence.
[0057] S230: According to the candidate abnormal time point, determine basic abnormal data from the log relative difference sequence, and determine the data abnormality rate of the candidate abnormal time point according to the basic abnormal data and the candidate abnormal data.
[0058] Among them, the basic abnormal data refers to the data corresponding to the candidate abnormal time point in the log relative difference sequence. The data abnormality rate refers to the ratio between the candidate abnormal data and the basic abnormal data at the same candidate abnormal time point. It should be noted that the larger the data abnormality rate, the higher the similarity between the candidate abnormal data and the basic abnormal data at the same candidate abnormal time point.
[0059] S240: Determine a target abnormal time point and an abnormal distance value of the target abnormal time point according to the data abnormality rate and a preset abnormality rate threshold.
[0060] The embodiment of the present invention does not impose any limitation on the size of the preset abnormality rate threshold, which can be set by a technician based on experience or needs, or determined repeatedly through a large number of experiments. The abnormal distance value refers to the size of the data abnormality rate deviation from the preset abnormality rate threshold. Exemplarily, the abnormal distance value is the ratio between the preset abnormality rate threshold and the data abnormality rate.
[0061] Specifically, the data anomaly rate is compared with a preset anomaly rate threshold, and based on the comparison result, a target anomaly time point and an anomaly distance value of the target anomaly time point are determined.
[0062] In an optional embodiment, the target abnormal time point and the abnormal distance value of the target abnormal time point are determined according to the data abnormality rate and the preset abnormality rate threshold, including: if the data abnormality rate is greater than the preset abnormality rate threshold, the corresponding candidate abnormal time point is determined as the target abnormal time point; the ratio between the preset abnormality rate threshold and the data abnormality rate of the target abnormal time point is used as the abnormal distance value of the target abnormal time point.
[0063] It can be understood that when the data anomaly rate is greater than the preset anomaly rate threshold, the corresponding candidate anomaly time point is used as the target anomaly time point, and then the ratio between the preset anomaly rate threshold and the data anomaly rate of the target anomaly time point is used as the anomaly distance value of the target anomaly time point, thereby improving the accuracy of the determined anomaly distance value.
[0064] Exemplarily, if the data anomaly rate is less than or equal to a preset anomaly rate threshold, it is prohibited to use the corresponding candidate anomaly time point as the target anomaly time point.
[0065] S250. Determine the target abnormal point weight corresponding to the corresponding target abnormal time point according to the abnormal distance value.
[0066] In an optional embodiment, determining the target abnormal point weight corresponding to the corresponding target abnormal time point according to the abnormal distance value includes: determining the distance difference between a preset abnormal distance threshold and the abnormal distance value; determining the candidate abnormal point weight according to the preset proportional factor and the distance difference; determining the target abnormal point weight of the target abnormal time point according to the candidate abnormal point weight and the preset weight threshold.
[0067] Among them, the embodiment of the present invention does not impose any limitation on the sizes of the preset abnormal distance threshold, the preset proportional factor and the preset weight threshold, which can be set by technicians according to experience or needs, or repeatedly determined through a large number of experiments.
[0068] Exemplarily, the preset abnormal distance threshold may be 1, the preset scale factor may be 15, and the preset weight threshold may be 10.
[0069] The distance difference refers to the difference between the preset distance threshold and the abnormal distance value. The candidate abnormal point weight refers to the weight of the log data at the target abnormal time point preliminarily determined based on the preset scale factor and the distance difference.
[0070] It should be noted that since the weight of the candidate outlier point may be greater than the preset weight threshold (such as 10), the weight of the target outlier point can be determined by taking the remainder of the preset weight threshold for the candidate outlier point weight, thereby avoiding the subsequent input of the target outlier point weight into the log anomaly detection model, which affects the performance of the log anomaly detection model.
[0071] It can be understood that by obtaining the candidate outlier weight according to the preset outlier distance threshold and the distance difference determined by the outlier distance value, and the preset proportional factor, and then processing the candidate outlier weight by the preset weight threshold to obtain the target outlier weight, the accuracy and applicability of the determined target outlier weight are improved.
[0072] S260: Input the current log data sequence, the preset normal point weights and the target abnormal point weights into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
[0073] The embodiment of the present invention provides a log anomaly detection scheme, which refines the target anomaly time point and the target anomaly point weight operation corresponding to the target anomaly time point according to the candidate anomaly time point, the candidate anomaly data and the log relative difference sequence into determining the basic anomaly data from the log relative difference sequence according to the candidate anomaly time point, and determining the data anomaly rate of the candidate anomaly time point according to the basic anomaly data and the candidate anomaly data; determining the target anomaly time point and the anomaly distance value of the target anomaly time point according to the data anomaly rate and the preset anomaly rate threshold; determining the target anomaly point weight corresponding to the corresponding target anomaly time point according to the anomaly distance value, thereby improving the determination mechanism of the target anomaly point weight. The above scheme improves the accuracy of the determined target anomaly point weight by determining the basic anomaly data corresponding to the candidate anomaly time point and the candidate anomaly data from the log relative difference sequence, determining the data anomaly rate, determining the target anomaly time point and the corresponding anomaly distance value according to the data anomaly rate and the preset anomaly rate threshold, and then determining the target anomaly point weight according to the anomaly distance value.
[0074] Embodiment 3
[0075] The embodiment of the present invention provides an optional example of log anomaly detection based on the above embodiment. It should be noted that for the part not described in detail in the embodiment of the present invention, reference can be made to the description of other embodiments.
[0076] In the existing technology, due to the strong correlation between log data and time, sequence-based deep learning modeling has gradually become the main method for log anomaly detection. The core of existing log anomaly detection is to use artificial intelligence algorithms to automatically analyze system logs to discover and locate faults. According to the data format fed into the detection model, the log anomaly detection algorithm model is divided into sequence model and frequency model, among which the sequence model can be divided into deep model and clustering model.
[0077] For example, although there are many methods for log anomaly detection, there are still some shortcomings and challenges. For example, traditional statistical methods are effective and faster for small and trace application scenarios, but they have obvious limitations when processing massive logs; machine learning methods, clustering and other algorithms in machine learning methods perform well in specific scenarios, but when processing large-scale, high-dimensional, and diverse log data, the computational complexity is high and the assumptions about the data are strong; deep learning methods, although the current deep learning methods can extract key information from logs well and introduce time series auxiliary analysis, they lack the reasonable application of time series, ignore the long-term dependence between logs and time, are difficult to train for large amounts of data, and have limited generalization capabilities.
[0078] The embodiment of the present invention proposes a new log sequence anomaly detection method based on a structural model of system logs using a deep learning method for anomaly detection and diagnosis, so as to solve the problems of insufficient application of current model time series and insufficient adaptability to complex data.
[0079] Exemplarily, the training process of the log anomaly detection model may include the following steps: data preprocessing, building a model framework, model training, selecting a loss function, and model evaluation. Specifically, data preprocessing, preprocessing the sample log data, such as cleaning, conversion, aggregation, and classification by type. Building a model framework, selecting a suitable deep learning model framework, such as selecting a model based on log sequence anomaly detection. Model training, using sample log data to train the log anomaly detection model; in the process of training the log anomaly detection model, a large amount of sample log data is used for training, and by learning the general rules and patterns of the logs, the log anomaly detection model can understand and process various types of logs. Selection of loss function, the loss function of the log anomaly detection model usually uses the cross entropy loss function, and the parameters of the model are updated through the back propagation algorithm to minimize the value of the loss function. Model evaluation, using the validation set to evaluate the performance of the trained log anomaly detection model. Common evaluation methods include accuracy, F1 score, root mean square error, etc.; according to the model evaluation results, the hyperparameters of the log anomaly detection model are adjusted, such as the learning rate, batch size, number of hidden units, etc.
[0080] Furthermore, after sufficient training and tuning, the log anomaly detection model is finally evaluated using the test dataset. The evaluation results will reflect the performance of the model in a real environment.
[0081] For example, see Figure 3 The training method of the log anomaly detection model shown includes:
[0082] S310, obtaining a sample log data sequence within two adjacent sample collection cycles, and determining a sample log relative difference sequence based on the sample log data sequence; wherein the sample log data sequence includes a current sample log data sequence and a previous sample log data sequence.
[0083] Among them, the current sample log data sequence refers to the sequence of log data in the current sample collection cycle. The current sample collection cycle refers to the cycle of collecting sample log data at the current moment. The previous sample log data sequence refers to the sequence of log data in the previous sample collection cycle. The previous sample collection cycle refers to the cycle of collecting sample log data at the previous moment.
[0084] The sample log relative difference sequence refers to the difference sequence between the current sample log data sequence and the previous sample log data sequence.
[0085] For example, data preparation involves collecting a large amount of sample log data of an application and performing preprocessing operations such as cleaning the sample log data. The time period used is T, and the current time is recorded as t0, and the time t0-T is recorded as t1. Among them, t1 represents the starting time point of the previous collection period; t0 represents the starting time point of the current collection period.
[0086] Exemplarily, a sample log relative difference sequence between a current sample log data sequence in a current sample collection cycle and a previous sample log data sequence in a previous sample collection cycle is calculated.
[0087] Specifically, the current sample log data sequence in the current sample collection cycle and the previous sample log data sequence in the previous sample collection cycle are obtained, and the sample log relative difference sequence is determined according to the current sample log data sequence and the previous sample log data sequence.
[0088] S320. Determine whether there is a log abnormality trend in the sample log relative difference sequence. If so, determine the candidate sample abnormal time point and the corresponding candidate sample abnormal data within the current sample collection cycle based on the current sample log data sequence and the previous sample log data sequence.
[0089] Among them, the candidate sample abnormal time point refers to the time point at which abnormal sample log data may exist in the current sample collection cycle. The candidate sample abnormal data refers to the abnormal absolute value corresponding to the candidate sample abnormal time point determined based on the current sample data waveform and the previous sample data waveform. The current sample data waveform refers to the fitted waveform determined based on the current sample log data sequence. The previous sample data waveform refers to the fitted waveform determined based on the previous sample log data sequence.
[0090] Exemplarily, a preset difference sequence algorithm is used to check and determine whether there is a log abnormality trend in the sample log relative to the difference sequence. If there is no log abnormality trend, a normal result is returned, otherwise the log abnormality trend processing step is entered.
[0091] Further, according to the current sample log data sequence, the current sample data waveform is fitted by the least square method, and similarly, the previous sample data waveform can be obtained; according to the current sample data waveform and the previous sample data waveform, the candidate sample abnormal data is obtained. It should be noted that the starting point of the abnormal trend can be fitted with the abnormal segment waveform by the least square method to obtain the abnormal absolute value, that is, according to the judgment result of the log abnormal trend, at least part of the sample log data with the log abnormal trend is intercepted from the current sample log data sequence, and then according to at least part of the sample log data intercepted, waveform fitting is performed to obtain the current sample data waveform. Similarly, according to at least part of the sample log data intercepted from the previous sample log data sequence, the previous sample data waveform can be obtained; according to the current sample data waveform and the previous sample data waveform, the candidate sample abnormal data is determined.
[0092] Specifically, based on the preset difference sequence algorithm, determine whether there is a log abnormality trend in the sample log relative difference sequence. If so, based on the preset waveform fitting algorithm, determine the current sample data waveform corresponding to the current sample log data sequence and the previous sample data waveform corresponding to the previous sample log data sequence, and determine the candidate sample abnormal time point and the corresponding candidate sample abnormal data within the current sample collection cycle based on the current sample data waveform and the previous sample data waveform.
[0093] S330. Determine the target sample abnormal time point and the corresponding target sample abnormal point weight according to the candidate sample abnormal time point, the candidate sample abnormal data and the sample log relative difference sequence.
[0094] Among them, the target sample abnormal time point refers to the candidate sample abnormal time point where abnormal sample log data exists in the current sample collection cycle, that is, the target sample abnormal time point can be understood as the candidate sample abnormal time point that the log anomaly detection model needs to focus on. The target sample abnormal point weight refers to the weight corresponding to the sample log data at the target sample abnormal time point.
[0095] Exemplarily, based on the candidate sample abnormal time point, the sample abnormality basic data is determined from the sample log relative difference sequence, and the sample data abnormality rate of the candidate sample abnormal time point is determined based on the sample abnormality basic data and the candidate sample abnormal data; based on the sample data abnormality rate and a preset abnormality rate threshold, the target sample abnormal time point and the sample abnormality distance value of the target sample abnormal time point are determined; based on the sample abnormality distance value, the target sample abnormal point weight corresponding to the corresponding target sample abnormal time point is determined.
[0096] Among them, the sample anomaly basic data refers to the data corresponding to the candidate sample anomaly time point in the sample log relative difference sequence. The sample data anomaly rate refers to the ratio between the candidate sample anomaly data and the sample anomaly basic data at the same candidate sample anomaly time point.
[0097] The sample anomaly distance value refers to the deviation of the sample data anomaly rate from the preset anomaly rate threshold. Exemplarily, the sample anomaly distance value is the ratio between the preset anomaly rate threshold and the sample data anomaly rate.
[0098] Exemplarily, if the sample data abnormality rate is greater than a preset abnormality rate threshold, the corresponding candidate sample abnormality time point is determined to be the target sample abnormality time point; the ratio between the preset abnormality rate threshold and the sample data abnormality rate of the target sample abnormality time point is used as the sample abnormality distance value of the target sample abnormality time point.
[0099] Exemplarily, if the abnormality rate of the sample data is less than or equal to a preset abnormality rate threshold, it is prohibited to use the corresponding candidate sample abnormal time point as the target sample abnormal time point.
[0100] Specifically, it is determined whether the abnormal rate of the sample data is greater than the preset abnormal rate threshold. If it is not greater than the preset abnormal rate threshold, the normal result is returned. Otherwise, the actual deviation of the abnormal rate of the sample data from the preset abnormal rate threshold is calculated to obtain the sample abnormal distance value θ ws The larger the sample anomaly distance value, the more similar the data waveforms at the same time point in two adjacent sample collection cycles are, so the log anomaly detection model does not need to pay too much attention to the information at the aforementioned time points; the smaller the sample anomaly distance value, the greater the data waveform deviation is, so the log anomaly detection model needs to focus on the messages at the aforementioned time points.
[0101] For example, if the waveform of the sample log data at time t0 is not significantly different from the waveform of the sample log data at time t1 (which is equal to or close to the preset abnormal rate threshold, i.e., 0.9≤θ ws≤1), then the sample log data at time t0 is considered to have no obvious anomaly, so the log anomaly detection model does not need to pay too much attention to it, and the weight α of time t0 is given to 1; if the waveform of the sample log data at time t0 is different from that of the sample log data at time t1 (that is, 0.5≤θ ws <0.9), the sample log data at time t0 is considered abnormal, and the log anomaly detection model needs to pay attention to the message at this moment; if the waveform of the sample log data at time t0 is different from that of the sample log data at time t1 (that is, 0≤θ ws <0.5), it is considered that the sample log data at time t0 has obvious anomalies, and the log anomaly detection model needs to focus on the messages at this moment.
[0102] Exemplarily, a sample distance difference between a preset anomaly distance threshold and a sample anomaly distance value is determined; a candidate sample anomaly point weight is determined based on a preset proportional factor and the sample distance difference; and a target sample anomaly point weight at a target sample anomaly time point is determined based on the candidate sample anomaly point weight and a preset weight threshold.
[0103] The sample distance difference refers to the difference between the preset distance difference and the sample abnormal distance value. The candidate sample abnormal point weight refers to the weight of the sample log data at the target sample abnormal time point preliminarily determined based on the preset scale factor and the sample distance difference.
[0104] For example, according to the solution to the linear interpolation problem, a preset scaling factor μ is determined. A slightly larger factor (for example, μ = 15) is selected and 1-θ ws (i.e., the sample distance difference) multiplied by the preset scaling factor μ to obtain θ ws The result of the linear transformation, i.e., the candidate sample outlier weight, is taken as the remainder of 10 to ensure that the value of the target sample outlier weight does not exceed 10, i.e., the sample weight (normal point weight or target sample outlier weight) α (0<α≤10, default.1).
[0105] S340: Input the current sample log data sequence of the current sample collection period, the preset normal point weights and the target sample abnormal point weights into the constructed log anomaly detection model.
[0106] S350. For any current sample time point within the current sample collection cycle, determine the data forgetting operator, data updating operator, and data output operator of the current sample time point based on the sample log data and the corresponding sample weight of the current sample time point, and determine the current sample sub-result of the current sample time point based on the data forgetting operator, data updating operator, and data output operator.
[0107] The current sample time point refers to any time point within the current sample collection cycle. The sample log data refers to the sample log data corresponding to the current sample time point in the current sample log data sequence. The sample weight can be a normal point weight or a target sample abnormal point weight. Exemplarily, the sample weight corresponds to the sample log data at the current sample time point.
[0108] Among them, the data forgetting operator can be used to determine the forgotten information in the log anomaly detection model. The data updating operator can be used to determine the updated information in the log anomaly detection model. The data output operator can be used to determine the information output to the hidden state in the log anomaly detection model.
[0109] Among them, the current sample sub-result refers to the result obtained by the log anomaly detection model for the sample log data at the current sample time point through anomaly detection. Exemplarily, the current sample sub-result can be that the current sample is normal or the current sample is abnormal. The current sample is normal means that the sample log data at the current sample time point is detected, and the prediction result obtained is normal. The current sample is abnormal means that the sample log data at the current sample time point is detected, and the prediction result obtained is abnormal.
[0110] It should be noted that the structure of the log anomaly detection model in the embodiment of the present invention is divided into a forget gate, an input gate, a state update and an output gate.
[0111] Exemplarily, for any current sample time point within the sample collection cycle, the data forgetting operator of the current sample time point is determined based on the forgetting parameter matrix, the forgetting bias parameter, the previous sample result corresponding to the previous sample time point of the current sample time point, the sample log data of the current sample time point and the corresponding sample weight in the log anomaly detection model.
[0112] Among them, the forgetting parameter matrix refers to the parameter matrix of the forgetting gate in the log anomaly detection model. The forgetting bias parameter refers to the bias parameter of the forgetting gate in the log anomaly detection model. The previous sample time point refers to the moment before the current sample time point. The previous sample result refers to the detection result of the sample log data at the previous sample time point.
[0113] For example, the main function of the forget gate is to forget the information in the previous state at the current sample time point t, and to determine how much of the state at the previous moment needs to be retained to the current moment. The forget gate receives the current input and the state at the previous time, and generates a vector between 0 and 1 through the sigmoid function (i.e., S-type logic function). The data forgetting operator can be determined by the following formula:
[0114] f t =σ(W f ·[h t-1 ,αχt ]+b f );
[0115] Among them, f t represents the data forgetting operator; σ represents the activation function; W f represents the forgetting parameter matrix; h t-1 represents the result of the previous sample; α represents the sample weight at the current sample time point; χ t Indicates the sample log data at the current sample time point; b f It should be noted that the above formula shows the output h at the previous moment. t-1 and the current data input χ t Get f through the forget gate t The process is adjusted by introducing the importance level α.
[0116] Exemplarily, the data update operator at the current sample time point and the current default state at the current sample time point are determined based on the input parameter matrix, input bias parameter, previous sample result, sample log data at the current sample time point and corresponding sample weight in the log anomaly detection model.
[0117] Among them, the input parameter matrix refers to the parameter matrix of the input gate in the log anomaly detection model. The input bias parameter refers to the bias parameter of the input gate in the log anomaly detection model. The current default state refers to the temporary state of the log anomaly detection model at the current sample time point.
[0118] For example, after the forget gate is executed, the input gate determines which of the input information at the current sample time point should be added to the memory cell. The log anomaly detection model uses sigmoid and tanh functions (hyperbolic tangent function) to process the current input, generate a vector of new information, and add it to the memory cell state, so that only the information associated with the current context is updated. The data update operator can be determined by the following formula:
[0119] i t =σ ( W i ·[h t-1 ,(10-α)χ t ]+b i );
[0120] Among them, i t represents the data update operator; W i represents the input parameter matrix; b i Indicates the input bias parameter.
[0121] Furthermore, the current default state at the current sample time point is determined by the following formula:
[0122]
[0123] in, Indicates the current default state at the current sample time point; W c represents the weight matrix for calculating the current default state, that is, the parameter matrix for calculating the current default state; b c Represents the corresponding bias term, that is, the bias parameter for calculating the current default state.
[0124] It should be noted that the above two formulas show the output h at the previous moment t-1 and the current data input χ t , get i through the input gate t , and obtain the current temporary state through the unit state The process introduces the importance α and passes the input value χ t To determine the update that needs to be made in the cell state. The state that needs to be updated can be determined by the input gate.
[0125] Exemplarily, the current unit state is determined according to the previous unit state of the log anomaly detection model at the previous sample time point, the current default state at the current sample time point, the data forgetting operator, and the data updating operator. The previous unit state refers to the unit state of the log anomaly detection model at the previous sample time point. The current unit state refers to the unit state of the log anomaly detection model at the current sample time point.
[0126] For example, after being processed by the forget gate and the input gate, the previous unit state C is obtained. t-1 , forget gate output f t , input gate output i t And the output of the status The forget gate filters out the old information that does not need to be believed, and the input gate writes the new information, thus completing the information transmission at the current time. Update the cell state at time t-1 to obtain the current cell state C t The current cell state can be determined by the following formula:
[0127]
[0128] Among them, C t Indicates the current unit status; C t-1 Indicates the previous unit state; f t ·C t-1 Indicates forgetting of the previous unit state; Indicates an update to the current default state.
[0129] Exemplarily, the data output operator is determined according to the output parameter matrix, output bias parameter, previous sample sub-result, sample log data at the current sample time point and corresponding sample weight in the log anomaly detection model; the current sample sub-result at the current sample time point is determined according to the data output operator and the current unit state. The output parameter matrix refers to the parameter matrix of the output gate. The output bias parameter refers to the bias parameter of the output gate.
[0130] For example, the output gate determines the state of the current time, that is, the output value o of the log anomaly detection t After being processed by the output gate, the information of the memory cell state is selectively output to the state and used as the input of the next time, thus completing the information transmission of the current time, h t That is, the output of the model based on the output value and the current unit state. Exemplarily, the current sample sub-result at the current sample time point can be determined by the following formula:
[0131] o t =σ(W o ·[h t-1 ,αχ t ]+b o );
[0132] h t =o t tanh(C t );
[0133] Among them, t represents the data output operator; W0 represents the output parameter matrix; b0 represents the output bias parameter; h t Indicates the current sample sub-result at the current sample time point.
[0134] Furthermore, the new cell state (i.e., C t ) and the new hidden state (i.e. h t ) is transmitted to the next moment.
[0135] S360: Generate a current sample detection result including current sample sub-results at each current sample time point.
[0136] The current sample detection result refers to a set of current sample sub-results obtained by detecting sample log data at different current sample time points by the log anomaly detection model.
[0137] S370: Train the log anomaly detection model according to the current sample detection result and the actual sample detection result corresponding to the current sample collection period.
[0138] The actual detection result of the sample refers to the actual label of the sample log data at different current sample time points in the current sample log data sequence. Exemplarily, the actual label can be actually normal or actually abnormal. Actual normal means that the sample log data at a certain current sample time point is actually normal. Actual abnormal means that the sample log data at a certain current sample time point actually has an abnormality.
[0139] The training scheme of the log anomaly detection model provided by the embodiment of the present invention determines the data forgetting operator, the data updating operator and the data output operator according to the sample log data at the current sample time point and the corresponding sample weight, and then determines the current sample sub-result, and then determines the current sample detection result according to the current sample sub-result. The log anomaly detection model is trained through the current sample detection result and the actual sample detection result, thereby improving the accuracy of training the log anomaly detection model.
[0140] The technical solution provided by the embodiment of the present invention determines the importance of the log information at the current moment by comparing the data waveform, and the logic of judging the validity of the information is enhanced, which is more suitable for the characteristics of the huge amount of data in monitoring and other scenarios; secondly, the above technology is combined and applied to the log anomaly detection model, which improves the accuracy of the model in detecting log anomaly information, while greatly reducing the amount of training data.
[0141] It should be noted that the information collected in the embodiments of the present invention is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse; if the user chooses to refuse, the expert decision-making process will be entered.
[0142] Embodiment 4
[0143] Figure 4 This is a structural diagram of a log anomaly detection device provided by Embodiment 4 of the present invention. This embodiment is applicable to the case of performing anomaly detection on a log data sequence. The method can be performed by a log anomaly detection device, which can be implemented in software and / or hardware and can be configured in an electronic device that carries the log anomaly detection function.
[0144] like Figure 4 As shown, the device includes: a relative difference sequence determination module 410, a candidate abnormal data determination module 420, a target abnormal point weight determination module 430 and a log detection result determination module 440. Among them,
[0145] The relative difference sequence determination module 410 is used to obtain the periodic log data sequence within two adjacent collection periods, and determine the log relative difference sequence according to the periodic log data sequence; wherein the periodic log data sequence includes the current log data sequence and the previous log data sequence;
[0146] The candidate abnormal data determination module 420 is used to determine whether the log relative difference sequence has a log abnormal trend, and if so, determine the candidate abnormal time point and the corresponding candidate abnormal data in the current collection period according to the current log data sequence and the previous log data sequence;
[0147] A target abnormal point weight determination module 430 is used to determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence;
[0148] The log detection result determination module 440 is used to input the current log data sequence, the preset normal point weight and the target abnormal point weight into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
[0149] An embodiment of the present invention provides a log anomaly detection solution, which obtains periodic log data sequences within two adjacent collection cycles, and determines a log relative difference sequence according to the periodic log data sequences; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence; determines whether the log relative difference sequence has a log anomaly trend, and if so, determines candidate anomaly time points and corresponding candidate anomaly data within the current collection cycle according to the current log data sequence and the previous log data sequence; determines a target anomaly time point and a target anomaly point weight corresponding to the target anomaly time point according to the candidate anomaly time point, the candidate anomaly data and the log relative difference sequence; inputs the current log data sequence, the preset normal point weight and the target anomaly point weight into a trained log anomaly detection model, and obtains a current log detection result corresponding to the current log data sequence. The above scheme determines the log relative difference sequence according to the current log data sequence and the previous log data sequence to determine whether there is a log abnormal trend. If there is a log abnormal trend, the candidate abnormal time points and the corresponding candidate abnormal data in the current acquisition cycle are determined, and then the target abnormal time point and the corresponding target abnormal point weight are determined from the candidate abnormal time points. The current log data sequence, the preset normal point weight and the target abnormal point weight are input into the log anomaly detection model to obtain the current log detection result, and the target abnormal time point is determined from each time point associated with the current log data sequence, and the corresponding target abnormal point weight is determined for the log data at the target abnormal time point. At the same time, the current log data sequence, the normal point weight or the target abnormal point weight corresponding to each time point is input into the log anomaly detection model, and the current log detection result is output, so that the log anomaly detection model can focus on the log data at the target abnormal time point corresponding to the target abnormal point weight, thereby improving the accuracy of the current log detection result, that is, improving the accuracy of the log anomaly detection.
[0150] Optionally, the target outlier weight determination module 430 includes:
[0151] a data anomaly rate determination unit, configured to determine basic anomaly data from the log relative difference sequence according to the candidate anomaly time point, and determine the data anomaly rate of the candidate anomaly time point according to the basic anomaly data and the candidate anomaly data;
[0152] An abnormal distance value determination unit, used to determine a target abnormal time point and an abnormal distance value of the target abnormal time point according to the data abnormality rate and a preset abnormality rate threshold;
[0153] The outlier weight determination unit is used to determine the target outlier weight corresponding to the corresponding target outlier time point according to the outlier distance value.
[0154] Optionally, the abnormal distance value determining unit is specifically used to:
[0155] If the data anomaly rate is greater than the preset anomaly rate threshold, determining the corresponding candidate anomaly time point as the target anomaly time point;
[0156] The ratio between the preset abnormal rate threshold and the data abnormal rate at the target abnormal time point is used as the abnormal distance value of the target abnormal time point.
[0157] Optionally, the outlier weight determination unit is specifically configured to:
[0158] Determine a distance difference between a preset abnormal distance threshold and the abnormal distance value;
[0159] Determining the weight of the candidate outlier point according to the preset scale factor and the distance difference;
[0160] The target abnormal point weight of the target abnormal time point is determined according to the candidate abnormal point weight and a preset weight threshold.
[0161] Optionally, the candidate abnormal data determination module 420 includes:
[0162] A data waveform determining unit, used to determine a current data waveform corresponding to the current log data sequence, and a previous data waveform corresponding to the previous log data sequence;
[0163] The candidate abnormal data determining unit is used to determine the candidate abnormal time point and the corresponding candidate abnormal data in the current acquisition cycle according to the current data waveform and the previous data waveform.
[0164] Optionally, the log anomaly detection model is trained based on the following devices, including:
[0165] A current sample sub-result determination module is used to determine, for any current sample time point within the current sample collection period, a data forgetting operator, a data updating operator and a data output operator at the current sample time point according to the sample log data at the current sample time point and the corresponding sample weight, and determine the current sample sub-result at the current sample time point according to the data forgetting operator, the data updating operator and the data output operator;
[0166] A current sample detection result determination module, used to generate a current sample detection result including current sample sub-results at each current sample time point;
[0167] The model training module is used to train the log anomaly detection model according to the current sample detection result and the actual sample detection result corresponding to the current sample collection period.
[0168] The log anomaly detection device provided in the embodiment of the present invention can execute the log anomaly detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each log anomaly detection method.
[0169] In the technical solution of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of the periodic log data sequence, preset abnormality rate threshold, preset abnormality distance threshold, preset proportional factor and preset weight threshold, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0170] According to an embodiment of the present invention, the present invention also provides an electronic device, a readable storage medium and a computer program product.
[0171] Embodiment 5
[0172] Figure 5 It is a structural diagram of an electronic device for implementing a log anomaly detection method provided by Embodiment 5 of the present invention. Electronic device 510 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0173] like Figure 5 As shown, the electronic device 510 includes at least one processor 511, and a memory connected to the at least one processor 511 in communication, such as a read-only memory (ROM) 512, a random access memory (RAM) 513, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 511 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 512 or the computer program loaded from the storage unit 518 to the random access memory (RAM) 513. In the RAM 513, various programs and data required for the operation of the electronic device 510 can also be stored. The processor 511, the ROM 512, and the RAM 513 are connected to each other via a bus 514. An input / output (I / O) interface 515 is also connected to the bus 514.
[0174] A number of components in the electronic device 510 are connected to the I / O interface 515, including: an input unit 516, such as a keyboard, a mouse, etc.; an output unit 517, such as various types of displays, speakers, etc.; a storage unit 518, such as a disk, an optical disk, etc.; and a communication unit 519, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 519 allows the electronic device 510 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0175] The processor 511 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 511 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 511 executes the various methods and processes described above, such as the log anomaly detection method.
[0176] In some embodiments, the log anomaly detection method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 518. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 510 via the ROM 512 and / or the communication unit 519. When the computer program is loaded into the RAM 513 and executed by the processor 511, one or more steps of the log anomaly detection method described above may be performed. Alternatively, in other embodiments, the processor 511 may be configured to perform the log anomaly detection method in any other appropriate manner (e.g., by means of firmware).
[0177] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0178] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0179] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0180] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0181] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0182] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0183] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0184] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A log anomaly detection method, characterized in that: include: Acquire a periodic log data sequence within two adjacent collection periods, and determine a log relative difference sequence based on the periodic log data sequence; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence; Determine whether the log relative difference sequence has a log abnormality trend, and if so, determine candidate abnormal time points and corresponding candidate abnormal data within the current collection period according to the current log data sequence and the previous log data sequence; Determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence; The current log data sequence, the preset normal point weights and the target abnormal point weights are input into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
2. The method according to claim 1, characterized in that The determining, according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence, a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point includes: According to the candidate abnormal time point, basic abnormal data is determined from the log relative difference sequence, and according to the basic abnormal data and the candidate abnormal data, the data abnormality rate of the candidate abnormal time point is determined; Determine a target abnormal time point and an abnormal distance value of the target abnormal time point according to the data abnormal rate and a preset abnormal rate threshold; According to the abnormal distance value, the target abnormal point weight corresponding to the corresponding target abnormal time point is determined.
3. The method according to claim 2, characterized in that The step of determining a target abnormal time point and an abnormal distance value of the target abnormal time point according to the data abnormality rate and a preset abnormality rate threshold comprises: If the data anomaly rate is greater than the preset anomaly rate threshold, determining the corresponding candidate anomaly time point as the target anomaly time point; The ratio between the preset abnormal rate threshold and the data abnormal rate at the target abnormal time point is used as the abnormal distance value of the target abnormal time point.
4. The method according to claim 2, characterized in that: Determining the target abnormal point weight corresponding to the corresponding target abnormal time point according to the abnormal distance value includes: Determine a distance difference between a preset abnormal distance threshold and the abnormal distance value; Determining the weight of the candidate outlier point according to the preset scale factor and the distance difference; The target abnormal point weight of the target abnormal time point is determined according to the candidate abnormal point weight and a preset weight threshold.
5. The method according to claim 1, characterized in that The determining, according to the current log data sequence and the previous log data sequence, candidate abnormal time points and corresponding candidate abnormal data within the current collection cycle includes: Determining candidate abnormal time points and corresponding candidate abnormal data within a current collection period according to the current log data sequence and the previous log data sequence includes: Determine a current data waveform corresponding to the current log data sequence, and a previous data waveform corresponding to the previous log data sequence; According to the current data waveform and the previous data waveform, candidate abnormal time points and corresponding candidate abnormal data within the current acquisition cycle are determined.
6. The method according to any one of claims 1 to 5, characterized in that The log anomaly detection model is trained based on the following methods, including: For any current sample time point within the current sample collection cycle, determine the data forgetting operator, data updating operator and data output operator at the current sample time point according to the sample log data and the corresponding sample weight at the current sample time point, and determine the current sample sub-result at the current sample time point according to the data forgetting operator, the data updating operator and the data output operator; Generate a current sample detection result including current sample sub-results at each current sample time point; The log anomaly detection model is trained according to the current sample detection result and the actual sample detection result corresponding to the current sample collection period.
7. A log anomaly detection device, characterized in that: include: A relative difference sequence determination module is used to obtain a periodic log data sequence within two adjacent collection periods, and determine a log relative difference sequence based on the periodic log data sequence; wherein the periodic log data sequence includes a current log data sequence and a previous log data sequence; A candidate abnormal data determination module is used to determine whether the log relative difference sequence has a log abnormal trend, and if so, determine the candidate abnormal time point and the corresponding candidate abnormal data in the current collection period according to the current log data sequence and the previous log data sequence; A target abnormal point weight determination module is used to determine a target abnormal time point and a target abnormal point weight corresponding to the target abnormal time point according to the candidate abnormal time point, the candidate abnormal data and the log relative difference sequence; The log detection result determination module is used to input the current log data sequence, the preset normal point weight and the target abnormal point weight into the trained log anomaly detection model to obtain the current log detection result corresponding to the current log data sequence.
8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a log anomaly detection method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a log anomaly detection method as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the steps of the log anomaly detection method according to any one of claims 1 to 6 are implemented.