A method and device for detecting log anomalies

By clustering and matching real-time cluster logs to generate tag trees, combined with baseline monitoring and sequential detection models, the problem of long troubleshooting time and cumbersome process of log exception detection methods in the existing technology is solved, and rapid abnormality detection and root cause positioning of massive large data cluster logs are achieved.

CN114647558BActive Publication Date: 2025-07-18JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210173675.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-07-18
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The troubleshooting time of log abnormality detection methods in the prior art is slow and the troubleshooting process is cumbersome, especially in the case of massive large data cluster logs, it is difficult to cover the missed coverage and take into account the sudden increase in the number of logs.

Method used

By clustering analysis of real-time cluster logs, the tag tree is generated, and matched with the log template library, the log exception category is determined, the baseline monitoring model and sequential detection model are used for abnormal detection, and combined with root cause positioning technology, real-time aggregation analysis and abnormal detection of massive large data cluster logs are realized.

Benefits of technology

It reduces the workload of manual troubleshooting, simplifies the troubleshooting process, improves the troubleshooting efficiency, and can promptly detect log abnormalities and quickly locate the root causes of problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114647558B_ABST
    Figure CN114647558B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for detecting log anomalies. The method includes: performing clustering analysis on the acquired real-time cluster logs to generate a corresponding label tree; matching the label tree with a log template library to determine the log template that matches the label tree and the corresponding log anomaly category, and saving the real-time cluster logs into the corresponding log template according to the log anomaly category, where the log template library includes a plurality of log templates, and each log template has a corresponding log anomaly category; performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result. It can realize the aggregation analysis of the real-time cluster logs of massive big data, and then perform anomaly detection on each type of real-time cluster logs to determine the detection result, reducing the workload of manual troubleshooting and simplifying the troubleshooting process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for detecting abnormal logs, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] For the abnormal detection of cluster logs, it is a relatively common detection technology for computer clusters to monitor the cluster logs and discover problems in a timely manner.

[0003] In the prior art, the detection of cluster logs depends on rule scripts written by operation and maintenance engineers according to experience. In the face of a large amount of big data cluster logs (reaching hundreds of millions per day), the existing methods will have omissions in coverage and it is difficult to take into account the abnormal situation of a sudden increase in the number of logs continuously in a certain period. Therefore, the fault troubleshooting time of the existing log abnormal detection method is slow and the fault troubleshooting process is cumbersome. Summary of the Invention

[0004] The present disclosure provides a method and apparatus for detecting abnormal logs, an electronic device, and a non-transitory computer-readable storage medium, so as to solve the technical problems of slow fault troubleshooting time and cumbersome fault troubleshooting process in the existing log abnormal detection method.

[0005] The present disclosure provides a method for detecting abnormal logs, including:

[0006] Performing clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0007] Matching the label tree with a log template library to determine a log template and a corresponding log abnormal category that match the label tree, and saving the real-time cluster logs according to the log abnormal category to the corresponding log template, where the log template library includes a plurality of log templates, and each log template has a corresponding log abnormal category;

[0008] Performing abnormal detection on the real-time cluster logs of different log abnormal categories to determine a detection result.

[0009] According to the method for detecting abnormal logs provided by an embodiment of the present disclosure, the method for generating the log template library includes:

[0010] Obtaining historical cluster logs;

[0011] Generating an initial label tree based on the historical cluster logs;

[0012] Building an initial template tree, training the initial template tree based on the initial label tree to generate a template, and generating a log template library from the template;

[0013] Perform secondary clustering on the templates, and label the corresponding log anomaly categories for each category of templates.

[0014] According to the method for log anomaly detection provided by an embodiment of the present disclosure, the method further includes:

[0015] In the case where the label tree does not match the log template library, calculate the similarity between the unmatched real-time cluster logs and the historical cluster logs corresponding to each existing log anomaly category, and determine the log anomaly category corresponding to the unmatched real-time cluster logs;

[0016] Based on the unmatched real-time cluster logs and their corresponding log anomaly categories, perform an incremental training task on the log template library to obtain an updated log template library.

[0017] According to the method for log anomaly detection provided by an embodiment of the present disclosure, perform anomaly detection based on the real-time cluster logs of different log anomaly categories, and determine the detection result, including:

[0018] Convert the real-time cluster logs and historical cluster logs corresponding to each log anomaly category into time series metrics;

[0019] Input the time series metrics into the baseline monitoring model, and output the anomaly prediction values corresponding to each log anomaly category;

[0020] Wherein, the anomaly prediction values include: mean change of time series metrics, jitter frequency change, detection of peaks and valleys, and drop ratio values.

[0021] According to the method for log anomaly detection provided by an embodiment of the present disclosure, perform anomaly detection based on the real-time cluster logs of different log anomaly categories, and determine the detection result, including:

[0022] Determine the proportion of the real-time cluster logs of different log anomaly categories, and determine the first detection result according to the proportion.

[0023] According to the method for log anomaly detection provided by an embodiment of the present disclosure, perform anomaly detection based on the real-time cluster logs of different log anomaly categories, and determine the detection result, including:

[0024] Input the real-time cluster logs of different log anomaly categories into the sequential detection model, and output the second detection result.

[0025] According to the method for log anomaly detection provided by an embodiment of the present disclosure, after determining the detection result, the method further includes:

[0026] Generate a log metric time series curve for each cluster and a total log metric time series curve according to the time series metrics corresponding to different cluster logs;

[0027] Compare the change trend of the log metric time series curve of each cluster with the change trend of the total log metric time series curve;

[0028] If the change trends are consistent, based on the proportion of the real-time cluster logs of the cluster in different log anomaly categories, determine the log anomaly category with a relatively large proportion as the main log anomaly category, and based on the main log anomaly category, perform root cause localization to determine the machine identifier with anomalies in the cluster.

[0029] The present disclosure provides an apparatus for log anomaly detection, including:

[0030] A clustering module for performing clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0031] A matching module for matching the label tree with a log template library to determine the log template and the corresponding log anomaly category that match the label tree, and saving the real-time cluster logs into the corresponding log template according to the log anomaly category, where the log template library includes a plurality of log templates, and each log template has a corresponding log anomaly category;

[0032] A detection module for performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine a detection result.

[0033] The present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for log anomaly detection as described in any one of the above are implemented.

[0034] The present disclosure further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for log anomaly detection as described in any one of the above are implemented.

[0035] The method and apparatus for log anomaly detection provided by the present disclosure cluster the cluster logs to generate a label tree, and match it with a log template library to determine the log template and the corresponding log anomaly category that match the label tree. Thus, problems in the big data cluster can be actively discovered from the perspective of cluster logs through online real-time matching, and aggregation analysis of real-time cluster logs of massive big data can be realized. Furthermore, anomaly detection is performed on each type of real-time cluster logs to determine a detection result, reducing the workload of manual troubleshooting and simplifying the troubleshooting process. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions in the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0037] Figure 1 is one of the schematic structural diagrams of the device for detecting log anomalies provided by the present disclosure;

[0038] Figure 2 is a schematic diagram of the tag tree generated after clustering the obtained real-time cluster logs provided by the present disclosure;

[0039] Figure 3 is a schematic diagram of a log template library provided by the present disclosure;

[0040] Figure 4 is the second schematic structural diagram of the device for detecting log anomalies provided by the present disclosure;

[0041] Figure 5 is the third schematic structural diagram of the device for detecting log anomalies provided by the present disclosure;

[0042] Figure 6 is the fourth schematic structural diagram of the device for detecting log anomalies provided by the present disclosure;

[0043] Figure 7 is the schematic structural diagram of the device for detecting log anomalies provided by the present disclosure;

[0044] Figure 8 is the schematic structural diagram of the electronic device provided by the present disclosure. Detailed implementation manners

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments in the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the embodiments of the present disclosure.

[0046] For the methods in the prior art, there are still other problems. For example, it is very difficult to detect abnormal events through rule matching of big data logs. Some error types belong to normal business error reports and are not system failures. Therefore, relevant rules are often not configured. However, under specific conditions, they may no longer be normal business error reports, which poses a security risk to the normal operation of the business. Moreover, it is very difficult for manual rule scripts to fully cover these scenarios.

[0047] To solve the technical defects existing in the prior art, an embodiment of the present disclosure discloses a method for log anomaly detection. Refer to Figure 1 , including:

[0048] Step 101: Perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree.

[0049] In this embodiment, instead of matching each real-time cluster log with the log template library, a corresponding label tree is generated after clustering analysis, and then the label tree is matched with the log template library, thereby converting the problem of troubleshooting complex big data cluster logs into the problem of matching the label tree with the log template, which is beneficial to improving the efficiency of troubleshooting and reducing the troubleshooting time.

[0050] Refer to Figure 2 , Figure 2 which is the label tree generated after clustering the obtained real-time cluster logs.

[0051] Step 102: Match the label tree with the log template library to determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category.

[0052] wherein, the log template library includes multiple log templates, and each log template has a corresponding log anomaly category.

[0053] It should be noted that the method for generating the log template library includes the following steps S21 to S24:

[0054] S21: Obtain historical cluster logs.

[0055] S22: Generate an initial label tree based on the historical cluster logs.

[0056] S23: Build an initial template tree, train the initial template tree based on the initial label tree to generate templates, and generate a log template library from the templates.

[0057] S24: Perform secondary clustering on the templates, and label the corresponding log anomaly category for each type of template.

[0058] Through the above steps S21 to S24, a log template library is generated for matching with the tag tree generated by the implementation cluster logs. See Figure 3 , Figure 3 which shows a log template library of this embodiment.

[0059] Among them, the log exception categories of the log template library include 6: MemStore data flushing operation delay, GC memory recycling occurred, heap memory usage rate exceeded the maximum quota, the cluster was slow in processing a certain table operation, the data block size exceeded the quota resulting in cache failure, and the connection to the zookeeper server timed out. After the log exception categories are determined, they generally will not change again, and only the real-time cluster logs need to be matched to each log exception category.

[0060] After successful matching, the real-time cluster logs are saved to each log exception category of the log template.

[0061] Step 103: Perform anomaly detection based on the real-time cluster logs of different log exception categories to determine the detection result.

[0062] Among them, the dimensions of anomaly detection can be various. For example, perform anomaly detection and time series detection on the real-time cluster logs corresponding to each log exception category, perform anomaly detection on the proportion of the real-time cluster logs of each log exception category, perform root cause location on the logs generated by different clusters, and so on.

[0063] The method for log anomaly detection provided by the present disclosure clusters the cluster logs to generate a tag tree, matches it with the log template library, determines the log template and the corresponding log exception category that match the tag tree, so as to actively discover problems in the big data cluster from the perspective of cluster logs through online real-time matching, and can realize the aggregation analysis of real-time cluster logs of massive big data, and then perform anomaly detection on each type of real-time cluster logs to determine the detection result, reducing the workload of manual troubleshooting and simplifying the troubleshooting process.

[0064] Furthermore, in the case where the tag tree and the log template library do not match, the method can further perform incremental learning on the log template library using the unmatched real-time cluster logs to expand the log template library.

[0065] Specifically, see Figure 4 , the method of the embodiment of the present disclosure includes steps 401 to 402:

[0066] 401: Calculate the similarity between the unmatched real-time cluster logs and the historical cluster logs corresponding to each existing log exception category to determine the log exception category corresponding to the unmatched real-time cluster logs.

[0067] In this embodiment, it is necessary to calculate the text similarity between each real-time cluster log and the historical cluster log to determine the log anomaly category corresponding to each unmatched real-time cluster log.

[0068] 402. Based on the unmatched real-time cluster logs and their corresponding log anomaly categories, perform an incremental training task on the log template library to obtain an updated log template library.

[0069] This embodiment can not only perform template training on historical cluster logs to generate a log template library with a wide coverage, and push it online for real-time matching through an algorithm based on a tag tree, but also continue to perform incremental training on the log template library for unmatched real-time cluster logs, so that the updated log template library has a wide coverage. The real-time matching algorithm based on the tag tree has higher timeliness than regular expressions, and can achieve aggregated analysis of massive big data cluster logs, reducing the workload of manual troubleshooting. And after saving the real-time cluster logs to the log template, problems can be discovered in advance by monitoring the log template, and losses can be stopped in time.

[0070] Further, step 103 includes:

[0071] Convert the real-time cluster logs and historical cluster logs corresponding to each log anomaly category into time series metrics;

[0072] Input the time series metrics into the baseline monitoring model to output the anomaly prediction value corresponding to each log anomaly category;

[0073] Among them, the anomaly prediction value includes: the mean change of the time series metric, the jitter frequency change, the detection of peaks and valleys, and the drop ratio value.

[0074] In this embodiment, the baseline monitoring model can be a DeepAR model, including an encoder and a decoder. Through the baseline monitoring model, the predicted values for a future period of time can be obtained, such as the predicted values for the next 10 minutes.

[0075] This model is based on the autoregressive principle. The true value of the previous moment is used as the feature input to the encoder network at the current moment, and the predicted value of the previous moment is used as the feature input to the decoder network for time series prediction. The effect of the DeepAR model is evaluated by calculating the root mean square error (RMSE). Using 120 time steps (the time step can be selected as 10s, 30s, 1min) of data to predict 100 time steps of time series data, the evaluation results show that the DeepAR model is suitable for intelligent baseline prediction after quantifying the metrics of big data cluster logs. By adapting the upper and lower limits, it can better identify anomalies in big data cluster logs.

[0076] Further, step 103 includes: determining the proportion of real-time cluster logs of different log anomaly categories, and determining a first detection result according to the proportion, so as to realize taking the angles of different log anomaly categories as considerations for anomaly detection.

[0077] Further, anomaly detection of the logical order of logs can also be realized. Step 103 includes: inputting the real-time cluster logs of different log anomaly categories into an order detection model, and outputting a second detection result.

[0078] In this embodiment, the order detection model can be a time series model of CNN+LSTM, so as to realize using the template sequence attribute of the logs to convert the log anomaly detection problem into a multi-classification problem and perform anomaly detection on the log logical order.

[0079] When using the CNN+LSTM model, the input is the time series data after indexing the log templates of 6 classes for a period of time.

[0080] When training the CNN+LSTM model, the input data for LSTM is: 128*6*5. Taking one sample as an example, the data input at each time step (a total of 5 time steps) is 30*1, and an output of 6*1 is obtained. That is, the data structure of the output of LSTM before concat is a 6*1 matrix. 128 is the batch size, that is, the number of samples selected for one training. By performing multiple trainings, all the data can be traversed.

[0081] The data input to CNN is: 1*5*6, the convolutional kernel: a 3*6 matrix, and the feature maps: 128. The convolutional kernel matrix is gradually swept over the input data, multiplied and added at the corresponding positions, and 0 is used to pad the input data at the same time, obtaining 128 feature maps of 5*6.

[0082] Optionally, after determining the detection result, refer to Figure 5 , the method further includes the following steps 501 to 503:

[0083] Step 501, generate a log metric time series curve for each cluster and a total log metric time series curve according to the time series metrics corresponding to different cluster logs.

[0084] Step 502, compare the change trend of the log metric time series curve of each cluster with the change trend of the total log metric time series curve.

[0085] Step 503, if the change trends are consistent, based on the proportion of the real-time cluster logs of the cluster in different log anomaly categories, determine the log anomaly category with a larger proportion as the main log anomaly category, and perform root cause location based on the main log anomaly category to determine the machine identifier with anomalies in the cluster.

[0086] In this embodiment, a root cause location detection model is configured for the log template metrics of important classes, so that the sudden increase in the total amount of logs can be quickly detected. Through multi-dimensional drill-down analysis, it is located that the time series curve of the log metrics of a certain cluster is consistent with the total amount change. Combining with the analysis of the trend change of the template ratio of this cluster, the specific machine identifiers that cause problems in the big data cluster are located. The log volume of the big data cluster has the characteristic of trend change, and there is a multi-dimensional drill-down relationship between the log and the template. Quickly locating the combination of machine identifiers that cause problems in the big data cluster dimension can solve the problems of slow problem troubleshooting and difficult root cause location.

[0087] Specifically, the root cause location detection model of this embodiment constructs evaluation indicators, screens the element set of the root cause, determines the preliminary search space, uses the reinforcement learning search method to search for the set with the highest multi-dimensional root cause possibility, and corrects the final root cause. The principle of root cause correction: The attribute combination with a larger potential score is more likely to be the root cause. When two element sets have the same potential score, the one with fewer elements wins.

[0088] An embodiment of the present disclosure also provides a method for log anomaly detection, see Figure 6 , including:

[0089] Step 601: Perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree.

[0090] Step 602: Match the label tree with the log template library, and determine whether the label tree matches the log template library. If it matches, execute Step 603; if it does not match, execute Step 604.

[0091] Step 603: Determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category.

[0092] Among them, the log template library includes multiple log templates, and each log template has a corresponding log anomaly category.

[0093] Step 604: Calculate the similarity between the unmatched real-time cluster logs and the historical cluster logs corresponding to each existing log anomaly category, and determine the log anomaly category corresponding to the unmatched real-time cluster logs; based on the unmatched real-time cluster logs and their corresponding log anomaly categories, perform an incremental training task on the log template library to obtain an updated log template library, and return to execute Step 602.

[0094] Step 605: Perform anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result.

[0095] Among them, there are various methods for anomaly detection:

[0096] In the first case, convert the real-time cluster logs and historical cluster logs corresponding to each log anomaly category into time series metrics; input the time series metrics into the baseline monitoring model, and output the anomaly prediction values corresponding to each log anomaly category.

[0097] In this embodiment, the anomaly prediction values include: mean change of time series metrics, jitter frequency change, detection of spikes and valleys, and drop ratio values.

[0098] In the second case, determine the proportion of real-time cluster logs of different log anomaly categories, and determine the first detection result according to the proportion.

[0099] In the third case, input the real-time cluster logs of different log anomaly categories into the sequential detection model, and output the second detection result.

[0100] Step 606: Generate the log metric time series curve of each cluster and the total log metric time series curve according to the time series metrics corresponding to different cluster logs.

[0101] Step 607: Compare the change trend of the log metric time series curve of each cluster with the change trend of the total log metric time series curve.

[0102] Step 608: If the change trends are consistent, based on the proportion of the real-time cluster logs of this cluster in different log anomaly categories, determine the log anomaly category with a larger proportion as the main log anomaly category, and based on the main log anomaly category, perform root cause localization to determine the machine identifier with anomalies in the cluster.

[0103] The embodiments of the present disclosure convert the complicated process of troubleshooting big data cluster logs into a form of template comparison, which can achieve fast clustering of logs and analysis of log categories from a global perspective. The log template library is trained and generated by the FT-Tree method. By matching the log templates in real time online to generate the time series metrics of log categories, log anomaly points are discovered, problems are discovered in advance, and the problems of slow troubleshooting time for big data cluster failures, difficult passive troubleshooting of big data cluster problems, and cumbersome troubleshooting processes are solved.

[0104] In addition, configure the baseline monitoring model with the metrics after quantifying the key monitoring templates. By comparing historical data, a continuous abnormal sudden increase in quantity during this period can be found, and faults can be hit in advance.

[0105] Again, in this embodiment, through the online incremental learning log template method, based on a sliding window, feature extraction can be performed on the logs to find the correlation between templates and detect anomalies in newly added log pattern combinations; using the time series attribute of the logs, the log anomaly detection problem is converted into a multi-classification problem, and an anomaly detection is performed on the logical order of the logs by training a sequential detection model (CNN+LSTM); after an anomaly is found, the associated alarms are analyzed by drilling down into the fault details to assist the operation and maintenance personnel in analyzing the specific root cause of the fault.

[0106] The embodiment of the present disclosure further includes a device for log anomaly detection. Refer to Figure 7 , including:

[0107] A clustering module 701, configured to perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0108] A matching module 702, configured to match the label tree with a log template library to determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category, where the log template library includes multiple log templates, and each log template has a corresponding log anomaly category;

[0109] A detection module 703, configured to perform anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result.

[0110] Optionally, the device further includes a historical template library generation module, configured to:

[0111] Obtain historical cluster logs;

[0112] Generate an initial label tree based on the historical cluster logs;

[0113] Build an initial template tree, train the initial template tree based on the initial label tree to generate templates, and generate a log template library from the templates;

[0114] Perform secondary clustering on the templates, and label each category of the templates with the corresponding log anomaly category.

[0115] Optionally, the device further includes:

[0116] A similarity calculation module, configured to calculate the similarity between the unmatched real-time cluster logs and the historical cluster logs corresponding to each existing log anomaly category in the case where the label tree does not match the log template library, and determine the log anomaly category corresponding to the unmatched real-time cluster logs;

[0117] An update module, configured to perform an incremental training task on the log template library based on the unmatched real-time cluster logs and their corresponding log anomaly categories, to obtain an updated log template library.

[0118] Optionally, the detection module 703 is specifically configured to:

[0119] Convert the real-time cluster logs and historical cluster logs corresponding to each log anomaly category into time series metrics;

[0120] Input the time series metrics into a baseline monitoring model, and output anomaly prediction values corresponding to each log anomaly category;

[0121] Wherein, the anomaly prediction values include: mean change of time series metrics, jitter frequency change, detection of spikes and valleys, and drop ratio values.

[0122] Optionally, the detection module 703 is specifically configured to: determine the proportion of real-time cluster logs of different log anomaly categories, and determine a first detection result according to the proportion.

[0123] Optionally, the detection module 703 is specifically configured to: input the real-time cluster logs of different log anomaly categories into a sequential detection model, and output a second detection result.

[0124] Optionally, the apparatus further includes:

[0125] A curve generation module, configured to generate a log metric time series curve for each cluster and a total log metric time series curve according to the time series metrics corresponding to different cluster logs after determining the detection result;

[0126] A trend comparison module, configured to compare the change trend of the log metric time series curve of each cluster with the change trend of the total log metric time series curve;

[0127] A root cause location module, configured to, if the change trends are consistent, based on the proportion of the real-time cluster logs of the cluster in different log anomaly categories, determine the log anomaly category with a larger proportion as the main log anomaly category, and perform root cause location based on the main log anomaly category to determine the machine identifier with anomalies in the cluster.

[0128] The apparatus for log anomaly detection provided by the embodiments of the present disclosure clusters cluster logs to generate a label tree, matches the label tree with a log template library, determines the log template and the corresponding log anomaly category that match the label tree, so as to actively discover problems in a big data cluster from the perspective of cluster logs through online real-time matching, can achieve aggregation analysis of real-time cluster logs of massive big data, and further perform anomaly detection on each type of real-time cluster logs to determine the detection result, reducing the workload of manual troubleshooting and simplifying the troubleshooting process.

[0129] Figure 8 illustrates a schematic diagram of the physical structure of an electronic device, as Figure 8 shown. The electronic device may include: a processor 801, a communications interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communications interface 802, and the memory 803 communicate with each other through the communication bus 804. The processor 801 may call the logical instructions in the memory 803 to execute the method for log anomaly detection, including:

[0130] Performing clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0131] Matching the label tree with the log template library to determine the log template and the corresponding log anomaly category that match the label tree, and saving the real-time cluster logs to the corresponding log template according to the log anomaly category. Among them, the log template library includes multiple log templates, and each log template has a corresponding log anomaly category;

[0132] Performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result.

[0133] In addition, when the logical instructions in the above-mentioned memory 803 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0134] On the other hand, the present disclosure also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for log anomaly detection provided by the above-mentioned various methods, including:

[0135] Perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0136] Match the label tree with the log template library to determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category, where the log template library includes multiple log templates, and each log template has a corresponding log anomaly category;

[0137] Perform anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result.

[0138] In another aspect, the present disclosure also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is implemented to execute the method for log anomaly detection provided above, including:

[0139] Perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree;

[0140] Match the label tree with the log template library to determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category, where the log template library includes multiple log templates, and each log template has a corresponding log anomaly category;

[0141] Perform anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result.

[0142] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A method for detecting log anomalies, characterized in that, Including: Performing clustering analysis on the obtained real-time cluster logs to generate a corresponding tag tree; wherein, the tag tree is a tree structure formed by concatenating the words in the real-time cluster logs in sequence; Matching the tag tree with a log template library to determine the log template and the corresponding log anomaly category that match the tag tree, and saving the real-time cluster logs into the corresponding log template according to the log anomaly category, wherein the log template library includes multiple log templates, and each log template has a corresponding log anomaly category; Performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result; The method for generating the log template library includes: Obtaining historical cluster logs; Generating an initial tag tree based on the historical cluster logs; Constructing an initial template tree, training the initial template tree based on the initial tag tree to generate templates, and generating a log template library from the templates; Performing secondary clustering on the templates, and labeling each category of templates with a corresponding log anomaly category.

2. The method for detecting abnormal logs according to claim 1, wherein, The method further includes: In the case where the tag tree does not match the log template library, calculating the similarity between the unmatched real-time cluster logs and the historical cluster logs corresponding to each existing log anomaly category to determine the log anomaly category corresponding to the unmatched real-time cluster logs; Performing an incremental training task on the log template library based on the unmatched real-time cluster logs and their corresponding log anomaly categories to obtain an updated log template library.

3. The method for log anomaly detection according to claim 1, wherein Performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result, including: Converting the real-time cluster logs and historical cluster logs corresponding to each log anomaly category into time series metrics; Inputting the time series metrics into a baseline monitoring model to output an anomaly prediction value corresponding to each log anomaly category; Wherein, the anomaly prediction value includes: mean change of time series metrics, jitter frequency change, detection of spikes and valleys, and drop ratio value.

4. The method for detecting abnormal logs according to claim 1, wherein Performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result, including: Determining the proportion of the real-time cluster logs of different log anomaly categories, and determining a first detection result according to the proportion.

5. The method for detecting abnormal logs according to claim 1, wherein, Performing anomaly detection based on the real-time cluster logs of different log anomaly categories to determine the detection result, including: Inputting the real-time cluster logs of different log anomaly categories into a sequential detection model to output a second detection result.

6. The method for log anomaly detection according to claim 1, wherein After determining the detection result, the method further includes: Generating a log metric time series curve for each cluster and a total log metric time series curve according to the time series metrics corresponding to different cluster logs; Comparing the change trend of the log metric time series curve of each cluster with the change trend of the total log metric time series curve; If the change trends are consistent, based on the proportion of the real-time cluster logs of the cluster in different log anomaly categories, determining the log anomaly category with the largest proportion as the main log anomaly category, and performing root cause localization based on the main log anomaly category to determine the machine identifier with anomalies in the cluster.

7. A device for detecting log anomalies, characterized in that, Including: A clustering module, which is used to perform clustering analysis on the obtained real-time cluster logs to generate a corresponding label tree; wherein, the label tree is a tree structure formed by concatenating the words in the real-time cluster logs in the order of appearance; A matching module, which is used to match the label tree with a log template library, determine the log template and the corresponding log anomaly category that match the label tree, and save the real-time cluster logs to the corresponding log template according to the log anomaly category, wherein the log template library includes multiple log templates, and each log template has a corresponding log anomaly category; A detection module, which is used to perform anomaly detection on the real-time cluster logs of different log anomaly categories to determine the detection result; The device further includes a historical template library generation module, which is used for: Obtaining historical cluster logs; Generating an initial label tree based on the historical cluster logs; Building an initial template tree, training the initial template tree based on the initial label tree to generate templates, and generating a log template library from the templates; Performing secondary clustering on the templates, and labeling each type of template with a corresponding log anomaly category.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for log anomaly detection according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for log anomaly detection according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Log template extraction method based on online hierarchical clustering

    CN109981625A

  • System operation log monitoring method and device, electronic equipment and storage medium

    CN113760645A