Abnormal Log Detection Method, Device, Electronic Device and Storage Medium

Through the abnormal log detection method of multi-algorithm fusion, logistic regression and catboost algorithm weighted fusion are used to determine the weight based on the number of domain keywords, solving the reliability and misjudgment problems of abnormal log detection in the existing technology, and achieving efficient and accurate abnormal log recognition.

CN115048345BActive Publication Date: 2025-07-22CHINA MOBILE GROUP JIANGSU +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110251137.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-08
Publication Date
2025-07-22
Estimated Expiration
2041-03-08

AI Technical Summary

Technical Problem

The existing abnormal log detection methods have poor reliability and high error rates. Methods based on human experience and manual rules cannot accurately identify abnormal logs, resulting in misjudgment and high maintenance costs.

Method used

A variety of anomaly detection algorithms (such as logistic regression and catboost) are used to detect log features, and the abnormal detection results are obtained through weighted fusion. The weight is determined based on the number of domain keywords contained in the log to be detected, and the model is trained based on the log features of the sample log and its exception labels.

Benefits of technology

It improves the accuracy and reliability of abnormal log detection, reduces manpower and material investment, adapts to large-scale operation and maintenance scenarios, and reduces the misjudgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048345B_ABST
    Figure CN115048345B_ABST
Patent Text Reader

Abstract

The present invention provides an abnormal log detection method, device, electronic device and storage medium. The method includes: determining the log features of the log to be detected; based on an anomaly detection model, performing anomaly detection on the log features under multiple different algorithms, and performing weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected; the anomaly detection model is trained based on the log features of sample logs and their anomaly labels. The method, device, electronic device and storage medium provided by the present invention determine the weight of weighted fusion by applying the number of domain keywords included in the log to be detected, so as to fuse the detection results obtained by performing anomaly detection under multiple different algorithms, realize anomaly detection with multi-algorithm fusion, and thus ensure the accuracy and reliability of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of operation and maintenance business support, and particularly to an abnormal log detection method, device, electronic device and storage medium. Background Art

[0002] With the advent of the 5G (the 5th generation mobile communication), Internet of Things and big data era, the enterprise information system has witnessed an explosive growth. The operation and maintenance logs are ever-changing, and the operation and maintenance pressure faced by operation and maintenance personnel is increasing.

[0003] In the network operation and maintenance center, operation and maintenance engineers have to face thousands of log data every day. The traditional operation and maintenance methods are based on human experience to determine whether the logs are abnormal, and a simple method based on fixed rules for log abnormality determination.

[0004] However, abnormal logs cannot be accurately identified based on human experience, and misjudgments often occur in the method based on manual rules, leading to other problems. In addition, for the determination system based on manual rules, a large amount of human and material resources are required for maintenance costs. Summary of the Invention

[0005] The present invention provides an abnormal log detection method, device, electronic device and storage medium to solve the problems of poor reliability and high error rate of the existing abnormal log detection methods.

[0006] The present invention provides an abnormal log detection method, including:

[0007] Determine the log features of the log to be detected;

[0008] Based on the anomaly detection model, perform anomaly detection on the log features under multiple different algorithms, and perform weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weights of the weighted fusion are determined based on the number of domain keywords included in the log to be detected;

[0009] The anomaly detection model is trained based on the log features of the sample logs and their anomaly labels.

[0010] According to an abnormal log detection method provided by the present invention, the weights of the weighted fusion are determined according to the following steps:

[0011] Based on the total number of log word segments included in the sample logs, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected, determine the algorithm weights of each algorithm;

[0012] Determine the weight of the weighted fusion based on the algorithm weights of each algorithm and the accuracy of anomaly detection under each algorithm.

[0013] According to an anomaly log detection method provided by the present invention, the multiple different algorithms include a logistic regression algorithm and a catboost algorithm;

[0014] Determining the algorithm weights of each algorithm based on the total number of log segments included in the sample log, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected includes:

[0015] Determine the algorithm weights of each algorithm based on the following formula:

[0016]

[0017]

[0018] In the formula, μ 逻辑回归 and μ catboost are the algorithm weights of the logistic regression algorithm and the catboost algorithm respectively, N = n + k, n is the total number of log segments included in the sample log, k is the total number of preset domain keywords, and p is the number of domain keywords included in the log to be detected.

[0019] According to an anomaly log detection method provided by the present invention, the anomaly label is determined based on the following steps:

[0020] Determine the log features of each sample log;

[0021] Cluster the log features of each sample log to obtain multiple log clusters;

[0022] Obtain the manual labels of each log cluster, and use the manual labels as the anomaly labels of the sample logs corresponding to the log features within the corresponding log cluster.

[0023] According to an anomaly log detection method provided by the present invention, determining the log features of the log to be detected includes:

[0024] Perform text segmentation on the log to be detected to obtain each segment of the log to be detected;

[0025] Perform domain keyword matching on each segment of the log to be detected, and determine the keyword vector based on the matching result;

[0026] Construct the log features of the log to be detected based on the word vectors of each segment of the log to be detected and the keyword vector.

[0027] An abnormal log detection method provided by the present invention, wherein the word vectors of each word segment are determined based on word frequency and inverse document frequency index.

[0028] An abnormal log detection method provided by the present invention, before performing text segmentation on the log to be detected, further includes:

[0029] Removing abnormal symbols and extracting Chinese information from the log to be detected;

[0030] Based on the log text after removing abnormal symbols and the extracted Chinese information, reconstruct the log to be detected.

[0031] The present invention also provides an abnormal log detection device, including:

[0032] A feature extraction unit, configured to determine the log features of the log to be detected;

[0033] An abnormal detection unit, configured to perform abnormal detection on the log features under multiple different algorithms based on an abnormal detection model, and perform weighted fusion on the detection results under multiple different algorithms to obtain an abnormal detection result, where the weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected;

[0034] The abnormal detection model is trained based on the log features of sample logs and their abnormal labels.

[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps of any one of the above-mentioned abnormal log detection methods are implemented.

[0036] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned abnormal log detection methods are implemented.

[0037] The abnormal log detection method, device, electronic device, and storage medium provided by the present invention determine the weight of weighted fusion by using the number of domain keywords included in the log to be detected, so as to fuse the detection results obtained by performing abnormal detection under multiple different algorithms, and realize abnormal detection with multi-algorithm fusion, thereby ensuring the accuracy and reliability of abnormal detection. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 is one of the schematic flowcharts of the abnormal log detection method provided by the present invention;

[0040] Figure 2 is the schematic flowchart of the weight determination method provided by the present invention;

[0041] Figure 3 is the schematic flowchart of the abnormal detection model construction method provided by the present invention;

[0042] Figure 4 is the schematic structural diagram of the abnormal detection model provided by the present invention;

[0043] Figure 5 is the schematic structural diagram of the abnormal log detection device provided by the present invention;

[0044] Figure 6 is the schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0046] Regarding the log anomaly detection work in the network operation and maintenance system, it is impossible to accurately identify abnormal logs based on human experience, and the method based on artificial rules often results in misjudgments, leading to other problems. In addition, for the judgment system based on artificial rules, a large amount of human and material maintenance costs need to be invested. By using machine learning algorithms to implement abnormal logs, not only can the accuracy be improved, but also a wide range of operation and maintenance scenarios can be covered, thereby saving a large amount of human and material costs.

[0047] The improvement of the accuracy of abnormal log recognition also has an obvious improvement effect on the operation and maintenance work efficiency and can save a large amount of human and material resources for the enterprise. However, due to the complex and changeable log data, the accuracy of traditional single machine learning algorithms is relatively low. Therefore, the embodiments of the present invention provide an abnormal log detection method.

[0048] Figure 1 is one of the flow schematic diagrams of the abnormal log detection method provided by the present invention. As Figure 1 shown, the method includes:

[0049] Step 110, determining the log features of the log to be detected.

[0050] Specifically, the log to be detected is the log that needs to be detected for abnormal logs. The log to be detected can be the log directly pulled from the network operation and maintenance system or the log after data preprocessing. The embodiments of the present invention do not make specific limitations on this.

[0051] The logs referred to here can include the standard output and text logs of the network operation and maintenance system. The standard output includes the information output through STDOUT and STDERR, including the text content redirected to the standard output, etc. The text logs include the internal log files that are not redirected to the standard output. The above logs may be mounted through the host directory or mounted to the remote storage center.

[0052] The log feature is the vector representation of the corresponding log. The log feature is used to characterize the features of the corresponding log. For example, the log feature may include the word vectors of each word segment in the corresponding log, which can reflect whether there are keywords in the corresponding log and the number of keywords. The embodiments of the present invention do not make specific limitations on this.

[0053] Step 120, based on the anomaly detection model, performing anomaly detection on the log features under multiple different algorithms, and performing weighted fusion on the detection results under multiple different algorithms to obtain the anomaly detection result. The weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected;

[0054] The anomaly detection model is trained based on the log features of the sample logs and their anomaly labels.

[0055] The anomaly detection model here integrates multiple anomaly detection algorithms, so it can make up for the problem of unsatisfactory accuracy when using a single algorithm. By detecting the log features of the log to be detected based on multiple different anomaly detection algorithms, multiple detection results can be obtained, and each detection result corresponds to an anomaly detection algorithm. On this basis, the anomaly detection model also needs to summarize and fuse the detection results under each anomaly detection algorithm to obtain the final anomaly detection result.

[0056] Further, when the anomaly detection model aggregates and fuses multiple detection results, it can adopt the method of weighted summation, and the weights of the weighted summation can be determined according to the number of domain keywords included in the log to be detected. The domain keywords referred to here are the log keywords in the pre-set network operation and maintenance field, such as "fault", "Error", "peak", "exceed", etc.

[0057] Considering that different anomaly detection algorithms have their own focuses. For example, the regularization feature in the logistic regression algorithm has high computational efficiency for large matrices, and the catboost algorithm has high accuracy for data with categorical features and continuous features. It is possible to judge which specific algorithm to use for anomaly detection may have higher accuracy according to whether there are prominent features in the log features of the input log to be detected.

[0058] At this time, the number of domain keywords included in the log to be detected can be used as an index to judge the prominence degree of the features reflected by the log features. The higher the number of domain keywords included in the log to be detected, the more prominent the features reflected by the log features. The accuracy of using an anomaly detection algorithm biased towards categorical features may be higher than that of using an anomaly detection algorithm biased towards large matrix calculations. Correspondingly, the weight of the detection result obtained by the anomaly detection algorithm biased towards categorical features can be increased, so as to ensure the reliability of the finally output anomaly detection result.

[0059] Before executing step 120, an anomaly detection model can also be pre-trained. The training of the anomaly detection model can be achieved through the following steps: First, collect a large number of sample logs, determine the log features of each sample log, and obtain the anomaly labels of the sample logs through manual annotation. On this basis, train the initial model containing multiple anomaly detection algorithms based on the log features and anomaly labels of the sample logs, and use the trained initial model as the anomaly detection model.

[0060] The method provided by the embodiment of the present invention determines the weights of weighted fusion by applying the number of domain keywords included in the log to be detected, so as to fuse the detection results obtained by anomaly detection under multiple different algorithms, realize anomaly detection with multi-algorithm fusion, and thus ensure the accuracy and reliability of anomaly detection.

[0061] Based on the above embodiment, Figure 2 is a flowchart of the weight determination method provided by the present invention. As Figure 2 shown, the weights of the weighted fusion are determined based on the following steps:

[0062] Step 210: Determine the algorithm weights of each algorithm based on the total number of log word segments included in the sample log, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected.

[0063] Specifically, the total number of log word segments included in the sample log can be understood as the total number of word segments that can be recognized during the abnormal log detection process, and the preset domain keywords are included in the log word segments. By combining the total number of log word segments included in the sample log, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected, the prominence degree of the features represented by the log features of the log to be detected can be measured, and then the algorithm weights of each algorithm, that is, the adaptation degree of each algorithm applied to the abnormal detection of the log to be detected, can be determined.

[0064] Step 220: Determine the weights of the weighted fusion based on the algorithm weights of each algorithm and the accuracy rates of abnormal detection under each algorithm.

[0065] Specifically, when determining the weights of the weighted fusion of the detection results under each algorithm, not only the adaptation degree of each algorithm applied to the abnormal detection of the log to be detected needs to be considered, but also the accuracy rate of each algorithm itself when applied to abnormal detection needs to be considered, and these two can be comprehensively measured. For example, the product of the algorithm weight of any algorithm and the accuracy rate of abnormal detection under this algorithm can be used as the weight of the weighted fusion of the detection results under this algorithm.

[0066] Furthermore, the accuracy rates of abnormal detection under each algorithm can be obtained by testing each abnormal detection sub-model after completing the training of the abnormal detection sub-models corresponding to each algorithm.

[0067] Based on any of the above embodiments, the multiple different algorithms include the logistic regression algorithm and the CatBoost algorithm.

[0068] Specifically, the CatBoost algorithm has a high accuracy rate for data with categorical features and continuous features. Therefore, in the case where the log features of the log to be detected contain more features of domain keywords, the precision priority of the CatBoost algorithm based on this part of the matrix of the log features can be correspondingly increased.

[0069] Correspondingly, step 210 specifically includes:

[0070] Determine the algorithm weights of each algorithm based on the following formula:

[0071]

[0072]

[0073] In the formula, μ 逻辑回归 and μcatboost They are the algorithm weights of the logistic regression algorithm and the CatBoost algorithm respectively. N = n + k, where n is the total number of log word segments in the sample log, k is the total number of preset domain keywords, p is the number of domain keywords in the log to be detected, and 0 < p < k. Here, when calculating μ catboost , the weight of the detection result under the CatBoost algorithm can be increased by 2 * p to avoid the influence caused by the weights of the other algorithms dragging down the overall prediction result.

[0074] Based on any of the above embodiments, in step 120, the weighted fusion of the detection results under multiple different algorithms to obtain the anomaly detection result includes:

[0075] Determine the anomaly detection result based on the following formula:

[0076]

[0077] where S1 is the classification accuracy of the logistic regression algorithm, and S2 is the classification accuracy of the CatBoost algorithm.

[0078] For example, the classification results of anomaly detection of sample data using the logistic regression and CatBoost algorithms respectively are as follows:

[0079]

[0080] Based on the results in the above table, the classification accuracies S1 and S2 corresponding to the logistic regression and CatBoost algorithms can be calculated (taking the data in the above table as an example, S1 = S2 = 0.57).

[0081] When there are 5 log data (m = 5), n = 50, k = 30, N = 80, S1 = S2 = 0.57, and the number p of domain keywords in each log data is: 25, 10, 0, 5, 0, the corresponding calculation results and prediction results are shown in the following table:

[0082]

[0083] Applying the supervised learning method for anomaly detection, an obvious disadvantage is that it requires manual annotation of abnormal logs in advance. In the actual operation and maintenance scenarios of enterprises, operation and maintenance engineers need to face a large number of machines every day, and it is very difficult to batch annotate the logs generated by the machines. And in the actual annotation process, the annotation results are greatly affected by the experience of on-site actual operation and maintenance personnel and the on-site environment. Colleagues with richer work experience have a lower risk of missing abnormal logs generated by machines, while colleagues with less work experience are not sure about abnormal logs and have the risk of missing and misjudging, and cannot achieve an ideal judgment effect.

[0084] Regarding this problem, based on any of the above embodiments, the abnormal labels are determined based on the following steps:

[0085] Determine the log features of each sample log;

[0086] Cluster the log features of each sample log to obtain multiple log clusters;

[0087] Obtain the manual labels of each log cluster, and use the manual labels as the abnormal labels of the sample logs corresponding to each log feature within the corresponding log cluster.

[0088] Specifically, after collecting a large number of sample logs, the log features of each sample log can be determined respectively. On this basis, unsupervised clustering can be performed on the log features of the sample logs to obtain multiple log clusters. Each log cluster can include multiple similar log features. By clustering a large number of sample logs, the complexity of log discrimination can be reduced. Subsequently, combined with human experience, one or more clusters of the clustering results can be determined as abnormal, thereby realizing batch annotation of sample logs and greatly reducing the workload of manual labeling.

[0089] Further, clustering the log features of each sample log can be implemented by the K-means algorithm. The specific implementation steps are as follows:

[0090] Assume that the log features of each sample log are represented as {x (1) ,..., x (m)}, and each log feature x (i) ∈ R n . k clustering centroids μ1, μ2,..., μ k ∈ R n can be randomly selected from them.

[0091] Subsequently, repeat the following process until convergence:

[0092] {For each example i, calculate the class it should belong to

[0093]

[0094] For each class j, recalculate the centroid of this class:

[0095]

[0096] The number of basic logs is m, and after clustering, n categories are generated, that is, n log clusters, where n << m. Usually, n is not greater than 100. Based on this process, by clustering a large number of logs, the complexity of log discrimination is reduced, and the subsequent workload of manual labeling is greatly reduced.

[0097] Based on any of the above embodiments, step 110 includes:

[0098] Perform text tokenization on the log to be detected to obtain each token of the log to be detected;

[0099] Perform domain keyword matching on each token of the log to be detected, and determine a keyword vector based on the matching result;

[0100] Construct the log feature of the log to be detected based on the word vectors of each token of the log to be detected and the keyword vector.

[0101] Specifically, in the process of generating the log feature, by performing domain keyword matching on each token in the log to be detected, determine the domain keywords included in the log to be detected as the matching result, thereby construct a keyword vector based on the situation of the domain keywords included in the log to be detected, and when constructing the log vector, add the keyword vector.

[0102] In the keyword vector referred to here, "0" and "1" can be used to indicate whether the corresponding domain keyword exists in the log to be detected. Assume that there are 30 domain keywords preset in total, then the keyword vector can be a vector with a length of 30, where each bit corresponds to the presence or absence of a domain keyword.

[0103] Based on any of the above embodiments, the word vectors of each token are determined based on the word frequency and inverse document frequency index of each token.

[0104] Specifically, for a given document, the term frequency (TF) refers to the frequency of a given term in the document. Further, the term frequency is a normalization of the term count to prevent it from biasing towards long documents. For the term t i in a specific document, its term frequency can be expressed as:

[0105]

[0106] In the formula, n i,j is the number of occurrences of the term t i in the document d j , and ∑kn k,j is the sum of the number of occurrences of all terms in the document d j .

[0107] The inverse document frequency (IDF) is a measure of the general importance of a term. The IDF of a specific term can be obtained by dividing the total number of documents by the number of documents containing the term and then taking the logarithm of the resulting quotient:

[0108]

[0109] Among them, |D| is the total number of files in the corpus, and |{j:t i ∈d j}| is the number of files containing the word t i (i.e., the number of files where n i,j ≠0). If the word is not in the corpus, it will cause the dividend to be zero. Therefore, generally, 1 + |{j:t i ∈d j}| is used.

[0110] Corresponding to the embodiments of the present invention, the log to be detected can be regarded as a given file, so as to calculate the word frequency of each word segment therein, and a corpus is constructed based on all logs including the log to be detected and all sample logs to calculate the inverse document frequency index of each word segment. On this basis, the word vector of each word segment can be obtained through the following formula:

[0111] tfidf i,j = tf i,j ×idf i

[0112] In the formula, tfidf i,j , that is, the TFIDF weight of each word segment, can be directly used as the word vector of each word segment.

[0113] For example, the following table shows 3 logs:

[0114]

[0115] After performing text word segmentation on the above 3 logs respectively, the word segmentation results of the 3 logs are obtained:

[0116] On this basis, calculate the TFIDF weights of each word segment therein respectively, so as to obtain the corresponding word vector matrix:

[0117]

[0118]

[0119] On this basis, combined with the keyword vector, the log features of the 3 logs can be obtained:

[0120] Log ID ...... Dedicated For Among which Background Processing Platform ...... Fault Error ...... Peak 1 ...... 0.2049 0.2049 0.2049 0.2049 0.2049 0.2049 ...... 1 0 ...... 1 2 ...... 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 ...... 0 0 ...... 1 3 ...... 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 ...... 0 1 ...... 0

[0121] Based on any of the above embodiments, in step 110, before performing text word segmentation on the log to be detected, it further includes:

[0122] Extract abnormal symbols and Chinese information from the log to be detected;

[0123] Based on the log text after removing abnormal symbols and the extracted Chinese information, reconstruct the log to be detected.

[0124] Specifically, removing abnormal symbols can filter out abnormal symbols in the log. For example, it can be filtered by regular expressions. The regular expression is as follows:

[0125] [^A-Za-z0-9]

[0126] The log text after removing abnormal symbols only includes English uppercase and lowercase letters and numbers.

[0127] Considering that a large amount of useful information in the log exists in Chinese, in the embodiments of the present invention, Chinese information is extracted to retain the Chinese in the log for subsequent combination with other information. For example, Chinese information can be extracted through regular expressions. The regular expression is as follows:

[0128] [u4e00-u9fa5]

[0129] On this basis, the log text after removing abnormal symbols and the extracted Chinese information can be recombined as the log to be detected after data cleaning.

[0130] Based on any of the above embodiments, Figure 3 is a schematic flowchart of the method for constructing an anomaly detection model provided by the present invention. As Figure 3 shown, the method includes the following steps:

[0131] S1. Sample log collection:

[0132] The logs output by the network operation and maintenance system can be uniformly processed by the logging worker. Different logging workers will write the standard output to different destinations. Here, by default, the logging worker will write the logs to the host file in json format.

[0133] Further, the collected sample logs can be divided into the following three types:

[0134] One is the node log file disk write, which can be mounted to the host directory in the form of volumeMounts. For this log information, it can be directly collected by the Fluentd agent in the form of DanemonSet.

[0135] Another is the log file dumped to the remote log storage center, which can be mounted in the form of volumeMounts.

[0136] Another type is the non-persistent log file, which can be obtained by collecting this type of log information. It is necessary to immerse into the specific internal part, locate the log storage path and then collect it. The collection scheme is: run Fluentd in the Sidecar mode in the corresponding internal part, which has strong invasiveness.

[0137] S2. Sample log data processing:

[0138] The data processing here can be specifically implemented by ElasticSearch. The processing steps specifically include data cleaning, adjusting the log format, and processing the log format into the form of index plus content.

[0139] S3. Sample log feature extraction:

[0140] First, abnormal symbols and Chinese information can be extracted from each sample log, and each sample log can be reconstructed based on the log text after removing abnormal symbols and the Chinese information obtained by extraction.

[0141] On this basis, use the jieba word segmentation tool to divide each sample log into phrase forms and use " / " as the delimiter. Calculate the TFIDF weights of each word segment in each sample log as the word vectors of each word segment, and combine the keyword vectors determined by the results of domain keyword matching for each word segment of each sample log to construct the log features of each sample log.

[0142] S4. Sample log clustering:

[0143] Cluster the log features of each sample log to obtain multiple log clusters.

[0144] S5. Batch labeling of sample logs:

[0145] Combined with human experience, one or more clusters of the clustering results can be determined as abnormal, thereby realizing the batch annotation of sample logs.

[0146] S6. Multi-algorithm training and combination:

[0147] Under normal circumstances, assuming that the error rates of the base classifiers are independent of each other, then according to the Hoeffding inequality, the error rate of integrating multiple classifiers is:

[0148]

[0149] Among them, h i are the respective base classifiers, and the error rate of the base classifier is ε.

[0150] The above formula shows that as the number T of sub-classifiers in the ensemble algorithm increases, the error rate of the ensemble will decrease exponentially and eventually tend to zero.

[0151] Based on the above theory, supervised algorithms such as logistic regression and CatBoost are used to perform weighted fusion on the above algorithms:

[0152]

[0153] Among them, F j (X) is the j-th log combination classifier, is the weight of the i-th algorithm for the j-th log; is the i-th sub-classifier for the j-th log, and S i is the accuracy of each sub-classifier. In the embodiments of the present invention, K = 2, which respectively represent logistic regression and CatBoost classifiers, and X ∈ R n .

[0154] Using the above weighted fusion method to combine each sub-model such as logistic regression and CatBoost together can improve the algorithmic ability.

[0155] Figure 4 is a schematic structural diagram of the anomaly detection model provided by the present invention. As Figure 4 shown, the anomaly detection model may include logistic regression and CatBoost, as well as a combination module F(X) of the detection results of the two. Figure 4 The new logs in

[0156] represent the log features of the logs to be detected. The label output after passing through F(X) is the anomaly detection result of the logs to be detected.

[0157]

[0158]

[0159]

[0160] In the formula, TP, that is, True Positive, represents the number of samples that are actually 0 and predicted to be 0; FN, that is, False Negative, represents the number of samples that are actually 0 and predicted to be 1; FP, that is, False Positive, represents the number of samples that are actually 1 and predicted to be 0.

[0161] The model effects of the anomaly detection model when using logistic regression and catboost alone, as well as when fusing logistic regression and catboost with weights, are shown in the following table:

[0162] Algorithm Name Precision Recall F1-score Logistic Regression 0.92 0.93 0.92 Catboost 0.93 0.95 0.93 Weighted Fusion 0.96 0.95 0.95

[0163] It can be seen that the anomaly detection model based on the weighted fusion scheme has better model effects.

[0164] Based on any of the above embodiments, Figure 5 is a schematic structural diagram of the anomaly log detection device provided by the present invention, as Figure 5 shown, the device includes:

[0165] A feature extraction unit 510, configured to determine the log features of the log to be detected;

[0166] An anomaly detection unit 520, configured to perform anomaly detection on the log features under multiple different algorithms based on an anomaly detection model, and perform weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weights of the weighted fusion are determined based on the number of domain keywords included in the log to be detected;

[0167] The anomaly detection model is trained based on the log features of the sample log and its anomaly labels.

[0168] The device provided by the embodiments of the present invention determines the weights of the weighted fusion by using the number of domain keywords included in the log to be detected, so as to fuse the detection results obtained by performing anomaly detection under multiple different algorithms, realize anomaly detection with multi-algorithm fusion, and thus ensure the accuracy and reliability of anomaly detection.

[0169] Based on any of the above embodiments, the device further includes a weight determination unit, configured to:

[0170] Determine the algorithm weights of each algorithm based on the total number of log word segments included in the sample log, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected;

[0171] Determine the weights of the weighted fusion based on the algorithm weights of each algorithm and the accuracy rates of anomaly detection under each algorithm.

[0172] Based on any of the above embodiments, the multiple different algorithms include a logistic regression algorithm and a catboost algorithm;

[0173] The weight determination unit is specifically configured to:

[0174] Determine the algorithm weights of each algorithm based on the following formula:

[0175]

[0176]

[0177] Wherein, μ 逻辑回归 and μ catboost are the algorithm weights of the logistic regression algorithm and the catboost algorithm respectively, N = n + k, n is the total number of log segmentations included in the sample log, k is the total number of preset domain keywords, and p is the number of domain keywords included in the log to be detected.

[0178] Based on any of the above embodiments, the device further includes a labeling unit for:

[0179] Determine the log features of each sample log;

[0180] Cluster the log features of each sample log to obtain multiple log clusters;

[0181] Obtain the manual labels of each log cluster, and use the manual labels as the anomaly labels of the sample logs corresponding to each log feature within the corresponding log cluster.

[0182] Based on any of the above embodiments, the feature extraction unit 510 is used for:

[0183] Perform text segmentation on the log to be detected to obtain each segmentation of the log to be detected;

[0184] Perform domain keyword matching on each segmentation of the log to be detected, and determine the keyword vector based on the matching result;

[0185] Based on the word vectors of each segmentation of the log to be detected and the keyword vector, construct the log feature of the log to be detected.

[0186] Based on any of the above embodiments, the word vectors of each segmentation are determined based on the word frequency and the inverse document frequency index.

[0187] Based on any of the above embodiments, the feature extraction unit 510 is further used for:

[0188] Remove abnormal symbols and extract Chinese information from the log to be detected;

[0189] Based on the log text after removing abnormal symbols and the extracted Chinese information, reconstruct the log to be detected.

[0190] Figure 6 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute an exception log detection method, which includes: determining the log features of the log to be detected; based on an anomaly detection model, performing anomaly detection on the log features under multiple different algorithms, and performing weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected; the anomaly detection model is trained based on the log features of sample logs and their anomaly labels.

[0191] In addition, when the logic instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0192] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the anomaly log detection method provided by the above-mentioned various methods. The method includes: determining the log features of the log to be detected; based on an anomaly detection model, performing anomaly detection on the log features under multiple different algorithms, and performing weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected; the anomaly detection model is trained based on the log features of sample logs and their anomaly labels.

[0193] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-provided exception log detection method, which includes: determining the log features of the log to be detected; based on an anomaly detection model, performing anomaly detection on the log features under multiple different algorithms, and performing weighted fusion on the detection results under multiple different algorithms to obtain an anomaly detection result, where the weights of the weighted fusion are determined based on the number of domain keywords included in the log to be detected; and the anomaly detection model is trained based on the log features of sample logs and their anomaly labels.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. An abnormal log detection method, characterized in that, Including: Determine the log features of the log to be detected; Based on the anomaly detection model, perform anomaly detection on the log features under multiple different algorithms, and perform weighted fusion on the detection results under multiple different algorithms to obtain the anomaly detection result. The weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected; The anomaly detection model is trained based on the log features of the sample logs and their anomaly labels; The weight of the weighted fusion is determined based on the following steps: Based on the total number of log word segments included in the sample logs, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected, determine the algorithm weights of each algorithm; Based on the algorithm weights of each algorithm and the accuracy of anomaly detection under each algorithm, determine the weight of the weighted fusion.

2. The abnormal log detection method according to claim 1, wherein The multiple different algorithms include the logistic regression algorithm and the catboost algorithm; The determining the algorithm weights of each algorithm based on the total number of log word segments included in the sample logs, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected includes: Determine the algorithm weights of each algorithm based on the following formula: Where, μ 逻辑回归 and μ catboost are the algorithm weights of the logistic regression algorithm and the catboost algorithm respectively, N = n + k, n is the total number of log word segments included in the sample log, k is the total number of preset domain keywords, and p is the number of domain keywords included in the log to be detected.

3. The abnormal log detection method according to claim 1, characterized in that The anomaly label is determined based on the following steps: Determine the log features of each sample log; Cluster the log features of each sample log to obtain multiple log clusters; Obtain the manual labels of each log cluster, and use the manual labels as the anomaly labels of the sample logs corresponding to the log features within the corresponding log cluster.

4. The abnormal log detection method according to any one of claims 1 to 3, characterized in that The determining the log features of the log to be detected includes: Perform text word segmentation on the log to be detected to obtain each word segment of the log to be detected; Perform domain keyword matching on each word segment of the log to be detected, and determine the keyword vector based on the matching result; Based on the word vectors of each word segment of the log to be detected and the keyword vector, construct the log features of the log to be detected.

5. The abnormal log detection method according to claim 4, wherein The word vectors of each word segment are determined based on the word frequency and the inverse document frequency index.

6. The abnormal log detection method according to claim 4, wherein Before performing the text word segmentation on the log to be detected, it further includes: Perform abnormal symbol removal and Chinese information extraction on the log to be detected; Based on the log text after abnormal symbol removal and the extracted Chinese information, reconstruct the log to be detected.

7. An abnormal log detection device, characterized in that Including: A feature extraction unit for determining the log features of the log to be detected; An anomaly detection unit for performing anomaly detection on the log features under multiple different algorithms based on the anomaly detection model, and performing weighted fusion on the detection results under multiple different algorithms to obtain the anomaly detection result. The weight of the weighted fusion is determined based on the number of domain keywords included in the log to be detected; The anomaly detection model is trained based on the log features of the sample logs and their anomaly labels; Wherein, the device further includes a weight determination unit for: Based on the total number of log word segments included in the sample logs, the total number of preset domain keywords, and the number of domain keywords included in the log to be detected, determine the algorithm weights of each algorithm; Based on the algorithm weights of each algorithm and the accuracy of anomaly detection under each algorithm, determine the weight of the weighted fusion.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the abnormal log detection method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the abnormal log detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Log anomaly detection method and device based on deep learning, terminal and medium

    CN110347547A

  • Webpage log attack information detection method, system and device and readable storage medium

    CN110830483A