An abnormal location determination method, device, electronic device, medium and program product

By building a collection of row templates and an unsupervised detection model, the problem of inconsistent log forms in cloud computing is solved, and the automated abnormal positioning of network cloud logs is realized, which improves the efficiency and accuracy of abnormal detection.

CN116192619BActive Publication Date: 2025-07-11ASIAINFO TECH CHINA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310035888.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-11
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

In the field of cloud computing, the log formats generated by network cloud devices are not unified, resulting in the inability to effectively perform preprocessing and abnormal analysis, and rely heavily on expert experience and cannot achieve fast and effective abnormal positioning.

Method used

By extracting log templates from the original log to build a row template collection, using an unsupervised detection model for exception detection, determining the exception row template based on the matching relationship between keywords and common fields, and positioning the target log rows in the original log to achieve log analysis and exception positioning with inconsistent forms.

Benefits of technology

Preprocessing and de-expert exception detection of inconsistent network cloud logs is realized, and target log rows carrying exception information are quickly positioned, improving the automation and efficiency of exception positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192619B_ABST
    Figure CN116192619B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an anomaly localization method, apparatus, device, medium, and program product, which relates to the field of cloud computing. The method includes: after obtaining the original log, constructing a set of line templates according to the original log, where each line template in the set of line templates is in the same order as the corresponding log line in the original log; performing anomaly detection on the set of line templates through an unsupervised detection model, and in the case where the detection result is abnormal, determining an abnormal line template through the matching relationship between the keywords characterizing the anomaly and the common fields carried by the line template; finally, locating the target log line corresponding to the abnormal line template in the original log. When the original log is a network cloud log, the solution shown in the embodiment of the present application realizes de-expertized anomaly localization for network cloud logs with inconsistent formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology. Specifically, this application relates to an abnormal location method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] On the one hand, in the field of cloud computing, general computing, storage, and network hardware devices are decomposed into various virtual resources (i.e., network cloudified hardware devices, simply referred to as network cloud devices), and users can obtain various network cloud devices according to the needs of applications to achieve cloud computing. Therefore, network cloud operation and maintenance for network cloud devices emerged as the times require.

[0003] On the other hand, the main direction in the operation and maintenance method for devices is abnormal detection and root cause analysis for logs. Currently, there are several related technologies for log abnormal detection and root cause analysis, such as methods based on machine learning, methods based on association mining, and methods based on statistics.

[0004] Although the above methods can achieve abnormal detection and root cause analysis of logs, there are still the following problems:

[0005] (1) Heavily rely on expert experience. For example, it is necessary for operation and maintenance experts to analyze log data and output certain knowledge based on experience, and then perform log analysis on the premise of the output knowledge.

[0006] (2) Unable to be applied to the analysis of distributed logs. Existing solutions mainly target single logs or multi-dimensional logs first. Since the forms of single logs or multi-dimensional logs are relatively fixed, they can be quickly preprocessed and then log analysis can be performed. However, the network cloud logs generated by each network cloud device have the problem of inconsistent log forms and cannot be effectively preprocessed, and it is even more difficult to carry out log analysis.

[0007] Therefore, the abnormal analysis of network cloud logs has become a technical problem that urgently needs to be solved. Summary of the Invention

[0008] The purpose of the embodiments of this application is to provide an abnormal location method, device, equipment, medium, and program product, aiming to solve one of the above technical problems. To achieve this purpose, the embodiments of this application provide the following several technical solutions.

[0009] On the one hand, the embodiments of this application provide an abnormal location method, and this method includes:

[0010] Construct a set of line templates from the log templates extracted from the original log. Each line template in the set of line templates is in the same order as the corresponding log line in the original log. Perform anomaly detection on the set of line templates based on an unsupervised detection model. If the detection result indicates an anomaly, obtain the keywords used to represent the anomaly, and determine the abnormal line template according to the matching relationship between the common fields carried by the line template and the keywords. Locate the target log line corresponding to the abnormal line template in the original log.

[0011] Optionally, the number of unsupervised detection models is at least two. Performing anomaly detection on the set of line templates based on an unsupervised detection model includes:

[0012] Input the set of line templates into each unsupervised detection model in turn, and obtain the results output by each unsupervised detection model. Determine the corresponding weight of each unsupervised detection model. Determine the detection result according to the respective weights and the output results of at least two unsupervised detection models.

[0013] Optionally, determining the abnormal line template according to the matching relationship between the common fields carried by the line template and the keywords includes:

[0014] Determine the matching relationship between the common fields carried by the line template and at least one keyword. If the matching relationship is that the common fields carried by the line template include at least one keyword, determine the corresponding line template as the abnormal line template.

[0015] Optionally, obtaining the keywords used to represent the anomaly includes at least one of the following:

[0016] Extract key fields from the set of line templates based on a preset algorithm, and determine the key fields as the keywords used to represent the anomaly. Obtain historical keywords and determine the historical keywords as the keywords used to represent the anomaly.

[0017] Optionally, before locating the target log line corresponding to the abnormal line template in the original log, it further includes:

[0018] For each abnormal line template, determine the similarity between the abnormal line template and other line templates in the set of line templates except for all the abnormal line templates. If it is determined that the similarity between any other line template and the abnormal line template is greater than a preset threshold, then use any other line template as a new abnormal line template.

[0019] Optionally, determining the similarity between the abnormal line template and other line templates in the set of line templates except for all the abnormal line templates includes:

[0020] Determine the similarity between the common fields of the abnormal line template and the common fields of other line templates as the similarity between the abnormal line template and other line templates.

[0021] Optionally, locate the target log line corresponding to the exception line template in the original log, including:

[0022] Determine at least one order of the exception line template in the line template set; locate the target log line corresponding to each order in the original log.

[0023] On the other hand, an embodiment of the present application provides an exception location device, which includes:

[0024] A construction module, configured to construct a line template set through log templates extracted from the original log, and each line template in the line template set is in the same order as the corresponding log line in the original log.

[0025] A detection module, configured to perform exception detection on the line template set based on an unsupervised detection model.

[0026] A determination module, configured to, if the detection result is an exception, obtain keywords for characterizing the exception, and determine the exception line template according to the matching relationship between the common fields carried by the line template and the keywords.

[0027] A location module, configured to locate the target log line corresponding to the exception line template in the original log.

[0028] On yet another aspect, an embodiment of the present application provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of an exception location method provided by an embodiment of the present application.

[0029] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of an exception location method provided by an embodiment of the present application.

[0030] An embodiment of the present application further provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the steps of an exception location method provided by an embodiment of the present application.

[0031] The beneficial effects brought by the technical solution provided by the embodiment of the present application are:

[0032] An embodiment of the present application provides an anomaly location method, which is applied to log analysis and anomaly location of network cloud logs with inconsistent formats. Specifically, after obtaining the original logs, a set of line templates is constructed according to the line templates extracted from the original logs, and each line template in the set of line templates is in the same order as the corresponding log line in the original logs; the set of line templates is detected for anomalies through an unsupervised detection model, and when the detection result indicates an anomaly, the anomalous line template is determined through the matching relationship between the keywords indicating the anomaly and the common fields carried by the line template; finally, the target log line corresponding to the anomalous line template is located in the original logs, and the target log line is the log line carrying the anomaly information. This method constructs a set of line templates from the line templates extracted from the original logs, realizing the preprocessing of the original logs with inconsistent formats; then, anomaly detection is performed through an unsupervised model, achieving the purpose of de-expertization detection; finally, the target log line corresponding to the anomalous log line is located in the original logs, that is, by locating the target log line carrying the anomaly information, the purpose of fast anomaly location is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0034] Figure 1 It is a schematic flow chart of an anomaly location method provided by an embodiment of the present application;

[0035] Figure 2 It is a schematic diagram of the scenario of anomaly location provided by an embodiment of the present application;

[0036] Figure 3 It is a schematic structural diagram of an anomaly location device provided by an embodiment of the present application;

[0037] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The following describes the embodiments of the present application with reference to the drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0039] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the technical field of the present invention, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0040] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0041] Continuing with the background technology, in the field of cloud computing, any hardware device can be incorporated into the network as long as it can be used as a virtual resource and be called as a network cloud device. And when needed, an application can call these virtual resources to use cloud computing services. Then, there is the following scenario: when using cloud computing services at the current moment, this batch of "network cloud devices" may be called, and at the next moment, the next batch of "hardware devices" may be called. Therefore, the forms of the network cloud logs generated by the hardware devices called by the application are not unified.

[0042] The processing processes of the methods based on machine learning, the methods based on association mining, and the methods regarding statistics existing in the prior art can be referred to the following content.

[0043] Among them, the processing process of the method based on machine learning: By collecting the original log data, screening the first abnormal range of the original log data according to the first log template (a template created according to expert experience), and obtaining the second log based on the first abnormal range; performing sliding window segmentation on the second log to determine the abnormal points of the log; or trimming the original log data to obtain the second log data containing error information, and extracting key fields (including time, error log summary, etc.) for the second log data, then performing clustering generalization on the second log data, constructing a backtracking tree model in a predetermined order, and combining the expert experience knowledge graph (the expert knowledge graph has covered more than 90% of the known error types) to locate the abnormal information.

[0044] Among them, the processing process of the algorithm based on association mining includes: by using the association mining algorithm, instantaneously performing association rule mining on the newly added original log data in real time to obtain infrequent items, frequent items, and association rules; performing association analysis on the log data in combination with multiple metrics of the information system to quickly locate system failures.

[0045] Among them, the processing process of the method based on statistics includes: performing event sequence regression analysis on all leaf (attribute KPI) dimension combinations of the original multi-dimensional log to obtain the actual value and predicted value of each leaf dimension combination, and further obtaining the deviation score through the actual value and predicted value; clustering all leaf dimension combinations in the multi-dimensional log through the deviation score to obtain all classes affected by the same root cause; finally, performing heuristic root cause search on each class respectively.

[0046] The above-mentioned method based on machine learning is obviously quite dependent on expert experience. The algorithm based on association mining is for the anomaly analysis of a single log, and the method based on statistics is for the anomaly analysis of multi-dimensional logs. That is to say, the solutions shown in the prior art either rely quite on expert experience or are for single logs or multi-dimensional logs. In the face of the massive and non-uniformly formatted logs generated by network cloud devices, the above three solutions cannot achieve fast and effective anomaly location.

[0047] The embodiment of the present application provides an anomaly location method, which can solve the above technical problems. Among them, this method is applied to a terminal device that can store a large amount of logs. Among them, this method preprocesses the original logs by constructing a set of row templates from the log templates extracted from the original logs; then, based on an unsupervised detection model, anomaly detection is performed on the set of row templates, realizing a de-expertized anomaly detection process; in the case where the detection result is abnormal, by matching the keywords characterizing the anomaly with the common fields of the row templates, the abnormal row templates are screened out, and the target log rows corresponding to the abnormal row templates are located in the original logs. By locating the target log rows carrying abnormal information, the final anomaly location is realized.

[0048] Optionally, the method provided by the embodiment of the present application can be implemented as an independent application program or a functional module / plugin of an application program. For example, this application program can be a dedicated anomaly location application program or an application program with anomaly location function, and through this application program, the purpose of anomaly location can be achieved.

[0049] Optionally, the original log is a log obtained from various network cloud devices, which may be a cloud server, an intelligent gateway, a cloud storage device, a cloud speaker, etc. Optionally, the original log may specifically be any of the following types of files: a file with a.csv suffix, a text with a.txt suffix, a file with a.doc / .ppt / .xls, etc. suffix. It should be noted that the embodiments of the present application do not limit the type of the original log.

[0050] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments may refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0051] Figure 1 A flowchart of an exception location method is shown. Among them, the method includes steps S110 to S140.

[0052] S110, construct a set of line templates through the log templates extracted from the original log, and the order of each line template in the set of line templates is the same as the order of the corresponding log lines in the original log.

[0053] Optionally, before executing S110, obtain the stored original log from at least one network cloud device. Among them, the original log may be a log file or multiple log files.

[0054] Among them, before constructing a set of line templates for the original log, extract all the line templates in the original log. Among them, after extracting the line templates, assign a template identifier to each line template according to the order of the extracted line templates. Each line template consists of common information and specific information identified by wildcards. Among them, when performing the extraction action,

[0055] Optionally, after receiving the original log, add a line identifier to each log line according to the order of each log line in the original log. The line identifier is unique and can identify the log file where the corresponding log line is located and the arrangement order in the log file.

[0056] To understand the above process more clearly, the embodiments of the present application also provide an example of the original log. In this example, the original log includes N log lines, and the identifiers of each log line are log line 1, log line 2... log line N in sequence.

[0057] Original log:

[0058] connected to 10.0.0.1. / / Log line 1

[0059] Connected to 192.168.0.1. / / Log line 2

[0060] Hex number 0xDEADBEAF. / / Log line 3

[0061] User davidoh logged in. / / Log line 4

[0062] User eranr logged in. / / Log line 5

[0063] ……

[0064] Line templates extracted from the original log include:

[0065] ID = 1 : size = 2 : connected to <*>[

[0066] ID = 2 : size = 1 : Hex number <*>[

[0067] ID = 3: size = 2: user <*> logged in

[0068] ……

[0069] In the line templates shown in this example, the ID is the identifier of each line template; the size is the total number of log lines corresponding to the line template in the original log.

[0070] In this example, for the line template with ID = 1, "connected to" is the common information, and <*> is a wildcard. Among them, for log line 1, the value at "<*>" is "10.0.0.1", and for log line 2, the value at "<*>" is "192.168.0.1".

[0071] Optionally, by constructing a set of line templates from the log templates extracted from the original log, it can include:

[0072] Determine the line template corresponding to each log line in the original log; replace each log line with the corresponding line template to obtain a set of line templates.

[0073] Or,

[0074] Construct a set of template identifiers based on the template identifiers of the line templates corresponding to each log line in the original log; create a set of line templates according to the line templates corresponding to each template identifier in the set of template identifiers.

[0075] Continuing with the above example, the line templates identified by ID=1, ID=2, and ID=3 are extracted from the original log. The corresponding template identifiers for each log line in the original log are: 1, 1, 2, 3, and 3. Therefore, the set of template identifiers includes: 1, 1, 2, 3, and 3. Further, the line template set can be constructed from this set of template identifiers as follows:

[0076] connected to<*>. / / Log line 1

[0077] connected to<*>. / / Log line 2

[0078] Hex number<*>. / / Log line 3

[0079] user<*>logged in. / / Log line 4

[0080] user<*>logged in. / / Log line 5

[0081] ……

[0082] Among them, since there is a logical order in each line template of the line template set, this line template set can be understood as a line template sequence.

[0083] Optionally, for the original log with a large amount of data, the original log can be first segmented according to a preset unit data volume to obtain multiple data blocks with a data volume less than or equal to the preset unit data volume; then, a line template set corresponding to each data block is constructed, and anomaly detection and anomaly localization are performed for each line template set.

[0084] S120, perform anomaly detection on the line template set based on an unsupervised detection model.

[0085] S130, if the detection result indicates an anomaly, obtain the keywords used to represent the anomaly, and determine the anomalous line template according to the matching relationship between the common fields carried by the line template and the keywords;

[0086] S140, locate the target log line corresponding to the anomalous line template in the original log.

[0087] An anomaly localization method provided by an embodiment of this application constructs a line template set from the line templates extracted from the original log, realizing the preprocessing of the complex original log; then, an anomaly judgment is performed through an unsupervised model, achieving the purpose of de-expertized anomaly judgment; finally, the target log line corresponding to the anomalous log line is located in the original log, that is, the final anomaly localization is realized by locating the target log line carrying the anomaly information.

[0088] Since existing anomaly analysis methods rely heavily on expert experience and cannot achieve automated analysis. However, when the volume of original logs is large and faced with the non-uniform original logs generated by network cloud devices, relying on expert experience becomes inefficient. Therefore, the embodiments of the present application provide alternative embodiments to solve this technical problem.

[0089] In an alternative embodiment, the number of unsupervised detection models in step S210 is at least two. Optionally, the at least two unsupervised detection models can be selected from the following models: log clustering model; principal component analysis model; invariant mining model; isolation forest model.

[0090] Among them, when the log clustering model processes the set of row templates, it first vectorizes the set of row templates so that each row template corresponds to a vector; secondly, it performs hierarchical clustering on the row templates in the set of row templates to generate multiple normal classes and abnormal classes, and performs the following operations for each class: selects the representative vector of the class, and matches the representative vector with other row templates in the set of row templates. If the distance between other row templates and the representative vector is less than a preset threshold, then add the other row templates to the set where the representative vector is located until all representative vectors are matched; if the distance is always greater than the preset threshold, then create a new class and use the row template as the new representative vector. Finally, the set of row templates is divided into multiple sets according to different "classes". Statistics the proportion of the set corresponding to the abnormal class in the set of row templates, and determines this proportion as the result.

[0091] Among them, when the principal component analysis template processes the set of row templates, it first processes the set of row templates into a matrix to be processed; secondly, projects the high-dimensional data in the matrix to be processed onto a new coordinate system composed of k principal components (i.e., k dimensions, k is less than the original dimension). Calculate the variances of each of the k principal components, and determine the principal component dimension from the k principal components. If the variance on the principal component dimension is greater than a preset threshold, it can be determined that the set of row templates is abnormal, otherwise, it is determined that the set of row templates is normal.

[0092] Among them, when the invariant mining algorithm processes the set of row templates, it first estimates the invariant space of the set of row templates through singular value decomposition. Secondly, uses a brute-force search algorithm to find invariants from the set of row templates. Verifies each row template by defining a threshold, and determines the row templates that exceed the threshold as abnormal. Finally, statistics the proportion of the row templates determined to be abnormal, and uses this proportion as the final result.

[0093] Among them, when the isolation forest model processes the set of line templates, it preprocesses the set of line templates so that each line template serves as a leaf node, and then constructs one or more isolation trees based on the set of line templates. Calculate the distance from the root node to each leaf node on each isolation tree, and regard the leaf node farther from the root node as an abnormal leaf node. Finally, count the proportion of all abnormal leaf nodes, and use this proportion as the final result.

[0094] In this embodiment, anomaly detection is performed on the set of line templates based on the unsupervised detection model, which may specifically include the following steps Sa1 to Sa3.

[0095] Sa1, input the set of line templates into each unsupervised detection model in sequence, and obtain the results output by each unsupervised detection model.

[0096] Sa2, determine the corresponding weight of each unsupervised detection model.

[0097] Sa3, determine the detection result according to the respective corresponding weights and the output results of at least two unsupervised detection models.

[0098] In one example, the above four models are used as the unsupervised detection models for performing anomaly detection. The set of line templates is input into each detection model, and the results output by each detection model are obtained.

[0099] Next, the following formula 1 can be used to synthesize the four results to obtain the final detection result.

[0100]

[0101] Among them, k i is the weight of each model, and s i is the result predicted by each model. After weighted averaging, the final detection result can be obtained.

[0102] In one implementation manner of this embodiment, a preset anomaly threshold is obtained, and the size of the detection result 5 and the anomaly threshold are compared. If the anomaly threshold is greater than the preset anomaly threshold, it is determined that the detection result represents an anomaly; otherwise, it is determined that the detection result represents normal.

[0103] Among them, for the detection result representing normal, the log analysis process can be ended. For the detection result representing an anomaly, the anomaly location process needs to be further executed.

[0104] After determining that the original log is abnormal through the unsupervised detection model, the embodiments of the present application also provide embodiments to illustrate how to perform anomaly location.

[0105] In an optional embodiment, after determining that the detection result characterizes an anomaly, keywords for characterizing the anomaly are obtained.

[0106] Optionally, obtaining the keywords for characterizing the anomaly includes at least one of the following methods:

[0107] Extract keyword fields from the set of row templates based on a preset algorithm, and determine the keyword fields as the keywords for characterizing the anomaly; obtain historical keywords, and determine the historical keywords as the keywords for characterizing the anomaly.

[0108] Optionally, the preset algorithm can be the TF-IDF (term frequency–inverse document frequency) algorithm. Among them, the principle of this algorithm is: if a

[0109] certain word appears frequently in an article and rarely appears in other articles, then it is considered that this word or phrase has good category discrimination ability and is suitable for classification. Among them, TF is the term frequency,

[0110] representing the probability of a term appearing in the text; IDF is the inverse document frequency, representing a measure of the general importance of a term. Further, taking the common fields of all row templates in the set of row templates as the extraction range, use the TF-IDF algorithm to extract keyword fields from this extraction range and use them as keywords.

[0111] Optionally, the preset algorithm can be the TextRank algorithm. The TextRank algorithm is a ranking algorithm for keyword extraction and document summarization, improved from the PageRank algorithm for ranking the importance of web pages by Google. It can extract keywords using the co-occurrence information (semantics) between words within a document. It can extract keywords and keyword groups from a given text and extract key sentences of the text using an extractive automatic summarization method. Further, taking the common fields of all row templates in the set of row templates as the extraction range, use the TextRank algorithm to extract keyword fields from this extraction range and use them as keywords.

[0112] It should be noted that the specific usage processes of the TF-IDF algorithm and the TextRank algorithm can refer to related technologies. For the sake of simplicity of description, they will not be elaborated here.

[0113] In an implementation manner of this embodiment, the abnormal row templates are determined according to the matching relationship between the common fields carried by the row templates and the keywords, which may specifically include steps Sb1 to Sb2.

[0114] Sb1, determine the matching relationship between the common fields carried by the row templates and at least one keyword.

[0115] Specifically, for each line template in the line template set, perform the determination operation shown in Sb1.

[0116] 0Sb2, if the matching relationship is that the common fields carried by the line template include at least one keyword, determine the corresponding line template as an abnormal line template.

[0117] Optionally, add the abnormal line templates that meet the above matching relationship in the line template set to the abnormal template set.

[0118] In an example, the obtained keywords characterizing abnormalities may include: "Errno", 5 "Failed", and "Exp". The following 6 abnormal line templates as shown in Table 1 below were searched for in the original log.

[0119] Abnormal line templates.

[0120] Table 1

[0121]

[0122]

[0123] In Table 1, the line template with ID = 8 represents a reset connection error, and this template includes the keyword "Errno", so it can be used as an abnormal line template.

[0124] Some abnormal line templates can be determined from the line template set through keywords. However, for

[0125] network cloud devices, since the forms of logs generated by each device are not unified, the log lines with abnormal 5 information generated by them may not all include keywords. Therefore, there may be abnormal line templates in the line template set that are not

[0126] identified. To solve this technical problem, the embodiments of the present application also provide a possible implementation.

[0127] In an optional embodiment, before locating the target

[0128] log line corresponding to the abnormal line template in the original log, it further includes:

[0129] 0For each abnormal line template, determine the similarity between this abnormal line template and other line templates in the line template set except for all the abnormal

[0130] line templates; if it is determined that the similarity between any other line template and this abnormal line template is greater than a preset threshold, then use any other line template as a new abnormal line template.

[0131] Optionally, after determining the new abnormal row template, the new abnormal row template may be added to the abnormal template set.

[0132] Optionally, the similarity between the abnormal row template and other row templates can be calculated by the edit distance algorithm. Among them, the edit distance algorithm is often used to measure the similarity between two sequences. For example, between two words word1 (such as kitten) and word2 (such as sitting), the conversion from word1 to

[0133] The minimum number of single-character editing operations required for another word word2; finally, it is counted that the minimum number of single-character editing operations required to convert 0 "kitten" to "sitting" is 3 steps, as follows:

[0134] Next steps Sc1 to Sc3.

[0135] Sc1, kitten→sitten(substitution of "s"for"k") / / replace

[0136] Sc2, sitten→sittin(substitution of "i" for "e") / / replace

[0137] Sc3, sittin→sitting(insertion of "g" at the end) / / Increase 5 The edit distance algorithm defines three character editing operations, such as insertion, deletion, and replacement.

[0138] For the use of the distance editing algorithm, please refer to the relevant technology, and for the sake of simplicity, it will not be repeated here.

[0139] In one example, taking the abnormal row template with ID=8 in Table 1 as an example, the similarity between other templates in the row template set and the abnormal row template is calculated, and finally two other row templates have high similarity and can be used as new abnormal row templates.

[0140] Table 2

[0141]

[0142] In Table 2, the other row template with ID=11 indicates a timeout connection error. Although the row template does not contain the above keywords such as "Errno", "Failid" and "Exp", it is difficult to be determined as an abnormal row template through the above keyword screening mode. However, it can be determined as a new abnormal row template by comparing the similarity with the determined abnormal row template (such as the abnormal row template with ID=8), thus realizing the generalization operation of the abnormal row template.

[0143] After determining the abnormal line template through the above embodiments, abnormal location can be performed according to the abnormal line template next.

[0144] In an alternative embodiment, locating target log lines corresponding to the abnormal line template in the original log may specifically include:

[0145] Determine at least one order of the abnormal line template in the line template set; locate the target log line corresponding to each order in the original log.

[0146] Continuing with the above example, the line template set constructed for the original log is as follows:

[0147] connected to<*>. / / Log line 1

[0148] connected to<*:>. / / Log line 2

[0149] Hex number<*X:>. / / Log line 3

[0150] user<*>logged in. / / Log line 4

[0151] user<*>logged in. / / Log line 5

[0152]

[0153] Failed to compute_task_migrate_server:No valid host was found.There are not enough hosts available. / / Log line 9

[0154] Failed to compute_task_migrate_server:No valid host was found.There are not enough hosts available. / / Log line 10

[0155] connected to<*:>. / / Log line 11

[0156] AMQP server on controller:<*>is unreachable:[Errno 104]Connection reset by peer.Trying again in<*>seconds. / / Log line 12

[0157] AMQP server on controller: <*> is unreachable: [Errno 104] Connection reset by peer. Trying again in <*> seconds. / / Log line 13

[0158] AMQP server on controller: <*> is unreachable: [Errno 104] Connection reset by peer. Trying again in <*> seconds. / / Log line 14

[0159] AMQP server on controller: <*> is unreachable: [Errno 104] Connection reset by peer. Trying again in <*> seconds. / / Log line 15

[0160] …

[0161] In this example, the following log lines can be determined as the target log lines: Log line 9, Log line 10, Log line 12, Log line 13, Log line 14, and Log line 15.

[0162] An exception location method shown in an embodiment of this application can be applied to perform exception analysis on various different types of logs. To more clearly understand the effect of this method, an embodiment of this application also provides an example for demonstration, and the implementation process of this example can refer to Figure 2 the process schematic diagram shown.

[0163] This example includes steps S1001 to S1004 in total.

[0164] S1001, Data acquisition.

[0165] Specifically, obtain network cloud logs (i.e., the original logs in the above embodiment) from multiple network cloud devices. Since the types of network cloud devices are different, the forms of the provided network cloud logs are not unified.

[0166] S1002, Data preprocessing.

[0167] Specifically, extract line templates for all network cloud logs. Among them, the network cloud logs can refer to the "original logs" shown in the above embodiment, and the extracted line templates can be the line templates shown in the above embodiment.

[0168] Further, replace each log line in the network cloud log with the corresponding line template to obtain the log to be processed corresponding to each network cloud log.

[0169] Among them, if the data volume of each log to be processed is too large, each log to be processed can be split and processed into units of logs to be processed with a data volume not greater than a preset threshold (that is, the set of line templates shown in the above embodiments). Perform step S1003 for each unit log to be processed.

[0170] S1003, use an unsupervised detection model to determine whether the unit log to be processed is abnormal.

[0171] Specifically, at least one unsupervised detection model can be called to determine whether the unit log to be processed is abnormal.

[0172] If the detection result is negative, end.

[0173] If the detection result is positive, perform step S1004.

[0174] S1004, perform root cause location according to the root cause set.

[0175] Before performing S1004, it also includes: obtaining the log to be processed obtained by the processing of S1002; extracting keywords from the common information of all line templates of the log to be processed; and determining whether to list the line template as a root cause template (that is, the abnormal line template in the above embodiments) according to whether each line template includes the keyword; combining all the root cause templates to obtain a root cause set (that is, the abnormal template set in the above embodiments).

[0176] Determine the target log line in the network cloud log corresponding to each root cause template in the root cause set. Among them, the target log line is the log line carrying abnormal information.

[0177] So far, the location of the abnormal cause is completed.

[0178] Figure 3 An abnormal location device 300 is shown. The device 300 includes the following modules:

[0179] A construction module 310, configured to construct a set of line templates by using log templates extracted from the original log, and the order of each line template in the set of line templates is the same as that of the corresponding log line in the original log.

[0180] A detection module 320, configured to perform abnormal detection on the set of line templates based on an unsupervised detection model.

[0181] A determination module 330, configured to, if the detection result is abnormal, obtain keywords for characterizing the abnormality, and determine abnormal line templates according to the matching relationship between the common fields carried by the line templates and the keywords.

[0182] A positioning module 340, configured to locate a target log line corresponding to an abnormal line template in the original log.

[0183] Optionally, the number of unsupervised detection models is at least two; in the abnormal detection of the line template set based on the unsupervised detection model, the detection module 320 is specifically configured to:

[0184] Input the line template set into each unsupervised detection model in sequence, and obtain the results output by each unsupervised detection model; determine the weights corresponding to each unsupervised detection model; and determine the detection result according to the weights and the results output by at least two unsupervised detection models respectively.

[0185] Optionally, in determining the abnormal line template according to the matching relationship between the common fields carried by the line template and the keywords, the determination module 330 is specifically configured to:

[0186] Determine the matching relationship between the common fields carried by the line template and at least one keyword; if the matching relationship is that the common fields carried by the line template include at least one keyword, determine the corresponding line template as an abnormal line template.

[0187] Optionally, in obtaining the keywords for characterizing abnormalities, the determination module 330 specifically performs at least one of the following processes:

[0188] Extract keyword fields from the line template set based on a preset algorithm, and determine the keyword fields as the keywords for characterizing abnormalities; obtain historical keywords, and determine the historical keywords as the keywords for characterizing abnormalities.

[0189] Optionally, before the determination module 330 locates the target log line corresponding to the abnormal line template in the original log, it is further used for

[0190] For each abnormal line template, determine the similarity between the abnormal line template and other line templates in the line template set except all the abnormal line templates; if it is determined that the similarity between any other line template and the abnormal line template is greater than a preset threshold, then use any other line template as a new abnormal line template.

[0191] Wherein, determining the similarity between the abnormal line template and other line templates in the line template set except all the abnormal line templates specifically includes: determining the similarity between the common fields of the abnormal line template and the common fields of other line templates as the similarity between the abnormal line template and other line templates.

[0192] Optionally, in locating the target log line corresponding to the abnormal line template in the original log, the positioning module 340 is specifically configured to:

[0193] Determine at least one order of the abnormal line template in the line template set; locate the target log line corresponding to each said order in the original log.

[0194] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the methods according to the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details will not be repeated here.

[0195] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the abnormal location method. Compared with the related art, it can achieve: de-expertized abnormal location for network cloud logs with inconsistent formats.

[0196] In an alternative embodiment, an electronic device is provided, as Figure 4 shown. Figure 3 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0197] The processor 4001 may be a CPU (Central Processing Unit, central processing unit), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 4001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0198] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 3 only a thick line is used to represent it in Figure 3 , but it does not mean that there is only one bus or one type of bus.

[0199] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.

[0200] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0201] Among them, the electronic device includes but is not limited to: a server.

[0202] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0203] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0204] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than that shown in the drawings or described in words.

[0205] It should be understood that although the flowchart of the embodiment of this application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiment of this application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiment of this application does not limit this.

[0206] The above are only alternative implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.

Claims

1. An abnormal location method, characterized in that The method includes: Constructing a set of line templates through log templates extracted from the original log, where each line template in the set of line templates is in the same order as the corresponding log line in the original log; Performing anomaly detection on the set of line templates based on an unsupervised detection model; If the detection result indicates an anomaly, obtaining keywords for characterizing the anomaly, and determining an anomalous line template according to the matching relationship between the common fields carried by the line template and the keywords; Locating a target log line corresponding to the anomalous line template in the original log; The number of the unsupervised detection models is at least two; performing anomaly detection on the set of line templates based on the unsupervised detection model includes: sequentially inputting the set of line templates into each unsupervised detection model, and obtaining the results output by each unsupervised detection model; determining the corresponding weight of each unsupervised detection model; and determining the detection result according to the respective weights and the output results of at least two unsupervised detection models; Before locating the target log line corresponding to the anomalous line template in the original log, it further includes: for each anomalous line template, determining the similarity between the anomalous line template and other line templates in the set of line templates except all the anomalous line templates; if it is determined that the similarity between any other line template and the anomalous line template is greater than a preset threshold, then taking any other line template as a new anomalous line template.

2. The method according to claim 1, wherein Determining the anomalous line template according to the matching relationship between the common fields carried by the line template and the keywords includes: Determining the matching relationship between the common fields carried by the line template and at least one keyword; If the matching relationship is that the common fields carried by the line template include at least one keyword, determining the corresponding line template as the anomalous line template.

3. The method according to claim 1, wherein Obtaining the keywords for characterizing the anomaly includes at least one of the following: Extracting key fields from the set of line templates based on a preset algorithm, and determining the key fields as the keywords for characterizing the anomaly; Obtaining historical keywords, and determining the historical keywords as the keywords for characterizing the anomaly.

4. The method according to claim 1, wherein Determining the similarity between the anomalous line template and other line templates in the set of line templates except all the anomalous line templates includes: Determining the similarity between the common fields of the anomalous line template and the common fields of the other line templates as the similarity between the anomalous line template and the other line templates.

5. The method according to any one of claims 2-3, characterized in that, Locating the target log line corresponding to the anomalous line template in the original log includes: Determining at least one order of the anomalous line template in the set of line templates; Locating the target log line corresponding to each order in the original log.

6. An abnormal location device, characterized in that, The device includes: A construction module for constructing a set of line templates through log templates extracted from the original log, where each line template in the set of line templates is in the same order as the corresponding log line in the original log; A detection module for performing anomaly detection on the set of line templates based on an unsupervised detection model; A determination module, configured to, if the detection result is abnormal, obtain keywords for characterizing the abnormality, and determine an abnormal line template according to the matching relationship between the common fields carried by the line template and the keywords; A positioning module, configured to locate a target log line corresponding to the abnormal line template in the original log; The number of the unsupervised detection models is at least two; the detection module is further configured to sequentially input the line template set into each unsupervised detection model, and obtain the results output by each unsupervised detection model; determine the weights corresponding to each unsupervised detection model; and determine the detection result according to the weights and the results output by at least two unsupervised detection models respectively; Before locating a target log line corresponding to the abnormal line template in the original log, the apparatus is further configured to, for each abnormal line template, determine the similarity between the abnormal line template and other line templates in the line template set except for all the abnormal line templates; if it is determined that the similarity between any other line template and the abnormal line template is greater than a preset threshold, use any other line template as a new abnormal line template.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Template-oriented Word2vec-based log exception detection method and device

    CN111459964A

  • Log anomaly detection method and device based on Word2Vec and electronic equipment

    CN113377607A

  • Abnormal log identification method and device, electronic equipment and storage medium

    CN115480994A