A method for parsing multi-source heterogeneous log data based on semantic enhancement

Through regular matching and the construction of template tree structures, combined with semantic vector processing, the problem of variable information being ignored in multi-source heterogeneous log data analysis is solved, the analysis speed and accuracy are improved, and the automated analysis of intelligent operation and maintenance systems is supported.

CN116341513BActive Publication Date: 2025-07-18NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310271716.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-07-18
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

The existing multi-source heterogeneous log data analysis methods mainly obtain log templates through similarity measurements, ignoring the variable information in the log, resulting in a reduction in parsing accuracy.

Method used

The regular matching method is used to preprocess, define the template tree structure and build the template tree, replace variables with one-to-one words corresponding to semantics, and use the hierarchy of the template tree and semantic vectors to split and merge the template to improve the parsing accuracy.

Benefits of technology

It has achieved minimal human intervention, improved log parsing speed and accuracy, reduced the time and cost of system upgrades and fault locations, provided valuable data sets for intelligent operation and maintenance systems, and promoted automated analysis and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116341513B_ABST
    Figure CN116341513B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for parsing multi-source heterogeneous log data based on semantic enhancement. First, a regular matching method is adopted for preprocessing heterogeneous log data, including presetting regular expressions to match common variables and using words with one-to-one semantic correspondence to replace these variables, so as to uniformly retain the important data parts in the log statements. Then, a template tree structure is defined and the template tree is constructed. The time for constructing and searching the template tree is saved by fixing the height of the template tree, and it is set that each layer of nodes of the template tree carries corresponding information to reduce the time required for template matching. Finally, template splitting and merging are performed to further improve the accuracy of the log parsing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data parsing, and particularly relates to a method for parsing multi-source heterogeneous log data based on semantic enhancement. Background Art

[0002] Multi-source heterogeneous log data has characteristics such as being unstructured and having a wide variety of types. It details the running information of multi-source heterogeneous systems and can help operation and maintenance personnel better monitor the system status, thereby detecting system anomalies. With the reform and upgrade of computer systems, the traditional method relying on matching rules and manual detection is not applicable. Therefore, parsing multi-source heterogeneous log data is an essential link for realizing anomaly detection of multi-source heterogeneous logs, and the accuracy of log parsing will directly affect the accuracy of anomaly detection. Therefore, a method for parsing multi-source heterogeneous log data is needed to convert the unstructured multi-source heterogeneous log data into a structured form and prepare data for subsequent steps of anomaly detection.

[0003] Parsing multi-source heterogeneous log data is a process of transforming log data from an unstructured form to a structured form while obtaining log template information. The existing methods for parsing multi-source heterogeneous log data mainly obtain log templates through similarity measurement, but the process of merging templates based on similarity often ignores variable information in the logs, thereby reducing the accuracy of log parsing. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies of the prior art and provide a method for parsing multi-source heterogeneous log data based on semantic enhancement. This method not only extracts variable data in the multi-source heterogeneous log structure, but also splits and merges templates according to the semantic information of the variable data, so that the finally obtained log template information is more accurate.

[0005] The present invention is achieved through the following technical solutions:

[0006] A method for parsing multi-source heterogeneous log data based on semantic enhancement, comprising the following steps:

[0007] Step 1, preprocess heterogeneous log data by using regular matching, including presetting regular expressions to match common variables and replacing these variables with words with one-to-one semantics;

[0008] Step 2, define a template tree structure and construct a template tree;

[0009] Step 2.1: Define the template tree structure. The first layer of the template tree only stores a root node without data information. The second layer is the length node, and the data stored in the node is the number of words in the log statement after regular matching. The third layer is the prefix node, and the data stored in the node is the prefix expression of the log statement after regular matching. The prefix expression consists of the first n / 2 words in the log statement, where n is the total number of words in the log statement. The fourth layer is the leaf node, and the data stored in the node is the log cluster information. This log cluster contains m log templates.

[0010] Step 2.2: Construct the template tree according to the template tree structure defined in Step 2.1, including the following steps:

[0011] Step 2.21: Find or create the length node of the second layer of the template tree according to the length of the log statement after regular matching.

[0012] Step 2.22: Find or create the prefix node of the third layer of the template tree according to the prefix expression of the log statement after regular matching.

[0013] Step 2.23: Judge whether the log statement matches the log template information according to the information of the first three layers of nodes. If the match is successful, the log statement will be added to the log template set of the log cluster. If the match fails, a log cluster containing log template information will be created based on the target log statement and added to the leaf node.

[0014] Step 3: Template splitting and merging. Among them, for template merging, the log templates of the same log cluster are merged by using the wildcard replacement method. For template splitting, the word vectors of the words in the log statement are represented by Word2vec, and then the similarity of the log statements in the same template is calculated according to the Pearson linear correlation coefficient. If it is less than 0, the original template will be split.

[0015] In the above technical solution, in Step 1, the variables to be replaced include: IP variable, digital variable, time variable.

[0016] In the above technical solution, in Step 1, it also includes: using regular expressions to locate all special characters in the original log data, and replacing special characters with a single space; reducing multiple consecutive spaces to one, and reducing multiple consecutive identical replaced words to one.

[0017] In the above technical solution, in Step 2.21, according to the number of words in the target log statement, traverse the nodes of the second layer of the template tree. If the search is successful, it means that the operation of matching the length node is completed. If the search fails, a new length node needs to be created according to the number of words.

[0018] In the above technical solution, in step 2.22, according to the prefix expression of the target log statement, traverse the third-level nodes of the template tree, and judge whether the matching is successful according to the similarity calculation. If the matching is successful, it means that the search is successful. If the matching fails, a prefix node needs to be created according to the prefix expression.

[0019] In the above technical solution, in step 2, the edit distance is used to calculate the similarity between the prefix node and the log template.

[0020] In the above technical solution, in step 3, the Pearson linear correlation coefficient formula is as follows:

[0021]

[0022] Among them, X j represents the word vector of the words in the log statement, and Y j represents the word vector of the statement to be matched. The range of the calculation result is Person ∈ [-1, 1]. If it is less than 0, it indicates negative correlation, and the template needs to be split. If it is greater than 0, it indicates positive correlation.

[0023] The present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed, the steps of the above method are realized.

[0024] The advantages and beneficial effects of the present invention are as follows:

[0025] Compared with the existing log parsing methods, the method of the present invention has the advantages of minimizing human intervention, fast parsing speed, and accurate parsing results. Due to the characteristics of diverse types and different structures of multi-source heterogeneous log data, traditional methods need to set corresponding matching rules for each type of log. However, the method proposed by the present invention first conducts a preliminary screening through regular matching, which can uniformly retain the important data parts in the log statements. Secondly, it saves the time for constructing and searching the template tree by fixing the height of the template tree. Then, it sets that each layer of nodes of the template tree carries corresponding information to reduce the time required for template matching. Finally, based on semantic vectors, the log templates are merged and split, thereby further improving the accuracy of the log parsing results.

[0026] The beneficial effects of the present invention are mainly reflected in the following aspects: on the one hand, it reduces the time and labor costs required for system architecture upgrade and fault location, and at the same time improves the time efficiency and parsing accuracy of processes such as log analysis and processing; on the other hand, it provides a valuable data set for the intelligent operation and maintenance system in the field of machine learning, thereby promoting the realization of engineering applications such as automated analysis and detection. Description of the Drawings

[0027] Figure 1It is the flowchart of the steps of the method for parsing multi-source heterogeneous log data based on semantic enhancement of the present invention.

[0028] Figure 2 It is the data diagram of the processing result after the regular matching process of multi-source heterogeneous log data Trace1 and Trace2.

[0029] Figure 3 It is the schematic structural diagram of the template tree of the present invention.

[0030] Figure 4 It is the flowchart for constructing the template tree.

[0031] Figure 5 It is the flowchart for template merging.

[0032] Figure 6 It is the flowchart for template splitting.

[0033] Figure 7 It is an example of template merging and splitting.

[0034] For those of ordinary skill in the art, without creative efforts, other relevant drawings can be obtained according to the above drawings. Detailed implementation manners

[0035] In order to enable those in the technical field to better understand the solution of the present invention, the technical solution of the present invention will be further described below in conjunction with specific embodiments.

[0036] A method for parsing multi-source heterogeneous log data based on semantic enhancement, see the attached Figure 1 , including the following steps:

[0037] Step 1, Regular matching

[0038] The original log data usually consists of a constant part and a variable part. The constant part generally adopts a fixed structure, and the variable part is various parameters, such as numbers, times, variables or special symbols such as "-", ">". The log parsing idea is to remove the variable part in the log data and only retain the constant part, and then perform operations such as integration and encoding on it. Among them, the variable part usually needs to be processed with the help of the expert knowledge of the actual application scenario to effectively improve the accuracy of mining the log template; the processing operations on the variable part mainly include direct removal and indirect replacement. The present invention uses the method of regular matching to perform preprocessing on heterogeneous log data. Specifically, it includes the following steps:

[0039] Step 1.1, First, use regular expressions to locate all special characters in the original log data, and then use a single space to replace the special characters.

[0040] Step 1.2, secondly, preset some regular expressions to match common variables (such as IP, numbers, time, etc.), and use words with one-to-one semantic correspondence to replace these variables (such as replacing the number "23" with the word "NUMBER"), because these variables still have corresponding semantic information when analyzing the entire log data. If the wildcard "<*>" is used for unified replacement, the relationships of all variables will be blurred, which may affect the accuracy of the sentence vector representation.

[0041] Step 1.3, finally, reduce multiple consecutive spaces to one, and reduce multiple consecutive identical replaced words to one, ensuring that the processed log statement still retains the original semantic information.

[0042] Figure 2 Shows the processing results after the regular matching process of multi-source heterogeneous log data Trace1 and Trace2. It can be seen from Figure 2 that regular matching performs a good replacement on the Trace1 and Trace2 data, and to a greater extent retains the structure and semantics of the original data. Most of the existing Drain log parsing methods are limited to the length of the log statement, that is, two statements with the same semantics and structure may be divided into two different log templates due to different lengths. In addition, most of the existing Spell log parsing algorithms use wildcards to uniformly replace variables in the variable part, and the replaced statements may lose some key information, resulting in incomplete or incorrect semantic information. The present invention uses words with semantic information to replace variables by category in the regular matching process, and merges multiple consecutive special symbols and keywords, so it can effectively alleviate the above two problems to a certain extent.

[0043] Step 2, construct a template tree

[0044] After the regular matching process, the variable part in the original log data has been completed. Next, it is necessary to generate a log template. The core of log template generation lies in the construction of the template tree; the present invention will specifically describe the template tree structure and the process of constructing the template tree.

[0045] Step 2.1, define the template tree structure

[0046] Before the start of the log parsing process, the template tree is an empty tree structure with a root node. As log data is input, new nodes are created to update the information of each layer of nodes in the tree, and finally a template tree is formed.

[0047] The structure of the template tree constructed by the present invention is as Figure 3As shown, the height of the entire template tree is set to a fixed height h = 4. Then, the time complexity of traversing the entire template tree is o(nlogn), that is, the parsing efficiency in the log parsing process is determined by the tree height h. In the template tree, only a root node without data information is saved in the first layer of the tree; the second layer is the length node, and the data saved in the node is the number of words in the log statement after regular matching; the third layer is the prefix node, and the data saved in the node is the prefix expression of the log statement after regular matching. The prefix expression consists of the first n / 2 words of the log statement, where n is the total number of words in the log statement; the fourth layer is the leaf node, and the data saved in the node is the log cluster information (represented as (LTi, LCi)). This log cluster contains m (m >= 1) log templates (represented as LTj).

[0048] The concepts of log cluster and log template are as follows:

[0049] The log sequence is represented as LS = {L1, L2, L3, ···, Ln}, where LS is the log sequence output in chronological order, n is the length of the log sequence, Li represents a log in the log sequence, and i ∈ [1, n]; the set of log templates corresponding to the log sequence is represented as LT = {LT1, LT2, LT3, ···, LTm}, m is the total number of log templates, LTj represents a generated corresponding log template, j ∈ [1, m], and m ∈ [1, n].

[0050] The concept of a log cluster is that the log texts clustered into the same log cluster in the log sequence are similar. The higher the similarity between two logs, the higher the probability of being divided into the same log cluster; on the contrary, the lower the similarity between two logs, the higher the probability of being divided into different log clusters. According to the concept of the log template, the log statements divided into the same log cluster have the same log template, and the log template LTj is used as the identifier of the log cluster. The set of log clusters is defined as:

[0051] SetLC = {(LT1, LC1)(LT2, LC2), ···, (LTm, LCm)};

[0052] LCi = {L1, L2, L3, ···, Ln};

[0053] Among them, the set of log clusters consists of many log clusters (LTi, LCi). LTi represents the log template, LCi represents the log sequence divided into this log cluster, and Lj represents a log in this log sequence, where i ∈ [1, m] and j ∈ [1, n].

[0054] Step 2.2, constructing the template tree

[0055] According to the template tree structure defined in Step 2.1, the process of constructing the template tree in this step is as follows Figure 4As shown, it includes the following steps.

[0056] Step 2.21: Find or create the second-layer length node of the template tree according to the length of the log statement after regular matching. Specifically: according to the number of words in the target log statement, traverse the second-layer nodes of the template tree. If the search is successful, it means that the matching length node operation is completed; if the search fails, a length node needs to be newly created according to the number of words.

[0057] Step 2.22: Find or create the third-layer prefix node of the template tree according to the prefix expression of the log statement after regular matching. Specifically: according to the prefix expression of the target log statement, traverse the third-layer nodes of the template tree, and judge whether the matching is successful according to the similarity calculation. If the matching is successful, it means that the search is successful; if the matching fails, a prefix node needs to be newly created according to the prefix expression.

[0058] Step 2.23: Judge whether the log statement matches the log template information according to the information of the first three layers of nodes. If the matching is successful, the log statement will be added to the log template set of the log cluster; if the matching fails, a log cluster containing log template information will be created based on the target log statement and added to the leaf node.

[0059] In the process of constructing the template tree above, both Step 2.22 and Step 2.23 involve the matching process, that is, the prefix expression matching in Step 2.22 and the log template matching in Step 2.23. The present invention uses similarity calculation to implement the above two matching processes.

[0060] Both the prefix expression and the log template are expressions obtained by abbreviating the original log statement, so as to represent the structural information of the entire log statement. Therefore, the edit distance (Levenshtein) is used to calculate the similarity Sim between the prefix node and the log template. The formula is defined as follows:

[0061]

[0062] When using this formula to calculate the similarity of the prefix node, fi represents the prefix expression of the target log statement (taking the first half of the target log statement as the prefix expression), si is the prefix node to be matched, Leven is the edit distance similarity calculation function, and Len(f) and Len(s) represent the number of words in the target statement and the number of words in the prefix node to be matched respectively.

[0063] When calculating the similarity of log templates using this formula, fi represents the prefix node information where the target log statement is located (i.e., the prefix expression stored in the prefix node where the target log statement matches successfully or is created successfully), s is the child node of this prefix node (i.e., the log template information stored in the leaf node to be matched), Leven is the edit distance similarity calculation function, and Len(f) and Len(s) represent the number of words in the prefix node and the number of words in the log template in the leaf node to be matched, respectively.

[0064] Among them, the similarity calculation result Sim ∈ (0, 1). If the result value of Sim is closer to 1, it indicates a higher similarity between the two; conversely, if it is closer to 0, it indicates a lower similarity or even dissimilarity between the two.

[0065] Step 3, Template Splitting and Merging

[0066] After the processing of the above steps, the log templates obtained through template tree construction may often have differences in semantic expression from the original log data. Therefore, it is necessary to perform semantic splitting and merging on these log templates. Although the variable part in the log data has been changed during the regular matching stage, the present invention uses words with corresponding semantics for replacement, which not only retains the original structure but also strengthens the semantic information.

[0067] For template merging, the present invention uses the method of wildcard replacement to merge the log templates of the same log cluster. The template merging process is as Figure 5 shown.

[0068] For template splitting, the present invention uses Word2vec to represent the word vectors of the words in the log statement, and then calculates the similarity of the log statements in the same template according to the Pearson correlation coefficient. If it is less than 0, the original template is split. Among them, the Pearson correlation coefficient formula is as follows:

[0069]

[0070] Among them, X j represents the word vector of the word in the log statement, Y j represents the word vector of the word in the statement to be matched. The range of the calculation result Person ∈ [-1, 1]. If it is less than 0, it indicates negative correlation, and the template needs to be split; if it is greater than 0, it indicates positive correlation.

[0071] Figure 6 shows the process of template splitting. The purpose of template splitting is to split the log statements with opposite semantics into different templates by comparing the log statements in the same log template.

[0072] After the above steps, the present invention merges similar templates, and the log messages with opposite semantics are reclassified into different log templates. For example Figure 7 As shown, Log1-3 should be classified into the same log template Template according to the template tree construction process. However, the semantic information of Log1 and Log3 is quite different. The algorithm will reclassify them and generate two sub-templates, corresponding to Template1-1 and Template1-2 respectively, that is, the process of splitting the log template is completed.

[0073] The above is an exemplary description of the present invention. It should be noted that without departing from the core of the present invention, any simple deformation, modification or equivalent replacement that can be made by those skilled in the art without creative labor falls within the protection scope of the present invention.

Claims

1. A method for parsing multi-source heterogeneous log data based on semantic enhancement, characterized in that Including the following steps: Step 1, preprocess heterogeneous log data by means of regular matching, including , presetting regular expressions to match common variables, and using words with one-to-one semantics to replace these variables; Step 2, define the template tree structure and construct the template tree; Step 2.1, define the template tree structure. The first layer of the template tree only stores a root node without data information; The second layer is the length node, and the data stored in the node is the number of words in the log statement after regular matching; The third layer is the prefix node, and the data stored in the node is the prefix expression of the log statement after regular matching. The prefix expression is composed of the first n / 2 words in the log statement, where n is the total number of words in the log statement; The fourth layer is the leaf node, and the data stored in the node is the log cluster information. This log cluster contains m log templates; Step 2.2, construct the template tree according to the template tree structure defined in Step 2.1, including the following steps: Step 2.21: Search for or create the length node of the second layer of the template tree according to the length of the log statement after regular matching; Step 2.22: Search for or create the prefix node of the third layer of the template tree according to the prefix expression of the log statement after regular matching; Step 2.23: Judge whether the log statement matches the log template information according to the information of the first three layers of nodes. If the matching is successful, the log statement will be added to the log template set of the log cluster; if the matching fails, a log cluster containing log template information will be created based on the target log statement and added to the leaf node; Step 3, template splitting and merging; among them, for template merging, use the wildcard replacement method to merge the log templates of the same log cluster; for template splitting, use Word2vec to represent the word vectors of the words in the log statement, and then calculate the similarity of the log statements in the same template according to the Pearson linear correlation coefficient. If it is less than 0, the original template will be split.

2. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, wherein: In Step 1, the variables to be replaced include: IP variables, numeric variables, and time variables.

3. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, wherein: In Step 1, it also includes: using regular expressions to locate all special characters in the original log data, and using a single space to replace special characters; reducing multiple consecutive spaces to one, and reducing multiple consecutive identical replaced words to one.

4. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, characterized in that: In Step 2.21, according to the number of words in the target log statement, traverse the nodes of the second layer of the template tree. If the search is successful, it means that the operation of matching the length node is completed; if the search fails, a length node needs to be newly created according to the number of words.

5. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, wherein: In Step 2.22, according to the prefix expression of the target log statement, traverse the nodes of the third layer of the template tree, and judge whether the matching is successful according to the similarity calculation. If the matching is successful, it means that the search is successful; if the matching fails, a prefix node needs to be newly created according to the prefix expression.

6. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, wherein: In Step 2, the edit distance is used to calculate the similarity between the prefix node and the log template.

7. The method for parsing multi-source heterogeneous log data based on semantic enhancement according to claim 1, wherein: In Step 3, the formula for the Pearson linear correlation coefficient is as follows: Among them, X j represents the word vector of the word in the log statement, and Y j represents the word vector of the statement to be matched. The range of the calculation result is Person ∈ [-1, 1]. If it is less than 0, it indicates negative correlation, and the template needs to be split. If it is greater than 0, it indicates positive correlation.

8. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed, it implements the steps of the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge graph construction method for log data

    CN112579707A

  • Method and device for classifying logs of data center

    CN114968933A