Log template generation method, log marking method and system

By building a log template tree and extracting keywords using TextRank algorithm for clustering, the problem of inefficient processing of massive log data in modern software systems is solved, and automated and accurate log template extraction and marking is realized, which is suitable for various log files.

CN120337891APending Publication Date: 2025-07-18HEFEI DAZHIHUI CAIHUI DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465215.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In modern software systems, the processing of massive log data of hundreds of millions of yuan is inefficient and lacks universality, and traditional classification methods are complex and difficult to achieve generalization.

Method used

By segmenting the training set logs, a template tree with hierarchical structure is constructed, the template tree branches are merged with similarity, keywords are extracted and clustered using the TextRank algorithm, and a general log template is generated to achieve automated marking.

Benefits of technology

It improves the accuracy and universality of log template extraction, reduces manual participation, improves log processing efficiency and system scalability, and is suitable for various log files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337891A_ABST
    Figure CN120337891A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of log file processing, in particular to a log template generation method and system and a log marking method and system. The method comprises the following steps: firstly, carrying out word segmentation on logs in a training set, taking segmented words as leaf nodes, establishing a template tree reflecting log contents through a hierarchical structure, and combining template tree branches with similarity; taking a path from the root node to the last segmented word of the log as a log template and extracting a keyword list; and clustering the keyword categories, and for each category, forming a general template of logs in the category by combining all keywords in the category. The method aims at automatically extracting and marking the hundred million-level log template, meanwhile, the universality of the template is improved, and the problem of massive log data processing in a modern software system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of log file processing, and in particular to a log template generating method, a log marking method and a system. Background Art

[0002] With the development of information technology, modern software systems are becoming increasingly complex, and the amount of log data generated has also increased dramatically. Billions of massive logs make it extremely difficult and time-consuming to manually locate anomalies. At the same time, due to the various log formats and styles, traditional classification methods require the configuration of complex rules and regular expressions, and the classification efficiency is low and not universal. This not only increases the workload, but also makes it difficult to implement a universal log processing solution. Summary of the invention

[0003] In order to overcome the defects of slow log processing efficiency and poor versatility in the above-mentioned prior art, the present invention proposes a log template generation method, which aims to automatically extract and mark hundreds of millions of log templates, while improving their versatility and solving the problem of massive log data processing in modern software systems.

[0004] The present invention proposes a log template generation method, which first performs word segmentation on logs in a training set, takes the word segmentation as a leaf node, establishes a template tree that reflects the log content through a hierarchical structure, and merges the template tree branches based on similarity; takes the path from the root node to the last word segmentation of the log as the log template and extracts a keyword list; clusters the keyword categories, and for each category, combines all the keywords in the category to form a general template for the log in the category.

[0005] Preferably, the first-level leaf nodes of the template tree are used to mark the number of word segments of the log; the word segments of each log are connected under the leaf nodes corresponding to the number of word segments.

[0006] Preferably, the construction of the template tree includes the following steps:

[0007] Construct a root node, and make the child nodes of the root node as the first-level leaf nodes; the first-level leaf nodes are used to mark the number of word segmentations in the log; make the logs in the training set as the logs to be processed in turn;

[0008] For the logs to be processed, the number of segmented words is counted; if there is a corresponding node for the first-layer leaf node, a branch expressing the segmented word sequence of the logs to be processed is generated under the node; if there is no corresponding node for the first-layer leaf node, a new first-layer leaf node is added to mark the number of segmented words for the logs to be processed, and a branch expressing the segmented word sequence of the logs to be processed is generated under the node;

[0009] On the template tree, the child nodes of the same parent node are different from each other. On the branches connected below each first-layer leaf node, the corresponding log word segments are arranged in a parent-child node relationship according to their sorting in the log.

[0010] Preferably, the keyword extraction of the log template adopts the TextRank algorithm.

[0011] Preferably, the logs in the training set are log data after log preprocessing; the log preprocessing includes: data cleaning, format unification, and variable replacement.

[0012] A log marking method based on a general log template proposed by the present invention first obtains a template tree and a clustering result of the training set by using the log template generation method described above; then obtains the log to be marked and extracts word segments; obtains the path corresponding to the log to be marked in the template tree according to the number and sorting of the word segments; extracts a keyword list from this path, and classifies the keyword list based on the known clustering result, and obtains the general template of the class where the keyword list is located as the marking result of the log to be marked.

[0013] Preferably, the method for obtaining the path corresponding to the log to be marked in the template tree includes the following steps:

[0014] Search for or add a first-layer leaf node corresponding to the number of word segments of the log to be marked in the template tree as the target node;

[0015] Take the target node as the parent node, and search for or add a child node corresponding to the first word segment of the log to be marked below the parent node and record it as the new target node;

[0016] Loop through this step, traverse the word segments of the log to be marked until the word segments of the log to be marked are inserted into the template tree in order in a parent-child node relationship; then take the path from the root node to the leaf node where the last word segment of the log to be marked is located as the path corresponding to the log to be marked.

[0017] Preferably, if the path of the log to be marked coincides with an existing path on the template tree, directly obtain the general template of the class where this existing path is located as the marking result of the log to be marked;

[0018] If the clustering result of the keyword list of the log to be marked is a new class, then construct a general template for the log to be marked based on the keyword list.

[0019] Preferably, the template tree is updated in real time in combination with the path of the log to be marked.

[0020] A log marking system based on a general log template proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the log marking method based on the general log template.

[0021] The advantages of the present invention are as follows:

[0022] (1) The log template generation method proposed by the present invention, through log word segmentation, regards each word segment as a leaf node of the template tree, and establishes a hierarchical structure to reflect the relationship between log contents. By combining similarities, the leaf nodes of different logs are merged to reduce the number of branches of the template tree and improve the accuracy of template extraction.

[0023] (2) Based on keyword information, the present invention uses a clustering algorithm to analyze and mark the template tree, thereby realizing the automatic classification and annotation of templates.

[0024] (3) By combining multiple methods, including similarity calculation, syntax tree construction, etc., the accuracy of log template extraction is ensured, and the possibility of human errors is reduced. The present invention does not depend on specific application scenarios or log formats, has good adaptability and scalability, and can be applied to various types of log files. Moreover, it can process massive data, realize fast log template extraction and marking, and greatly improve the efficiency of data analysis.

[0025] (4) The log marking method based on a general log template proposed by the present invention classifies logs based on the training results of the general template on the training set, and uses the general template as the log extraction result for marking; the present invention improves the accuracy and generality of log template extraction through intelligent means, and solves the problem of low processing efficiency of massive log data. Through the automatic marking of log templates, the scalability and maintainability of the system are further enhanced.

[0026] (5) By automatically clustering and extracting templates for logs, the manual participation is significantly reduced, and the efficiency of log processing is remarkably improved. Compared with traditional classification methods that require configuring complex rules and regular expressions, this method realizes keyword extraction and automatic classification based on a word segmentation tree-like relationship graph, simplifies the process of log classification through an intelligent processing flow, and reduces the workload. The log processing solution given by the present invention has extremely high generality and applicability.

[0027] (6) Through word segmentation processing and similarity calculation, the present invention improves the accuracy of log extraction and ensures the effectiveness of anomaly detection. By constructing a template tree and saving, managing, and updating the template tree, the efficient storage and retrieval of log templates are realized.

[0028] (7) The present invention combines various technical means such as data preprocessing, word segmentation, template construction, keyword extraction, and clustering analysis to form a complete process for automatic extraction and marking of log templates, which is applicable to various application scenarios that require efficient management and analysis of log data. Description of the Drawings

[0029] Figure 1 It is a flowchart of a method for generating a log template;

[0030] Figure 2 It is a flowchart of a method for constructing a template tree;

[0031] Figure 3 It is a flowchart of a log marking method based on a general log template;

[0032] Figure 4 It is a schematic diagram of a template tree constructed based on a training set in the embodiment;

[0033] Figure 5 It is a schematic diagram of the updated template tree in the embodiment. Detailed implementation manners

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] Referring to Figure 1 , a method for generating a log template proposed by the present invention includes the following steps:

[0036] S1. Obtain historical logs to form a training set, and extract the word segmentation of each log in the training set.

[0037] In this embodiment, first, log embedding is performed on the logs from the target program (that is, log recording instructions are added at key positions in the program code) to ensure that all necessary running information can be captured. These log files are then sent to the Kafka message queue. This step ensures the real-time nature and high-throughput processing ability of the data.

[0038] Then, preliminary cleaning and formatting processing are performed on the collected log files, including but not limited to removing useless characters such as invalid spaces and line breaks, and using variable replacement for fixed statements such as IP addresses and email addresses to reduce the complexity of the logs. This step helps to reduce noise in subsequent analysis, improve data quality and processing efficiency.

[0039] The processed logs are formed into a training set and word segmentation is performed. Specifically, when implementing, Jieba word segmentation or other suitable Chinese word segmentation tools can be used to segment the log content, converting the continuous text sequence into individual word segments (tokens). This process is the basis for constructing the log template. Through accurate word segmentation, the specific meaning and structure of the logs can be understood more accurately.

[0040] S2. Construct a template tree by combining word segmentation of logs, and use the path of the template tree as the log template; each word segmentation on the template tree is used as a leaf node, and whether to merge or generate a new branch is determined by calculating the similarity between adjacent nodes. This method not only improves the accuracy of log template extraction but also effectively reduces the overall complexity of the template tree. Each newly generated template will be saved in the database, and the corresponding template file will be saved for subsequent query and analysis.

[0041] As Figure 2 shown, the construction of the template tree specifically includes the following sub-steps;

[0042] S21. Construct a root node, and make the child nodes of the root node be the first-layer leaf nodes; the first-layer leaf nodes are used to mark the log length, that is, the number of word segmentations;

[0043] S22. For the log to be processed, count the number of word segmentations; determine whether there is a corresponding node in the first-layer leaf nodes;

[0044] If not, add a leaf node marking the length of the log to be processed in the first-layer leaf nodes, and then execute step S23;

[0045] If yes, execute step S23;

[0046] S23. Set the statistical value T, and the initial value of T is 1; make the first-layer leaf node marking the length of the log to be processed be the initial value of the parent node;

[0047] S24. Determine whether there is a node marking the T-th word segmentation of the log to be processed in the child nodes of the parent node;

[0048] If not, add a node marking the T-th word segmentation of the log to be processed under the parent node, and then execute step S25;

[0049] If yes, execute step S25;

[0050] S25. Make the node marking the T-th word segmentation of the log to be processed in the child nodes of the parent node be the target child node, and determine whether the value of T is greater than or equal to the number of word segmentations of the log to be processed;

[0051] If not, update T to T + 1, update the parent node to the target child node, and then return to step S24;

[0052] If yes, use the branch from the root node to the target child node as the log template of the log to be processed;

[0053] S26. Traverse all logs in the training set and execute steps S22 - S25 to obtain a template tree containing all log templates;

[0054] S3. For each log template in the template tree, extract a keyword list. Keywords are crucial for understanding the core meaning of the log template and represent the most critical information points in each template.

[0055] In specific implementation, a keyword algorithm such as TextRank is used to extract keywords to form the keyword list serving as the log template tag; TextRank extracts keywords by constructing a relationship weight graph and calculating the importance of each word, which can better understand the core meaning of the log template.

[0056] The TextRank algorithm is a graph-based ranking algorithm for automatic text summarization and keyword extraction. It draws on Google's PageRank algorithm, represents sentences or words in the text as nodes in a graph, and constructs a weight graph using the connection relationships between the nodes. By iteratively calculating the weights between the nodes, the TextRank algorithm can sort the nodes according to the importance of the nodes.

[0057] S4. Cluster the tags. For each class, combine all the keywords within the class (i.e., the keyword set of all keyword lists within the class) to form an expression template, which serves as the general template for the logs of this class.

[0058] Refer to Figure 3 , a log marking method based on a general log template proposed in this embodiment of the present invention includes the following steps:

[0059] St1. Obtain the log to be marked and extract word segments, using the number of word segments as the length of the log to be marked.

[0060] In specific implementation, the log to be marked is the data result after preprocessing the original log data.

[0061] In specific implementation, Kafka can be used as a message queue to achieve efficient and stable log data collection. First, log points are inserted into the target program (i.e., log recording instructions are added at key positions in the program code) to ensure that all necessary running information can be captured; then, through configuration, the producer sends the original log data from different sources to the Kafka cluster, and the consumer program receives these original log data and preprocesses them as the log to be marked for subsequent marking.

[0062] Log preprocessing mainly performs format unification and variable substitution to ensure that different types of logs can be processed consistently and reduce interference in subsequent processing. Specifically, first, the input log is parsed to remove redundant characters such as spaces, and fixed variables such as IP addresses and email addresses are replaced. The specified fixed variables can be identified and replaced with strings of specific patterns using regular expressions or other methods.

[0063] St2. Determine whether there is a node corresponding to the length of the log to be marked among the first-layer leaf nodes of the template tree;

[0064] If not, add a leaf node marking the length of the log to be marked among the first-layer leaf nodes, and then execute step St3;

[0065] If yes, execute step St3;

[0066] St3. Set the statistical value T, and the initial value of T is 1; let the first-layer leaf node marking the length of the log to be marked be the initial value of the parent node;

[0067] St4. Determine whether there is a node marking the T-th word segment of the log to be marked among the child nodes of the parent node;

[0068] If not, add a node marking the T-th word segment of the log to be marked under the parent node, and then execute step St5;

[0069] If yes, execute step St5;

[0070] St5. Let the node marking the T-th word segment of the log to be marked among the child nodes of the parent node be the target child node, and determine whether the value of T is greater than or equal to the number of word segments of the log to be marked;

[0071] If not, let T be updated to T + 1, update the parent node to the target child node, and then return to step St4;

[0072] If yes, use the branch from the root node to the target child node as the log template of the log to be marked;

[0073] St6. Use the keyword algorithm to extract the keyword list of the log template, and classify the keyword list based on the known clustering results, and obtain the general template of the class where the keyword list is located as the marking result of the log to be marked.

[0074] Specifically, let the clustering result of the training set be denoted as the original class, and the minimum value of the known clustering results in step St6 is the original class;

[0075] If it is classified into the original class in St6, use the general template of the original class as the diary marking result;

[0076] If it is classified into a new class, for the new class, combine the keywords within the class to generate the label of the class as the diary marking result, and the label is used as the general template of the class.

[0077] In this embodiment, a clustering algorithm is used to further analyze and classify the extracted keyword list. By constructing a tree-like relationship graph and performing clustering calculations, the marking results of each group of similar log templates are finally determined. This automated marking method greatly improves the processing efficiency and accuracy, enabling a large amount of log data to be marked and classified quickly and accurately.

[0078] Specifically, assume that the word segments extracted from two logs, log 1 and log 2, on the training set are "java lang null" and "java lang exception" respectively.

[0079] Thus, as Figure 4 shown, when constructing the template tree, first establish the root node. Then, for log 1 in the training set, the number of word segments is counted as 3. A node marking the log length of 3 is set as the parent node at the first-level leaf node; a child node marking the word segment "java" is set below the parent node, a child node marking the word segment "lang" is set below the "java" node, and a child node marking the word segment "null" is set below the "lang" node. Thus, the path of log 1 in the training set is obtained: "root node - log length 3 - java - jlang - jnull";

[0080] For log 2 in the training set, the number of word segments is counted as 3. The first-level leaf node marking the log length of 3 already exists, and this node marking the log length of 3 is used as the parent node; there is a child node marking the word segment "java" below the parent node, and there is a child node marking the word segment "lang" below the "java" node. There is no "exception" in the child nodes of the "lang" node, so a child node marking the word segment "exception" is set below the "lang" node. Thus, the path of log 2 in the training set is obtained: "root node - log length 3 - java - jlang - exception".

[0081] As Figure 5 shown, assume that the word segments of the log to be marked are "receive from", and the number of word segments is 2; since there is no node marking the log length of 2 in the first-level leaf nodes of the training set template tree, a node marking the log length of 2 is set as the parent node at the first-level leaf node; a child node marking the word segment "receive" is set below the parent node, and a child node marking the word segment "from" is set below the "receive" node; thus, the path of this log to be marked is obtained: "root node - log length 2 - receive - from";

[0082] Suppose there is another word segmentation of the log to be tagged as "get from", and its word segmentation count is 2. Since the first-layer leaf nodes of the updated template tree have added nodes with a tagged log length of 2, the node with a tagged log length of 2 is used as the parent node. Since the child nodes of the parent node do not have "get", a child node is set under the parent node to tag the word segmentation "get". Since the child nodes of the "get" node do not have "from", a child node is set under the "get" node to tag the word segmentation "from". Thus, the path of the log to be tagged is obtained: "root node - log length 2 - get - from".

[0083] Of course, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0084] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0085] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.

Claims

1. A method for generating a log template, characterized in that, First, tokenize the logs in the training set. Using the tokens as leaf nodes, build a template tree that reflects the log content through a hierarchical structure, and merge the branches of the template tree by combining similarities. Take the path from the root node to the last token of the log as the log template and extract the keyword list. Cluster the keyword categories. For each category, combine all the keywords within the category to form a common template for the logs within the category.

2. The log template generation method according to claim 1, wherein The first-layer leaf nodes of the template tree are used to label the number of tokens in the logs. The tokens of each log are connected below the leaf nodes corresponding to the number of tokens.

3. The log template generation method according to claim 2, wherein The construction of the template tree includes the following steps: Build the root node, and make the child nodes of the root node be the first-layer leaf nodes. The first-layer leaf nodes are used to mark the number of tokens in the logs. Let the logs in the training set be the logs to be processed in turn. For the log to be processed, count the number of tokens. If there is a corresponding node among the first-layer leaf nodes, generate a branch representing the token sequence of the log to be processed below that node. If there is no corresponding node among the first-layer leaf nodes, add a new first-layer leaf node that marks the number of tokens of the log to be processed, and generate a branch representing the token sequence of the log to be processed below that node. On the template tree, the child nodes of the same parent node are different from each other. On the branches connected below each first-layer leaf node, the corresponding log tokens are arranged in a parent-child node relationship according to their order in the log.

4. The log template generation method according to claim 1, wherein The keyword extraction of the log template adopts the TextRank algorithm.

5. The log template generation method according to claim 1, characterized in that The logs in the training set are the log data after log preprocessing. Log preprocessing includes: data cleaning, format unification, and variable substitution.

6. A log tagging method based on a general log template using the log template generation method according to any one of claims 1-5, characterized in that First, use the log template generation method described in any one of claims 1-5 to obtain the template tree and the clustering results of the training set. Then obtain the log to be labeled and extract the tokens. Obtain the path corresponding to the log to be labeled in the template tree according to the number of tokens and the sorting. Extract the keyword list from this path, and classify the keyword list based on the known clustering results. Obtain the common template of the class where the keyword list is located as the labeling result of the log to be labeled.

7. The log tagging method based on a general log template according to claim 6, characterized in that, The method of obtaining the path corresponding to the log to be labeled in the template tree includes the following steps: Find or add a first-layer leaf node corresponding to the number of tokens of the log to be labeled in the template tree as the target node. Take the target node as the parent node, and find or add a child node corresponding to the first token of the log to be labeled below the parent node and record it as the new target node. Loop through this step, traverse the tokens of the log to be labeled until the tokens of the log to be labeled are inserted into the template tree in order in a parent-child node relationship. Then take the path from the root node to the leaf node where the last token of the log to be labeled is located as the path corresponding to the log to be labeled.

8. The log marking method based on a general log template according to claim 7, wherein, If the path of the log to be labeled coincides with an existing path on the template tree, directly obtain the common template of the class where the existing path is located as the labeling result of the log to be labeled. If the clustering result of the keyword list of the log to be labeled is a new class, build a common template for the log to be labeled based on the keyword list.

9. The log tagging method based on a general log template according to claim 6 or 7 or 8, characterized in that, Update the template tree in real time in combination with the path of the log to be labeled.

10. A log marking system based on a general log template, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. The processor is connected to the memory and is used to execute the computer program to implement the log tagging method based on a general log template as described in claim 6 or 7 or 8 or 9.