Method for generating log template, log analysis method and log analysis system

By building a parallel storage framework through a hash table library, the efficiency and accuracy issues of log parsing methods when processing large amounts of logs are solved, and efficient parallel processing and accurate log template generation are achieved.

CN120743872APending Publication Date: 2025-10-03SAMSUNG SEMICON CHINA RES & DEV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900607.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing log parsing methods based on prefix tree structures are difficult to process in parallel, resulting in increased running time when processing a large number of logs and inability to effectively parse and generate templates.

Method used

A parallel storage framework is constructed using a hash table library. Log entries are preprocessed, divided into buckets and sub-buckets, and a high-performance hash table library is used to store and match log templates. Multi-core processors and distributed computing resources are used for parallel processing.

Benefits of technology

It improves the speed and efficiency of log parsing, ensures accuracy and efficiency when processing massive logs at the TB or even PB level, and overcomes the defects of the prefix tree structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743872A_ABST
    Figure CN120743872A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating a log template, a log analysis method and a log analysis system. The method for generating the log template comprises the following steps: preprocessing a plurality of log entries to obtain preprocessed log entries expressed in a regularization manner; dividing the preprocessed log entry into a plurality of buckets based on the length of the preprocessed log entry; based on the first K lexical elements of the preprocessed log entries, the log entries in each bucket are divided into a plurality of sub-buckets, and K is an integer larger than or equal to 2; generating a sub-bucket log template set for each sub-bucket; storing the length of the log entry in each sub-bucket and the first K-1 lexical elements as keys of a hash table of each sub-bucket, and storing a sub-bucket log template set of each sub-bucket as a value of the hash table; and combining the hash tables of the plurality of sub-buckets into a hash table library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and more specifically, to a method for generating a log template, a log parsing method, and a log parsing system. Background Art

[0002] Log parsing is the process of extracting log templates and corresponding parameters from raw logs (or log entries). Parsed logs can be used for downstream tasks such as anomaly detection, fault prediction, and fault analysis.

[0003] Existing log parsing methods typically use a prefix tree structure to store intermediate and final results during the log parsing process. However, log parsing methods based on prefix trees struggle with parallel processing. When the number of logs is large, the runtime of these methods increases significantly, potentially preventing efficient parsing and template generation. Summary of the Invention

[0004] This Summary is provided to introduce in simplified form a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key features and / or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0005] According to one aspect of the present disclosure, a method for generating log templates is provided, the method comprising: preprocessing a plurality of log entries to obtain preprocessed log entries of regularized expressions; dividing the preprocessed log entries into a plurality of buckets based on the length of the preprocessed log entries; dividing the log entries in each bucket into a plurality of sub-buckets based on the first K tokens of the preprocessed log entries, where K is an integer greater than or equal to 2; generating a sub-bucket log template set for each sub-bucket; storing the length and the first K-1 tokens of the log entries in each sub-bucket as keys of a hash table for each sub-bucket, and storing the sub-bucket log template set of each sub-bucket as values ​​of the hash table; and merging the hash tables of the plurality of sub-buckets into a hash table library.

[0006] The step of preprocessing the multiple log entries may include: preprocessing the multiple log entries using a delimiter set and a regular expression set, wherein the method may further include: sampling the multiple log entries; and obtaining a delimiter set and a regular expression set based on the sampling results.

[0007] The step of sampling the multiple log entries may include: performing multiple samplings on the multiple log entries in parallel, and combining multiple sampling subsets obtained through the multiple samplings to obtain a sampled log set, wherein, for each sampling: selecting a first number of log entries from the multiple log entries as anchor log entries to obtain an anchor log set; randomly selecting a second number of log entries from the multiple log entries as candidate log entries to obtain a candidate log set; when the similarity between a candidate log entry in the candidate log set and each anchor log entry in the anchor log set meets a predetermined condition, adding the candidate log entry to the sampling subset.

[0008] The similarity may be a Jaccard distance, wherein when the Jaccard distance between the candidate log entry and each anchor log entry is within a predetermined interval, the candidate log entry is added to the sampling subset, wherein the first number of log entries are typical log entries among the plurality of log entries.

[0009] The step of obtaining a delimiter set and a regular expression set based on the sampling result may include: determining the delimiter set based on the frequency of occurrence of candidate delimiters in the sampled log set; and generating a regular expression set based on the sampled log set.

[0010] The step of determining the delimiter set may include: adding the candidate delimiter to the delimiter set when a ratio of the number of log entries including the candidate delimiter in the sampled log set to the total number of log entries in the sampled log set is greater than a first threshold; otherwise, discarding the candidate delimiter.

[0011] The step of generating a set of regularized expressions may include: identifying dynamic variables in log entries in the sampled log set based on heuristic rules; identifying static variables in log entries in the sampled log set based on the identified dynamic variables; and replacing the dynamic variables with placeholders.

[0012] The step of preprocessing the multiple log entries using a delimiter set and a regular expression set may include: identifying a delimiter for each of the multiple log entries based on the delimiter set; separating each of the multiple log entries into multiple tokens using the identified delimiter to obtain a log expression formatted as tokens; identifying dynamic variables in the log expression using the regular expression set; identifying static variables in the log expression based on the identified dynamic variables in the log expression; and replacing the dynamic variables with placeholders to generate preprocessed log entries.

[0013] The step of dividing the pre-processed log entries into a plurality of buckets based on the lengths of the pre-processed log entries may include: calculating the length of each of the pre-processed log entries; and clustering log entries having the same length into the same bucket.

[0014] The step of dividing the log entries in each bucket into multiple sub-buckets based on the first K tokens of the preprocessed log entries may include: for each bucket, if the first K tokens of the first log entry are the same as the first K tokens of the second log entry, clustering the first log entry and the second log entry into the same sub-bucket, where the size of K is inversely proportional to the number of sub-buckets.

[0015] The step of dividing the log entries in each bucket into multiple sub-buckets based on the first K words of the preprocessed log entries may also include: calculating the number of log entries in each sub-bucket of the multiple sub-buckets belonging to one bucket; when the first K-1 words of the log entries in the first sub-bucket are the same as the first K-1 words of the log entries in the second sub-bucket, and the sum of the number of log entries in the first sub-bucket and the number of log entries in the second sub-bucket is less than a second threshold, merging the first sub-bucket and the second sub-bucket into one sub-bucket.

[0016] The step of generating a sub-bucket log template set for each sub-bucket may include: for each sub-bucket, determining the first log entry in the sub-bucket as a template in the sub-bucket log template set; calculating the distance between each of the templates in the sub-bucket log template set and each of the remaining second log entries in the sub-bucket; when the distance is greater than a third threshold, adding the second log entry to the sub-bucket log template set.

[0017] According to one aspect of the present disclosure, a log parsing method is provided, comprising: preprocessing a target log entry to obtain a preprocessed target log entry of a regularized expression; forming a hash key using the length of the preprocessed target log entry and the first K-1 tokens, where K is an integer greater than or equal to 2; obtaining a corresponding sub-bucket log template set from a hash table library based on the hash key; selecting a template that matches the target log entry based on a similarity between the preprocessed target log entry and each log template in the sub-bucket log template set; and obtaining parsing information from the target log entry using the selected template.

[0018] The step of selecting a template that matches the target log entry may include: calculating the Jaccard distance between the preprocessed target log entry and each log template in the obtained sub-bucket log template set to generate a similarity list; selecting a minimum value in the similarity list; and determining the log template corresponding to the minimum value as the template that matches the target log entry.

[0019] The step of pre-processing the target log entry may include: pre-processing the target log entry using a delimiter set and a regularization expression set, wherein the delimiter set and the regularization expression set are obtained based on sampling a plurality of training sample logs.

[0020] According to one aspect of the present disclosure, a log parsing system is provided, comprising: a memory configured to store a hash table library; a preprocessing module configured to preprocess a target log entry to obtain a preprocessed target log entry of a regularized expression; a parsing module configured to form a hash key using the length of the preprocessed target log entry and the first K-1 tokens, where K is an integer greater than or equal to 2; obtain a corresponding sub-bucket log template set from the hash table library based on the hash key; select a template that matches the target log entry based on a similarity between the preprocessed target log entry and each log template in the sub-bucket log template set; and obtain parsing information from the target log entry using the selected template.

[0021] The parsing module can be configured to: calculate the Jaccard distance between the preprocessed target log entry and each log template in the obtained sub-bucket log template set to generate a similarity list; select the minimum value in the similarity list; and determine the log template corresponding to the minimum value as the template that matches the target log entry.

[0022] The preprocessing module may be configured to preprocess the target log entry using a delimiter set and a regular expression set.

[0023] The separator set and the regularization expression set may be obtained based on sampling a plurality of training sample logs.

[0024] According to one aspect of the present disclosure, a non-transitory computer-readable storage medium stores instructions, which, when executed by a processor, causes the processor to perform a log parsing method according to an example embodiment of the present disclosure.

[0025] Additional aspects and / or advantages of the present inventive concepts will be set forth in part in the description which follows and, in part, will be apparent from the description, and / or may be learned by practice of various exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings.

[0027] Figure 1 is a diagram illustrating an example of a log entry.

[0028] Figure 2is a flowchart illustrating a method of generating a log template according to an example embodiment of the present disclosure.

[0029] Figure 3 is a flowchart illustrating a method of generating a log template according to an example embodiment of the present disclosure.

[0030] Figure 4 is a diagram illustrating a log parsing method according to an exemplary embodiment of the present disclosure.

[0031] Figure 5 is a pseudo code illustrating a log sampling method according to an example embodiment of the present disclosure.

[0032] Figure 6 is a pseudo code illustrating a method of generating a log template according to an exemplary embodiment of the present disclosure.

[0033] Figure 7 is a flowchart illustrating a log parsing method according to an example embodiment of the present disclosure.

[0034] Figure 8 This is a table showing the accuracy indicators of various log analysis methods.

[0035] Figure 9 is a diagram illustrating the running time of a log parsing method according to an exemplary embodiment of the present disclosure and a comparative example.

[0036] Figure 10 is a diagram illustrating a log parsing system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] The following detailed description is provided to help the reader gain a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be clear after understanding the disclosure of the present application. For example, except for operations that must occur in a specific order, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as will be clear after understanding the disclosure of the present application. In addition, for greater clarity and brevity, descriptions of features known in the art may be omitted.

[0038] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein have been provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein that will be apparent upon understanding the disclosure of this application.

[0039] The structural or functional description of the examples disclosed herein below is intended only for the purpose of describing the examples, and the examples may be implemented in various forms. The examples are not intended to be limiting, but rather to encompass various modifications, equivalents, and substitutes within the scope of the claims.

[0040] Although the terms "first" or "second" are used to explain various components, the components are not limited to the terms. These terms should only be used to distinguish one component from another. For example, within the scope of the rights according to the concept of the present disclosure, a "first" component may be referred to as a "second" component, or similarly, a "second" component may be referred to as a "first" component.

[0041] It will be understood that when a component is referred to as being “connected to” another component, it can be directly connected or coupled to the other component or intervening components may be present.

[0042] As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. It should also be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of stated features, integers, steps, operations, elements, components, or combinations thereof, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0043] Unless otherwise defined, all terms used herein (including technical or scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the examples belong. It will also be understood that, unless expressly defined as such herein, terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense.

[0044] Hereinafter, examples will be described in detail with reference to the accompanying drawings. With regard to reference numerals assigned to elements in the drawings, it should be noted that like elements will be denoted by like reference numerals, and redundant descriptions thereof will be omitted.

[0045] Figure 1 is a diagram illustrating an example of a log entry.

[0046] Figure 1 The table in Figure 1 shows four typical log entries generated by a software system. Each log entry may include an identifier (ID) and a log message. Figure 1 Four types of log entries are shown, but the types and numbers of log entries are not limited thereto.

[0047] Log parsing methods are a crucial link between raw log entries and downstream analysis. They involve using delimiters to accurately separate log entries, employing regular expressions to distinguish dynamic and static variables within the raw logs, replacing dynamic variables with placeholders while preserving static variables, and finally clustering multiple similar processed logs into unique templates.

[0048] Existing log parsing methods typically utilize prefix tree structures to store intermediate and final results during the log parsing process. The fundamental concept behind prefix tree structures is to leverage large storage capacity to reduce execution time. However, when faced with massive log volumes, such as terabytes (TB) or even petabytes (PB), hardware resource limitations significantly restrict the storage and operational capabilities of prefix tree structures. Furthermore, as a typical tree data structure consisting of a root node and multiple layers of child nodes, prefix tree structures inherently lack parallel processing capabilities. Consequently, when processing massive log volumes, such as TB or even PB, the runtime of log parsing methods based on prefix tree structures increases exponentially. This can lead to inefficiencies in parsing and generating templates.

[0049] With the development of artificial intelligence (AI) technology, log analysis systems that utilize AI algorithm models have attracted significant attention within the industry. Among these AI-based log analysis systems, mainstream log parsing methods can be divided into two categories: semantic-based deep learning methods and grammar-based heuristic methods.

[0050] Semantic-based deep learning methods (particularly those utilizing large language models (LLMs)) parse logs by exploring the inherent semantic relationships within them. However, these methods require computing embeddings for the log information. Applying deep learning, or even LLMs, to derive these embeddings is often time-consuming. Consequently, semantic-based methods are less effective when dealing with massive log volumes, typically terabytes or even petabytes.

[0051] In industrial scenarios involving large-scale log parsing, grammar-based heuristic methods demonstrate significant advantages due to their high computational efficiency. In real-world industrial operations, raw logs are continuously generated in a streaming fashion. Since grammar-based heuristic methods typically utilize batch offline clustering techniques (e.g., the K-means algorithm), they have limitations when processing logs that are updated in real time. To ensure real-time computational efficiency, grammar-based heuristic methods often ignore semantic relationships between logs. As a result, grammar-based heuristic methods are relatively less accurate than semantic-based methods.

[0052] Therefore, it is desirable to provide a method that can efficiently process a large number of log entries while ensuring the accuracy of log parsing.

[0053] This paper utilizes a high-performance hash table library to build a parallel storage framework, effectively replacing the prefix tree structure commonly used in log parsing methods. The hash table-based log template generation and log parsing methods are designed to fully utilize multi-core processors and distributed computing resources, thereby improving the speed and efficiency of log parsing methods.

[0054] Figure 2 is a flowchart illustrating a method of generating a log template according to an example embodiment of the present disclosure.

[0055] Reference Figure 2 , the method of generating a log template according to the present disclosure may include operations 210 to 260 .

[0056] In operation 210 , a plurality of log entries may be preprocessed to obtain preprocessed log entries of regularized expressions.

[0057] The plurality of log entries may be pre-processed using a predetermined set of delimiters and a predetermined set of regularization expressions.

[0058] Log entries often contain a large number of special symbols (e.g., spaces, commas, quotation marks, semicolons, colons, equal signs, brackets, etc.) These symbols can be defined as delimiters, and delimiters can be used to divide each log entry into multiple tokens.

[0059] In one example, a delimiter set may be defined by a user, and a delimiter for each of the plurality of log entries may be identified based on the delimiter set. Each of the plurality of log entries may be separated into a plurality of tokens using the identified delimiter to obtain a log expression formatted as tokens.

[0060] A regular expression set may be used to identify dynamic variables in a log expression, a placeholder <*> may be used to replace the dynamic variables, and the remaining variables in the log expression may be identified as static variables, thereby obtaining a preprocessed log.

[0061] In operation 220, the pre-processed log entries may be divided into a plurality of buckets based on their lengths. In one example, the length of each of the pre-processed log entries (e.g., the number of tokens in each log entry) may be calculated, and log entries with the same length may be clustered into the same bucket.

[0062] In operation 230 , the log entries in each bucket may be further divided into a plurality of sub-buckets based on the first K tokens of the pre-processed log entries.

[0063] When multiple log entries are divided into multiple buckets, each bucket can be operated on independently and in parallel. For example, when processing a large number of log entries, multi-core processors and / or distributed computing resources can be used to perform operations on each bucket in parallel. For each bucket, if the first K tokens of one log entry are the same as the first K tokens of another log entry, the two log entries are clustered into the same sub-bucket.

[0064] Because sub-buckets that do not contain sufficient log information will result in poor performance in generating log templates, in some cases, sparse sub-buckets can be further merged.

[0065] In one example, the number of log entries in each of the multiple sub-buckets included in a bucket may be calculated. and Share the same first K-1 tokens, and the sum of the sizes of the two sub-buckets (e.g., the sum of the number of log entries included in the two sub-buckets) is less than a predetermined threshold , then these two sub-buckets can be merged into a new sub-bucket .

[0066] In operation 240, a sub-bucket log template set may be generated for each sub-bucket. Each sub-bucket can be operated in parallel and independently. For each sub-bucket, the first log entry in the sub-bucket can be The template in the sub-bucket log template set is determined. Then, each template in the sub-bucket log template set can be calculated. Each of the remaining log entries in the sub-bucket For example, the Jaccard distance can be used to measure similarity. When the distance between the sub-bucket template and the log entry (for example, the Jaccard distance) When the value is greater than the predetermined threshold sim, the log entry can be Stored as a new template in the sub-bucket log template collection.

[0067] In operation 250, the length of the log entry in each sub-bucket and the first K-1 tokens may be stored as keys in a hash table, and the sub-bucket log template set for each sub-bucket may be stored as values ​​in the hash table. A key-value pair may be generated for each sub-bucket, where the key may be defined as: Hash(log length & first K-1 tokens), and the value may correspond to the corresponding sub-bucket log template set. .

[0068] In operation 260 , the hash tables of the plurality of sub-buckets may be merged into a hash table base HT.

[0069] In industrial scenarios, log entries contain a large amount of special symbols and semi-structured content generated by various systems and software. Without customized processing for these unique special symbols and semi-structured content, the effectiveness of log parsing methods is reduced. For example, commonly used delimiter sets and / or regular expression combinations are unable to reflect all delimiters / regular expressions that appear in massive log entries, leading to errors in log segmentation and variable identification.

[0070] This disclosure provides a method for generating log templates based on sampling. By performing multiple sampling operations, the difference between the sampled log distribution and the global log distribution can be reduced, thereby creating a precise set of delimiters and regularization expressions, thereby improving the accuracy of log parsing.

[0071] Figure 3 is a flowchart illustrating a method of generating a log template according to an example embodiment of the present disclosure.

[0072] Reference Figure 3 , the method of generating a log template according to the present disclosure may include operations 310 to 360 .

[0073] In operation 310 , sampling may be performed on a plurality of log entries. Multiple sampling operations may be performed on the plurality of log entries in parallel, and a plurality of sampling subsets obtained through the multiple sampling operations may be combined to obtain a sampled log set.

[0074] In one example, S sampling operations may be performed on multiple log entries in parallel. For each sampling operation, a first number of log entries may be selected from the multiple log entries as anchor log entries to obtain an anchor log set, and a second number of log entries may be randomly selected from the multiple log entries as candidate log entries to obtain a candidate log set. A similarity may be calculated between each candidate log entry in the candidate log set and each anchor log entry in the anchor log set. When the similarity meets a predetermined condition, the candidate log entry may be included in the sampling subset.

[0075] An anchor log entry can be a typical log entry among multiple log entries. For example, based on expert experience, log type, log generation source, etc., a user can select 100 different log entries as anchor log entries to form an anchor log set. In addition, 0.01% of the total number of log entries can be randomly selected each time to form a candidate log set. .

[0076] Each candidate log entry in the candidate log set can be calculated With each anchor log entry in the anchor log collection The following equation (1) shows the calculation of candidate log entries based on the Jaccard distance Log entries with anchor The similarities between them.

[0077] (1) In equation (1), represents the Jaccard distance between the anchor log entry and the candidate log entry, i represents the index of the anchor log, and j represents the index of the candidate log entry.

[0078] When the similarity between the candidate log entry and the anchor log entry satisfies a predetermined condition, the candidate log entry may be included in the sampling subset, otherwise, the candidate log entry may be discarded.

[0079] In order to make the sampled candidate log entries reflect the distribution characteristics of global log entries as much as possible and maintain a balance between consistency and diversity among the sampled candidate log entries, an interval defined by an upper bound (UB) and a lower bound (LB) can be determined. The upper bound ensures diversity among the selected log entries, and the lower bound ensures consistency among the selected log entries. Therefore, if the Jaccard distance is within the interval [LB, UB], the candidate log entry can be selected. Included in the sampling subset Otherwise, discard .

[0080] Multiple sampling subsets obtained through S sampling can be combined to obtain a sampling log set .

[0081] In operation 320 , a delimiter set and a regularization expression set may be obtained based on the sampling result.

[0082] In one example, a set of separators is determined based on the frequency of candidate separators appearing in a set of sampled logs. V , then traverse the candidate set V Each delimiter in , calculation includes The number of log entries is The ratio between the total number of log entries in .if If is greater than the threshold T, then Put into the delimiter set μ Otherwise, the delimiter is discarded.

[0083] Generate a set of regularized expressions based on a set of sampled logs R In one example, a sample log set may be identified based on heuristic rules. Dynamic variables in log entries in the; based on the identified dynamic variables, the sampled log set can be identified Static variables in log entries in ; and dynamic variables can be replaced with placeholders.

[0084] In one example, a specific pattern such as a memory address, IP address, file path, email address, etc. After identifying the dynamic variables, the remaining elements in the log entry can be identified as static variables. Then, the dynamic variables can be replaced with the placeholder <*> to obtain the set of regular expressions R .

[0085] Next, in operation 330, the delimiter set obtained in operation 320 may be used to μ and a set of regular expressions R Preprocessing is performed on a plurality of log entries to obtain preprocessed log entries.

[0086] Figure 3 Operations 340 to 380 in the Figure 2 Operations 220 to 260 in FIG. 2 correspond to each other one by one. Therefore, repeated descriptions are omitted.

[0087] Figure 4 is a diagram illustrating a log parsing method according to an exemplary embodiment of the present disclosure. Figure 5 is a pseudo code illustrating a log sampling method according to an example embodiment of the present disclosure. Figure 6 is a pseudo code illustrating a method of generating a log template according to an exemplary embodiment of the present disclosure.

[0088] Figure 4 The process of generating a delimiter set and a regular expression set through a bounded multiple sampling module, generating a hash table library through a two-stage log template generation module, and outputting the log parsing results is shown below. Figure 5 and Figure 6 To describe Figure 4 .

[0089] In operation 410 , a sampling log set may be generated using bounded multiple samplings.

[0090] Reference Figure 5 , S samplings can be performed on multiple log entries in parallel, and multiple sampling subsets obtained through S samplings are combined to obtain a sampled log set (see Figure 5 1 to 9 in the .

[0091] In operation 420, a set of separators may be generated based on statistical analysis of the sampled log set (see Figure 5 10 to 13 in the .

[0092] In operation 430, a set of regularized expressions may be generated from the sampled log set based on heuristic rules (see Figure 5 14 in the .

[0093] In operation 440 , the original log may be processed into tokens using the delimiter set and the regularization expression set obtained in operations 420 to 430 to obtain processed log entries of regularized expressions.

[0094] In operation 450, the plurality of log entries may be divided into a plurality of buckets according to the length of each of the processed log entries (see Figure 6 1 in the .

[0095] In operation 460, the log entries in each bucket may be further divided into a plurality of sub-buckets based on the first K words (see Figure 6 2 to 4 in ).

[0096] In operation 470, when the log information of the sub-bucket is insufficient, the two sub-buckets that share the first K-1 tokens can be merged into a new sub-bucket (see Figure 6 5 to 8 in the .

[0097] In operation 480, for each sub-bucket, a sub-bucket log template set of each sub-bucket may be generated based on the Jaccard distance and stored in a hash table (see Figure 6 9 to 16 in the .

[0098] After obtaining the hash table library, the hash table library can be used to parse the newly generated logs.

[0099] Figure 7 is a flowchart illustrating a log parsing method according to an example embodiment of the present disclosure.

[0100] In operation 710, the target log entry may be preprocessed to obtain a preprocessed target log entry of a regularized expression. A separator set obtained based on multiple samplings may be used. μ and a set of regular expressions R Preprocess the target log entries.

[0101] In operation 720 , a hash key may be formed using the length of the pre-processed target log entry and the first K−1 tokens.

[0102] In operation 730, the corresponding sub-bucket log template set can be obtained from the hash table library HT based on the hash key. .

[0103] In operation 740, a template matching the target log entry may be selected based on the similarity between the pre-processed target log entry and each log template in the sub-bucket log template set. In one example, the pre-processed target log entry may be calculated. Each log template in the obtained sub-bucket log template collection Jaccard distance between them to generate a similarity list A minimum value in the similarity list may be selected, and the log template corresponding to the minimum value may be determined as the template matching the target log entry.

[0104] In operation 750 , parsed information may be obtained from the target log entry using the selected template.

[0105] Figure 8 This is a table showing the accuracy indicators of various log analysis methods.

[0106] Because incorrectly parsed log templates can significantly degrade the overall performance of AI-based log analysis systems, evaluating the accuracy (or effectiveness) of log parsing methods is crucial. This evaluation employs two commonly used metrics: the Kalinsky-Harabaschi Index (CHI) (also known as the variance ratio criterion) and the Davis-Boulding Index (DBI). A higher CHI value indicates better clustering performance, meaning that data points are more dispersed between classes than within them. A lower DBI reflects better clustering.

[0107] Reference Figure 8 , compared with other log parsing methods in the prior art, the log parsing method according to the example embodiment of the present disclosure performs well in both evaluation indicators. It is worth noting that relative to the second-place Drain method, the log parsing method according to the example embodiment of the present disclosure shows a 20% enhancement in both indicators. Even more striking is that it far exceeds other recognized log parsing methods. Therefore, the log parsing method according to the example embodiment of the present disclosure can effectively meet the actual needs of log parsing in engineering environments.

[0108] Figure 9 is a diagram illustrating the running time of a log parsing method according to an exemplary embodiment of the present disclosure and a comparative example.

[0109] The log parsing method according to the exemplary embodiments of the present disclosure can be applied to multi-core processors and / or distributed resources. When the number of log entries to be processed is large, the log parsing method based on distributed parallel processing presents a significant parsing efficiency advantage.

[0110] Reference Figure 9, the log parsing method according to the example embodiment of the present disclosure can use 10 CPU cores to perform distributed processing. As the number of log entries increases (for example, from 9M to 24M+ entries), the log parsing method according to the example embodiment of the present disclosure shows a linear increase in running time, while the running time of the Drain method shows an exponential increase. When the number of log entries is 9 million, the parsing speed of the log parsing method according to the example embodiment of the present disclosure is 11.5 times that of the Drain method. When the number of log entries is 21 million, the parsing speed gap between the two widens to 27.6 times. It is worth noting that the Drain method encounters the structural limitations of its prefix tree design and crashes when the number of log entries exceeds 24 million, while the log parsing method according to the example embodiment of the present disclosure successfully completes the parsing in 2.9 minutes with this number of logs.

[0111] It can be seen that the log template generation method and log parsing method according to the example embodiments of the present disclosure can make full use of multi-core processors and distributed computing to achieve parallel processing, and use a high-performance hash table library to build a parallel storage framework, so that the solution proposed in the present disclosure has high processing efficiency while ensuring the accuracy of log parsing.

[0112] Figure 10 is a diagram illustrating a log parsing system according to an exemplary embodiment of the present disclosure.

[0113] Reference Figure 10 According to an exemplary embodiment of the present disclosure, the log parsing system 1000 may include a pre-processing module 1010 , a parsing module 1020 , and a memory 1030 .

[0114] The memory 1030 may include volatile memory and / or non-volatile memory. The memory 1030 may store instructions, applications, etc. for executing the method for generating log templates and the log parsing method. The memory 1030 may also store a hash table library for log parsing.

[0115] The preprocessing module 1010 may receive a log entry to be parsed. The preprocessing module 1010 may preprocess the log entry to obtain a preprocessed log entry with a regularized expression. For example, the preprocessing module 1010 may preprocess the log entry using a set of delimiters and a set of regularized expressions, where the set of delimiters and the set of regularized expressions are obtained by sampling a plurality of training sample logs.

[0116] Parsing module 1020 can parse the preprocessed log entry. Parsing module 1020 can use the length of the preprocessed log entry and the first K-1 tokens to form a hash key. Parsing module 1020 can refer to the hash table library stored in memory 1030 and obtain a corresponding sub-bucket log template set based on the hash key. Parsing module 1020 can select a template that matches the target log entry based on the similarity between the preprocessed log entry and each log template in the sub-bucket log template set. Parsing module 1020 can use the selected template to obtain parsed information from the log entry.

[0117] The method for generating log templates according to an exemplary embodiment of the present disclosure utilizes a hash table parallel framework to overcome the shortcomings of traditional prefix tree structures. Because each bucket or sub-bucket operates independently of the others, the hash table parallel framework facilitates concurrent execution during log template generation. Furthermore, during log template generation, a two-stage process of bucket and sub-bucket division allows for further refinement and adjustment of the generated sub-buckets, which improves the accuracy of the generated log template.

[0118] The log template generation method according to an exemplary embodiment of the present disclosure employs multiple bounded sampling, where an upper bound (UB) is used to adjust for diversity and a lower bound (LB) is used to ensure consistency. Therefore, compared to traditional random sampling techniques, which are fraught with uncertainty, bounded multiple sampling significantly reduces the difference between the sampled log distribution and the global log distribution, thereby improving overall sampling accuracy. Furthermore, the set of delimiters and regularized expressions obtained based on sampling can improve the accuracy of log template generation and log parsing.

[0119] The devices, units, modules, and other components described herein are implemented by hardware components. Examples of hardware components that can be used to perform the operations described herein include, where appropriate, controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described herein. In other examples, one or more of the hardware components that perform the operations described herein are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer may be implemented by one or more processing elements (such as logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond to and execute instructions in a defined manner to achieve a desired result). In one example, the processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer may execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) for performing the operations described herein. Hardware components can also access, manipulate, process, create, and store data in response to the execution of instructions or software. For simplicity, the singular term "processor" or "computer" may be used in the description of the examples described in this application, but in other examples, multiple processors or computers may be used, or the processor or computer may include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component, or two or more hardware components, may be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components may be implemented by one or more processors, or a processor and a controller, and one or more other hardware components may be implemented by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, may implement a single hardware component, or two or more hardware components. Hardware components may have any one or more of different processing configurations, examples of which include: a single processor, an independent processor, a parallel processor, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.

[0120] The methods for performing the operations described herein are performed by computing hardware (e.g., by one or more processors or computers) that is implemented to execute instructions or software as described above to perform the operations described herein performed by the methods. For example, a single operation, or two or more operations, may be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be performed by one or more processors, or a processor and a controller, and one or more other operations may be performed by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, may perform a single operation, or two or more operations.

[0121] The instructions or software for controlling a processor or computer to implement the hardware components and perform the methods described above may be written as a computer program, code segments, instructions, or any combination thereof to individually or collectively instruct or configure the processor or computer to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and methods described above. In one example, the instructions and / or software include machine code (such as machine code generated by a compiler) that is directly executed by the processor or computer. In another example, the instructions or software include high-level code that is executed by the processor or computer using an interpreter. A person of ordinary skill in the art or a programmer can easily write instructions and / or software based on the block diagrams and flow charts shown in the accompanying drawings and the corresponding descriptions in the specification, which disclose algorithms for performing the operations performed by the hardware components and methods described above.

[0122] The instructions or software for controlling a processor or computer to implement the hardware components and perform the methods described above, as well as any associated data, data files, and data structures, are recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random-access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random-access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R At least one of an LTH, BD-RE, Blu-ray or optical disc storage device, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a card memory (such as a multimedia card or a micro card (e.g., Secure Digital (SD) or Extreme Digital (XD))), a magnetic tape, a floppy disk, a magneto-optical data storage device, an optical data storage device, a hard disk, a solid state disk, and any other device, any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the instructions.

[0123] Although various example embodiments have been described, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the claims.

Claims

1. A method for generating a log template, comprising: Preprocessing the plurality of log entries to obtain preprocessed log entries of regularized expressions; Divide the preprocessed log entries into multiple buckets based on their lengths; Based on the first K tokens of the preprocessed log entries, the log entries in each bucket are divided into multiple sub-buckets, where K is an integer greater than or equal to 2; Generate a sub-bucket log template set for each sub-bucket; Store the length and first K-1 tokens of the log entry in each sub-bucket as the key of the hash table of each sub-bucket, and store the sub-bucket log template set of each sub-bucket as the value of the hash table; The hash tables of the multiple sub-buckets are merged into a hash table library.

2. The method according to claim 1, in, The step of preprocessing the plurality of log entries comprises: preprocessing the plurality of log entries using a set of separators and a set of regular expressions, The method further comprises: sampling the plurality of log entries; A delimiter set and a regularization expression set are obtained based on the sampling result.

3. The method according to claim 2, in, The step of sampling the plurality of log entries includes: performing a plurality of sampling operations on the plurality of log entries in parallel, and combining a plurality of sampling subsets obtained by the plurality of sampling operations to obtain a sampled log set. Among them, for each sample: selecting a first number of log entries from the plurality of log entries as anchor log entries to obtain an anchor log set; randomly selecting a second number of log entries from the plurality of log entries as candidate log entries to obtain a candidate log set; When the similarity between a candidate log entry in the candidate log set and each anchor log entry in the anchor log set satisfies a predetermined condition, the candidate log entry is added to the sampling subset.

4. The method according to claim 1, wherein The step of pre-processing the plurality of log entries using a set of delimiters and a set of regular expressions comprises: identifying a delimiter for each of the plurality of log entries based on a delimiter set; separating each of the plurality of log entries into a plurality of tokens using the identified delimiter to obtain a log expression formatted as tokens; Use a set of regular expressions to identify dynamic variables in log expressions; Identifying static variables in the log expression based on the identified dynamic variables in the log expression; and Use placeholders to replace dynamic variables to generate preprocessed log entries.

5. The method according to claim 1, wherein The steps of dividing the preprocessed log entries into a plurality of buckets based on the length of the preprocessed log entries include: Calculate the length of each of the preprocessed log entries; Cluster log entries with the same length into the same bucket.

6. The method according to claim 1, wherein The steps of dividing the log entries in each bucket into multiple sub-buckets based on the first K tokens of the preprocessed log entries include: For each bucket, if the first K tokens of the first log entry are the same as the first K tokens of the second log entry, the first log entry and the second log entry are clustered into the same sub-bucket. The size of K is inversely proportional to the number of sub-buckets.

7. The method according to claim 1, wherein The steps of generating a sub-bucket log template set for each sub-bucket include: for each sub-bucket, Determine the first log entry in the sub-bucket as a template in the sub-bucket log template set; Calculating a distance between each of the templates in the sub-bucket log template set and each of the remaining second log entries in the sub-bucket; When the distance is greater than a third threshold, the second log entry is added to the sub-bucket log template set.

8. A log parsing method, comprising: Preprocessing the target log entry to obtain a preprocessed target log entry of a regularized expression; Use the length of the preprocessed target log entry and the first K-1 tokens to form a hash key, where K is an integer greater than or equal to 2; Obtaining a corresponding sub-bucket log template set from a hash table library based on the hash key; Based on the similarity between the preprocessed target log entry and each log template in the sub-bucket log template set, a template matching the target log entry is selected; Gets parsed information from the target log entry using the selected template.

9. A log parsing system, comprising: A memory configured to store a hash table library; The preprocessing module is configured to: preprocess the target log entry to obtain a preprocessed target log entry of a regularized expression; The parsing module is configured as follows: Use the length of the preprocessed target log entry and the first K-1 tokens to form a hash key, where K is an integer greater than or equal to 2; Obtaining a corresponding sub-bucket log template set from a hash table library based on the hash key; Based on the similarity between the preprocessed target log entry and each log template in the sub-bucket log template set, a template matching the target log entry is selected; Gets parsed information from the target log entry using the selected template.

10. A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, cause the processor to perform the method according to claim 1.