Log analysis method and device, equipment, medium and program product
By using a large language model and preset prompt words to classify and correct log features, and by combining a character index tree to optimize log templates, the problem of poor reliability in log parsing using the large language model is solved, and efficient and accurate log parsing is achieved.
Patent Information
- Application Number
- CN202511007095.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-14
AI Technical Summary
Log parsing techniques based on large language models suffer from poor reliability of parsing results, lack effective correction methods, and existing techniques suffer from high computational complexity and low efficiency.
By using a large language model to parse the target logs, extract constant and variable features, and use preset prompt words to determine the category, the marked category in the initial log template is corrected, and the log template is optimized by combining character index trees and semantic analysis.
It improves the accuracy and reliability of log parsing results, reduces computational complexity, and enhances system performance and efficiency.
Smart Images

Figure CN120950933A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a log parsing method, apparatus, device, medium, and program product. Background Technology
[0002] In log parsing techniques based on Large Language Models (LLMs), each log is parsed only once, and there is a lack of effective correction methods when errors occur during parsing. However, it is not uncommon for LLMs to produce errors during parsing, and the relevant techniques do not have corresponding effective correction methods to handle these errors, resulting in poor reliability of log parsing results. Summary of the Invention
[0003] This application provides a log parsing method, apparatus, device, medium, and program product to solve the problem of poor reliability of related log parsing results.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a log parsing method, including:
[0006] The target log is parsed using a large language model to obtain the initial log template of the target log;
[0007] Feature extraction is performed on the constants and / or variables in the target log to obtain the target features;
[0008] The target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. The first preset prompt word is used to represent the category classification rule.
[0009] If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0010] Secondly, embodiments of this application provide a log parsing apparatus, including:
[0011] The parsing module is used to parse the target log using a large language model to obtain the initial log template of the target log;
[0012] The first extraction module is used to extract features from the constants and / or variables in the target log to obtain target features;
[0013] The judgment module is used to use the large language model and the first preset prompt word to judge the category of the target feature and obtain the category of the target feature. The first preset prompt word is used to represent the category judgment rule.
[0014] The modification module is used to modify the tag category when the category of the target feature is inconsistent with the tag category of the initial log template, so as to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0015] Thirdly, embodiments of this application provide an electronic device, including a processor, the processor being used for:
[0016] The target log is parsed using a large language model to obtain the initial log template of the target log;
[0017] Feature extraction is performed on the constants and / or variables in the target log to obtain the target features;
[0018] The target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. The first preset prompt word is used to represent the category classification rule.
[0019] If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0020] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the log parsing method described in the first aspect above.
[0021] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the log parsing method described in the first aspect above.
[0022] In a sixth aspect, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the log parsing method described in the first aspect above.
[0023] In this embodiment, a large language model is used to parse the target log to obtain an initial log template. Features are extracted from the constants and / or variables in the target log to obtain target features. The large language model and a first preset prompt word are used to determine the category of the target features, where the first preset prompt word represents the category determination rule. If the category of the target feature is inconsistent with the labeled category of the initial log template, the labeled category is changed to obtain the target log template, where the labeled category is the category of the target feature indicated by the initial log template. Thus, by using the large language model and the first preset prompt word to determine the category of the target feature, the labeled category in the initial log template can be validated and modified, thereby continuously optimizing and correcting the initial log template, which improves the accuracy and reliability of the log parsing results. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a log parsing method provided in an embodiment of this application;
[0026] Figure 2 This is a schematic diagram illustrating a category determination based on a large language model, provided in an embodiment of this application.
[0027] Figure 3 This is a flowchart illustrating a log template error correction method provided in an embodiment of this application;
[0028] Figure 4 This is a flowchart illustrating a category determination method provided in an embodiment of this application;
[0029] Figure 5 This is a schematic diagram illustrating semantic judgment based on a large language model, provided in an embodiment of this application.
[0030] Figure 6 This is a flowchart of a merged log template provided in an embodiment of this application;
[0031] Figure 7 This is a flowchart of an initial log parsing provided in an embodiment of this application;
[0032] Figure 8 This is a schematic diagram illustrating log parsing based on a large language model, provided in an embodiment of this application.
[0033] Figure 9 This is a schematic diagram of a character index tree provided in an embodiment of this application;
[0034] Figure 10 This is a schematic diagram of character matching provided in an embodiment of this application;
[0035] Figure 11 This is a flowchart of a character index tree matching method provided in an embodiment of this application;
[0036] Figure 12 This is a flowchart illustrating a matching log template provided in an embodiment of this application;
[0037] Figure 13 This is a flowchart of a log parsing method provided in an embodiment of this application;
[0038] Figure 14 This is a schematic diagram of the structure of a log parsing device provided in an embodiment of this application;
[0039] Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0041] For ease of understanding, the following describes some aspects of the embodiments of this application:
[0042] Log message parsing based on regular expressions is a well-known technology in the field of information technology. Log files typically consist of two parts: a log header and a log message. The log header contains strictly formatted information such as timestamps and log levels, while the log message contains a detailed description of the log and is the primary object of log parsing. The task of log parsing is to extract patterns and separate variables from the log message. Specifically, separating the log header and log message is relatively simple, as the log header's format is relatively fixed and can be easily extracted using predefined rules. However, the log message contains a large amount of unstructured or semi-structured data, which requires further parsing to extract useful information. Regular expressions, as a pattern matching tool, are applied in log message parsing. Through manually defined regular expressions, fixed patterns in log messages can be identified, and variables can be extracted from them. For example, a log message may contain specific event descriptions, user IDs, operation objects, and other information. Regular expressions can define matching patterns for this information, thereby separating variables from the log message.
[0043] The related technologies have the following technical problems:
[0044] (1) Limitations in handling misidentification: In the log parsing technology based on large language models, each log parsing process queries the large language model once. However, it is not uncommon for large language models to produce errors during the parsing process. When faced with the situation of erroneous parsing of large language models, the related technologies do not have corresponding effective correction methods to handle these errors, which affects the reliability of the parsing results.
[0045] (2) Limitations of merging existing templates: Related techniques use the Longest Common Subsequence (LCS) technique to merge the parsing results of the current large model with existing templates, but this method has a disconnect. LCS is based on statistical features, while the parsing results of large models are based on semantics. Using LCS similarity technology based on statistical features to determine whether to merge will weaken the high effectiveness of using large language models for semantic analysis to some extent.
[0046] (3) Limitations of regular expression-based matching: Related technologies mainly rely on template matching based on regular expressions. This method requires the current log to be matched against each predefined template one by one. This matching process is not only computationally complex, but also inefficient when processing large amounts of log data, which significantly affects system performance.
[0047] In this application embodiment, a log parsing method, apparatus, device, medium, and program product are proposed to solve the problem of poor reliability of related log parsing results.
[0048] See Figure 1 , Figure 1 This is a flowchart of a log parsing method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0049] Step 101: Use a large language model to parse the target log to obtain the initial log template of the target log.
[0050] In this step, the large language model described above can be used to automatically extract structured information from logs to obtain corresponding log templates. The target log can be the raw log data that needs to be parsed, such as system logs, application logs, or network logs. The initial log template can be a structured template generated after parsing by the large language model.
[0051] Step 102: Extract features from the constants and / or variables in the target log to obtain target features.
[0052] It is understood that the aforementioned target features may include multiple contents from the target log, which are used for subsequent category determination.
[0053] Step 103: Use the large language model and the first preset prompt word to determine the category of the target feature and obtain the category of the target feature. The first preset prompt word is used to represent the category determination rule.
[0054] In this step, the first preset prompt can be a predefined rule or instruction used to guide the large language model in classification, and can be determined based on predefined rules based on expert experience. For example, the first preset prompt could be "variables are usually dynamic values, while constants are usually fixed values".
[0055] The above category determination can be to determine whether the target feature belongs to a constant or a variable, that is, the category of the target feature includes constants and variables.
[0056] For example, Figure 2 This is a schematic diagram illustrating category determination based on a large language model, as provided in an embodiment of this application. Figure 2 As shown, by inputting prompt words into the large language model, the task content and rules for judging constants and variables are communicated. The log content and judgment fragments are input together into the large language model. The first preset prompt word can include: Given a log entry and some content, determine whether the content is a variable or a constant. You can make the judgment based on the following principles: 1. Variables are usually nouns, corresponding to specific objects or entities; 2. Variables are usually values dynamically populated into the log; 3. Constants usually carry important contextual information and define the interpretation of the log.
[0057] Step 104: If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0058] In this step, the aforementioned tag categories can be understood as the classification of target features in the initial log template. The aforementioned target log template is a modified structured template, and the modified tag classification is consistent with the categories used for category judgment by the large language model. For example, Figure 3 This is a flowchart of a log template error correction method provided in an embodiment of this application, such as... Figure 3 As shown, the initial log template is segmented into multiple fragments, and these fragments are pruned sequentially. The pruning process can involve filtering some fragments before using a large language model for category determination. Specifically, it can be done by determining whether the label categories corresponding to constants and variables in each fragment are correct based on preset rules. If the label categories are determined to be incorrect, the large language model is then used for judgment. This allows for a re-examination of multiple fragments of the initial log template, and the initial log template can be corrected using a correction module, thereby achieving self-optimization of the log template.
[0059] It should be noted that the Longest Common Subsequence (LCS) algorithm can also be used to identify the differences between the initial log template and the original log, record the different parts in the initial template and the original log, generate a difference list, and then correct the errors in the template based on the difference list, such as replacing redundant characters or adding missing characters, to generate a validated template.
[0060] In this embodiment, a large language model is used to parse the target log to obtain an initial log template. Features are extracted from the constants and / or variables in the target log to obtain target features. The large language model and a first preset prompt word are used to determine the category of the target features, where the first preset prompt word represents the category determination rule. If the category of the target feature is inconsistent with the labeled category of the initial log template, the labeled category is changed to obtain the target log template, where the labeled category is the category of the target feature indicated by the initial log template. Thus, by using the large language model and the first preset prompt word to determine the category of the target feature, the labeled category in the initial log template can be validated and modified, thereby continuously optimizing and correcting the initial log template, which improves the accuracy and reliability of the log parsing results.
[0061] Optionally, the step of using the large language model and the first preset prompt word to determine the category of the target feature and obtain the category of the target feature includes:
[0062] Determine whether the label category corresponding to the target feature is correct based on preset rules;
[0063] If the label category is determined to be incorrect, the target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature.
[0064] Specifically, the aforementioned preset rules can be predefined classification rules or logic used to quickly determine the category of target features. For example, since constants are usually fixed formats or specific keywords, while variables may contain dynamic content such as numbers and letters, the preset rules can be determined as follows: if the content of the target feature labeled as a variable only contains letters or Chinese characters, or if the target feature labeled as a constant conforms to the structure of common variables, such as IP addresses or paths, then it is judged that it may be misidentified. In the case of misidentification, a large language model is then used to determine the category.
[0065] For example, Figure 4 This is a flowchart of a category determination provided in an embodiment of this application, such as... Figure 4 As shown, when the target features labeled as variables contain only letters or Chinese characters, or when the target features labeled as constants contain only numbers, a large language model is needed for reflection, that is, to use a large language model for category determination.
[0066] In this implementation, the label category corresponding to the target feature is determined based on preset rules. If the label category is determined to be incorrect, the target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. This allows target features that conform to preset rules to be filtered out before the large language model is used for category classification, thereby improving the efficiency of log parsing.
[0067] Optionally, after changing the tag category to obtain the target log template, the method further includes:
[0068] The similarity between multiple candidate log templates and the target log template is calculated respectively, and the target candidate log template is determined. The target candidate log template is the log template with the highest similarity to the target log template among the multiple candidate log templates.
[0069] The target log template and the target candidate log template are semantically determined using the large language model and the second preset prompt word to determine whether the target log template and the target candidate log template are the same template. The second preset prompt word is used to characterize the judgment rule of the same template.
[0070] If the target log template and the target candidate log template are determined to be the same template, the target log template and the target candidate log template are merged to obtain a merged template.
[0071] Specifically, the aforementioned candidate log templates can be log templates continuously accumulated and generated by the log parsing system during the parsing process, i.e., currently stored log templates, or they can be a subset of candidate log templates selected from the stored log templates. The aforementioned similarity can be an indicator used to measure the structural or semantic similarity between two templates, specifically based on string matching, semantic analysis, or feature vector distance, etc. For example, the similarity could be Jaccard similarity.
[0072] The aforementioned second preset prompt can be a predefined rule or instruction used to guide the large language model in determining whether two templates are "the same template". For example, "If the structure, variable type and semantics of two templates are consistent, they are considered to be the same template".
[0073] For example, Figure 5 This is a schematic diagram illustrating semantic judgment based on a large language model, as provided in an embodiment of this application. Figure 5 As shown, the second preset prompt may include: Given two templates, you need to determine whether they should belong to the same template. You can make this judgment based on the following criteria: 1. If the templates have different contextual formats, then answer "No"; 2. If the difference between them is a preposition or conjunction, then answer "No"; 3. If the difference between them is a value filled in the log, then answer "Yes"; 4. If all the differences between them contain important contextual information, then answer "No".
[0074] It is understood that the above semantic determination refers to analyzing the semantic consistency between the target log template and the target candidate log template through a large language model, rather than just string matching.
[0075] It should be noted that before using the large language model for semantic determination, a certain pruning strategy can be adopted. That is, if two templates can cover each other, or the difference only exists in the recurring multiple elements, then there is no need to use the large language model for semantic determination, and the two templates can be directly identified as belonging to the same template.
[0076] The process of merging the target log template and the target candidate log template described above can be to extract the common and variable parts in the target log template and the target candidate log template to generate a unified template.
[0077] For example, Figure 6 This is a flowchart of a merged log template provided in an embodiment of this application, such as... Figure 6 As shown, during the log template merging process, the system compares the currently generated target log template with the stored historical templates, identifies the most similar template, and performs semantic analysis on the two templates using a large language model to determine the similarity between the two templates.
[0078] In this implementation, by calculating the similarity between multiple candidate log templates and the target log template, the target candidate log template that is closest to the target log template can be selected. Furthermore, a semantic determination based on a large language model is used to decide whether to merge the two templates. By performing semantic analysis and contextual understanding on the templates, the limitations of traditional character or word matching can be overcome, and the actual similarity between templates can be identified more accurately. This allows for continuous improvement of log templates, enhancing their versatility, as well as the accuracy and consistency of log parsing.
[0079] Optionally, before calculating the similarity between the multiple candidate log templates and the target log template, the method further includes:
[0080] Based on the acquired multiple preset log templates, an inverted index is constructed, which is used to indicate the position information of each word in the multiple preset log templates.
[0081] Lexical extraction is performed on the target log template to obtain the target vocabulary;
[0082] Based on the target vocabulary and the inverted index, multiple candidate log templates are determined. The candidate log templates are log templates that include the target vocabulary among the multiple preset log templates.
[0083] Specifically, the aforementioned preset log templates can be log templates continuously accumulated and generated by the log parsing system during the parsing process, that is, currently stored log templates. The aforementioned target vocabulary can be key components in the target log template, including multiple constants and variables.
[0084] In this implementation, an inverted index can be used to quickly locate all log templates containing the target words, thereby identifying multiple candidate log templates that include the target words, which can further improve the efficiency and accuracy of log parsing.
[0085] Optionally, the step of parsing the target log using a large language model to obtain an initial log template for the target log includes:
[0086] Obtain a preset log dataset, which includes multiple log template pairs, each of which includes a log and a template corresponding to the log.
[0087] Calculate the similarity between multiple logs in the preset log dataset and the target log;
[0088] The K logs with the highest similarity to the target log among the plurality of logs are identified as example logs, where K is an integer greater than 0;
[0089] The target log is parsed using the large language model and the example set to obtain an initial log template for the target log. The example set is used to instruct the large language model to learn and process the log. The example set includes the example log and the log template corresponding to the example log.
[0090] Specifically, the aforementioned preset log dataset can be a pre-constructed collection containing logs and corresponding log templates, where the log templates can be stored historical log templates. The aforementioned similarity can measure the degree of similarity between the target log and the logs in the preset log dataset, and can be determined based on lexical matching, semantic analysis, or structural similarity, etc.
[0091] It is understood that the above-mentioned parsing of the target log using the large language model and example set can be achieved by utilizing the context learning capability of the large language model and providing examples to parse the target log.
[0092] It's important to note that the example set was chosen to provide the large language model with sufficient reference information to correctly distinguish between constants and variables in the logs. The selected examples should provide as much information as possible so that the large language model can understand how to process new logs. In log parsing, some words may behave as constants or variables in different contexts. For the same content, it might be a constant in some logs and a variable in others. Therefore, the selection of the example set mentioned above considers not only the textual similarity between logs but also the different identifications of controversial words.
[0093] For example, the above similarity can be calculated based on the text similarity between logs and the degree of controversy for each word in the logs. The degree of controversy can be used to measure the ratio of the frequency of a word appearing as a constant to the frequency of a variable in a preset log dataset. The formula for calculating the above similarity is as follows:
[0094]
[0095] Where bow(e) is the bag-of-words for log e, bow(e′) is the bag-of-words for log e′, cont(w) is the degree of controversy of log content w, |S c |S represents the number of times the log content w is used as a constant. v | represents the number of times the log content w is used as a variable.
[0096] Based on the above calculation formula, select the K highest-scoring log template pairs from the preset log dataset to form an example set.
[0097] For example, Figure 7 This is a flowchart of an initial log parsing provided in an embodiment of this application, such as... Figure 7 As shown, log template pairs are obtained from publicly available log datasets to construct a preset log dataset, and an iterative sampling strategy is used to ensure the diversity of the selected logs. Then, the K log template pairs that are most similar to the target log are selected from the preset log dataset to form an example set. Finally, the example set is used to construct prompts, and context learning guides a large-scale language model to generate the initial template of the target log.
[0098] Figure 8 This is a schematic diagram illustrating log parsing based on a large language model, as provided in an embodiment of this application. Figure 8 As shown, when building a suggestion, logs and templates from the example set can be inserted into the suggestion, and then... <start>and <end>Tags limit the output range to avoid generating unnecessary content.
[0099] In this implementation, by selecting K log template pairs that are most similar to the target log from a preset log dataset to form an example set, and using the example set to guide the large language model to parse the target log, manual intervention can be reduced, thereby further improving the accuracy and efficiency of log parsing.
[0100] Optionally, the step of parsing the target log using a large language model to obtain an initial log template for the target log includes:
[0101] Generate a character index tree based on the acquired multiple preset log templates;
[0102] The target log is matched character by character using the character index tree;
[0103] If the target log fails to match, the target log is parsed using the large language model to obtain the initial log template of the target log.
[0104] Specifically, the character index tree described above can be used for efficient storage and matching of strings, where each node in the character index tree represents a character. For example, Figure 9 This is a schematic diagram of a character index tree provided in an embodiment of this application, such as... Figure 9 As shown, each node in the character index tree represents a log constant character or a wildcard, and the leaf nodes represent complete log templates.
[0105] Figure 10 This is a schematic diagram of character matching provided in an embodiment of this application, such as... Figure 10 As shown, matching constant characters only requires determining whether their contents are consistent. Wildcards often match more than a single character; normally, they will match multiple characters. When handling the matching of wildcards "<*>", the character types that have appeared before can be listed in a set. When a character type that is not in the set is matched, the wildcard matching ends, and the matching of subsequent characters continues. For example, the character types in the set are generated based on the types of characters that appear in the corresponding positions of the logs that generated the template. Only the following cases are considered: (1) 26 uppercase and lowercase English letters, (2) Chinese characters, (3) numbers, (4) various symbols. Among them, categories (1), (2), and (3) constitute three classes, while each symbol in (4) is a separate class.
[0106] It's important to note that the character index tree described above can be understood as a character-level prefix tree. Traditional prefix trees typically segment words based on fixed delimiters, leading to inconsistencies and insufficient flexibility when processing templates generated by large language models. For example, consider a log template "ready=<*>policy=<*>wakefulness=<*>wksummary=<*>uasummary=<*>". Because it uses spaces as delimiters, using a prefix tree for log matching would require sacrificing some template granularity and precision, simplifying the template to "<*><*><*><*><*>" (if a space-segmented segment contains <*>, the entire segment is simplified to <*>). This causes the template to lose all valid information and become unusable for log matching. Using this template for log matching would result in "over-matching" of log messages containing all four spaces.
[0107] The process of matching the target log character by character using the character index tree described above can be such that when matching the target log, all characters of the target log are used for matching in the character index tree, and the matching process is considered successful when the leaf node of the character index tree is reached.
[0108] For example, Figure 11 This is a flowchart of a character index tree matching method provided in an embodiment of this application, such as... Figure 11 As shown, firstly, the current character in the log is matched. If the match fails, the matching process terminates, and the template matching fails. If the match succeeds, the corresponding successor node is found based on the content of the child nodes of the current node and the next character in the log. Next, it is determined whether the successor node exists. If the successor node for the corresponding content does not exist (assuming the next character in the log is 'a', if the content of the child nodes of the current node is not 'a' or '<*>', then it is considered that there is no successor node), the match fails. If it exists, it is further determined whether the successor node is an internal node or a leaf node. If it is an internal node, the matching continues for the next character. If it is a leaf node, the match succeeds.
[0109] For example, Figure 12 This is a flowchart of a matching log template provided in an embodiment of this application, such as... Figure 12 As shown, when the target log is obtained for the first time, it is first checked whether it can be matched with an existing log template through the character index tree. If the match is successful, the matched template is directly used for parsing to reduce the number of calls to the large model, thereby avoiding repeated parsing of the same template, improving processing speed, and achieving efficient log parsing. If the match fails, the large language model is called to generate a log template to obtain a high-precision log parsing result.
[0110] Furthermore, after generating the target log template, the target log template can be updated in the character index tree.
[0111] In this implementation, a character index tree is generated based on multiple pre-defined log templates, and the target log is matched character by character using this character index tree. This allows for efficient parsing of the target log and flexible handling of different types of log templates, ensuring the accuracy and consistency of the log parsing results. Furthermore, compared to using regular expressions for log matching, this method improves matching efficiency.
[0112] For example, Figure 13 This is a flowchart of a log parsing method provided in an embodiment of this application, such as... Figure 13 As shown, log parsing includes five core steps: template matching, initial parsing, template correction, self-reflection, and template merging.
[0113] For a given log sequence, a streaming parsing process is performed, and the extracted log templates are recorded in a set containing several unique templates, each of which can correspond to multiple logs;
[0114] For each target log entry, we can try to match it with templates in the template set using a character index tree to improve efficiency. If the log entry matches a template, we can directly assign it to that template.
[0115] For the initial template, the target log is compared with the initial template through template validation to correct differences, thereby generating a corrected template;
[0116] For the correction template, the correctness of the identified constants and variables is reviewed using a large language model to reflect on itself, correct misidentifications, and generate improved templates.
[0117] For the improved template, the differences between the improved template and the most similar template in the set are analyzed using a large language model. The templates are then merged and used to further improve the template, and finally the parsing result is obtained.
[0118] See Figure 14 , Figure 14 This is a schematic diagram of the structure of a log parsing device provided in an embodiment of this application, as shown below. Figure 14 As shown, the log parsing device 1400 includes:
[0119] Parsing module 1401 is used to parse the target log using a large language model to obtain the initial log template of the target log;
[0120] The first extraction module 1402 is used to extract features from constants and / or variables in the target log to obtain target features;
[0121] The judgment module 1403 is used to perform category judgment on the target feature using the large language model and the first preset prompt word to obtain the category of the target feature, wherein the first preset prompt word is used to characterize the category judgment rule;
[0122] The modification module 1404 is used to modify the tag category when the category of the target feature is inconsistent with the tag category of the initial log template, so as to obtain a target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0123] Optionally, the determination module 1403 includes:
[0124] The first determining unit is used to determine whether the label category corresponding to the target feature is correct based on preset rules;
[0125] The judgment unit is used to determine the category of the target feature by using the large language model and the first preset prompt word when it is determined that the tag category is incorrect, so as to obtain the category of the target feature.
[0126] Optionally, the log parsing device 1400 further includes:
[0127] The first determining module is used to calculate the similarity between multiple candidate log templates and the target log template, and to determine the target candidate log template, wherein the target candidate log template is the log template with the highest similarity to the target log template among the multiple candidate log templates;
[0128] The second determining module is used to perform semantic determination on the target log template and the target candidate log template using the large language model and the second preset prompt word, so as to determine whether the target log template and the target candidate log template are the same template. The second preset prompt word is used to characterize the judgment rule of the same template.
[0129] The merging module is used to merge the target log template and the target candidate log template to obtain a merged template when it is determined that the target log template and the target candidate log template are the same template.
[0130] Optionally, the log parsing device 1400 further includes:
[0131] A construction module is used to construct an inverted index based on multiple preset log templates, wherein the inverted index is used to indicate the position information of each word in the multiple preset log templates;
[0132] The second extraction module is used to extract vocabulary from the target log template to obtain target vocabulary;
[0133] The third determining module is used to determine multiple candidate log templates based on the target vocabulary and the inverted index, wherein the candidate log templates are log templates that include the target vocabulary among the multiple preset log templates.
[0134] Optionally, the parsing module 1401 includes:
[0135] An acquisition unit is used to acquire a preset log dataset, the preset log dataset including multiple log template pairs, each of the multiple log template pairs including a log and a template corresponding to the log;
[0136] A calculation unit is used to calculate the similarity between multiple logs in the preset log dataset and the target log, respectively.
[0137] The second determining unit is used to determine the K logs with the highest similarity to the target log among the plurality of logs as example logs, where K is an integer greater than 0;
[0138] The first parsing unit is used to parse the target log using the large language model and the example set to obtain an initial log template of the target log. The example set is used to instruct the large language model to learn and process logs, and the example set includes the example logs and the log templates corresponding to the example logs.
[0139] Optionally, the parsing module 1401 includes:
[0140] The generation unit is used to generate a character index tree based on multiple pre-defined log templates.
[0141] The matching unit is used to perform character-by-character matching on the target log using the character index tree;
[0142] The second parsing unit is used to parse the target log using the large language model when the target log fails to match, so as to obtain the initial log template of the target log.
[0143] It should be noted that the log parsing device provided in this application embodiment is a device capable of executing the above-described log parsing method. Therefore, all implementation methods in the above-described log parsing method embodiments are applicable to this device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0144] For details, see Figure 15 As shown in the figure, this application embodiment also provides an electronic device, including a bus 1501, a transceiver 1502, an antenna 1503, a bus interface 1504, a processor 1505, and a memory 1506.
[0145] Processor 1505, used for:
[0146] The target log is parsed using a large language model to obtain the initial log template of the target log;
[0147] Feature extraction is performed on the constants and / or variables in the target log to obtain the target features;
[0148] The target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. The first preset prompt word is used to represent the category classification rule.
[0149] If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
[0150] exist Figure 15 In this document, a bus architecture (represented by bus 1501) is used. Bus 1501 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 1505 and memory represented by memory 1506. Bus 1501 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 1504 provides an interface between bus 1501 and transceiver 1502. Transceiver 1502 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 1505 is transmitted over a wireless medium via antenna 1503, which further receives data and transmits it to processor 1505.
[0151] Processor 1505 manages bus 1501 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 1506 can be used to store data used by processor 1505 during operation.
[0152] Optionally, the processor 1505 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0153] Optionally, the processor 1505 is specifically used for:
[0154] Determine whether the label category corresponding to the target feature is correct based on preset rules;
[0155] If the label category is determined to be incorrect, the target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature.
[0156] Optionally, the processor 1505 is also used for:
[0157] The similarity between multiple candidate log templates and the target log template is calculated respectively, and the target candidate log template is determined. The target candidate log template is the log template with the highest similarity to the target log template among the multiple candidate log templates.
[0158] The target log template and the target candidate log template are semantically determined using the large language model and the second preset prompt word to determine whether the target log template and the target candidate log template are the same template. The second preset prompt word is used to characterize the judgment rule of the same template.
[0159] If the target log template and the target candidate log template are determined to be the same template, the target log template and the target candidate log template are merged to obtain a merged template.
[0160] Optionally, the processor 1505 is also used for:
[0161] Based on the acquired multiple preset log templates, an inverted index is constructed, which is used to indicate the position information of each word in the multiple preset log templates.
[0162] Lexical extraction is performed on the target log template to obtain the target vocabulary;
[0163] Based on the target vocabulary and the inverted index, multiple candidate log templates are determined. The candidate log templates are log templates that include the target vocabulary among the multiple preset log templates.
[0164] Optionally, the processor 1505 is specifically used for:
[0165] Obtain a preset log dataset, which includes multiple log template pairs, each of which includes a log and a template corresponding to the log.
[0166] Calculate the similarity between multiple logs in the preset log dataset and the target log;
[0167] The K logs with the highest similarity to the target log among the plurality of logs are identified as example logs, where K is an integer greater than 0;
[0168] The target log is parsed using the large language model and the example set to obtain an initial log template for the target log. The example set is used to instruct the large language model to learn and process the log. The example set includes the example log and the log template corresponding to the example log.
[0169] Optionally, the processor 1505 is specifically used for:
[0170] Generate a character index tree based on the acquired multiple preset log templates;
[0171] The target log is matched character by character using the character index tree;
[0172] If the target log fails to match, the target log is parsed using the large language model to obtain the initial log template of the target log.
[0173] It should be noted that the electronic device provided in this application embodiment is a device capable of executing the above-described log parsing method. Therefore, all implementation methods in the above-described log parsing method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0174] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described log parsing method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0175] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described log parsing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0176] This application also provides a computer program product, including computer instructions. When executed by a processor, these computer instructions implement the various processes of the above-described log parsing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0177] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0178] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0179] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.< / end> < / start>
Claims
1. A log parsing method, characterized in that, include: The target log is parsed using a large language model to obtain the initial log template of the target log; Feature extraction is performed on the constants and / or variables in the target log to obtain the target features; The target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. The first preset prompt word is used to represent the category classification rule. If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
2. The method according to claim 1, characterized in that, After changing the tag category to obtain the target log template, the method further includes: The similarity between multiple candidate log templates and the target log template is calculated respectively, and the target candidate log template is determined. The target candidate log template is the log template with the highest similarity to the target log template among the multiple candidate log templates. The target log template and the target candidate log template are semantically determined using the large language model and the second preset prompt word to determine whether the target log template and the target candidate log template are the same template. The second preset prompt word is used to characterize the judgment rule of the same template. If the target log template and the target candidate log template are determined to be the same template, the target log template and the target candidate log template are merged to obtain a merged template.
3. The method according to claim 1, characterized in that, The step of using the large language model and the first preset prompt word to determine the category of the target feature, and obtaining the category of the target feature, includes: Determine whether the label category corresponding to the target feature is correct based on preset rules; If the label category is determined to be incorrect, the target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature.
4. The method according to claim 2, characterized in that, Before calculating the similarity between the multiple candidate log templates and the target log template, the method further includes: Based on the acquired multiple preset log templates, an inverted index is constructed, which is used to indicate the position information of each word in the multiple preset log templates. Lexical extraction is performed on the target log template to obtain the target vocabulary; Based on the target vocabulary and the inverted index, multiple candidate log templates are determined. The candidate log templates are log templates that include the target vocabulary among the multiple preset log templates.
5. The method according to claim 1, characterized in that, The process of parsing the target log using a large language model to obtain the initial log template of the target log includes: Obtain a preset log dataset, which includes multiple log template pairs, each of which includes a log and a template corresponding to the log. Calculate the similarity between multiple logs in the preset log dataset and the target log; The K logs with the highest similarity to the target log among the plurality of logs are identified as example logs, where K is an integer greater than 0; The target log is parsed using the large language model and the example set to obtain an initial log template for the target log. The example set is used to instruct the large language model to learn and process the log. The example set includes the example log and the log template corresponding to the example log.
6. The method according to claim 1, characterized in that, The process of parsing the target log using a large language model to obtain the initial log template of the target log includes: Generate a character index tree based on the acquired multiple preset log templates; The target log is matched character by character using the character index tree; If the target log fails to match, the target log is parsed using the large language model to obtain the initial log template of the target log.
7. A log parsing device, characterized in that, include: The parsing module is used to parse the target log using a large language model to obtain the initial log template of the target log; The first extraction module is used to extract features from the constants and / or variables in the target log to obtain target features; The judgment module is used to use the large language model and the first preset prompt word to judge the category of the target feature and obtain the category of the target feature. The first preset prompt word is used to represent the category judgment rule. The modification module is used to modify the tag category when the category of the target feature is inconsistent with the tag category of the initial log template, so as to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
8. An electronic device, characterized in that, Includes a processor, the processor being used for: The target log is parsed using a large language model to obtain the initial log template of the target log; Feature extraction is performed on the constants and / or variables in the target log to obtain the target features; The target feature is classified using the large language model and the first preset prompt word to obtain the category of the target feature. The first preset prompt word is used to represent the category classification rule. If the category of the target feature is inconsistent with the tag category of the initial log template, the tag category is changed to obtain the target log template, wherein the tag category is the category of the target feature indicated by the initial log template.
9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the log parsing method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the log parsing method as described in any one of claims 1 to 6.
11. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the log parsing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Log analysis method and device, equipment and storage medium
CN116822491A
Low-cost and zero-sample online log analysis method based on large language model
CN117407242A
Log template acquisition method, electronic equipment and storage medium
CN118093325A