Rule-based log analysis method, device, equipment, medium and product
Through the rule-based log analysis method, the training set and prompt guidance log analysis model output optimization rule sequences is solved, and the problem that large language models are difficult to capture the implicit knowledge of logs is improved, and the accuracy and generalization ability of log analysis are improved.
Patent Information
- Application Number
- CN202510300167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, log analysis methods based on large language models are difficult to capture the implicit knowledge and rules in the log, resulting in the model output that may contradict real-world knowledge, and rely on rote memorization and lack effective generalization ability.
A rule-based log analysis method is proposed. By creating a training set for characterizing different log analysis tasks and designing corresponding prompts, the initial log analysis model is guided to output rules corresponding to the log analysis task, and by optimizing the rule sequence and embedding the rule prompts, the model's inference and generalization capabilities are enhanced.
This method can understand log data more accurately, enhance the inference and generalization capabilities of the log analysis model, improve the accuracy of log analysis, and has cost control advantages when facing log data changes.
Smart Images

Figure CN119806965B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of log analysis, and in particular to a rule-based log analysis method, device, equipment, medium and product. Background Art
[0002] System logs provide critical runtime information, providing important debugging and troubleshooting information for developers, operators, and maintenance personnel. However, as systems become increasingly complex, analyzing large amounts of log data has become a major challenge. Large language models (LLMs) for natural language processing have powerful reasoning and interpretation capabilities and can effectively handle log analysis tasks.
[0003] Since LLM has difficulty capturing implicit knowledge and rules in logs, hallucinations occur when LLM produces outputs that appear reasonable but contradict real-world knowledge. The existing technology uses sample-based log analysis methods, which helps the model learn to identify patterns and anomalies in logs by providing specific samples, thereby reducing the generation of outputs that contradict the real world. However, this method causes LLM to rely on rote memorization and lacks effective generalization capabilities. Summary of the invention
[0004] The purpose of this application is to provide a rule-based log analysis method, device, equipment, medium and product that can accurately understand log data, thereby improving the generalization ability of the model.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a rule-based log analysis method, comprising:
[0007] Creating a training set for characterizing different log analysis tasks, and designing corresponding prompts for the different log analysis tasks, wherein the training set includes training log samples and training label samples, and each training log in the training log samples uniquely corresponds to a training label in the training label samples;
[0008] Inputting the training set and the prompt into an initial log analysis model, and guiding the initial log analysis model to output rules corresponding to the log analysis task through the prompt;
[0009] Prioritizing the rules to obtain an optimized rule sequence, and storing the optimized rule sequence in a rule base within the initial log analysis model;
[0010] Obtaining sample prompts and rule prompts, embedding the optimization rule sequence into the rule prompts, and obtaining final rule prompts;
[0011] Fine-tuning the parameters of the initial log analysis model by using a comparative preference optimization method to obtain a log analysis model;
[0012] The log to be analyzed is input into the log analysis model, the log analysis model is guided to output a final result through the final rule prompt, and the log to be analyzed is analyzed through the final result.
[0013] In a second aspect, the present application provides a rule-based log analysis device, comprising:
[0014] A creation module, used to create a training set for characterizing different log analysis tasks and design corresponding prompts for the different log analysis tasks, wherein the training set includes training log samples and training label samples, and each training log in the training log samples uniquely corresponds to a training label in the training label samples;
[0015] A guiding module, used for inputting the training set and the prompt into an initial log analysis model, and guiding the initial log analysis model to output a rule corresponding to the log analysis task through the prompt;
[0016] A storage module, used for prioritizing the rules to obtain an optimized rule sequence, and storing the optimized rule sequence in a rule base located in the initial log analysis model;
[0017] An embedding module, used for obtaining sample prompts and rule prompts, embedding the optimized rule sequence into the rule prompts, and obtaining final rule prompts;
[0018] A fine-tuning module, used to fine-tune the parameters of the initial log analysis model by using a comparative preference optimization method to obtain a log analysis model;
[0019] The output module is used to input the log to be analyzed into the log analysis model, guide the log analysis model to output a final result through the final rule prompt, and analyze the log to be analyzed through the final result.
[0020] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any one of the rule-based log analysis methods described above.
[0021] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the rule-based log analysis methods described above.
[0022] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any one of the rule-based log analysis methods described above.
[0023] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0024] Compared with the sample-based log analysis method used in the prior art, the present application embeds the optimized rule sequence into the rule prompt to obtain the final rule prompt, which can capture the potential rules in the log after the rule is called and understand the log data more accurately, thereby enhancing the reasoning and generalization capabilities of the log analysis model and improving the accuracy of log analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 This is an application environment diagram of a rule-based log analysis method in an embodiment of the present application;
[0027] Figure 2 A flowchart of a rule-based log analysis method provided in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of functional modules of a rule-based log analysis device provided in one embodiment of the present application;
[0029] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0031] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] The rule-based log analysis method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the training set to be processed to the server 104. After the server 104 receives the training set to be processed, for the training set to be processed, the server 104 designs corresponding prompts for different log analysis tasks, inputs the training set and the prompts into the initial log analysis model, and guides the initial log analysis model to output rules corresponding to the log analysis tasks through the prompts; prioritizes the rules to obtain an optimized rule sequence, and stores the optimized rule sequence in a rule library located in the initial log analysis model; obtains sample prompts and rule prompts, embeds the optimized rule sequence into the rule prompts, and obtains final rule prompts; further fine-tunes the parameters of the initial log analysis model through a comparison preference optimization method to obtain a log analysis model; inputs the log to be analyzed into the log analysis model, calls the final rule prompt corresponding to the log analysis task in the rule library according to the log analysis task represented by the log to be analyzed, and guides the log analysis model to output the final result through the final rule prompt. The server 104 can feed back the final result obtained to the terminal 102. In addition, in some embodiments, the rule-based log analysis method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the training set to be processed, or the server 104 can obtain the training set to be processed from the data storage system and process the training set to be processed.
[0033] The terminals may be, but are not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices and portable wearable devices. IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or a cloud server.
[0034] System logs are key information that records various events and states during system operation. They provide a window for developers, operators, and maintenance personnel to deeply understand the behavior patterns of the system and effectively debug when problems occur. With the continuous advancement of technology, systems are becoming more and more complex, which has led to a sharp increase in the amount of system log data. Large language models (LLMs) in the field of natural language processing have powerful reasoning and interpretation capabilities and can effectively handle log analysis tasks.
[0035] In practical applications, the system generates a massive amount of log data. For example, it has been reported that a cloud system can generate about 30GB to 50GB of log data per hour, which is equivalent to about 100 to 200 million lines of tracking logs. On the one hand, since the reasoning cost of LLM is very high, it is unrealistic to call a large language model for every log message; on the other hand, using a smaller language model is cheaper and more economical than a large-scale model. Most of the current LLM-based log analysis methods rely on the reasoning capabilities of large-scale models such as ChatGPT, and these methods often perform poorly on smaller language models.
[0036] In addition, LLM also faces the problem of "hallucination" when analyzing logs. This is mainly because LLM has difficulty capturing the knowledge and rules implicit in the logs. Hallucinations occur when LLM outputs information that seems reasonable but actually contradicts real-world knowledge. In order to enhance the reasoning ability of LLM and reduce the phenomenon of hallucinations, one approach is to adopt prompt engineering, that is, to clarify the goals and methods by providing examples. This method is called example-based prompts. Although example-based prompts help reduce hallucinations, their generalization ability is limited, especially when system logs are constantly evolving. As the system is updated, developers may modify source code and log statements, thereby changing the structure of log data. Example-based prompts cannot capture the underlying rules in the logs, which often leads to large language models relying on rote memorization, thereby limiting the model's ability to generalize new situations.
[0037] In order to solve the above problems, this application proposes a rule-based log analysis method that does not rely on specific examples and uses rule-based prompts to enhance the reasoning ability of small-scale language models. Compared with example-based prompts, rule-based prompts help language models better understand the nature of tasks and knowledge, thereby improving their generalization ability. Since this method allows the use of smaller and lower-cost language models to complete complex log analysis tasks, this method can not only cope with the constant changes in log data, but also has obvious advantages in cost control.
[0038] In an exemplary embodiment, Figure 2 As shown, a rule-based log analysis method is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps S201 to S206. Among them:
[0039] In step S201, a training set is created for characterizing different log analysis tasks, and corresponding prompts are designed for different log analysis tasks. The training set includes training log samples and training label samples. Each training log in the training log sample uniquely corresponds to a training label in the training label sample.
[0040] Specifically, due to the huge number of logs, generating rules from a large number of logs will be more reliable, but the cost of large-scale reasoning using GPT is high. Therefore, this application first samples the logs, that is, randomly selects a small number of logs from the original large-scale logs to create a dataset ,in and Represent log samples and label samples respectively, Represents a specific log analysis task. In this application, T represents a log parsing task or a log anomaly detection task. Then, the dataset Split into training set and test set , so that For example, 2000 logs are randomly sampled from the original one million logs, and then the 2000 logs are divided into a training set and a test set according to a ratio of 4:1, that is, the training set has 1600 logs and the test set has 400 logs.
[0041] Logging is a method of recording events that occur when a system, application, or other software is running. Log files usually contain but are not limited to the following:
[0042] Timestamp: Records the specific date and time when an event occurs.
[0043] Log level: such as DEBUG, INFO, WARNING, ERROR, CRITICAL, etc., indicating the severity or urgency of the log entry.
[0044] Message: Describes the details of an event, which may include error messages, system status updates, user actions, transactions, etc.
[0045] Source information: Information indicating the source of the log message, such as the file name, class name, function name, or process ID.
[0046] User information: If the log is related to user operations, it may contain user ID, user name, or other user identification information.
[0047] Environmental information: Information about the system environment or runtime environment, such as operating system type, software versions, configuration settings, etc.
[0048] State information: The state of an application or system at a specific point in time, which may include memory usage, resource utilization, etc.
[0049] The main purpose of logs is to help developers and system administrators monitor system operation status, debug problems, audit operations, and analyze performance. Training tags contain content after log parsing or log anomaly detection.
[0050] In step S202, the training set and the prompt are input into the initial log analysis model, and the prompt is used to guide the initial log analysis model to output rules corresponding to the log analysis task.
[0051] Specifically, the prompt includes a log parsing prompt and a log anomaly detection task prompt. The log analysis task includes a log parsing task and a log anomaly detection task. A log parsing prompt is designed for the log parsing task, and a log anomaly detection prompt is designed for the log anomaly detection task. In step S202, the initial log analysis model can select the LLM model. It should be noted that only in the rule generation related steps (such as step S202), the initial log analysis model can select the GPT-4o-mini model, because the GPT-4o-mini model can process a small amount of data, but the processing speed is fast.
[0052] In one embodiment, the log analysis task includes a log parsing task and a log anomaly detection task, the prompt includes a log parsing prompt and a log anomaly detection prompt, the rule includes a log parsing rule and a log anomaly detection rule, and step S202 includes the following sub-steps:
[0053] S2021. Input the training set and log parsing prompts used to characterize the log parsing task into the initial log analysis model, and output the log parsing rules.
[0054] S2022: Input the training set and log anomaly detection prompt used to characterize the log anomaly detection task into the initial log analysis model, and output the log anomaly detection rule.
[0055] Specifically, since the log analysis tasks disclosed in the present invention include log parsing and log anomaly detection, there are prompts corresponding to different tasks, namely log parsing prompts and log anomaly detection prompts, and obviously two rules will be generated, namely log parsing rules and log anomaly detection rules. Through log parsing prompts and log anomaly detection prompts, the model is guided to output rules related to the log analysis task based on the training sets that characterize the log parsing task and the log anomaly detection task. The goal of generating log parsing rules and log anomaly detection rules is to find a function , which maps logs and tags to rules , represents the original log sample, Represents a label sample. The rules are generated based on the pre-defined vocabulary space in the initial log analysis model. Represents the vocabulary space of the model, and the vocabulary space is predefined in the model. Log parsing rules and log anomaly detection rules are generated based on the predefined vocabulary space in the initial log analysis model. When performing log analysis, there will be a predefined vocabulary or dictionary first. This vocabulary contains key words, patterns, and structures that are essential for understanding specific system or application logs. When generating log parsing rules and log anomaly detection rules, these rules are created based on this predefined vocabulary space. By using a predefined vocabulary space, key information in the data set can be identified more accurately. Since the rules are customized according to the actual possible log format and content, with clear vocabulary and parsing rules, the system can classify and analyze the data in the data set more quickly without checking all possible variants line by line, thereby improving overall processing efficiency.
[0056] For example,
[0057] The log parsing prompt in step S2021 is in the following form:
[0058] You will get some log parsing examples, and you need to summarize some log parsing rules based on these examples. The format of the examples is "<original log>-><parsed log>". You must output the summarized rules in JSON format, including two keys "task" and "rules", where "task" is "log_parsing" and "rules" is a string list.
[0059] For example,
[0060] The log anomaly detection prompt in step S2022 is in the following form:
[0061] You will get some examples of log anomaly detection, and you need to summarize some log anomaly detection rules based on these examples. The format of the examples is "<original log>-><category>". You must output the summarized rules in JSON format, including two keys "task" and "rules", where "task" is "anomaly_detection" and "rules" is a list of strings.
[0062] Specifically, the form of the rule is one of natural language description, regular expression and code. Parsing rules can be expressed in multiple forms to adapt to different application scenarios and requirements. Parsing rules in the form of natural language description are defined in the way of expressing in human daily language, which is easy to understand. For example, "if the log contains the word 'ERROR', it is marked as an error log." Parsing rules in the form of regular expressions use mathematical or logical expressions to define rules, which are highly formalized and precise. For example: "ERROR in log_message → error_log". Parsing rules in the form of code are written in programming languages and can be directly executed in software.
[0063] For example,
[0064] The log parsing rules in step S2021 are:
[0065] {
[0066] "task": "log_parsing",
[0067] "rules": [
[0068] "Replace the numeric sequence with <*>",
[0069] "Replace the IP address with <*>",
[0070] "Replace the hostname with <*>",
[0071] "Size <number>'Replace with 'size<*>'",
[0072] "Write 'blk_ <number>'Replace with 'blk_<*>'",
[0073] "Replace the user sequence with 'user=<*>'",
[0074] "Change 'uid= <number>' and 'euid= <number>'replace with 'uid=<*>' and 'euid=<*>' respectively,
[0075] "For logs without variable parts, preserve their structure and text",
[0076] "Write a time indicator (such as ' <number>s') is replaced by '<*>s'",
[0077] "For statements that are like events or actions and contain numbers, replace the numbers with <*>",
[0078] "For logs indicating connections or disconnections, replace the specific address and port with <*>",
[0079] "Replace strings matching a specific description with <*> where the description is similar" ]
[0081] }
[0082] For example,
[0083] The log anomaly detection rule in step S2022 is:
[0084] {
[0085] "task": "anomaly_detection",
[0086] "rules": [
[0087] "If the log contains 'data storage interrupt', mark it as 'Anomalous'",
[0088] "If the log contains 'data TLB error interrupt', mark it as 'Anomalous'",
[0089] "If the log contains 'rts: kernel terminated for reason', mark it as 'Anomalous'",
[0090] "If the log contains 'ciod: failed to read message prefix', mark it as 'Anomalous'",
[0091] "If the log contains 'ciod: Error reading message prefix', mark it as 'Anomalous'",
[0092] "If the log contains 'status error: status=0x00', mark it as 'Anomalous'",
[0093] "If the log contains 'drive not ready for command', mark it as 'Anomalous'",
[0094] "If the log contains 'system epilog failed w / rc=1', mark it as 'Anomalous'" ]
[0096] }
[0097] In step S203, the rules are prioritized to obtain an optimized rule sequence, and the optimized rule sequence is stored in a rule library located in the initial log analysis model.
[0098] In one embodiment, "prioritizing rules" in step S203 includes the following sub-steps S2031-S2032:
[0099] S2031. In the process of the initial log analysis model outputting rules corresponding to the log analysis task, the number of occurrences of each rule is recorded.
[0100] S2032: Prioritize all the rules in descending order of the number of occurrences.
[0101] Specifically, the order in which rules are executed directly affects the reasoning logic and final results of the system. In order to ensure that the system can operate as expected and make the best choice when faced with multiple potentially applicable rules, the rules need to be clearly prioritized. This allows models like LLM to understand the relative importance of these rules based on the order in which they are ranked, and to prioritize higher priority rules as much as possible during the reasoning process.
[0102] In order to achieve this priority sorting, a counting-based method can be used. Whenever a rule is generated, it is counted. Then, the rules are sorted from large to small according to their counts, that is, the rules with larger counts are sorted higher. The core idea of this method is that the more times a rule is generated, it usually means that it appears more frequently in the system, or it is more important in solving a specific problem. Therefore, in this way, the priority of the rules can be effectively determined, so that the system can give priority to those more important or more frequently applicable rules during the reasoning process, and store the optimized rule sequence in the rule base located in the initial log analysis model, so that in the process of log analysis, the rules in the rule base can be called accordingly according to the tasks represented by the logs. The rules here include
[0103] It should be noted that the rules in step S203 include log parsing rules and log anomaly detection rules, and therefore, the optimization rule sequence also correspondingly includes a log parsing optimization rule sequence and a log anomaly detection optimization rule sequence.
[0104] In step S204, sample prompts and rule prompts are obtained, and the optimization rule sequence is embedded in the rule prompt to obtain a final rule prompt.
[0105] Specifically, since the log analysis task disclosed in the present invention includes a log parsing task and a log anomaly detection task, the sample prompt includes a log parsing sample prompt and a log anomaly detection sample prompt, and the rule prompt includes a log parsing rule prompt and a log anomaly detection rule prompt.
[0106] First, obtain sample prompts and rule prompts, and embed the optimization rule sequence in the database into the rule prompt. Since the optimization rule sequence in the database includes the log parsing optimization rule sequence and the log anomaly detection optimization rule sequence, the log parsing optimization rule sequence is embedded in the log parsing rule prompt, and the corresponding log parsing final rule prompt is obtained; the log anomaly detection optimization rule sequence is embedded in the log anomaly detection rule prompt, and the corresponding log anomaly detection final rule prompt is obtained. The final rule prompt in step S204 includes the log parsing final rule prompt and the log anomaly detection final rule prompt. Since the role of the log parsing optimization rule and the log anomaly detection optimization rule is to guide the model to capture the potential rules in the log and accurately understand the log data, and the role of the log parsing rule prompt and the log anomaly detection rule prompt is to guide the model to output the corresponding label samples according to different execution tasks. Therefore, the optimization rule sequence is embedded in the rule prompt, the purpose of which is to obtain the output obtained after the log is input into the model, which can follow the log parsing rule prompt and the log anomaly detection rule prompt, and can also follow the log parsing optimization rule sequence and the log anomaly detection optimization rule sequence at the same time.
[0107] For example,
[0108] The log parsing example prompt in the interpretation content of step S204 is:
[0109] You will be provided with some log parsing examples in JSON format, which you need to use to parse the incoming logs. You must output the parsed logs in JSON format, with two keys "task" and "parsed_log", where "task" is "log_parsing" and "parsed_log" is a string.
[0110] Example:
[0111] {cases}
[0112] For example,
[0113] The log anomaly detection sample prompt in the interpretation content of step S204 is:
[0114] You will get some log anomaly detection samples in JSON format, which you need to use to determine whether the incoming logs are abnormal. You must output the results in JSON format, including two keys "task" and "category", where "task" is "anomaly_detection" and "category" is "Normal" or "Anomalous".
[0115] Example:
[0116] {cases}
[0117] For example,
[0118] The log parsing final rule prompt in the interpretation content of step S204 is:
[0119] You will be given some log parsing rules in JSON format, which you need to use to parse the incoming logs. You must output the parsed logs in JSON format, with two keys "task" and "parsed_log", where "task" is "log_parsing" and "parsed_log" is a string.
[0120] Rules (here, rules represent a sequence of log parsing optimization rules):
[0121] {rules}
[0122] For example,
[0123] The final rule prompt for log anomaly detection in the interpretation content of step S204 is:
[0124] You will get some log anomaly detection rules in JSON format, and you need to use these rules to determine whether the incoming logs are abnormal. You must output the results in JSON format, including two keys "task" and "category", where "task" is "anomaly_detection" and "category" is "Normal" (normal) or "Anomalous" (abnormal).
[0125] Rules (here, rules represent a sequence of rules for optimizing log anomaly detection):
[0126] {rules}
[0127] In one embodiment, before step S205, the rule-based log analysis method further includes the following sub-steps S101-S104:
[0128] S101, inputting training log samples and sample prompts into an initial log analysis model, and guiding the initial log analysis model to output sample label samples through the sample prompts.
[0129] If the training log sample represents the log parsing task, the corresponding sample prompt is the log parsing sample prompt. If the training log sample represents the log anomaly detection task, the corresponding sample prompt is the log anomaly detection sample prompt. Different log analysis tasks will correspond to corresponding sample label samples. Here we regard it as a unified sample label sample.
[0130] S102: Input the training log samples and rule prompts into the initial log analysis model, and guide the initial log analysis model to output rule label samples through the rule prompts.
[0131] If the training log sample represents the log parsing task, the corresponding rule prompt is the log parsing rule prompt. If the training log sample represents the log anomaly detection task, the corresponding rule prompt is the log anomaly detection rule prompt. Different log analysis tasks will correspond to corresponding rule label samples. Here we regard it as a unified rule label sample.
[0132] S103: Compare the sample labels in the sample label sample with the corresponding training labels, filter out the mismatched sample labels and mark them as rejected labels.
[0133] S104: Compare the rule labels in the rule label sample with the corresponding training labels, screen out matching rule labels and mark them as accepted labels.
[0134] Specifically, the training log samples The sample prompt obtained in step S204 is used as the input of the initial log analysis model to generate a sample label sample . The training log samples The rule prompts designed in step S204 are used as the input of the initial log analysis model to generate rule label samples .
[0135] The samples whose sample labels generated based on the sample prompts do not match the true labels of the selected samples are regarded as rejected label samples, that is, The samples whose rule labels generated based on the rule prompts match the true labels of the selected samples are taken as the accepted label samples, that is, , and thus the training set is expanded to .
[0136] In step S205, the parameters of the initial log analysis model are fine-tuned by using a comparison preference optimization method to obtain a log analysis model.
[0137] In one embodiment, the comparative preference optimization method includes calculating a preference function and calculating a behavior cloning function, and the part of "fine-tuning the parameters of the initial log analysis model by the comparative preference optimization method" in step S205 includes the following sub-steps S2051-S2053:
[0138] S2051. Calculate a preference function based on the rejection label and the acceptance label.
[0139] S2052. Calculate the behavior cloning function based on the accepted label.
[0140] S2053. Combine the preference function and the behavior cloning function to construct a total loss function, and use the total loss function to fine-tune the parameters of the initial log analysis model.
[0141] Specifically, the rejected label samples are , the accepted label sample is , the contrastive preference optimization (CPO) method is used to fine-tune the parameters of the model, thereby optimizing the initial log analysis model. For example, the number of parameters of LLM is about 8 billion, and LLM models with a number of parameters of about 8 billion are usually defined as small models. The calculation of CPO includes two items: preference function calculation and behavior cloning (BC) function calculation. Both preference function and behavior cloning function are loss functions.
[0142] Preference function Calculated by the following formula:
[0143] ;
[0144] Among them, the preference function To optimize model parameters ; Indicates that the parameter is Model; Indicates that from the training set The expected value of the sample in represents a log sample, Indicates rejection of label samples, Indicates acceptance of labeled samples; Indicates that given input Under the condition of The logarithmic probability of Indicates that given input Under the condition of The logarithmic probability of It is a hyperparameter, usually set to 0.5, which is used to balance the weights of the two logarithmic probability terms; is the Sigmoid function, which is defined as:
[0145] ;
[0146] The output range of the Sigmoid function is (0, 1), that is, for any real number x, the value of the Sigmoid function is between 0 and 1. The Sigmoid function has the following characteristics:
[0147] When the value of x is large, will be very small, so Close to 1;
[0148] When the value of x is small, will be very large, so Close to 0.
[0149] The input of the Sigmoid function is ,if The value of is very large, resulting in the output Close to 1, which means accepting labeled samples Relative to the rejection label sample The model considers it more likely to appear. This indicates that the model is more inclined to choose This result is "pulling" the ideal output. Therefore, by pulling the ideal output closer, the model's tendency to the correct or better solution is increased, that is, the model learns to be more inclined to generate the ideal output. ability.
[0150] On the contrary, if The value of is very small, resulting in the output Close to 0, which means It is only a possible choice for the model, but not the best choice, which means that the model should reduce its inclination towards such output, that is, "push away" the undesirable output. Therefore, by pushing away the undesirable output, the model's inclination towards undesirable output is reduced, that is, the model is prevented from generating outputs that do not conform to the rules or have poor effects. Through the pull-in and push-out mechanism, it helps to optimize the model's learning process and improve its performance and generalization ability.
[0151] BC Function Calculated by the following formula:
[0152] ;
[0153] The BC (Behavioral Cloning) function maximizes The probability of achieving Do not deviate from the distribution of accepted samples, thus ensuring the stability of the model optimization process.
[0154] Finally, the total loss function is a combination of the preference function and the behavior cloning function. The total loss function is calculated by the following formula:
[0155] ;
[0156] in, Represents a hyperparameter used to control the strength of the BC regularization term, usually set to 0.1.
[0157] By adding the BC term, the model can maintain the stability and generalization ability of the original task while learning the preference, and prevent the model from overfitting to the preference label. By combining the preference function with the behavior cloning function, the total loss function not only considers adjusting the model output according to the preference when optimizing the model, making the model output more accurate, but also ensures that the model will not deviate from the distribution of the accepted label samples while learning the preference, thus ensuring the stability of the model optimization process. This combination improves the performance and reliability of the model.
[0158] In step S206, the log to be analyzed is input into the log analysis model, and the final rule prompt corresponding to the log analysis task in the rule library is called according to the log analysis task represented by the log to be analyzed. The log analysis model is guided to output the final result through the final rule prompt, and the log to be analyzed is analyzed through the final result.
[0159] The final results include parsed logs and logs after anomaly detection, and the logs are analyzed accordingly through the parsed logs and logs after anomaly detection.
[0160] For example,
[0161] The log before parsing is "generating core.2473";
[0162] For the log parsing task, according to the final rule prompt, the parsed log is output, that is, the final result is as follows:
[0163] { "task": "log_parsing",
[0164] "parsed_log": "generating core.<*>"}
[0165] Finally, a regular expression is used to extract the final answer from the output (i.e. generating core.<*>).
[0166] In one embodiment, the rule-based log analysis method further includes the following sub-steps A1-A3:
[0167] A1. Get the test set ,The test set contains test log samples and test label samples. Each test log in the test log sample uniquely corresponds to a test label in the test label sample.
[0168] A2. Input the test log sample into the log analysis model to obtain the output result.
[0169] A3. Compare the output results with the test label samples and determine the accuracy of the log analysis model based on the comparison results.
[0170] The comparison results of rule prompts and sample prompts on the log parsing task are compared. The comparison results are shown in Table 1:
[0171] Table 1
[0172]
[0173] GA (Grouping Accuracy) is used to evaluate the ability to correctly group logs belonging to the same template. GA is defined as the ratio of the number of correctly grouped logs to the total number of logs. A log is considered correctly grouped if and only if the log template is consistent with the same group of logs in the real situation.
[0174] FGA (F1 score group accuracy) is a template-level metric that focuses on the proportion of correctly grouped templates rather than individual logs. is the actual correct number of templates in the real situation, The number of templates generated for the log parser. If is the number of templates correctly parsed by the log parser, and defines the grouping accuracy (PGA) as , the recall rate of group accuracy (RGA) is , and then calculate FGA as their harmonic mean, that is, .
[0175] FTA (F1 score template accuracy) is a template-level metric calculated based on the proportion of correctly identified templates. It is calculated as the harmonic mean of the precision and recall of the template accuracy. The difference is that a template is considered correctly identified if and only if the log messages of the parsed template share the same true template and all the tokens of the template are exactly the same as the tokens of the true template.
[0176] As shown in Table 1, for the log parsing task, in the three categories (GA, FGA, FTA), the accuracy of rule prompts is higher than that of sample prompts.
[0177] The comparison results of rule prompts and sample prompts on the log anomaly detection task are compared. The comparison results are shown in Table 2:
[0178] Table 2
[0179]
[0180] Precision (P) represents the percentage of abnormal logs correctly detected among all logs predicted to be abnormal, that is, .
[0181] Recall (R) represents the percentage of correctly detected abnormal logs among all the logs that are actually abnormal, that is, .
[0182] The F1 score (F1) represents the harmonic mean of precision and recall, i.e. .
[0183] A true positive (TP) means that the log is actually abnormal and the model also detects it as abnormal; a false negative (FN) means that the log is actually normal, but the model mistakenly detects it as abnormal; a false positive (FP) means that the log is actually abnormal, but the model mistakenly detects it as normal.
[0184] As shown in Table 2, for the log anomaly detection task, in the three categories (precision, recall, and F1 score), the accuracy of the rule prompt is higher than the accuracy of the sample prompt. Combining Tables 1 and 2, it can be seen that this application constructs rule prompts and then embeds the optimized rule sequence into the rule prompts to obtain the final rule prompts, which can accurately capture the potential rules in the logs and understand the log data more accurately, thereby enhancing the reasoning and generalization capabilities of the log analysis model and improving the accuracy of log analysis.
[0185] In practical applications, the format (e.g., the arrangement of information recorded in a log) or the content (e.g., the specific type of information recorded) of a log may change over time or as the system is updated. Since rule-based log analysis methods rely on rules to analyze logs, when the log format or content changes, these changes can be adapted by updating these rules. This may include adding new rules to handle new log formats, or modifying existing rules to correctly parse the changed content. In traditional machine learning methods, if the format or content of the input data changes, it is usually necessary to retrain the model with new data to ensure that the model can understand and process the new data format or content. However, in rule-based methods, since the rules can be updated independently of the model training process, there is no need to retrain the entire model. The model has already learned how to parse and understand the logs based on the rules, and it is only necessary to ensure that the rule set is consistent with the new log format or content.
[0186] Based on the same inventive concept, the embodiment of the present application also provides a device for implementing the rule-based log analysis involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more rule-based log analysis device embodiments provided below can refer to the limitations of the rule-based log analysis method above, and will not be repeated here.
[0187] In an exemplary embodiment, Figure 3 As shown, a rule-based log analysis device is provided, comprising:
[0188] A creation module 310 is used to create a training set for representing different log analysis tasks and design corresponding prompts for the different log analysis tasks, wherein the training set includes training log samples and training label samples, and each training log in the training log samples uniquely corresponds to a training label in the training label samples;
[0189] A guiding module 320, configured to input the training set and the prompt into an initial log analysis model, and guide the initial log analysis model to output a rule corresponding to the log analysis task through the prompt;
[0190] A storage module 330 is used to prioritize the rules to obtain an optimized rule sequence, and store the optimized rule sequence in a rule base located in the initial log analysis model;
[0191] An embedding module 340 is used to obtain sample prompts and rule prompts, embed the optimization rule sequence into the rule prompts, and obtain a final rule prompt;
[0192] A fine-tuning module 350, configured to fine-tune the parameters of the initial log analysis model by using a comparison preference optimization method to obtain a log analysis model;
[0193] The output module 360 is used to input the log to be analyzed into the log analysis model, call the final rule prompt corresponding to the log analysis task in the rule base according to the log analysis task represented by the log to be analyzed, and guide the log analysis model to output the final result through the final rule prompt, and analyze the log to be analyzed through the final result.
[0194] As an optional implementation, the log analysis task includes a log parsing task and a log anomaly detection task, a log parsing prompt is designed for the log parsing task, and a log anomaly detection prompt is designed for the log anomaly detection task; the guiding module 320 is specifically used to:
[0195] Inputting the training set and the log parsing prompt used to characterize the log parsing task into the initial log analysis model, and outputting log parsing rules;
[0196] The training set used to characterize the log anomaly detection task and the log anomaly detection prompt are input into the initial log analysis model, and a log anomaly detection rule is output.
[0197] As an optional implementation, in the aspect of prioritizing the rules, the module 330 is stored, specifically for:
[0198] In the process of the initial log analysis model outputting the rules corresponding to the log analysis task, recording the number of occurrences of each rule;
[0199] All the rules are prioritized in descending order of the number of occurrences.
[0200] As an optional implementation, the rule is in a form of a natural language description, a regular expression, and a code form; the rule is generated based on a vocabulary space predefined in the initial log analysis model.
[0201] As an optional implementation, the rule-based log analysis device further includes:
[0202] Inputting the training log sample and the sample prompt into the initial log analysis model, and guiding the initial log analysis model to output a sample label sample through the sample prompt;
[0203] Inputting the training log sample and the rule prompt into the initial log analysis model, and guiding the initial log analysis model to output a rule label sample through the rule prompt;
[0204] Compare the sample labels in the sample label sample with the corresponding training labels, filter out the sample labels that do not match and mark them as rejected labels;
[0205] The rule labels in the rule label samples are compared with the corresponding training labels, and matching rule labels are screened out and marked as accepted labels.
[0206] As an optional implementation, the comparison preference optimization method includes calculating the preference function and calculating the behavior clone function, and the fine-tuning module 350 is specifically used for:
[0207] calculating the preference function based on the rejection labels and the acceptance labels;
[0208] Calculating the behavior cloning function based on the accepted label samples;
[0209] The preference function and the behavior cloning function are combined to construct a total loss function, and the parameters of the initial log analysis model are fine-tuned by the total loss function.
[0210] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store rule-based log analysis data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a rule-based log analysis method is implemented.
[0211] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0212] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0213] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0214] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0216] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0217] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0218] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0219] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.< / number> < / number> < / number> < / number> < / number>
Claims
1. A rule-based log analysis method, characterized in that: The rule-based log analysis method includes: Creating a training set for characterizing different log analysis tasks, and designing corresponding prompts for the different log analysis tasks, wherein the training set includes training log samples and training label samples, and each training log in the training log samples uniquely corresponds to a training label in the training label samples; Inputting the training set and the prompt into an initial log analysis model, and guiding the initial log analysis model to output a rule corresponding to the log analysis task through the prompt; Prioritizing the rules to obtain an optimized rule sequence, and storing the optimized rule sequence in a rule base within the initial log analysis model; Obtaining sample prompts and rule prompts, embedding the optimization rule sequence into the rule prompts, and obtaining final rule prompts; Fine-tuning the parameters of the initial log analysis model by using a comparative preference optimization method to obtain a log analysis model; Input the log to be analyzed into the log analysis model, call the final rule prompt corresponding to the log analysis task in the rule base according to the log analysis task represented by the log to be analyzed, guide the log analysis model to output a final result through the final rule prompt, and analyze the log to be analyzed through the final result; The log analysis task includes a log parsing task and a log anomaly detection task, the prompt includes a log parsing prompt and a log anomaly detection prompt, and the rule includes a log parsing rule and a log anomaly detection rule; the inputting the training set and the prompt into the initial log analysis model, and guiding the initial log analysis model to output a rule corresponding to the log analysis task through the prompt, comprises: Inputting the training set and the log parsing prompt used to characterize the log parsing task into the initial log analysis model, and outputting log parsing rules; Inputting the training set used to characterize the log anomaly detection task and the log anomaly detection prompt into the initial log analysis model, and outputting a log anomaly detection rule; The prioritizing the rules comprises: In the process of the initial log analysis model outputting the rules corresponding to the log analysis task, recording the number of occurrences of each rule; All the rules are prioritized in descending order of the number of occurrences.
2. The rule-based log analysis method according to claim 1, characterized in that: The form of the rule is one of a natural language description form, a regular expression form and a code form; the rule is generated based on a vocabulary space predefined in the initial log analysis model.
3. The rule-based log analysis method according to claim 1, characterized in that: Before the step of fine-tuning the parameters of the initial log analysis model by using the comparison preference optimization method, the rule-based log analysis method further includes: Inputting the training log sample and the sample prompt into the initial log analysis model, and guiding the initial log analysis model to output a sample label sample through the sample prompt; Inputting the training log sample and the rule prompt into the initial log analysis model, and guiding the initial log analysis model to output a rule label sample through the rule prompt; Compare the sample labels in the sample label sample with the corresponding training labels, filter out the sample labels that do not match and mark them as rejected labels; The rule labels in the rule label samples are compared with the corresponding training labels, and matching rule labels are screened out and marked as accepted labels.
4. The rule-based log analysis method according to claim 3, characterized in that: The comparative preference optimization method includes a calculation preference function and a calculation behavior cloning function. The comparative preference optimization method is used to fine-tune the parameters of the initial log analysis model, including: calculating the preference function based on the rejection labels and the acceptance labels; Calculating the behavior cloning function based on the acceptance label; The preference function and the behavior cloning function are combined to construct a total loss function, and the parameters of the initial log analysis model are fine-tuned through the total loss function.
5. A rule-based log analysis device, characterized in that: The rule-based log analysis device comprises: A creation module, used to create a training set for characterizing different log analysis tasks and design corresponding prompts for the different log analysis tasks, wherein the training set includes training log samples and training label samples, and each training log in the training log samples uniquely corresponds to a training label in the training label samples; A guiding module, used for inputting the training set and the prompt into an initial log analysis model, and guiding the initial log analysis model to output a rule corresponding to the log analysis task through the prompt; A storage module, used for prioritizing the rules to obtain an optimized rule sequence, and storing the optimized rule sequence in a rule base located in the initial log analysis model; An embedding module, used for obtaining sample prompts and rule prompts, embedding the optimized rule sequence into the rule prompts, and obtaining final rule prompts; A fine-tuning module, used to fine-tune the parameters of the initial log analysis model by using a comparative preference optimization method to obtain a log analysis model; An output module, used for inputting the log to be analyzed into the log analysis model, calling the final rule prompt corresponding to the log analysis task in the rule base according to the log analysis task represented by the log to be analyzed, and guiding the log analysis model to output a final result through the final rule prompt, and analyzing the log to be analyzed through the final result; The log analysis task includes a log parsing task and a log anomaly detection task, the prompt includes a log parsing prompt and a log anomaly detection prompt, and the rule includes a log parsing rule and a log anomaly detection rule; the guiding module is specifically used to: Inputting the training set and the log parsing prompt used to characterize the log parsing task into the initial log analysis model, and outputting log parsing rules; Inputting the training set used to characterize the log anomaly detection task and the log anomaly detection prompt into the initial log analysis model, and outputting a log anomaly detection rule; In terms of prioritizing the rules, the storage module is specifically used to: In the process of the initial log analysis model outputting the rules corresponding to the log analysis task, recording the number of occurrences of each rule; All the rules are prioritized in descending order of the number of occurrences.
6. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the rule-based log analysis method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the rule-based log analysis method described in any one of claims 1 to 4 are implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the rule-based log analysis method described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Error log processing method and device, electronic equipment and readable storage medium
CN117874236A
Model increment training method and system and model applied to log anomaly detection
CN118228801A
Log analysis method and device based on large language model
CN119416770A