Dynamic adaptive log analysis method and system based on local large language model
By deploying large language models locally to generate and verify log parsing rules, and combining the dynamic adaptive log parsing method of the traditional log parsing engine, the problem of traditional log parsing methods relying on manual writing rules and poor cross-domain adaptability is solved, and efficient, secure and real-time log parsing effects are achieved.
Patent Information
- Application Number
- CN202510178789.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional log analysis methods rely on manual writing rules, which are costly and have poor cross-domain adaptability. The cloud-based large language model log analysis solution has poor real-time and data leakage risks, especially in areas such as finance and medical care that require high security and real-time.
The dynamic adaptive log analysis method based on the local large language model is adopted to generate and verify the parsing rules through the locally deployed large language model, combine the traditional log analysis engine to realize real-time structured log analysis, and improve the analysis coverage and timeliness through dynamic rule expansion and injection modules.
It significantly reduces the cost of manual writing rules, improves the accuracy and adaptability of log analysis, ensures data security, improves the timeliness and efficiency of log analysis, and is suitable for areas with high requirements for security and real-timeness.
Smart Images

Figure CN120106044A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of log analysis, and in particular to a dynamic adaptive log analysis method and system based on a local large language model. Background Art
[0002] Traditional log parsing methods such as Drain and Spell are highly dependent on manually written parsing rule expressions, which not only leads to high labor costs, but also poor cross-domain adaptability. In recent years, although log parsing solutions based on large language models (LLMs) (including LogGPT) do not require manual rule writing, they usually rely on cloud services and have problems such as poor real-time performance and data leakage risks. In particular, these solutions are not applicable to fields such as finance and medical care that have high requirements for security and real-time performance.
[0003] To this end, those skilled in the art have proposed a dynamic adaptive log parsing method and system based on a local large language model to solve the problems raised by the background technology. Summary of the invention
[0004] 1. Technical problems to be solved
[0005] A dynamic log parsing method based on a locally deployed large language model (LLM) and a traditional log parsing method is provided. The locally deployed large language model (LLM) is used to generate parsing rules, verify them, and dynamically expand the rule base, thereby solving the problems of high cost of manual rule writing and poor dynamic adaptability. At the same time, data security and log parsing timeliness are guaranteed, and the cost of log parsing can be well controlled.
[0006] 2. Technical solution
[0007] The overall framework includes three core modules: rule generation and verification module based on local large language model (LLM), real-time log parsing module, and dynamic rule expansion and injection, realizing the closed loop of "sampling-generation-parsing-expansion". The overall framework is shown in the figure below. Figure 1 .
[0008] Main module function description:
[0009] 1. Local Large Language Model (LLM) rule generation and verification module
[0010] (1) Layered weighted sampling: assign sampling weights according to log levels (ERROR>WARN>INFO), giving priority to covering key logs;
[0011] (2) Additional sampling: If the proportion of unresolved logs exceeds the threshold (including the unresolved rate exceeding 5%), additional sampling is triggered, and cluster analysis is combined to dynamically supplement samples to ensure sample coverage;
[0012] (3) Generate rule parsing prompts. The structured prompts guide LLM to generate standardized regular expressions.
[0013] (4) Rule generation: Submit the sampled logs and prompts to the local large language model, generate log parsing rules, and put them into the rule cache pool after rule verification.
[0014] 2. Log real-time analysis module
[0015] (1) According to the parsing rules, use the traditional log parsing engine (including enhanced Driven) to parse the logs one by one and output structured parsed logs;
[0016] (2) Mark unparsed logs. When the unparsed log rate triggers a threshold, dynamic rule expansion and injection are performed.
[0017] 3. Dynamic rule expansion and injection module
[0018] (1) When the log parsing engine detects a new log template, it puts the new log into the cache queue and performs deduplication operations, waiting for the dynamic rules to be triggered to avoid blocking real-time parsing.
[0019] (2) When the dynamic rule extension is triggered, the rule generator is asynchronously called to generate parsing rules. After automatic verification (including when the coverage rate is >90%), the rules are dynamically injected into the rule cache pool and persistence pool so that the log parsing engine can perform real-time log parsing on the new log template.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] 1. The present invention realizes the automatic generation and verification of log parsing rules through a rule generation and verification module based on a local large language model, significantly reducing the cost of manually writing rules while improving the accuracy and adaptability of the rules; thereby solving the problem that traditional log parsing methods are highly dependent on manual labor and have poor cross-domain adaptability.
[0022] 2. The present invention realizes active learning and dynamic rule expansion of unparsed logs through a dynamic rule expansion and injection module; when the unparsed log rate exceeds a preset threshold, the system will asynchronously generate new rules and dynamically inject them into the rule cache pool without manual intervention, thereby improving the coverage and timeliness of log parsing.
[0023] 3. The present invention adopts a locally deployed large language model to avoid the risk of data leakage and improve the efficiency of real-time analysis, thereby making the present invention have broad application prospects in fields such as finance and medical care that have high requirements for security and real-time performance.
[0024] 4. The present invention combines the advantages of traditional log parsing engines and local large language models to realize a closed-loop process of "sampling-generation-parsing-expansion", effectively improving the efficiency and accuracy of log parsing; at the same time, the present invention also provides functions such as rule cache pool and persistence pool, further enhancing the stability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is the overall framework diagram of the system of the present invention;
[0026] Figure 2 Active learning and rule dynamic expansion flow chart for the present invention;
[0027] Figure 3 A flow chart is generated for the rules of the present invention. DETAILED DESCRIPTION
[0028] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0029] Embodiment: The present invention provides a dynamic adaptive log parsing method based on a local large language model, comprising:
[0030] Step 1: Layered weighted sampling step, assigning sampling weights according to log levels, giving priority to covering key logs;
[0031] Step 2: Local Large Language Model (LLM) rule generation and verification step, guiding the locally deployed LLM to generate regular expressions for log parsing through structured prompts, and perform rule verification;
[0032] Step 3: real-time log parsing step, using a traditional log parsing engine to parse the logs one by one according to the generated rules and output structured logs;
[0033] Step 4: Active learning and dynamic rule expansion step. When the unparsed log rate exceeds the preset threshold, asynchronous rule generation is triggered. After verification, new rules are dynamically injected into the rule cache pool without manual intervention.
[0034] From the above, it can be seen that this method prioritizes key logs through stratified weighted sampling, uses the local large language model (LLM) to automatically generate and verify log parsing rules, combines with the traditional log parsing engine to achieve real-time structured parsing of logs, and automatically triggers rule expansion and injection when the unparsed log rate exceeds the preset threshold without manual intervention; thereby significantly reducing the cost of manual rule writing, improving the accuracy and adaptability of log parsing, while ensuring data security, and improving the timeliness and efficiency of log parsing, which is particularly suitable for fields with high requirements for security and real-time performance.
[0035] Specifically, the specific steps of step 1 are as follows:
[0036] Step 1.1: Layered weighted sampling: weighted sampling of historical logs according to log levels, such as 50% for ERROR logs, 30% for WARNING logs, and 20% for INFO logs. Weighted sampling can also be performed based on whether the logs contain specified keywords.
[0037] Step 1.2: Log preprocessing: Perform preliminary preprocessing on log samples, such as removing blank lines, duplicate logs, illegal characters (including garbled characters, unresolvable symbols), etc., and unify timestamps and log levels to facilitate subsequent operations;
[0038] Step 1.3: Dynamically adjust the sampling volume. If the number of a certain type of logs is insufficient, redistribute the remaining sampling volume in proportion. For example, if the number of ERROR-level logs is small, the total number of samples collected may be insufficient, so the remaining sampling volume is redistributed in proportion.
[0039] Step 1.4: Sample supplementation, dynamically monitor the real-time log parsing failure rate. If the log parsing failure rate exceeds the threshold (including 50%), additional sampling is performed to supplement the collected samples.
[0040] From the above, we can see that by finely controlling the log sampling ratio and preprocessing process, the priority coverage of key logs and the acquisition of high-quality samples are ensured; it not only improves the pertinence and efficiency of log parsing, but also can dynamically adjust the sampling volume and supplement samples when the log distribution is uneven or the parsing failure rate is high, effectively enhancing the representativeness and completeness of the samples, thereby further improving the accuracy and adaptability of subsequent log parsing, laying a solid foundation for the entire dynamic adaptive log parsing method.
[0041] More specifically, in the stratified weighted sampling step, in order to more accurately control the sampling ratio of logs at different levels, a weight distribution formula is introduced; for example, let the weights of the ERROR, WARNING, and I NFO log levels be w, respectively. error 、w warning 、w info , and satisfies w error +w warning +w info =1; then the number of samples is n i It is expressed as:
[0042] n i =N×w i ;
[0043] Where N is the total number of samples, i represents the log level; by adjusting the weight w i , to flexibly control the sampling quantity of logs of different levels.
[0044] Specifically, the specific steps of step 2 are as follows:
[0045] Step 2.1: Design a prompt template to clarify the requirements for generating rules. Design a large language model (LLM) prompt, such as "Task: Generate log parsing rules, requirements: - Use <*> to mark dynamic variables (including IP, numbers, - Generate regular expressions, use brackets to capture variables, example input: User 192.168.1.1 logged in, example output: User <*> logged in", and perform preliminary verification on the large language model (LLM) according to the designed prompts to obtain the best results. The large language model (LLM) can also automatically generate prompts, such as "Please generate prompts to implement the task of generating log parsing rules, such as 'Task: Generate log parsing rules, requirements: - Use <*> to mark dynamic variables (including IP, numbers, - Generate regular expressions, use brackets to capture variables'";
[0046] Step 2.2: Generate rules, set log submission batches to avoid submitting information that exceeds the single submission limit of the large language model (LLM), merge the prompt template and single batch log information and submit them to the large language model (LLM), and generate a list of candidate parsing rules;
[0047] Step 2.3: Rule verification: Use historical logs to verify candidate rules and set verification match rate thresholds (including 90%). Keep rules with a match rate greater than the threshold and perform dynamic variable verification to ensure that variable parts, such as IP addresses and timestamps, are correctly captured.
[0048] Step 2.4: Conflict detection: Check for overlapping rules. When multiple rules match the same log, the longest matching rule is retained first.
[0049] Step 2.5: Exception handling. After steps 2.1 to 2.4, if no rules are left or the Large Language Model (LLM) fails to generate rules, a default matching rule is provided for temporary matching, such as replacing numbers and IP addresses with <*>, and marking the log for manual review for subsequent optimization;
[0050] Step 2.6: Clean up the rules in the rule cache pool and dynamically clean up the parsing rules that have not been used for a long time in the rule cache pool based on the principle of most recent and most frequently used.
[0051] From the above, we can see that by designing refined prompt templates to guide the local large language model (LLM) to generate log parsing rules, and performing strict rule verification, conflict detection and exception handling, it is ensured that the generated parsing rules are both in line with the requirements and efficient and reliable; in addition, by dynamically cleaning up long-term unused rules in the rule cache pool, the performance and stability of the system are effectively maintained; it not only realizes the automatic generation and verification of log parsing rules, significantly reducing the cost of manual intervention, but also improves the accuracy and adaptability of log parsing by continuously optimizing the rule base, providing strong support for the fast and accurate parsing of log data.
[0052] More specifically, in the LLM rule generation and verification step, in order to guide LLM to generate high-quality parsing rules, a loss function is designed to evaluate the difference between the generated rules and the expected goals; the loss function can be constructed based on indicators such as the accuracy, recall rate or F1 score of rule matching; for example, let y be the true label, For the labels predicted by the rules generated by LLM, the loss function L can be expressed as:
[0053]
[0054] By minimizing the loss function, the quality of rules generated by LLM can be optimized.
[0055] Specifically, the specific steps of step 3 are as follows:
[0056] Step 3.1: Real-time log parsing. The log parsing engine (including Dra in) parses the logs one by one in real time according to the parsing rules in the rule cache pool and outputs structured logs (template + parameters).
[0057] Step 3.2: New template detection. If parsing fails, mark the log as "unparsed", calculate its hash value to remove duplicates, and put it into the cache queue so that the rule generator can be called later to generate new rules.
[0058] From the above, we can see that the log parsing engine applies the parsing rules in the rule cache pool in real time to parse the logs one by one, realizing the rapid structured processing of log data; at the same time, the new template detection mechanism can timely discover and mark unparsed logs, and put them into the cache queue after deduplication by calculating the hash value, providing new learning samples for subsequent rule generators; it not only ensures the real-time and accuracy of log parsing, but also can dynamically adapt to newly emerging log templates, and continuously improve the coverage and efficiency of log parsing by continuously learning and expanding the parsing rule base, providing a strong guarantee for the comprehensive parsing and effective utilization of log data.
[0059] Specifically, the specific steps of step 4 are as follows:
[0060] Step 4.1: Active learning: When the unparsed log rate exceeds a certain threshold (including 5%), or the waiting time for unparsed logs exceeds a threshold (including 30 minutes), the clustered results are asynchronously submitted to the rule generator to generate new extraction rules;
[0061] Step 4.2: Same as steps 2.3 to 2.5, perform rule verification, conflict detection and exception handling on the new parsing rules;
[0062] Step 4.3: Dynamic expansion, the new rules processed in step 4.2 are dynamically injected into the rule cache pool and the rule persistence pool.
[0063] From the above, we can see that the unparsed log situation is dynamically monitored through the active learning mechanism. When the unparsed rate or waiting time exceeds the preset threshold, the asynchronous generation of new rules is triggered; the newly generated rules are dynamically injected into the rule cache pool and persistence pool after strict verification, conflict detection and exception handling; it not only realizes the timely response and processing of unparsed logs, but also significantly improves the coverage and accuracy of log parsing by continuously expanding and optimizing the parsing rule library; in addition, the dynamic rule injection mechanism does not require human intervention, ensuring the automatic operation and efficiency of the system, and providing strong support for the intelligence and adaptability of log parsing.
[0064] More specifically, in the active learning and dynamic rule expansion steps, a clustering algorithm can be used to effectively cluster unparsed logs and trigger asynchronous rule generation. The core of the clustering algorithm is to define a reasonable distance metric and a cluster center selection strategy. For example, in the K-means algorithm, the goal is to minimize the sum of squares within the cluster:
[0065]
[0066] Where k is the number of clusters, C i is the i-th cluster, μ i is the center point of the i-th cluster;
[0067] By iteratively optimizing cluster centers and cluster assignments, efficient log clustering and new rule generation triggering can be achieved.
[0068] Furthermore, the effects of the dynamic adaptive log parsing method based on the local large language model of the embodiment and the current traditional log parsing method (comparative example) are compared to obtain the following table:
[0069]
[0070]
[0071] As can be seen from the above table, the dynamic adaptive log parsing method based on the local large language model is superior to the traditional log parsing method in terms of rule generation, adaptability, parsing efficiency, parsing accuracy, rule update mechanism, system stability and security. It can significantly reduce the cost of manual intervention, improve the efficiency and accuracy of log parsing, and provide strong support for the enterprise's log management and data analysis.
[0072] Dynamic and adaptive log parsing system based on local large language model, such as Figures 1 to 3 As shown, the dynamic adaptive log parsing method based on the local large language model includes:
[0073] A stratified weighted sampling module is used to assign sampling weights according to log levels and trigger additional sampling; the stratified weighted sampling module triggers clustering-based additional sampling when the unparsed log exceeds the threshold; the stratified weighted sampling module can dynamically adjust the sampling volume when the collected sample logs are less than the predetermined sample volume;
[0074] LLM rule generation and verification module, the LLM rule generation and verification module includes an LLM rule generator, which is used to guide the locally deployed LLM to generate and verify log parsing rules through structured prompts; the LLM rule generator automatically generates log parsing rules using a locally deployed large language model (LLM); the LLM rule generator includes a rule cache pool, which reuses historical rules through edit distance matching;
[0075] The log real-time parsing module is used to parse the logs in real time according to the rules in the rule cache pool;
[0076] The active learning and dynamic rule expansion module is used to asynchronously generate new rules and dynamically inject them into the rule cache pool when the unparsed log rate exceeds a preset threshold; the active learning and dynamic expansion mechanism includes actively learning new log parsing rules; the active learning and dynamic expansion mechanism includes dynamically injecting new rules into the log parsing engine without restarting the system.
[0077] As can be seen from the above, the system achieves high automation and intelligence of log parsing by integrating four modules: hierarchical weighted sampling, LLM rule generation and verification, real-time log parsing, and active learning and dynamic rule expansion. The system can intelligently allocate sampling weights according to the log level to ensure the priority processing of key logs, and effectively deal with the problem of uneven log distribution by additional sampling and dynamic adjustment of sampling volume. The LLM rule generation and verification module uses the locally deployed large language model to automatically generate and verify parsing rules, improves the accuracy and adaptability of the rules, and optimizes rule reuse through the rule cache pool and edit distance matching mechanism. The real-time log parsing module ensures the rapid structured processing of log data, while the active learning and dynamic rule expansion module can automatically trigger the generation and injection of new rules when the unparsed log rate exceeds the preset threshold without manual intervention. This not only significantly reduces the cost of manual rule writing and improves the efficiency and accuracy of log parsing, but also enhances the adaptability and stability of the system, providing strong support for the in-depth analysis and utilization of log data.
[0078] Working principle: First, the hierarchical weighted sampling module is used to intelligently collect log data, and sampling weights are assigned according to the log level. When necessary, additional sampling is triggered and the sampling volume is dynamically adjusted. Then, the LLM rule generation and verification module uses the locally deployed LLM to guide the generation and verification of log parsing rules through structured prompts. These rules are stored in the rule cache pool for reuse. The log real-time parsing module parses the logs in real time according to the rules in the rule cache pool and outputs structured logs. Finally, the active learning and dynamic rule expansion module monitors the situation of unparsed logs. When the unparsed log rate exceeds the preset threshold, new rules are asynchronously generated and dynamically injected into the rule cache pool, so that continuous optimization and expansion of rules can be achieved without human intervention. The entire system realizes a high degree of automation and intelligence in log parsing, which significantly improves the efficiency and accuracy of log parsing.
[0079] In summary, the present invention has the following beneficial effects:
[0080] Improved efficiency: Compared with traditional log parsing methods, the log parsing rules of this method are automatically generated without manual intervention, which significantly improves parsing efficiency.
[0081] Dynamic adaptability: Compared with traditional log parsing methods, this method can dynamically cover the parsing of new log templates without manual intervention, and can automatically adapt to new log templates, significantly improving log parsing coverage.
[0082] Enhanced security: Combined with the locally deployed Large Language Model (LLM), it prevents data leakage, improves real-time parsing efficiency, and meets compliance requirements in the financial, medical and other fields.
[0083] Reduced maintenance costs: Compared with traditional methods, this method automatically generates log parsing rules, significantly reducing human maintenance costs. Compared with other methods based on cloud-based large language models, this method does not need to transfer data to the cloud-based large language model (LLM), and only uses the large language model (LLM) to extract rules from sample logs, which can significantly reduce parsing costs.
[0084] The embodiment of the present application provides an electronic device, which is applicable to the above-mentioned dynamic adaptive log parsing method based on a local large language model, including:
[0085] Memory, used to protect computer programs and data;
[0086] Processor, used to run system programs.
[0087] The embodiment of the present application provides a computer storage medium, which is applicable to the above-mentioned dynamic adaptive log parsing method based on the local large language model, and performs hierarchical confidentiality management on the above-mentioned system and data in accordance with confidentiality management requirements.
[0088] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a system or a computer program product. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0089] The present application is described with reference to the flowcharts and / or block diagrams of the devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0090] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0092] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0093] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0094] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include transitory media such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, commodity or device including the elements.
[0096] The embodiments of the present invention are provided for the purpose of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A dynamic adaptive log parsing method based on a local large language model, characterized in that: include: Step 1: Assign sampling weights according to log levels, giving priority to covering key logs; Step 2: Use structured prompts to guide the locally deployed LLM to generate regular expressions for log parsing and perform rule verification; Step 3: Use the traditional log parsing engine to parse the logs one by one according to the generated rules and output structured logs; Step 4: When the unparsed log rate exceeds the preset threshold, asynchronous rule generation is triggered, and new rules are dynamically injected into the rule cache pool after verification without manual intervention.
2. The dynamic adaptive log parsing method based on the local large language model as claimed in claim 1, characterized in that: The specific steps of step 1 are as follows: Step 1.1: Layered weighted sampling: weighted collection of historical logs based on log levels and whether they contain specified keywords; Step 1.2: Log preprocessing: preprocess the log samples and unify the timestamp and log level to facilitate subsequent operations; Step 1.3: Dynamically adjust the sampling volume. If the number of logs of a certain type is insufficient, redistribute the remaining sampling volume in proportion; Step 1.4: Supplement samples and dynamically monitor the real-time log parsing failure rate.
3. The dynamic adaptive log parsing method based on the local large language model as claimed in claim 1, characterized in that: The specific steps of step 2 are as follows: Step 2.1: Design prompt templates, clarify the requirements for generating rules, and design large language model (LLM) prompts; Step 2.2: Generate rules, set log submission batches to avoid submitting information that exceeds the single submission limit of the large language model, merge the prompt template and single batch log information and submit them to the large language model to generate a list of candidate parsing rules; Step 2.3: Rule verification: Use historical logs to verify candidate rules and set a verification match rate threshold. Keep rules with a match rate greater than the threshold and perform dynamic variable verification to ensure that variable parts, such as IP addresses and timestamps, are captured correctly. Step 2.4: Conflict detection: Check for overlapping rules. When multiple rules match the same log, the longest matching rule is retained first. Step 2.5: Exception handling. After steps 2.1 to 2.4, if no rules are left or the large language model fails to generate rules, a default matching rule is provided for temporary matching, and the log is marked for manual review for subsequent optimization. Step 2.6: Clean up the rules in the rule cache pool and dynamically clean up the parsing rules that have not been used for a long time in the rule cache pool based on the principle of most recent and most frequently used.
4. The dynamic adaptive log parsing method based on the local large language model as claimed in claim 1, characterized in that: The specific steps of step 3 are as follows: Step 3.1: Real-time log parsing: the log parsing engine parses logs one by one in real time according to the parsing rules in the rule cache pool and outputs structured logs; Step 3.2: New template detection. If parsing fails, mark the log as "unparsed", calculate its hash value to remove duplicates, and put it into the cache queue to generate new rules.
5. The dynamic adaptive log parsing method based on the local large language model as claimed in claim 1, characterized in that: The specific steps of step 4 are as follows: Step 4.1: Active learning: When the unparsed log rate exceeds a certain threshold, or the waiting time of unparsed logs exceeds a threshold, the clustering results are asynchronously submitted to the rule generator to generate new extraction rules; Step 4.2: Same as steps 2.3 to 2.5, perform rule verification, conflict detection and exception handling on the new parsing rules; Step 4.3: Dynamic expansion, the new rules processed in step 4.2 are dynamically injected into the rule cache pool and the rule persistence pool.
6. Dynamic adaptive log parsing system based on local large language model, characterized by: The method for dynamically adapting log parsing based on a local large language model according to claims 1 to 7 comprises: Hierarchical weighted sampling module, used to assign sampling weights according to log levels and trigger additional sampling; LLM rule generation and verification module, which is used to guide the locally deployed LLM to generate and verify log parsing rules through structured prompts; The log real-time parsing module is used to parse the logs in real time according to the rules in the rule cache pool; Active learning and dynamic rule expansion module, used to asynchronously generate new rules and dynamically inject them into the rule cache pool when the unparsed log rate exceeds the preset threshold.
7. The dynamic adaptive log parsing system based on the local large language model as claimed in claim 6, characterized in that: The stratified weighted sampling module triggers cluster-based additional sampling when the unparsed log exceeds the threshold; The stratified weighted sampling module can dynamically adjust the sampling amount when the collected sample logs are less than the predetermined sample amount.
8. The dynamic adaptive log parsing system based on the local large language model as claimed in claim 6, characterized in that: The LLM rule generator automatically generates log parsing rules using a locally deployed large language model; The LLM rule generator includes a rule cache pool and reuses historical rules through edit distance matching.
9. The dynamic adaptive log parsing system based on the local large language model as claimed in claim 1, characterized in that: The active learning and dynamic expansion mechanism includes actively learning new log parsing rules; The active learning and dynamic expansion mechanism includes dynamic injection of new rules into the log parsing engine without restarting the system.
Citation Information
Cited By
Data analysis rule matching method and device, storage medium and electronic equipment
CN120658810A
Intelligent data processing system and method integrating table cleaning and relation graph
CN121524186A
Website data analysis method and device based on large model technology
CN121681912A