Operation and maintenance log online analysis and detection method and device based on large model, equipment and medium

By adopting large-model-based online analysis and detection methods in operation and maintenance log analysis, the problems of low efficiency and poor detection accuracy of operation and maintenance log analysis are solved, and more efficient and accurate abnormal detection is achieved.

CN119988147AInactive Publication Date: 2025-05-13SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510459285.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency and poor detection accuracy in operation and maintenance log analysis, especially when facing complex and changeable abnormal patterns, false alarms and missed alarms often occur.

Method used

The online analysis and detection method of operation and maintenance logs based on large models is adopted, and the initial log is preprocessed and regular matching is performed, and the log template is generated and adjusted by preset large language models to achieve accurate log analysis and exception detection.

Benefits of technology

It improves the analysis efficiency of operation and maintenance logs and the accuracy of abnormal detection, and can capture complex abnormal patterns more accurately, reducing false alarms and missed reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988147A_ABST
    Figure CN119988147A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance log online analysis and detection method and device based on a large model, equipment and a medium, and relates to the field of operation and maintenance log paragraption.The method comprises the steps that in a target operation and maintenance system, obtained processed logs and all first regular expressions are subjected to regular matching; when the matching fails, determining a newly added log template based on a preset large language model and the processed log, and respectively determining the similarity between the newly added log template and each log template in a preset log template set; combining the log template with the similarity greater than a preset threshold and the newly added log template to obtain a target log template, and obtaining a target log by utilizing the target log template and the processed log; adjusting a preset large language model by using a log test set and an abnormal log example determined from the target log; and performing anomaly detection on logs generated in the system by utilizing the adjusted model, and displaying detected anomaly information on a target user interface. Therefore, the accuracy of log anomaly detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of operation and maintenance log analysis, and in particular to a large model-based online analysis and detection method, device, equipment and medium for operation and maintenance logs. Background Art

[0002] In today's digital age, the scale and complexity of various information systems are constantly increasing, and operation and maintenance work has become increasingly important. As a key data source for recording the system's operating status and operation process, operation and maintenance logs play a vital role in ensuring the stable operation of the system. However, existing technologies have many shortcomings in operation and maintenance log analysis.

[0003] On the one hand, the main methods of operation and maintenance log analysis are rule-based, statistical, and machine learning. For rule-based analysis, once a new anomaly exceeds the scope of the rule, it is difficult to identify and has poor flexibility. Statistical analysis is easily interfered by data noise and has poor detection effect on complex atypical anomalies. Machine learning-based analysis requires a large amount of high-quality data training, which is costly and has poor model interpretability. On the other hand, most anomaly detection algorithms for operation and maintenance logs are based on simple rule matching or statistical models, which make it difficult to accurately capture complex and changeable anomaly patterns. Faced with diverse system failures and security threats, these methods often have false positives and false negatives.

[0004] Therefore, how to improve the analysis efficiency and detection accuracy of operation and maintenance logs is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a large model-based online analysis and detection method, device, equipment and medium for operation and maintenance logs, which can improve the analysis efficiency of operation and maintenance logs and the accuracy of anomaly detection. The specific scheme is as follows: In a first aspect, the present application provides an online analysis and detection method for operation and maintenance logs based on a large model, comprising: In the target operation and maintenance system, the obtained initial log is preprocessed, and the obtained processed log is matched with each first regular expression; the first regular expression is a regular expression obtained by converting a log template in a preset log template set; When the match fails, a new log template is determined based on the preset large language model and the processed log, and the similarity between the new log template and each log template in the preset log template set is determined respectively; Merging the log templates with similarity greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and parsing the processed log using the target log template to obtain a target log; Determine a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjust the preset large language model using the log test set, the abnormal log example, and the model adjustment prompt information; The adjusted model is used to perform anomaly detection on the logs generated in the target operation and maintenance system, and the detected anomaly information is displayed on the target user interface.

[0006] Optionally, the preprocessing of the obtained initial log includes: Acquire an initial log from each log source, and determine a corresponding log format based on each log source, so as to determine a corresponding second regular expression using the log format; Using the second regular expression to remove the time field, the log level field, and the thread ID field in the initial log to complete the data removal operation; The preset content in the initial log after the data elimination operation is completed is extracted using the second regular expression to complete the data extraction operation.

[0007] Optionally, the determining a new log template based on the preset large language model and the processed log includes: Determine a wildcard format and a placeholder format of a template variable based on the processed log to determine a log template output format; Generate an initial log template based on the processed log and the log template output format by using a preset large language model; Converting the initial log template into a third regular expression, and performing regular matching on the third regular expression and the processed log to obtain a corresponding matching result; If the matching result indicates that the match is successful, the initial log template is determined as a newly added log template; If the matching result indicates that the matching fails, the non-alphabetic and non-numeric characters in the initial log template and the processed log are determined as separators, and the initial log template and the processed log are respectively split using the separators to obtain a template word sequence and a log word sequence; Determine the common words in the template word sequence and the log word sequence, and determine the wildcard of the initial log template as the target wildcard; The newly added log template is determined based on the common words, the target wildcard, and the log word sequence.

[0008] Optionally, the respectively determining the similarity between the newly added log template and each log template in the preset log template set includes: Determine the longest common subsequence based on the newly added log template and each log template in the preset log template set; The similarity is determined using the longest common subsequence, the newly added log template, and each log template in the preset log template set.

[0009] Optionally, after respectively determining the similarity between the newly added log template and each log template in the preset log template set, the method further includes: Determine a similarity judgment result; If the similarity judgment result indicates that there is a log template with the similarity greater than the preset similarity threshold in the preset log template set, the log template with the similarity greater than the preset similarity threshold is output from the preset log template set, so as to merge the log template with the similarity greater than the preset similarity threshold and the newly added log template; If the similarity judgment result indicates that there is no log template with the similarity greater than the preset similarity threshold in the preset log template set, the newly added log template is added to the preset log template set.

[0010] Optionally, after obtaining the target log template, the step further includes: Determine a first log matching range of the target log template, and determine a second log matching range of each log template in the preset log template set; Compare the first log matching range with each of the second log matching ranges, and obtain a comparison result; If the comparison result shows that there is a log template whose second log matching range is included in the first log matching range in the preset log template set, then removing the log template whose second log matching range is included in the first log matching range from the preset log template set; The target log template is added to the preset log template set.

[0011] Optionally, the adjusting the preset large language model by using the log test set, the abnormal log example and the model adjustment prompt information includes: Determine the abnormality detection requirement, the input information of the target user terminal, and the functional mechanism information of the preset large language model as model adjustment prompt information; The pre-trained model of the preset large language model is adjusted using the log test set, the abnormal log example and the model adjustment prompt information.

[0012] In a second aspect, the present application provides an online analysis and monitoring device for operation and maintenance logs based on a large model, comprising: A log matching module, used to pre-process the acquired initial log in the target operation and maintenance system, and perform regular matching between the obtained processed log and each first regular expression; the first regular expression is a regular expression obtained by converting a log template in a preset log template set; A similarity determination module, used for determining a new log template based on a preset large language model and the processed log when the match fails, and determining the similarity between the new log template and each log template in the preset log template set; A log parsing module, used to merge the log templates with similarities greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and parse the processed log using the target log template to obtain a target log; A model adjustment module, used to determine a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjust the preset large language model using the log test set, the abnormal log example and the model adjustment prompt information; The anomaly detection module is used to use the adjusted model to perform anomaly detection on the logs generated in the target operation and maintenance system, and display the detected anomaly information on the target user interface.

[0013] In a third aspect, the present application provides an electronic device, including: Memory, used to store computer programs; The processor is used to execute the computer program to implement the aforementioned large model-based online analysis and detection method for operation and maintenance logs.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned large-model-based online analysis and detection method for operation and maintenance logs is implemented.

[0015] In this application, in the target operation and maintenance system, the obtained initial log is preprocessed, and the obtained processed log is matched with each first regular expression; the first regular expression is a regular expression obtained by converting the log template in the preset log template set; when the match fails, a new log template is determined based on the preset large language model and the processed log, and the similarity between the new log template and each log template in the preset log template set is determined respectively; the log templates in the preset log template set whose similarity is greater than the preset similarity threshold and the new log template are merged, and the target log template is obtained, and the processed log is parsed using the target log template to obtain the target log; a log test set and an abnormal log example are determined from the target log based on a preset natural language processing method, and the preset large language model is adjusted using the log test set, the abnormal log example and the model adjustment prompt information; the log generated in the target operation and maintenance system is detected for abnormality using the adjusted model, and the detected abnormal information is displayed on the target user interface. As can be seen from the above, in this application, in the target operation and maintenance system, the obtained initial log will first be preprocessed. After the preprocessing is completed, the processed log will be obtained, and then it will be matched with each first regular expression. The first regular expression here is obtained by converting the log template in the preset log template set. If the above matching operation fails, the newly added log template will be determined based on the preset large language model and the processed log. After the newly added log template is determined, the similarity between it and each log template in the preset log template set will be calculated respectively. After that, the log templates in the preset log template set whose similarity is greater than the preset similarity threshold and the newly added log template are merged to obtain the target log template. The processed log is then parsed using this target log template to finally obtain the target log. Next, the log test set and abnormal log examples are determined from the target log with the help of the preset natural language processing method. Then the preset large language model is adjusted using the log test set, abnormal log examples and model adjustment prompt information. Finally, the adjusted model is used to detect anomalies in the logs generated in the target operation and maintenance system. Once abnormal information is detected, it will be displayed on the target user interface. In this way, this application can improve the analysis efficiency of operation and maintenance logs and the accuracy of anomaly detection, thereby providing more comprehensive and accurate support for operation and maintenance work. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0017] Figure 1 This is a flow chart of an online analysis and detection method for operation and maintenance logs based on a large model disclosed in this application; Figure 2 A flowchart of a specific large-model-based online analysis and detection method for operation and maintenance logs disclosed in this application; Figure 3 This is a schematic diagram of the structure of an online analysis and detection device for operation and maintenance logs based on a large model disclosed in this application; Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] At present, for the analysis of operation and maintenance logs, on the one hand, the main methods of operation and maintenance log analysis are rule-based, statistical and machine learning. For rule-based analysis, once a new anomaly exceeds the scope of the rule, it is difficult to identify and has poor flexibility. However, statistical-based analysis is susceptible to data noise interference and has poor detection effect on complex atypical anomalies. Analysis based on machine learning requires a large amount of high-quality data training, which is costly and has poor model interpretability. On the other hand, most anomaly detection algorithms for operation and maintenance logs are based on simple rule matching or statistical models, which makes it difficult to accurately capture complex and changeable anomaly patterns. To this end, the present application provides an online analysis and detection method, device, equipment and medium for operation and maintenance logs based on a large model, which can improve the analysis efficiency of operation and maintenance logs and the accuracy of anomaly detection.

[0020] See also Figure 1 As shown, the embodiment of the present invention discloses an online analysis and detection method for operation and maintenance logs based on a large model, comprising: Step S11: In the target operation and maintenance system, the obtained initial log is preprocessed, and the processed log is matched with each first regular expression; the first regular expression is a regular expression obtained by converting a log template in a preset log template set.

[0021] In this embodiment, first, the initial log is obtained from each log source in the target operation and maintenance system, and the initial log is preprocessed. Since the log formats generated by different log sources are different, it is necessary to determine the corresponding log format based on each log source. Specifically, for each log source, its log format will be determined based on its own characteristics and specifications. For example, the log format of some log sources may have specific delimiters, field arrangement order and other features. By analyzing and identifying these features, the corresponding log format is determined.

[0022] After determining the log format of each log source, you can use these log formats to determine the corresponding second regular expression. The construction of the second regular expression is based on the characteristics of the log format, so as to identify and process the fields in the log. For example, for the time field in the log, its format may have a specific date and time representation. By constructing the corresponding second regular expression, you can accurately match and locate the time field.

[0023] Further, the time field, log level field and thread ID field in the initial log are eliminated by the second regular expression to complete the data elimination operation. In actual log analysis, the time field, log level field and thread ID field often have less impact on the core content analysis of the log, and the existence of these fields may increase the complexity of log processing. Therefore, by eliminating these fields from the initial log through the second regular expression, the amount of data can be reduced and the efficiency of subsequent processing can be improved.

[0024] After the data elimination operation is completed, the second regular expression is used to extract the preset content in the initial log after the data elimination operation is completed to complete the data extraction operation. The preset content refers to the part that is important for log analysis, such as key event descriptions and error information in the log. The remaining log content is scanned and matched by the second regular expression to extract the part that conforms to the preset content format. The processed log obtained in this way only contains information that is valuable for log analysis.

[0025] Finally, the processed log is matched with each first regular expression. The first regular expression is converted from a log template in the preset log template set. Its purpose is to compare the processed log with the existing log template to determine whether the processed log can match the existing log template. During the matching process, the processed log will be compared with each first regular expression in turn. If the match is successful, it means that the processed log has a similar structure and content to the corresponding log template, that is, the corresponding log template can be used to parse the processed log, thereby advancing the subsequent process steps.

[0026] Step S12: When the match fails, a new log template is determined based on the preset large language model and the processed log, and the similarity between the new log template and each log template in the preset log template set is determined respectively.

[0027] In this embodiment, when the processed log fails to match the first regular expressions in the preset log template set, it is necessary to generate a new log template with the help of the preset large language model. First, the wildcard format and the placeholder format of the template variables must be determined based on the processed log to clarify the output format of the log template. The wildcard format is used to represent any character or character combination that may appear in the log. Its determination needs to comprehensively consider the characteristics of the processed log and the needs of subsequent log analysis. The placeholder format of the template variable is to mark the parts of the log that may change, so that these variables can be accurately identified and processed in the subsequent log matching and parsing process. By conducting an in-depth analysis of the structure, content, and semantics of the processed log, the appropriate wildcard format and placeholder format of the template variable can be determined.

[0028] After the output format of the log template is determined, the initial log template is generated based on the processed log and the output format of the log template through the preset large language model. The preset large language model can generate an initial log template with a certain structure and semantics based on the input processed log and the specified output format. During the generation process, the large language model will perform semantic analysis, pattern recognition and other operations on the processed log, and generate the initial log template in combination with the requirements of the log template output format.

[0029] Further, the generated initial log template is converted into a third regular expression, and the third regular expression is matched with the processed log to obtain a corresponding matching result. If the matching result indicates a successful match, it means that the initial log template can fit the processed log well, and the initial log template can be determined as a new log template.

[0030] However, if the matching result indicates that the match fails, the initial log template needs to be corrected. Specifically, the non-alphabetic and non-numeric characters in the initial log template and the processed log are determined as separators, and these separators are used to split the initial log template and the processed log respectively to obtain the template word sequence and the log word sequence. The purpose of the splitting is to decompose the log and template into more fine-grained units for subsequent comparison and analysis.

[0031] Next, the common words in the template word sequence and the log word sequence are determined. These common words represent the common parts between the log and the template. At the same time, the wildcard of the initial log template is determined as the target wildcard. Finally, the newly added log template is determined based on the common words, the target wildcard, and the log word sequence. In the determination process, the common words and the target wildcard are retained, and the non-common parts are supplemented and adjusted according to the content in the log word sequence, so as to obtain a newly added log template after error correction.

[0032] After obtaining the newly added log template, it is necessary to determine its similarity with each log template in the preset log template set. First, the longest common subsequence is determined based on the newly added log template and each log template in the preset log template set. The longest common subsequence refers to the longest identical subsequence that appears in order but not necessarily consecutively in two sequences, which reflects the similarity between the two sequences.

[0033] The similarity is determined using the longest common subsequence, the newly added log template, and each log template in the preset log template set. By calculating the similarity, we can understand the similarity between the newly added log template and each log template in the preset log template set, which provides a basis for subsequent log template merging or adding operations. In addition, the formula for calculating the similarity is as follows: ; In the formula, Indicates the similarity between the newly added log template and a log template in the preset log template set. s Indicates a new log template. t Indicates a log template in the preset log template set. represents the length of the longest common subsequence, express s Length, express t Length, Represents the larger value of and.

[0034] In addition, after calculating the similarity, a similarity judgment result is determined, that is, it is determined whether there is a log template in the preset log template set whose similarity is greater than a preset similarity threshold. The preset similarity threshold is a pre-set standard used to measure whether the similarity between two log templates reaches a level that requires a merge operation.

[0035] In a specific implementation, if the similarity judgment result indicates that there are log templates with similarity greater than a preset similarity threshold in the preset log template set, then these log templates with similarity greater than the preset similarity threshold are output from the preset log template set so as to merge them with the newly added log template. The purpose of the merger is to reduce the number of log templates and improve the efficiency of log parsing.

[0036] In another specific implementation, if the similarity judgment result shows that there is no log template with a similarity greater than a preset similarity threshold in the preset log template set, the newly added log template is added to the preset log template set. In this way, the preset log template set can be continuously expanded to cover more log modes and types, thereby improving the accuracy and comprehensiveness of log analysis.

[0037] Step S13: merge the log templates with similarity greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and use the target log template to parse the processed log to obtain a target log.

[0038] In this embodiment, after determining that there are log templates with similarities greater than a preset similarity threshold in the preset log template set, it is necessary to merge these log templates with the newly added log template to obtain the target log template. The merging process needs to comprehensively consider the characteristics and structure of each template. When merging, the common parts of each template should be retained because the common parts represent the core features of the log pattern. For non-common parts, wildcard substitution is used. Through this merging operation, the number of log templates can be effectively reduced, the efficiency of log parsing can be improved, and at the same time, it is ensured that the log template can accurately reflect the pattern of the log.

[0039] After obtaining the target log template, use it to parse the processed log. By matching and parsing the processed log, the target log template can extract key information from the log and convert it into a target log that is easier to understand and analyze. During the parsing process, the target log template will identify the various fields and variables in the processed log according to its defined patterns and rules, thereby obtaining a structured target log.

[0040] In addition, in order to further optimize the preset log template set, it is necessary to determine the first log matching range of the target log template and the second log matching range of each log template in the preset log template set. The log matching range refers to the range and characteristics of the logs that the log template can match. When determining the first log matching range, the structure of the target log template, the use of wildcards, and the characteristics of the log pattern represented by the template should be considered. Similarly, for each log template in the preset log template set, its second log matching range also needs to be accurately determined. This helps to compare and screen the log templates later to ensure the efficiency and accuracy of the preset log template set.

[0041] Furthermore, in this embodiment, the first log matching range and each second log matching range are compared to obtain a comparison result. During the comparison process, the inclusion relationship, overlap degree, etc. between each log matching range should be carefully analyzed. By comparing, it can be found which log templates in the preset log template set have matching ranges covered by the target log template, and which log templates have overlapping or complementary relationships with the target log template.

[0042] If the comparison results show that there are log templates in the preset log template set whose second log matching range is included in the first log matching range, it means that the log patterns represented by these log templates have been covered by the target log template, and their existence will increase the complexity of the log template set and reduce the efficiency of log parsing. Therefore, these log templates whose second log matching range is included in the first log matching range are removed from the preset log template set. This can make the preset log template set more streamlined and improve the speed and accuracy of log parsing. After removing the included log templates, the target log template is added to the preset log template set. In this way, the preset log template set is updated and optimized, and can better adapt to the ever-changing log patterns, providing more accurate and efficient support for subsequent online log analysis and detection.

[0043] Step S14: determining a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjusting the preset large language model using the log test set, the abnormal log example and the model adjustment prompt information.

[0044] After obtaining the target log, you need to use the preset natural language processing method to extract the log test set and abnormal log examples from it. The preset natural language processing method covers a variety of technologies, such as text classification, cluster analysis, keyword extraction, etc.

[0045] Specifically, for the determination of the log test set, the target logs can be classified first. By analyzing the log's subject, source, format and other features, it can be divided into different categories. Representative log samples are selected from each category to form a log test set. These samples should be able to fully reflect the various situations of the target log, including normal logs and possible abnormal log situations. The determination of abnormal log examples can be assisted by more in-depth analysis. Cluster analysis can be used to cluster abnormal logs with similar characteristics to find out the typical patterns. At the same time, keyword extraction technology is used to identify keywords that frequently appear in abnormal logs. These keywords often represent the core information of abnormal events. Combining these methods, representative abnormal logs are selected from the target log as examples, and these examples will be used for subsequent adjustments to the preset large language model. In addition, the log test set and abnormal log examples can also be determined manually.

[0046] Next, in this embodiment, the anomaly detection requirements, the input information of the target user end, and the functional mechanism information of the preset large language model are determined as model adjustment prompt information. The anomaly detection requirements clarify the conditions and standards that the large language model needs to meet when performing log anomaly detection, such as the definition of anomalies, the accuracy requirements of detection, etc. The input information of the target user end reflects the user's specific needs and concerns, which may include concerns about certain specific types of anomalies, specific requirements for log analysis results, etc. Finally, the functional mechanism information of the preset large language model helps to understand the working principle and characteristics of the model so that adjustments can be made more targeted.

[0047] Furthermore, after obtaining the log test set, the abnormal log examples, and the model adjustment prompt information, the pre-trained model of the preset large language model can be adjusted. The log test set is input into the pre-trained model, and the model is asked to analyze and process these logs. During the processing, the model will classify and judge the logs according to its own algorithms and parameters, and try to identify the abnormal logs. At the same time, the abnormal log examples are used as references to let the model learn the characteristics and patterns of abnormal logs. By comparing with the abnormal log examples, the model can better understand the manifestation of abnormalities, thereby improving the accuracy of anomaly detection. The model adjustment prompt information plays a guiding role in this process. It tells the model the focus and direction to pay attention to when making adjustments, such as how to better analyze the input information of the target user end while meeting the requirements of abnormal detection. Based on these prompt information, the model will adjust its own parameters and continuously optimize its performance.

[0048] In the process of adjusting the model, the performance of the model can be continuously evaluated. Part of the data in the log test set can be used as a validation set to evaluate the effect of the model adjustment. If the performance of the model does not meet the expected requirements, it is necessary to further adjust the parameters or optimize the method until the model can accurately detect abnormal information in the log and meet the anomaly detection requirements and the input information of the target user.

[0049] Step S15: Use the adjusted model to perform anomaly detection on the logs generated in the target operation and maintenance system, and display the detected anomaly information on the target user interface.

[0050] In this embodiment, the adjusted model has higher accuracy and adaptability, and can perform efficient anomaly detection on logs generated in real time in the target operation and maintenance system. When a new log is generated, the model will quickly analyze it and determine whether the log is abnormal based on the characteristic patterns of normal and abnormal logs learned previously.

[0051] The detected anomaly information is then collated and formatted for clear and intuitive display on the target user interface. The user interface presents the anomaly information in an easy-to-understand manner, such as highlighting the anomaly log, providing detailed anomaly descriptions and suggested treatment measures. In this way, operation and maintenance personnel can obtain anomaly information in a timely manner and quickly take corresponding measures to ensure the stable operation of the target operation and maintenance system.

[0052] As can be seen from the above, in the target operation and maintenance system, in this application, the initial log obtained will first be preprocessed. After the preprocessing is completed, the processed log will be obtained, and then it will be regularly matched with each first regular expression. The first regular expression here is obtained by converting the log template in the preset log template set. If the above matching operation fails, the newly added log template will be determined based on the preset large language model and the processed log. After determining the newly added log template, the similarity between it and each log template in the preset log template set will be calculated respectively. After that, the log templates in the preset log template set whose similarity is greater than the preset similarity threshold and the newly added log template are merged to obtain the target log template. Then use this target log template to parse the processed log and finally obtain the target log. Next, the log test set and abnormal log examples are determined from the target log with the help of the preset natural language processing method. Then the preset large language model is adjusted using the log test set, abnormal log examples and model adjustment prompt information. Finally, the adjusted model is used to perform abnormal detection on the logs generated in the target operation and maintenance system. Once abnormal information is detected, it will be displayed on the target user interface. In this way, this application can improve the analysis efficiency of operation and maintenance logs and the accuracy of anomaly detection, thereby providing more comprehensive and accurate support for operation and maintenance work.

[0053] Combine the following Figure 2 The schematic diagram shown specifically illustrates the technical solution of the embodiment of the present application.

[0054] Specifically, in the operation and maintenance system, a large amount of operation and maintenance logs are generated from a wide range of sources. These logs come from various facilities such as servers, network equipment, storage systems, etc. The first step of this solution is to perform preprocessing operations on the initial logs collected from various sources. This step improves the quality of log parsing and reduces the number of log templates through operations such as log format verification, data filtering and cleaning, so as to obtain processed logs.

[0055] Next, the processed log is matched with the first regular expression converted from the preset log template set. If the match fails, the preset large language model is started, and the initial log template is generated in combination with the processed log, and the initial log template and the processed log are matched. When the match fails, the separator is determined from the initial log template and the processed log, and the separator is used to split the initial log template and the processed log respectively, and the common words in the obtained template word sequence and the log word sequence are determined, and the wildcard of the initial log template is determined as the target wildcard. The above processing process of the initial log template is the error correction process. After the initial log template is corrected, the corrected initial log template is determined as the newly added log template, and the similarity between the newly added log template and each template in the preset log template set is calculated.

[0056] Furthermore, the log templates with similarity greater than the preset threshold are merged with the newly added log template to generate the target log template. In addition, the target log template needs to be overwritten, that is, to check whether the target log template can overwrite or be overwritten by an original template in the preset log template set, and remove the overwritten log template from the preset log template set. Then, the target log template is matched with the processed log. After the match is successful, the processed log is parsed using the target log template to obtain the target log.

[0057] Then, log files are exported from the existing operation and maintenance system, and the logs are labeled as normal / abnormal through manual or traditional natural language processing methods, thereby forming a log abnormality training set (also known as a log test set), and a set of representative abnormal log message examples are prepared. Next, the log abnormality training set is input into the preset large language model, so that the preset large language model is supervised and fine-tuned according to the preset time, and abnormal log message examples are added to the prompt (i.e., the prompt information for the model) of the preset large language model after supervised fine-tuning, so as to detect the target log based on the prompt and the preset large language model. Once an abnormality is detected, the abnormal result is immediately displayed on the target user interface.

[0058] In this way, this application can greatly improve the efficiency of operation and maintenance log analysis, accurately identify anomalies, and provide strong support for ensuring the stable and efficient operation of the operation and maintenance system.

[0059] Accordingly, see Figure 3 As shown, the embodiment of the present application provides an online analysis and monitoring device for operation and maintenance logs based on a large model, comprising: The log matching module 11 is used to pre-process the acquired initial log in the target operation and maintenance system, and perform regular matching between the processed log and each first regular expression; the first regular expression is a regular expression obtained by converting a log template in a preset log template set; A similarity determination module 12, for determining a new log template based on a preset large language model and the processed log when the match fails, and determining the similarity between the new log template and each log template in the preset log template set; The log parsing module 13 is used to merge the log templates with similarities greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and parse the processed log using the target log template to obtain a target log; A model adjustment module 14 is used to determine a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjust the preset large language model using the log test set, the abnormal log example and the model adjustment prompt information; The anomaly detection module 15 is used to use the adjusted model to perform anomaly detection on the logs generated in the target operation and maintenance system, and display the detected anomaly information on the target user interface.

[0060] As can be seen from the above, in the target operation and maintenance system, in this application, the initial log obtained will first be preprocessed. After the preprocessing is completed, the processed log will be obtained, and then it will be regularly matched with each first regular expression. The first regular expression here is obtained by converting the log template in the preset log template set. If the above matching operation fails, the newly added log template will be determined based on the preset large language model and the processed log. After determining the newly added log template, the similarity between it and each log template in the preset log template set will be calculated respectively. After that, the log templates in the preset log template set whose similarity is greater than the preset similarity threshold and the newly added log template are merged to obtain the target log template. Then use this target log template to parse the processed log and finally obtain the target log. Next, the log test set and abnormal log examples are determined from the target log with the help of the preset natural language processing method. Then the preset large language model is adjusted using the log test set, abnormal log examples and model adjustment prompt information. Finally, the adjusted model is used to perform abnormal detection on the logs generated in the target operation and maintenance system. Once abnormal information is detected, it will be displayed on the target user interface. In this way, this application can improve the analysis efficiency of operation and maintenance logs and the accuracy of anomaly detection, thereby providing more comprehensive and accurate support for operation and maintenance work.

[0061] In some specific implementations, the log matching module 11 specifically includes: An expression determination unit, used for obtaining an initial log from each log source, and determining a corresponding log format based on each log source, so as to determine a corresponding second regular expression using the log format; A data elimination unit, used to eliminate the time field, the log level field and the thread ID field in the initial log by using the second regular expression to complete the data elimination operation; The data extraction unit is used to extract preset content in the initial log after the data elimination operation is completed by using the second regular expression to complete the data extraction operation.

[0062] In some specific implementations, the similarity determination module 12 specifically includes: a format determination unit, configured to determine a wildcard format and a placeholder format of a template variable based on the processed log, so as to determine a log template output format; A template generating unit, configured to generate an initial log template based on the processed log and the log template output format by using a preset large language model; A log matching unit, configured to convert the initial log template into a third regular expression, and perform regular matching on the third regular expression and the processed log to obtain a corresponding matching result; A first template determining unit, configured to determine the initial log template as a newly added log template if the matching result indicates a successful match; A log splitting unit, configured to, if the matching result indicates that the matching fails, determine the non-alphabetic and non-numeric characters in the initial log template and the processed log as separators, and use the separators to split the initial log template and the processed log respectively to obtain a template word sequence and a log word sequence; A wildcard determination unit, used for determining the common words in the template word sequence and the log word sequence, and determining the wildcard of the initial log template as a target wildcard; The second template determination unit is used to determine the newly added log template based on the common words, the target wildcard and the log word sequence.

[0063] In some specific implementations, the similarity determination module 12 specifically includes: A sequence determination unit, configured to determine a longest common subsequence based on the newly added log template and each log template in the preset log template set; The similarity determination unit is used to determine the similarity by using the longest common subsequence, the newly added log template and each log template in the preset log template set.

[0064] In some specific implementations, the similarity determination module 12 further includes: A result determination unit, used to determine a similarity determination result; a template merging unit, configured to output the log template with the similarity greater than the preset similarity threshold from the preset log template set if the similarity judgment result indicates that the log template with the similarity greater than the preset similarity threshold exists in the preset log template set, so as to merge the log template with the similarity greater than the preset similarity threshold and the newly added log template; The first template adding unit is configured to add the newly added log template to the preset log template set if the similarity judgment result indicates that there is no log template with the similarity greater than a preset similarity threshold in the preset log template set.

[0065] In some specific implementations, the log parsing module 13 specifically further includes: a range determination unit, configured to determine a first log matching range of the target log template, and determine a second log matching range of each log template in the preset log template set; A range comparison unit, configured to compare the first log matching range with each of the second log matching ranges and obtain a comparison result; a template removing unit, configured to remove the log template whose second log matching range is included in the first log matching range from the preset log template set if the comparison result shows that there is a log template whose second log matching range is included in the first log matching range in the preset log template set; The second template adding unit is used to add the target log template to the preset log template set.

[0066] In some specific implementations, the model adjustment module 14 specifically includes: An information determination unit, used to determine the abnormality detection requirement, the input information of the target user terminal and the functional mechanism information of the preset large language model as model adjustment prompt information; A model adjustment unit is used to adjust the pre-trained model of the preset large language model by using the log test set, the abnormal log example and the model adjustment prompt information.

[0067] Furthermore, the present application also discloses an electronic device. Figure 4It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input and output interface 25 and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the online analysis and detection method of operation and maintenance logs based on a large model disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment can specifically be an electronic computer.

[0068] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0069] In addition, the memory 22 as a carrier for resource storage may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0070] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large model-based online analysis and detection method for operation and maintenance logs performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0071] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned large model-based online analysis and detection method for operation and maintenance logs disclosed above is implemented. The specific steps of the method can be referred to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.

[0072] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0073] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0074] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0075] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0076] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A large model-based online analysis and detection method for operation and maintenance logs, characterized in that: include: In the target operation and maintenance system, the obtained initial log is preprocessed, and the obtained processed log is matched with each first regular expression; The first regular expression is a regular expression obtained by converting a log template in a preset log template set; When the match fails, a new log template is determined based on the preset large language model and the processed log, and the similarity between the new log template and each log template in the preset log template set is determined respectively; Merging the log templates with similarity greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and parsing the processed log using the target log template to obtain a target log; Determine a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjust the preset large language model using the log test set, the abnormal log example, and the model adjustment prompt information; The adjusted model is used to perform anomaly detection on the logs generated in the target operation and maintenance system, and the detected anomaly information is displayed on the target user interface.

2. According to the large model-based online analysis and detection method for operation and maintenance logs according to claim 1, it is characterized in that: The preprocessing of the obtained initial log includes: Acquire an initial log from each log source, and determine a corresponding log format based on each log source, so as to determine a corresponding second regular expression using the log format; Using the second regular expression to remove the time field, the log level field, and the thread ID field in the initial log to complete the data removal operation; The preset content in the initial log after the data elimination operation is completed is extracted using the second regular expression to complete the data extraction operation.

3. The large model-based online analysis and detection method for operation and maintenance logs according to claim 1 is characterized in that: The determining of a new log template based on the preset large language model and the processed log includes: Determine a wildcard format and a placeholder format of a template variable based on the processed log to determine a log template output format; Generate an initial log template based on the processed log and the log template output format by using a preset large language model; Converting the initial log template into a third regular expression, and performing regular matching on the third regular expression and the processed log to obtain a corresponding matching result; If the matching result indicates that the match is successful, the initial log template is determined as a newly added log template; If the matching result indicates that the matching fails, the non-alphabetic and non-numeric characters in the initial log template and the processed log are determined as separators, and the initial log template and the processed log are respectively split using the separators to obtain a template word sequence and a log word sequence; Determine the common words in the template word sequence and the log word sequence, and determine the wildcard of the initial log template as the target wildcard; The newly added log template is determined based on the common words, the target wildcard, and the log word sequence.

4. The large model-based online analysis and detection method for operation and maintenance logs according to claim 1 is characterized in that: The respectively determining the similarity between the newly added log template and each log template in the preset log template set includes: Determine the longest common subsequence based on the newly added log template and each log template in the preset log template set; The similarity is determined using the longest common subsequence, the newly added log template, and each log template in the preset log template set.

5. The large model-based online analysis and detection method for operation and maintenance logs according to claim 1 is characterized in that: After respectively determining the similarity between the newly added log template and each log template in the preset log template set, the method further includes: Determine a similarity judgment result; If the similarity judgment result indicates that there is a log template with the similarity greater than the preset similarity threshold in the preset log template set, the log template with the similarity greater than the preset similarity threshold is output from the preset log template set, so as to merge the log template with the similarity greater than the preset similarity threshold and the newly added log template; If the similarity judgment result indicates that there is no log template with the similarity greater than the preset similarity threshold in the preset log template set, the newly added log template is added to the preset log template set.

6. The large model-based online analysis and detection method for operation and maintenance logs according to claim 1 is characterized in that: After obtaining the target log template, the method further includes: Determine a first log matching range of the target log template, and determine a second log matching range of each log template in the preset log template set; Compare the first log matching range with each of the second log matching ranges, and obtain a comparison result; If the comparison result shows that there is a log template whose second log matching range is included in the first log matching range in the preset log template set, then removing the log template whose second log matching range is included in the first log matching range from the preset log template set; The target log template is added to the preset log template set.

7. The online analysis and detection method for operation and maintenance logs based on a large model according to any one of claims 1 to 6, characterized in that: The adjusting the preset large language model by using the log test set, the abnormal log example and the model adjustment prompt information includes: Determine the abnormality detection requirement, the input information of the target user terminal, and the functional mechanism information of the preset large language model as model adjustment prompt information; The pre-trained model of the preset large language model is adjusted using the log test set, the abnormal log example and the model adjustment prompt information.

8. An online analysis and monitoring device for operation and maintenance logs based on a large model, characterized in that: include: A log matching module, used to pre-process the acquired initial log in the target operation and maintenance system, and perform regular matching between the obtained processed log and each first regular expression; the first regular expression is a regular expression obtained by converting a log template in a preset log template set; A similarity determination module, used for determining a new log template based on a preset large language model and the processed log when the match fails, and determining the similarity between the new log template and each log template in the preset log template set; A log parsing module, used to merge the log templates with similarities greater than a preset similarity threshold in the preset log template set and the newly added log template to obtain a target log template, and parse the processed log using the target log template to obtain a target log; A model adjustment module, used to determine a log test set and an abnormal log example from the target log based on a preset natural language processing method, and adjust the preset large language model using the log test set, the abnormal log example and the model adjustment prompt information; The anomaly detection module is used to use the adjusted model to perform anomaly detection on the logs generated in the target operation and maintenance system, and display the detected anomaly information on the target user interface.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the large model-based online analysis and detection method for operation and maintenance logs as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the online analysis and detection method for operation and maintenance logs based on a large model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Low-cost and zero-sample online log analysis method based on large language model

    CN117407242A

  • Log analysis method and device based on artificial intelligence, computer equipment and medium

    CN119201597A

  • Multi-feature log anomaly detection method and system based on log full semantics

    US20220405592A1