Log-based fault determination method and device, equipment and storage medium
By performing pattern analysis and correlation processing on the logs, the fault points of abnormal operation and maintenance objects are quickly identified, solving the limitations of log failure perception and time-consuming log analysis in the existing technology, and achieving efficient fault location and analysis.
Patent Information
- Application Number
- CN202510207284.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the log fault perception content is limited, and the log analysis of cross-operation and maintenance objects requires a lot of time, and the fault location efficiency is low.
By obtaining information abnormal alarms, determine the abnormal operation and maintenance objects, and form an operation and maintenance object combination based on the object association relationship. Using the historical log mode in the log mode library, the log mode of the operation and maintenance object combination is combined and marked, and the target mode of multiple categories is generated, and the target mode of each category is then fault determination is performed to accurately identify the fault points of the abnormal operation and maintenance object.
It significantly reduces the burden of manually reviewing logs, quickly locks the core information of the fault, and improves the efficiency of log correlation analysis across operations and maintenance objects.
Smart Images

Figure CN120144347A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data, and particularly to a method, apparatus, device, and storage medium for fault determination based on logs. Background Art
[0002] In the field of monitoring and operation and maintenance, ensuring the stable operation of the system and promptly discovering and resolving faults are the core tasks of operation and maintenance work. As an important monitoring method, log monitoring can reflect the running status, operating conditions, and abnormal fault information of the program in real time by collecting and analyzing the logs generated by the system or application.
[0003] Currently, the monitoring of logs generally adopts the method of full-scale collection, that is, all generated log data are centrally collected and then subjected to subsequent analysis and processing. During the analysis process, methods such as keyword screening and regular expression matching are commonly used to filter important information and upload it to the alarm platform. After the alarm, the operation and maintenance personnel trace back to the server based on the alarm information and consult the relevant period logs to determine the specific cause and solution of the fault.
[0004] However, the process of fault location requires manual analysis and judgment by the operation and maintenance personnel. In the case of complex faults involving multiple components and servers, a large amount of time is required for the correlation analysis of log data, which not only has low efficiency but also may lead to inaccurate fault location or omission of key information due to individual experience and skill differences, thus prolonging the time for fault troubleshooting and recovery. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for fault determination based on logs to solve the technical problems in the prior art that the content of log fault perception is limited, a large amount of time is required for the log analysis of cross-operation and maintenance objects, and the fault location efficiency is low.
[0006] In a first aspect, this application provides a method for fault determination based on logs, including:
[0007] Obtain an information anomaly alarm, and determine an abnormal operation and maintenance object according to the information anomaly alarm;
[0008] Determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship; the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object;
[0009] Perform combined marking processing on the log patterns of the operation and maintenance object combination according to the historical log patterns in the log pattern library to obtain target patterns of multiple categories; the log pattern library is generated according to the historical log information of all operation and maintenance objects;
[0010] Perform fault determination processing on the target patterns of multiple categories respectively to obtain the fault points of the abnormal operation and maintenance object.
[0011] In a second aspect, the present application provides a log-based fault determination device, including:
[0012] An acquisition module, configured to acquire an information anomaly alarm, and determine an abnormal operation and maintenance object according to the information anomaly alarm;
[0013] A determination module, configured to determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship; the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object;
[0014] A processing module, configured to perform combined marking processing on the log patterns of the operation and maintenance object combination according to the historical log patterns in the log pattern library to obtain target patterns of multiple categories; the log pattern library is generated according to the historical log information of all operation and maintenance objects;
[0015] The processing module is configured to perform fault determination processing on the target patterns of multiple categories respectively to obtain the fault points of the abnormal operation and maintenance object.
[0016] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0017] The memory stores computer execution instructions;
[0018] The processor executes the computer execution instructions stored in the memory to implement the log-based fault determination method as described in the first aspect and various possible implementation manners of the first aspect.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the log-based fault determination method as described in the first aspect and various possible implementation manners of the first aspect.
[0020] In a fifth aspect, the present application provides a program product, including a computer program, and when the computer program is executed by a processor, it implements the log-based fault determination method as described above.
[0021] The log-based fault determination method, device, equipment, and storage medium provided by this application capture information anomaly alerts and identify operation and maintenance objects with abnormal problems based on these alerts. Then, using the association relationships between operation and maintenance objects, operation and maintenance object combinations are formed. By referring to the log pattern library constructed from the historical logs of all operation and maintenance objects, these combinations are matched and classified with log patterns to generate target log patterns of multiple categories. Subsequently, fault diagnosis is performed on the target log patterns of each category to accurately identify the fault locations of abnormal operation and maintenance objects. This method greatly facilitates the log correlation analysis across operation and maintenance objects, significantly reduces the burden of manual log review, and can quickly lock in the core information of faults. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0023] Figure 1 Schematic flow of a log-based fault determination method provided by this application Figure 1 ;
[0024] Figure 2 Schematic flow of a log-based fault determination method provided by this application Figure 2 ;
[0025] Figure 3 Schematic flow of a log-based fault determination method provided by this application Figure 3 ;
[0026] Figure 4 Schematic flow of a log-based fault determination method provided by this application Figure 4 ;
[0027] Figure 5 Schematic flow of a log-based fault determination method provided by this application Figure 5 ;
[0028] Figure 6 Schematic structural diagram of a log-based fault determination device provided by this application;
[0029] Figure 7 Schematic structural diagram of a log-based fault determination equipment provided by this application.
[0030] Through the above drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions later. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to explain the concept of this application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0033] Furthermore, the present application involves big data analysis of user information (including but not limited to personal biometric characteristics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and uses artificial intelligence technology for automated decision-making. For technical solutions that make decisions having a significant impact on personal rights and interests based on the results of automated decision-making, corresponding operation entrances are provided for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, an expert decision-making process will be entered.
[0034] It should be noted that the log-based fault determination method, device, equipment, and storage medium provided in the present application can be used in the field of big data, and can also be used in any field other than big data. The application fields of the log-based fault determination method, device, equipment, and storage medium in the present application are not limited.
[0035] In the field of monitoring and operation and maintenance, ensuring the stable operation of the system and timely discovering and solving faults are the core tasks of operation and maintenance work. As an important monitoring method, log monitoring can reflect the running situation, running state, and abnormal fault information of the program in real time by collecting and analyzing the logs generated by the system or application. The logs contain rich runtime data. Especially when the system fails, the logs can record key information, key changes, or operations of the program, providing valuable clues for fault troubleshooting for operation and maintenance personnel.
[0036] Currently, the monitoring of logs generally adopts the method of full - volume collection, that is, all generated log data are centrally collected and then subjected to subsequent analysis and processing. During the analysis process, common methods include keyword screening, regular expression matching, etc. Through these means, important information in the logs is filtered out and sent to the alarm platform. When the alarm platform receives relevant alarm information, the operation and maintenance personnel will, based on the keywords, alarm time and other information provided in the alarm, trace back to the servers involved in the failure, further consult and analyze the log data for the corresponding time period to determine the specific cause of the failure and the solution.
[0037] However, the fault perception in the existing technology mainly relies on pre - configured keyword alarms. Although this method can effectively monitor and alarm for known specific problems, if the configured keywords are not triggered during a failure, the system cannot perceive the occurrence of the failure, resulting in potential failures that may not be discovered and processed in a timely manner, increasing the operation risk of the system. In terms of fault location, the existing technology mainly relies on the analysis and judgment of operation and maintenance personnel. Especially when facing complex faults across components and servers, operation and maintenance personnel need to manually correlate and analyze various log data, which is time - consuming and inefficient. In addition, manual analysis may also be limited by factors such as personal experience and skill level, resulting in inaccurate fault location or omission of key information, further prolonging the time for fault recovery.
[0038] In response to the above - mentioned problems, this application proposes a log - based fault determination method. By analyzing the patterns of logs, abnormal situations in the log information are discovered and information anomaly alarms are generated. When an anomaly alarm is generated, the abnormal operation and maintenance object is determined, and according to the pre - established association relationships of the operation and maintenance objects, all operation and maintenance objects associated with the abnormal operation and maintenance object are retrieved. The log information of all operation and maintenance objects is retrieved and matched and classified with the historical log patterns in the log pattern library. The classified patterns are sorted, and fault determination is carried out according to the sorting order to quickly locate the location of the fault. This method brings great convenience to the log correlation analysis across operation and maintenance objects, can greatly reduce the amount of logs for manual analysis, and quickly discover the key information of faults.
[0039] The specific application scenarios of this application are very extensive. For example, it can be used in the troubleshooting of servers. When a server experiences problems such as performance degradation, downtime, or abnormal restart, log monitoring can capture and send alarm messages in real time. Based on the received alarm messages, the system can retrieve relevant servers or components, classify and display their log information in patterns, and review relevant logs and diagnose the cause of the failure according to the pattern categories. It can also be used in business performance monitoring. For some key business processes or performance indicators, such as response time, transaction success rate, etc., log monitoring can record the changes of these indicators in real time. When these indicators show abnormal fluctuations, alarm messages can be sent in a timely manner, enabling staff to quickly resolve the problems and ensure the stable operation of the business.
[0040] The following will specifically describe the technical solutions of this application and how the technical solutions of this application solve the above technical problems with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0041] Figure 1 Flow schematic of a log-based fault determination method provided by an embodiment of this application Figure 1 As Figure 1 shown, the log-based fault determination method provided by this embodiment includes:
[0042] S101: Obtain an information anomaly alarm, and determine an abnormal operation and maintenance object according to the information anomaly alarm.
[0043] Among them, the abnormal operation and maintenance object specifically refers to those servers or components with alarms in the logs.
[0044] It can be understood that the system can monitor and analyze the log information generated by the operation and maintenance objects in real time. Some anomalies will be found during the analysis of the log information and their corresponding alarm messages will be generated. After obtaining the alarm messages, the system processes the alarm messages. First, it will find the log information corresponding to the information anomaly alarm, identify the abnormal operation and maintenance objects, and further process the abnormal operation and maintenance objects.
[0045] S102: Determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship.
[0046] Among them, the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object.
[0047] It is understandable that the object association relationship refers to a certain connection or dependency relationship existing among multiple operation and maintenance objects. These relationships can be physical connections (such as network connections, hardware interface connections), logical dependencies (such as an application calling a database), or business logic associations (such as a service depending on the normal operation of another service). The operation and maintenance object combination is a set containing multiple operation and maintenance objects, and these objects are directly or indirectly associated with the abnormal operation and maintenance objects. By determining this combination, the operation and maintenance team can more comprehensively understand the impact scope of abnormal situations, thereby more effectively conducting fault troubleshooting and recovery work.
[0048] S103: According to the historical log patterns in the log pattern library, perform combined marking processing on the log patterns of the operation and maintenance object combination to obtain target patterns of multiple categories.
[0049] Among them, the log pattern library is generated based on the historical log information of all operation and maintenance objects.
[0050] It is understandable that during fault location, the log patterns of all operation and maintenance object combinations and abnormal operation and maintenance objects will be retrieved. These log patterns will be automatically generated according to a preset method when the log information is generated for subsequent extraction during the analysis of the log information. There are multiple types of historical log patterns stored in the log pattern library. The log patterns of the operation and maintenance object combination can be divided into multiple categories according to the historical log patterns. There may be multiple identical log patterns in each category. If the log patterns of the operation and maintenance object combination and abnormal operation and maintenance objects do not exist in the log pattern library, the log patterns of the abnormal operation and maintenance objects and combinations should also be divided into a separate category, that is, the unknown log pattern, for the staff to discover unknown abnormal situations.
[0051] S104: Perform fault determination processing on the target patterns of multiple categories respectively to obtain the fault points of the abnormal operation and maintenance objects.
[0052] It is understandable that a category is the result of classifying target patterns with similar characteristics or attributes. By classifying the same patterns into one category, the characteristics of each fault pattern can be understood more clearly, thereby more effectively conducting fault determination. The category also includes unknown patterns. Unknown patterns refer to those fault log information patterns that have not been recognized or classified. Due to the complexity and diversity of the system, new fault patterns may continuously appear, so unknown patterns will occur. For unknown patterns, operation and maintenance personnel need to handle them more carefully in order to promptly identify and respond to new fault types. By analyzing and processing according to the patterns, the location of the fault can be determined, helping operation and maintenance personnel quickly locate the problem and take corresponding measures for repair.
[0053] A log - based fault determination method provided in this embodiment. This method obtains an information anomaly alert, and determines an abnormal operation and maintenance object according to the information anomaly alert. According to the association relationship between the abnormal operation and maintenance object and the object, an operation and maintenance object combination is determined. According to the historical log patterns in the log pattern library, combination marking processing is performed on the log patterns of the operation and maintenance object combination to obtain target patterns of multiple categories; the log pattern library is generated according to the historical log information of all operation and maintenance objects. Fault determination processing is respectively performed on the target patterns of multiple categories to obtain the fault points of the abnormal operation and maintenance objects. This method brings great convenience to the log correlation analysis across operation and maintenance objects, can greatly reduce the amount of logs for manual analysis, and quickly discovers key fault information.
[0054] Figure 2 The flow diagram of a log - based fault determination method provided in an embodiment of this application Figure 2 As Figure 2 shown, on the basis of the Figure 1 embodiment, a possible implementation manner of establishing an object association relationship and generating a log pattern library is described in detail, including:
[0055] S201: Obtain the association configuration information of multiple servers and components in the monitoring system.
[0056] Among them, the servers and components together constitute the operation and maintenance objects, and the association configuration information is used to indicate the association relationship between the components and the servers and the association relationship between the system and the servers.
[0057] It can be understood that the monitoring system is the system that needs to be maintained and detected. The system includes multiple servers and components. First, obtain the association configuration information of the servers and components. The association configuration information includes, but is not limited to, the network address (Internet Protocol Address, abbreviated as IP) of the server where the component is located, the server IP group of the cluster network system, the application access relationship, etc. The association configuration information can be stored in different tables in a relational database for convenient query.
[0058] S202: According to the association configuration information, establish an object association relationship among the monitoring system, the servers, and the components.
[0059] It can be understood that according to the association configuration information, the monitoring system, the servers, and the components are connected to the system, so that subsequently, the servers and components included in the monitoring system can be determined according to the monitoring system, and at the same time, the system where the servers and components are located can also be found according to the servers and components, providing a basis for data cross - component or cross - server association.
[0060] For example, take a container and its host machine as an example. The container has an independent set of IPs and ports. A host machine has an independent IP. One or more containers may be running on a host machine. The above three pieces of information (container IP, container port, host machine IP) can be stored in the container-host configuration table, that is, there is an object association relationship. During use, all containers running on a host machine can be queried through the host machine IP, and the host machine IP where the container is located can also be queried through the container IP and container port.
[0061] S203: Obtain the historical log information of the operation and maintenance object.
[0062] It can be understood that obtaining the historical log information means extracting the log information generated by the operation and maintenance object in the past. These log information record important data such as the running status, error reports, and operation records of the operation and maintenance object. Based on these data, the historical faults and log patterns of the operation and maintenance object can be determined, which are the basis for subsequent analysis and processing.
[0063] S204: Perform sentence segmentation on the historical log information to obtain multiple first segments.
[0064] It can be understood that since log information usually exists in the form of long texts or continuous streams, in order to facilitate subsequent analysis and processing, it needs to be segmented into multiple independent sentences or segments, that is, the first segments. It can be segmented according to sentence segmentation punctuation marks (, / ., / : / ? / ( / ), etc.). Each segment consists of multiple words, ensuring that each segment contains complete information and meaning.
[0065] S205: Perform matching and extraction processing on the first segments according to general elements to obtain multiple second segments.
[0066] Among them, the general elements include: IP (a number between 1 and 255 in four segments, connected by ·), time (T, in the format of year, month, day, hour, minute, second, etc.), independent string (S), integer data (N), etc.
[0067] It can be understood that the system will perform matching and extraction on the first segments according to the preset general elements. These general elements are common information in the log information. Through matching and extraction, key information such as time and objects can be obtained from the segments, which is convenient for subsequent statistical analysis. And the first segments after extracting the general elements are the second segments, which contain the key information in the log information.
[0068] S206: Perform structural analysis processing on the second segments to obtain the syntactic structure and structural semantics of the second segments.
[0069] Understandably, after obtaining the second language segment, the system will perform a deeper analysis on it, including syntactic structure analysis and semantic extraction. Syntactic structure analysis helps to understand the grammatical composition and sentence structure of the language segment, while semantic extraction can identify the key concepts and vocabulary in the language segment.
[0070] S207: Determine the proportion of each syntactic structure in the historical log information according to the syntactic structure.
[0071] Understandably, by counting the frequency and proportion of each grammatical structure in the entire log, we can understand which types of sentences or structures are more common in the log. We can mark the log patterns according to the meaning represented by the grammatical structure so that we can directly obtain the log information subsequently.
[0072] S208: Perform a similarity matching process on the structural semantics to obtain multiple historical log patterns, and save the historical log patterns to generate a log pattern library.
[0073] Among them, the similarity matching process refers to merging the second language segments with semantic similarity higher than the merging threshold into historical log patterns.
[0074] Understandably, the system will use the similarity calculation technology in natural language processing to perform a similarity matching process on the structural semantics in the second language segment. When the semantic similarity of two or more language segments exceeds the preset merging threshold, they will be merged into a historical log pattern. These patterns represent the common and repetitive content in the log information. By saving them as a log pattern library, it can provide an efficient reference and basis for subsequent log analysis and anomaly detection. At the same time, according to the structural proportion determined in the above steps, we can name each historical log pattern. For example, select the three keywords with the highest proportion as the name of the pattern to facilitate the classification and management of the pattern.
[0075] A log-based fault determination method provided by this embodiment. This method obtains the association configuration information between the server and components in the monitoring system, and constructs the object association relationship between the detection system, the server and the components. These information jointly define the operation and maintenance objects and their internal associations. Subsequently, the system collects the historical log information of the operation and maintenance objects, and performs sentence segmentation processing to obtain multiple first language segments. Using common element matching extraction, the second language segment is obtained, and its syntactic structure and structural semantics are further analyzed. According to the proportion of the syntactic structure, combined with the semantic similarity (merging the second language segments with similarity higher than the merging threshold), multiple historical log patterns are identified and saved to the log pattern library. The object association relationship established by this method can quickly extract relevant operation and maintenance objects, reducing the time generated by multiple extractions of operation and maintenance objects. At the same time, constructing a log pattern library facilitates the quick matching and analysis of subsequent logs, improving the stability of the system.
[0076] Figure 3 Schematic diagram of a log - based fault determination method provided by an embodiment of this application Figure 3 As shown Figure 3 in Figure 1 the figure, on the basis of the embodiment, a possible implementation manner of generating an information anomaly alarm is described in detail, including:
[0077] S301: Obtain the log information of the operation and maintenance object in real time.
[0078] It can be understood that in real - time log analysis and fault diagnosis, the system will perform real - time analysis and detection on the log information generated by the operation and maintenance object, so as to timely discover the faults of the operation and maintenance object indicated therein.
[0079] S302: Perform sentence segmentation and structure parsing processing on the log information to obtain the structural semantic meaning of the log information.
[0080] Among them, obtaining the structural semantic meaning of the log information according to the log information obtained in real time is the same as the process of processing the historical log information in steps S204 to S206, and will not be elaborated here.
[0081] It can be understood that the log information generated by each operation and maintenance object needs to be subjected to pattern analysis processing to obtain and save the log pattern, which can facilitate subsequent fault location based on the log, reduce the time for log parsing and analysis when an alarm occurs, quickly process the alarm information, and timely optimize the operation and maintenance object.
[0082] S303: Determine whether the structural semantic meaning matches the preset alarm policy. If so, execute step S304; if not, execute step S305.
[0083] S304: Generate an anomaly information alarm and generate a log pattern according to the structural semantic meaning.
[0084] S305: Generate a log pattern according to the structural semantic meaning.
[0085] It can be understood that the alarm policy is the first step in detecting the log, and the purpose is to identify whether the key information (i.e., the structural semantic meaning) in the log matches the preset alarm policy. The preset alarm policy is usually set based on business logic and specific conditions to identify potential problems or anomalies. The anomaly information alarm can be, for example, an alarm generated according to keyword and fixed - threshold detection. This alarm corresponds to the known anomaly information in the log.
[0086] Among them, generating the log pattern according to the structural semantic meaning in the above steps is the same as steps S207 to S208.
[0087] S306: Generate an unknown log alert when the log mode does not match the historical log modes in the log mode library.
[0088] It can be understood that after determining that there is no known abnormal information in the log information, the system will further match the generated log mode with the historical log modes in the log mode library. If the match fails, it means that an unknown log mode is encountered. The new log mode generally represents some newly emerging information in an abnormal operating state. At this time, the system will generate an unknown log alert to remind the staff that the log mode is unknown and to pay attention to and investigate possible anomalies.
[0089] Optionally, the abnormal alert also includes: a data interruption alert, and the data interruption alert includes: a log interruption alert and a process interruption alert. The following details the method of generating a data interruption alert:
[0090] When a data interruption occurs in the log information, perform a process check on the log information.
[0091] It can be understood that data interruption refers to the loss, delay, or abnormal interruption of log data, which may be caused by reasons such as network problems, system errors, or hardware failures. To determine the cause of the data interruption, the system will perform a process check, such as a heartbeat check. A heartbeat check is a commonly used monitoring method that confirms the running state of a process or service by periodically sending heartbeat signals.
[0092] When the process check indicates that the process is normal, generate a log interruption alert and generate a log mode based on the log information.
[0093] When the process check indicates that the process is abnormal, generate a process interruption alert to remind the staff that there is an anomaly in operation and maintenance.
[0094] It can be understood that if the process check indicates that the process is normal but the log data is still interrupted, the system will generate a log interruption alert. This may be due to problems with log generation or transmission. A log mode can be generated based on the determined log information and the log interruption can be reported. If the process check indicates that the process is abnormal, the system will generate a process interruption alert and remind the staff to perform operation and maintenance processing. This usually means that a serious problem has occurred in some part of the system or service and needs to be resolved immediately.
[0095] A log-based fault determination method provided in this embodiment obtains log information of an operation and maintenance object in real time, performs sentence segmentation and structure parsing processing on the log information to obtain the structural semantic meaning of the log information, determines whether the structural semantic meaning matches a preset alarm policy. If it matches, it is determined that there is an abnormal information alarm in the log information, and a log pattern is generated according to the structural semantic meaning. If it does not match, a log pattern is generated according to the structural semantic meaning. When the log pattern does not match the historical log pattern in the log pattern library, it is determined that there is an unknown log alarm in the log information. This method is more sensitive to the fault perception shown in the log, is not limited to known alarm types and key information, and can automatically perceive unknown abnormal situations.
[0096] Figure 4 It is a flow diagram of a log-based fault determination method provided in an embodiment of the present application. Figure 4 As Figure 4 shown, on the basis of the Figure 1 embodiment, a possible implementation manner for determining the fault point of the abnormal operation and maintenance object is described in detail, including:
[0097] S401: For any one of multiple log patterns of an abnormal operation and maintenance object and the combination of operation and maintenance objects, determine whether the log pattern matches the historical log pattern. If so, execute step S402. If not, execute step S403.
[0098] It can be understood that after obtaining the abnormal alarm, the event center in the system will retrieve the object association relationship and extract the log patterns of the associated operation and maintenance objects. Importantly, there may be operation and maintenance objects that have had information abnormal alarms among the associated operation and maintenance objects. For the log patterns of these operation and maintenance objects, the log patterns within a certain time range need to be extracted for the staff to analyze whether there are existing abnormalities caused by historical abnormalities. After extraction, the log patterns of the operation and maintenance objects are classified and saved for the staff to determine the fault according to the log pattern.
[0099] S402: Classify the log pattern to obtain a first target pattern.
[0100] S403: Save the log pattern to obtain a second target pattern.
[0101] It can be understood that if the log pattern matches the historical log pattern, that is, the information of the log pattern can be found in the log pattern library, it means that the pattern has been parsed in detail during the historical period, so it is saved as the first target pattern. If the log pattern does not match the historical log pattern, the log pattern is a new unparsed log pattern, which is saved as the second type of target pattern, and the staff can be reminded to pay attention in the future.
[0102] S404: Sort the target patterns according to the preset time sequence, category, and number of categories to obtain a pattern anomaly table.
[0103] Among them, the number of categories refers to the number of log patterns included in the same category.
[0104] It can be understood that the preset time sequence can be, for example, a reverse sort according to the generation time of the exception alarm, that is, starting from the generation time of the real-time exception alarm and tracing back to the generation time of the historical exception alarms existing in the operation and maintenance object combination. The category refers to the specific classification of the target log pattern. The system can comprehensively consider which patterns should be placed at the front based on these three situations and generate a pattern anomaly table. For example, if the log patterns of a certain category often have failures, this pattern can be ranked at the front, so that the staff can analyze this pattern first and quickly determine the failure. Or arrange multiple unknown patterns existing in the target patterns at the back, so that the staff can have enough time to analyze them. The arrangement order of the patterns in the pattern anomaly table can be adjusted according to the needs of the staff, not limited to only one sorting method.
[0105] S405: According to the pattern anomaly table, perform a fault determination process on the target patterns to obtain the fault points of the abnormal operation and maintenance objects.
[0106] It can be understood that, in accordance with the order in the pattern anomaly table, display the log information of each target pattern, and determine the fault location points in the log information, which can effectively reduce the number of log analyses. And since the patterns of the log information are the same, the same method can be used to determine the faults for the log information of the same pattern, improving the fault location efficiency.
[0107] A log-based fault determination method provided in this embodiment. This method determines whether the log pattern of the abnormal operation and maintenance object combination matches the historical log pattern. If it matches, classify the log pattern to obtain the first target pattern. If it does not match, save the log pattern to obtain the second target pattern. The first target pattern and the second target pattern constitute target patterns of multiple categories. Sort the target patterns according to the preset time sequence and the number of categories to obtain a pattern anomaly table; the number of categories refers to the number of log patterns included in the same category. According to the pattern anomaly table, perform a fault determination process on the target patterns to obtain the fault points of the abnormal operation and maintenance objects. This method divides the log patterns into multiple categories and locates the faults according to the categories, improving the fault location efficiency.
[0108] Figure 5 It is a flow schematic diagram of a log-based fault determination method provided in an embodiment of the present application. Figure 5 As Figure 5 shown, in Figure 1Based on the embodiments, a possible implementation manner of reusing fault events will be described in detail, including:
[0109] S501: Mark the abnormal information corresponding to the fault point to obtain a fault event.
[0110] It can be understood that in the post - mortem analysis session after fault recovery, it is necessary to mark the fault events. The marking elements include: associated logic marking, time marking, and object marking. The associated logic marking includes marking the abnormal alarm corresponding to the fault and the key log patterns. Marking the key log patterns can be part of the associated logic or a corresponding alarm policy can be configured. The time marking includes generating a fault time period based on the time matched by the abnormal alarm and the marked log patterns, and recording the duration of the fault. At the same time, the system also supports manual setting of the marking time period by the staff. The object marking includes determining the fault object scope (abnormal impact scope) based on the object configuration relationship and the associated logic marking, and also supports manual modification and adjustment of the operation and maintenance objects by the staff.
[0111] In addition to marking the above elements, supplementary descriptions of the fault (such as emergency operation description, fault root cause description, etc.) can also be added. The fault event after marking is composed of the data marked in the above aspects, forming elements such as the abnormal alarm associated logic of the fault event, the period and duration of the event, and the key - concerned impact scope. The relevant information will be stored in the system's fault library as a whole. In addition, it also includes the time when this fault occurred, as well as information such as the records of historical fault matching hits.
[0112] S502: If an information abnormal alarm is obtained again, match the information abnormal alarm according to the fault event to obtain a matching result.
[0113] It can be understood that in the case of an existing fault library, if an information abnormal alarm appears during the analysis of real - time log information, it can be matched in the fault library through the information abnormal alarm, and match the comprehensive information such as abnormal alarm, time interval, object association relationship, and associated logic with the fault events in the fault library.
[0114] S503: When the matching result indicates a successful match, display the abnormal information of the fault event so that the staff can determine the fault point of the information abnormal alarm.
[0115] It can be understood that if a fault event matches successfully or meets the matching requirements, relevant abnormal alarms and event information can be extracted and displayed on the display screen for reference by the staff to solve new faults. Among them, meeting the matching requirements can be that the similarity between the information abnormal alarm and the fault event in the fault library meets a preset value. For example, when the similarity between the information indicated by the information abnormal alarm and the fault event reaches 80%, the information of this fault event can be retrieved and displayed.
[0116] A fault determination method based on logs provided in this embodiment obtains a fault event by performing a marking process on abnormal information corresponding to a fault point. If an information abnormal alarm is obtained again, the information abnormal alarm is matched according to the fault event to obtain a matching result. When the matching result indicates a successful match, the abnormal information of the abnormal point is displayed so that the staff can locate the fault point of the information abnormal alarm. Through fault marking and fault matching, this method can effectively automatically associate and match valid information when a fault occurs, realize the precipitation and reuse of fault location and handling experience in the operation and maintenance process, and improve the operation and maintenance efficiency.
[0117] Figure 6 It is a schematic structural diagram of a fault determination device based on logs provided in this application. As Figure 6 shown, this application provides a fault determination device based on logs. The fault determination device 600 based on logs includes:
[0118] An acquisition module 601, configured to acquire an information abnormal alarm and determine an abnormal operation and maintenance object according to the information abnormal alarm;
[0119] A determination module 602, configured to determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship; the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object;
[0120] A processing module 603, configured to perform combined marking processing on the log patterns of the operation and maintenance object combination according to the historical log patterns in the log pattern library to obtain target patterns of multiple categories; the log pattern library is generated according to the historical log information of all operation and maintenance objects;
[0121] The processing module 603 is further configured to perform fault determination processing on the target patterns of multiple categories respectively to obtain the fault points of the abnormal operation and maintenance objects.
[0122] Optionally, the device further includes: a establishment module 604;
[0123] The obtaining module 601 is further configured to obtain the association configuration information between multiple servers and components in the monitoring system; wherein, the servers and the components together constitute the operation and maintenance objects, and the association configuration information is used to indicate the association relationship between the components and the servers and the association relationship between the system and the servers;
[0124] The establishing module 604 is configured to establish an object association relationship among the monitoring system, the servers, and the components according to the association configuration information.
[0125] Optionally, the obtaining module 601 is further configured to obtain the historical log information of the operation and maintenance objects;
[0126] The processing module 603 is further configured to perform sentence segmentation processing on the historical log information to obtain a plurality of first text segments;
[0127] The processing module 603 is further configured to perform matching and extraction processing on the first text segments according to common elements to obtain a plurality of second text segments;
[0128] The processing module 603 is further configured to perform merging and classification processing on the second text segments to obtain the historical log patterns of the historical log information, and save the historical log patterns to generate a log pattern library.
[0129] Optionally, the processing module 603 is further configured to perform structural analysis processing on the second text segments to obtain the syntactic structure and structural semantic meaning of the second text segments;
[0130] The determining module 602 is further configured to determine the proportion of each syntactic structure in the historical log information according to the syntactic structure;
[0131] The processing module 603 is further configured to perform similarity matching processing on the structural semantic meanings to obtain a plurality of historical log patterns, and the similarity matching processing refers to merging the second text segments with the semantic similarity higher than the merging threshold into historical log patterns.
[0132] Optionally, the device further includes: a judgment module 605 and a generation module 606;
[0133] The obtaining module 601 is further configured to obtain the log information of the operation and maintenance objects in real time;
[0134] The processing module 603 is further configured to perform sentence segmentation and structural analysis processing on the log information to obtain the structural semantic meaning of the log information;
[0135] The judgment module 605 is configured to judge whether the structural semantic meaning matches a preset alarm policy;
[0136] The generation module 606 is configured to generate an abnormal information alert when the structural semantic meaning matches the preset warning policy, and generate a log pattern according to the structural semantic meaning;
[0137] The generation module 606 is further configured to generate a log pattern according to the structural semantic meaning when the structural semantic meaning does not match the preset warning policy;
[0138] The generation module 606 is further configured to generate an unknown log alert when the log pattern does not match the historical log pattern in the log pattern library.
[0139] Optionally, the processing module 603 is further configured to perform a process check on the log information when a data interruption occurs in the log information;
[0140] The generation module 606 is further configured to generate a log interruption alert and generate a log pattern according to the log information when the process check indicates that the process is normal;
[0141] The generation module 606 is further configured to generate a process interruption alert to remind the staff that there is an abnormality in operation and maintenance when the process check indicates that the process is abnormal.
[0142] Optionally, the judgment module 605 is further configured to determine whether any one of the multiple log patterns of the abnormal operation and maintenance object and the operation and maintenance object combination matches the historical log pattern;
[0143] The processing module 603 is further configured to perform a classification process on the log pattern to obtain a first target pattern when the log pattern matches the historical log pattern;
[0144] The processing module 603 is further configured to perform a saving process on the log pattern to obtain a second target pattern when the log pattern does not match the historical log pattern.
[0145] Optionally, the processing module 603 is further configured to sort the target patterns according to a preset time sequence, category, and number of categories to obtain a pattern exception table; the number of categories refers to the number of log patterns included in the same category;
[0146] The processing module 603 is further configured to perform a fault determination process on the target pattern according to the pattern exception table to obtain the fault point of the abnormal operation and maintenance object.
[0147] Optionally, the processing module 603 is further configured to perform a marking process on the abnormal information corresponding to the fault point to obtain a fault event;
[0148] The processing module 603 is further configured to, if an information anomaly alarm is obtained again, perform matching processing on the information anomaly alarm according to the fault event to obtain a matching result;
[0149] The determining module 602 is further configured to, when the matching result indicates a successful match, display the anomaly information of the fault event so that the staff can determine the fault point of the information anomaly alarm.
[0150] The implementation principle and technical effect of the log-based fault determination device provided by the embodiments of the present application are similar to those of each part in the foregoing log-based fault determination method, and will not be elaborated here.
[0151] Figure 7 It is a schematic structural diagram of a log-based fault determination device provided by the present application. As Figure 7 shown, the present application provides a log-based fault determination device. The log-based fault determination device 700 includes: a receiver 701, a transmitter 702, a processor 703, and a memory 704.
[0152] The receiver 701 is configured to receive instructions and data;
[0153] The transmitter 702 is configured to send instructions and data;
[0154] The memory 704 is configured to store computer execution instructions;
[0155] The processor 703 is configured to execute the computer execution instructions stored in the memory 704 to implement each step performed by the method in the foregoing embodiments. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0156] Optionally, the foregoing memory 704 can be either independent or integrated with the processor 703.
[0157] When the memory 704 is independently provided, the electronic device further includes a bus for connecting the memory 704 and the processor 703.
[0158] The implementation principle and technical effect of the electronic device provided in this embodiment can be referred to the foregoing embodiments, and will not be elaborated here.
[0159] The embodiments of the present application further provide a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the method described in any of the foregoing embodiments is implemented.
[0160] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by the processor, the method described in any of the foregoing embodiments is implemented.
[0161] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0162] Furthermore, it should be noted that although the steps in the flowchart are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the direction of the arrows. Unless there is a clear description in this article, the execution of these steps has no strict sequence limitation, and these steps can be executed in other sequences. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution sequence of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0163] It should be understood that the above device embodiments are merely illustrative, and the devices of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0164] In addition, without special description, in each embodiment of this application, the functional units / modules can be integrated into one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.
[0165] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0166] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0167] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include the common general knowledge or conventional technical means in this technical field not disclosed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this application are pointed out by the following claims.
[0168] It should be understood that this application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
Claims
1. A log-based fault determination method, characterized in that: include: Obtaining an abnormal information alarm, and determining an abnormal operation and maintenance object based on the abnormal information alarm; Determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship; the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object; According to the historical log patterns in the log pattern library, the log patterns of the operation and maintenance object combination are subjected to combination marking processing to obtain target patterns of multiple categories; The log pattern library is generated based on the historical log information of all operation and maintenance objects; Fault determination processing is performed on the target modes of multiple categories respectively to obtain the fault points of the abnormal operation and maintenance objects.
2. The method according to claim 1, characterized in that Before obtaining the abnormal information alarm, the method further includes: Acquire association configuration information of multiple servers and components in the monitoring system; wherein the server and the component together constitute an operation and maintenance object, and the association configuration information is used to indicate the association relationship between the component and the server and the association relationship between the system and the server; An object association relationship is established between the monitoring system, the server and the component according to the association configuration information.
3. The method according to claim 2, characterized in that The method further comprises: Obtaining historical log information of the operation and maintenance object; Segmenting the historical log information to obtain a plurality of first paragraphs; Performing matching and extraction processing on the first language segment according to the common elements to obtain a plurality of second language segments; The second paragraphs are merged and classified to obtain a historical log pattern of the historical log information, and the historical log pattern is saved to generate a log pattern library.
4. The method according to claim 3, characterized in that The merging and classifying the second paragraphs to obtain the historical log mode of the historical log information includes: Performing structural analysis on the second paragraph to obtain the syntactic structure and structural meaning of the second paragraph; According to the syntactic structure, determining the proportion of each syntactic structure in the historical log information; A similarity matching process is performed on the structural word meaning to obtain a plurality of historical log patterns, wherein the similarity matching process refers to merging the second paragraphs whose word meaning similarity is higher than a merging threshold into the historical log pattern.
5. The method according to claim 1, characterized in that The abnormal alarm includes: abnormal information alarm and unknown log alarm. Before obtaining the abnormal information alarm, the method further includes: Obtain log information of operation and maintenance objects in real time; Performing sentence segmentation and structural analysis on the log information to obtain the structural meaning of the log information; Determining whether the structure word meaning matches the preset alarm strategy; In the case where the structural word meaning matches the preset alarm strategy, an abnormal information alarm is generated, and a log pattern is generated according to the structural word meaning; In the case where the structural word meaning does not match the preset alarm strategy, generating a log pattern according to the structural word meaning; In the event that the log pattern does not match the historical log pattern, an unknown log alert is generated.
6. The method according to claim 5, characterized in that The abnormal alarm further includes: a data interruption alarm, and the data interruption alarm includes: a log interruption alarm and a process interruption alarm. The method further includes: In the event that data interruption occurs in the log information, performing a process check on the log information; If the process check indicates that the process is normal, a log interruption alarm is generated, and a log mode is generated according to the log information; When the process check indicates that the process is abnormal, a process interruption alarm is generated to remind the operation and maintenance staff that there is an abnormality.
7. The method according to claim 1, characterized in that According to the historical log patterns in the log pattern library, the log patterns of the operation and maintenance object combination are combined and marked to obtain target patterns of multiple categories, including: For any one of the multiple log patterns of the abnormal operation and maintenance object and the combination of the operation and maintenance object, determine whether the log pattern matches the historical log pattern; In the case where the log pattern matches the historical log pattern, classifying the log pattern to obtain a first target pattern; In the case that the log mode does not match the historical log mode, the log mode is saved to obtain a second target mode.
8. The method according to claim 1, characterized in that The performing fault determination processing on the target modes of the plurality of categories respectively to obtain the fault point of the abnormal operation and maintenance object includes: The target patterns are sorted according to a preset time sequence, category and category quantity to obtain a pattern exception table; the category quantity refers to the number of log patterns included in the same category; According to the mode abnormality table, a fault determination process is performed on the target mode to obtain the fault point of the abnormal operation and maintenance object.
9. The method according to claim 1, characterized in that: After obtaining the fault point of the abnormal operation and maintenance object, the method further includes: Marking the abnormal information corresponding to the fault point to obtain a fault event; If an information abnormality alarm is obtained again, matching processing is performed on the information abnormality alarm according to the fault event to obtain a matching result; When the matching result indicates a successful match, the abnormal information of the fault event is displayed so that the staff can determine the fault point of the abnormal information alarm.
10. A fault determination device based on logs, characterized in that: include: An acquisition module is used to acquire an abnormal information alarm and determine an abnormal operation and maintenance object according to the abnormal information alarm; A determination module, used to determine an operation and maintenance object combination according to the abnormal operation and maintenance object and the object association relationship; the object association relationship is used to indicate the association relationship between multiple operation and maintenance objects, and the operation and maintenance object combination is used to indicate multiple operation and maintenance objects associated with the abnormal operation and maintenance object; A processing module, used to perform combined marking processing on the log patterns of the operation and maintenance object combination according to the historical log patterns in the log pattern library, so as to obtain target patterns of multiple categories; the log pattern library is generated according to the historical log information of all operation and maintenance objects; The processing module is used to perform fault determination processing on the target modes of multiple categories respectively to obtain the fault point of the abnormal operation and maintenance object.
11. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 9 when executed by a processor.
13. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 9 when being executed by a processor.
Citation Information
Cited By
Cloud management method and device of product, storage medium and electronic equipment
CN121724008A