Fault root cause positioning method and system, computer storage medium and program product

By building a troubleshooting logic tree and combining qualitative and quantitative analysis, the problem of fault root cause positioning in complex application systems is solved, and fast and accurate fault location and fault propagation link acquisition is achieved, which improves troubleshooting efficiency and system availability.

CN119961029APending Publication Date: 2025-05-09CHINA UNIONPAY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411474283.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In complex enterprise application systems, the fundamental cause of rapid and accurate positioning of system failure is a key link for operation and maintenance personnel, but the existing technology is difficult to effectively solve this problem, especially when facing problems such as massive monitoring data, false alarms, missed reports and alarm delays.

Method used

By building a troubleshooting logic tree, combined with qualitative and quantitative analysis, the architectural topology diagram and abnormal alarm information of the application system are obtained, potential fault points are identified, the fault probability and impact degree are calculated, and the root cause fault point and its fault propagation link are finally determined.

Benefits of technology

It realizes rapid and accurate positioning of the root causes of application system failures, obtains fault propagation links, significantly improves the efficiency of troubleshooting and resolution, reduces fault recovery time, and improves the stability and availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961029A_ABST
    Figure CN119961029A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a fault root cause positioning method, a fault root cause positioning system for implementing the method, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring an architecture topological graph and abnormal alarm information of an application system; constructing a troubleshooting logic tree based on the architecture topological graph and the abnormal alarm information, wherein the troubleshooting logic tree indicates potential fault propagation links among the configuration components in the architecture topological graph by utilizing a logic relationship among the nodes; performing qualitative analysis on the fault removal logic tree to identify potential fault points causing faults; performing quantitative analysis on the fault removal logic tree to calculate the fault probability and the influence degree of each potential fault point; and determining a root cause fault point according to results of the qualitative analysis and the quantitative analysis, and obtaining a fault propagation link of the root cause fault point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically to a fault root cause locating method, a fault root cause locating system implementing the method, a computer-readable storage medium, and a computer program product. Background Art

[0002] In today's enterprise environment, with the increasing logical complexity of application systems and the emphasis and investment of R&D and operation teams on high-availability architecture, the scale of computer equipment is showing a trend of continuous expansion. This development trend inevitably leads to frequent system anomalies and failures, which in turn leads to the diversification of equipment abnormality alarm types and a sharp increase in the number of alarms. In order to ensure that the system can maintain high availability and reliability at the service level, operation and maintenance personnel must respond quickly to various abnormal alarms and troubleshoot in a timely manner to minimize the impact on business operations. In this process, it is particularly important to quickly and accurately locate the root cause of the system failure, which is a key link for operation and maintenance personnel to take effective countermeasures and restore the normal operation of the system.

[0003] Although the general high-availability architecture of the system has improved the robustness of the system to a certain extent, and with the continuous improvement of monitoring tools, business anomalies or system failures can be handled more properly. However, this also puts higher demands on the company's data collection and monitoring system, and requires operation and maintenance personnel to have rich experience and professional knowledge in order to accurately identify the root cause of the failure from the massive monitoring data. This process is not only time-consuming, but also prone to errors. To make it more complicated, when the system abnormal alarm has problems such as false alarms, missed alarms, or alarm delays, it will become more difficult to locate the root cause of the system failure.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] In order to solve or at least alleviate one or more of the above problems, the embodiments of the present application provide a fault root cause location method, a fault root cause location system implementing the method, a computer-readable storage medium and a computer program product, which construct a troubleshooting logic tree and combine qualitative and quantitative analysis to achieve rapid and accurate location of the root cause of the application system fault, and at the same time obtain the fault propagation link, effectively improving the efficiency of fault detection and resolution.

[0006] According to a first aspect of the present application, a method for locating the root cause of a fault is provided, the method comprising the following steps: obtaining an architecture topology diagram and abnormal alarm information of an application system, wherein the architecture topology diagram indicates the association relationship and hierarchical structure between the configuration components of the application system, and the abnormal alarm information includes an identifier of the configuration component related to the fault and the level at which it is located, and the importance of the abnormality; constructing a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information, the troubleshooting logic tree using the logical relationship between each node to indicate the potential fault propagation link between each configuration component in the architecture topology diagram; performing a qualitative analysis on the troubleshooting logic tree to identify the potential fault point that causes the fault; performing a quantitative analysis on the troubleshooting logic tree to calculate the failure probability and impact degree of each potential fault point; and determining the root cause fault point based on the results of the qualitative analysis and the quantitative analysis, and obtaining the fault propagation link of the root cause fault point.

[0007] As an alternative or supplement to the above solution, the fault root cause location method according to an embodiment of the present application also includes: generating the architecture topology diagram based on the CMDB configuration management information of the application system and a preset traceability path.

[0008] As an alternative or supplement to the above scheme, in a fault root cause locating method according to an embodiment of the present application, based on the CMDB configuration management information of the application system and a preset traceability path, generating the architecture topology diagram includes: retrieving and obtaining the CMDB configuration model of the application system from a troubleshooting knowledge base; traversing and parsing the CMDB configuration model using a breadth-first search algorithm to determine the hierarchical structure and nodes at each level of the application system; supplementing the hierarchical structure and nodes at each level according to the traceability path stored in the troubleshooting knowledge base; and generating the architecture topology diagram based on the supplemented hierarchical structure and nodes at each level.

[0009] As an alternative or supplement to the above solution, in a fault root cause location method according to an embodiment of the present application, the tracing path indicates how to trace back from a bottom-level child node to a sub-level root node.

[0010] As an alternative or supplement to the above scheme, the fault root cause locating method according to an embodiment of the present application also includes: monitoring the operating status of each configuration component in the architecture topology diagram; and determining the abnormal alarm information based on the monitoring data and abnormal indicator information for each type of fault.

[0011] As an alternative or supplement to the above scheme, the method for locating the root cause of a fault according to an embodiment of the present application also includes: adding annotation information to the nodes in the architecture topology diagram based on the abnormal alarm information, wherein each node in the architecture topology diagram corresponds to a type of configuration component, and the annotation information includes the number of instances of this type of configuration component and the number of abnormal alarms.

[0012] As an alternative or supplement to the above solution, in a fault root cause locating method according to an embodiment of the present application, the triggering conditions of the method include: manual triggering, timed polling triggering, and alarm event triggering.

[0013] As an alternative or supplement to the above scheme, in a method for locating the root cause of a fault according to an embodiment of the present application, constructing a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information includes: comparing the abnormal alarm information of each root node in the architecture topology diagram, and determining the root node with the highest abnormal importance as the root event; decomposing the cause of the root event into sub-events step by step, wherein each sub-event indicates whether the operating status of the subordinate node of the root event is abnormal; and connecting the sub-events through logical AND or logical OR according to the association relationship and hierarchical structure between the configuration components in the architecture topology diagram.

[0014] As an alternative or supplement to the above scheme, in a fault root cause locating method according to an embodiment of the present application, a qualitative analysis is performed on the troubleshooting logic tree to identify potential fault points that cause the fault, including: using a fault tree analysis method to simplify the troubleshooting logic tree to identify multiple minimum cut sets that cause the fault, wherein each minimum cut set includes a minimum set of basic events that cause the fault.

[0015] As an alternative or supplement to the above solution, in a fault root cause location method according to an embodiment of the present application, each minimum cut set corresponds to a minimum troubleshooting subtree.

[0016] As an alternative or supplement to the above scheme, in a fault root cause locating method according to an embodiment of the present application, quantitative analysis of the troubleshooting logic tree includes: determining the critical importance of the sub-event based on the importance of the sub-event, the level of the configuration component associated with the sub-event, and the timestamp of the sub-event; determining the structural importance of the sub-event based on the ratio of the number of combinations in which the state of the sub-event is the same as the state of the root event to the total number of system states without considering the probability of occurrence of the sub-event; and determining the probabilistic importance of the sub-event based on the degree of change in the probability of occurrence of the root event caused by the change in the probability of occurrence of the sub-event.

[0017] As an alternative or supplement to the above solution, in the fault root cause location method according to an embodiment of the present application, sub-event X i The key importance of K (i) The calculation is as follows:

[0018] I K (i) = St(a(X i ),l(X i ),t(X i))

[0019] Among them, a(X i ) represents the sub-event X i The importance of l(X i ) is related to the sub-event X i The level at which the associated configuration component is located, t(X i ) is the sub-event X i The timestamp generated, function St() represents the time according to a(X i )、l(X i ), t(X i ) for sub-event X i Sort and rate the importance of.

[0020] As an alternative or supplement to the above solution, in the fault root cause location method according to an embodiment of the present application, sub-event X i The structural importance of I τ (i) The calculation is as follows:

[0021]

[0022] Among them, X i _1 indicates the sub-event X i The state is 1, ΣT(1,X i _1) represents the sub-event X i The number of combinations with the same status as the root event, X i _0The sub-event X i The state is 0, ΣT(1,X i _0) indicates the sub-event X i The state of the root event is 0 and the state of the root event is 1, where N is the total number of sub-events.

[0023] As an alternative or supplement to the above solution, in the fault root cause location method according to an embodiment of the present application, sub-event X i The probability importance I ρ (i) The calculation is as follows:

[0024] I ρ (i) = P(T,X i _1)-P(T,X i _0)

[0025] Among them, P(T,X i _1) represents the sub-event X i The probability of the root event occurring when it occurs; P(T,X i _0) indicates the sub-event X i The probability of the root event occurring when it does not occur.

[0026] As an alternative or supplement to the above scheme, in a fault root cause location method according to an embodiment of the present application, quantitative analysis of the troubleshooting logic tree also includes: calculating the weighted sum of the critical importance, the structural importance and the probability importance to determine the comprehensive importance of the sub-event.

[0027] As an alternative or supplement to the above solution, in a fault root cause location method according to an embodiment of the present application, determining the root cause failure point according to the results of qualitative analysis and quantitative analysis includes: determining the one or more sub-events with the highest comprehensive importance as the root cause failure point.

[0028] As an alternative or supplement to the above scheme, in a fault root cause locating method according to an embodiment of the present application, obtaining the fault propagation link of the root cause fault point includes: obtaining a minimum troubleshooting subtree including the root cause fault point; and determining the fault propagation link based on the minimum troubleshooting subtree.

[0029] As an alternative or supplement to the above scheme, the fault root cause location method according to an embodiment of the present application also includes: constructing a troubleshooting knowledge base, wherein the troubleshooting knowledge base includes one or more of the following items: CMDB configuration management information of each application system, the type of each type of fault, abnormal indicator information of each type of fault, the abnormal importance of each type of fault, the importance of each type of sub-event, the traceability path, and the processing logic of fault location; and updating the troubleshooting knowledge base according to historical fault root cause location results.

[0030] According to a second aspect of the present application, a fault root cause location system is provided, comprising: a memory; a processor; and a computer program stored in the memory and executable on the processor, wherein the execution of the computer program results in the following operations: obtaining an architecture topology diagram and abnormal alarm information of an application system, wherein the architecture topology diagram indicates the association relationship and hierarchical structure between configuration components of the application system, and the abnormal alarm information includes an identifier of a configuration component related to the fault and the level at which it is located, and an abnormal importance; constructing a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information, wherein the troubleshooting logic tree uses the logical relationship between each node to indicate a potential fault propagation link between each configuration component in the architecture topology diagram; performing a qualitative analysis on the troubleshooting logic tree to identify a potential fault point that causes the fault; performing a quantitative analysis on the troubleshooting logic tree to calculate the fault probability and impact degree of each potential fault point; and determining a root cause fault point based on the results of the qualitative analysis and the quantitative analysis, and obtaining the fault propagation link of the root cause fault point.

[0031] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes instructions, and the instructions, when run, execute any one of the fault root cause locating methods described in the first aspect of the present application.

[0032] According to a fourth aspect of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, any one of the fault root cause locating methods described in the first aspect of the present application is implemented.

[0033] According to one or more embodiments of the present application, the fault root cause location solution realizes the rapid and accurate location of the root cause of the application system fault by constructing a troubleshooting logic tree and combining qualitative and quantitative analysis, and obtains the fault propagation link, which effectively improves the efficiency of fault troubleshooting and resolution. Specifically, the solution obtains the architecture topology diagram and abnormal alarm information of the application system, so as to fully and accurately understand the association relationship and hierarchical structure between the configuration components of the system, as well as the specific manifestation of the fault, laying a solid foundation for subsequent analysis; then, by constructing a troubleshooting logic tree, it not only intuitively shows the possible propagation path of the fault, but also utilizes the logical relationship between components, greatly improving the accuracy and efficiency of fault analysis; through qualitative analysis, the solution can narrow the root cause location range, identify potential fault points, and reduce the time cost of blind troubleshooting; through further quantitative analysis, the failure probability and impact degree of each potential fault point can be determined, providing a basis for locating the final root cause fault point; finally, according to the comprehensive analysis results, the solution can accurately lock the root cause fault point and clearly display its fault propagation link, which helps operation and maintenance personnel to quickly locate the problem and formulate a repair plan, thereby effectively shortening the fault recovery time and improving the stability and availability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and / or other aspects and advantages of the present application will become clearer and easier to understand through the following description of various aspects in conjunction with the accompanying drawings, in which the same or similar units are represented by the same reference numerals. In the accompanying drawings:

[0035] Figure 1 is a schematic flow chart of a fault root cause location method 10 according to one or more embodiments of the present application;

[0036] Figure 2 is a schematic flow chart of a method 20 for generating an architecture topology diagram according to one or more embodiments of the present application;

[0037] Figure 3 A schematic diagram of an architecture topology diagram drawn according to one or more embodiments of the present application;

[0038] Figure 4A schematic diagram of a troubleshooting logic tree constructed according to one or more embodiments of the present application;

[0039] Figure 5 A schematic diagram of a minimum troubleshooting logic tree constructed according to one or more embodiments of the present application;

[0040] Figure 6 FIG. 6 is a schematic block diagram of a fault root cause location system 60 according to one or more embodiments of the present application. DETAILED DESCRIPTION

[0041] The description of the following specific embodiments is merely exemplary in nature and is not intended to limit the disclosed technology or the application and use of the disclosed technology. In addition, it is not intended to be bound by any express or implied theory presented in the aforementioned technical field, background technology or the following specific embodiments.

[0042] In the following detailed description of the embodiments, many specific details are set forth in order to provide a more thorough understanding of the disclosed technology. However, it is apparent to one of ordinary skill in the art that the disclosed technology can be practiced without these specific details. In other instances, well-known features are not described in detail to avoid unnecessarily complicating the description.

[0043] Terms such as "comprising" and "including" indicate that in addition to the units and steps directly and clearly stated in the specification, the technical solution of the present application does not exclude the situation of having other units and steps that are not directly or clearly stated. Terms such as "first" and "second" do not indicate the order of units in terms of time, space, size, etc., but are only used to distinguish between units.

[0044] Hereinafter, various exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings.

[0045] Referring to the accompanying drawings, Figure 1 It is a schematic flow chart of a fault root cause location method 10 according to one or more embodiments of the present application.

[0046] like Figure 1 As shown, in step S110, the architecture topology diagram and abnormal alarm information of the application system are obtained.

[0047] The triggering of the fault root cause location method 10 or step S110 can be based on internal factors (e.g., timed polling trigger, alarm event trigger) or external factors (e.g., manual trigger). Timed polling refers to regular checks on the system based on preset time intervals. In one possible implementation, when the set time point is reached, the system automatically triggers the fault root cause location method 10 or step S110 to obtain the architecture topology diagram and abnormal alarm information of the specific application system, and start fault analysis. Alarm event triggering refers to the generation of an alarm event as a trigger signal in response to the detection of an abnormality or a fault reaching a preset alarm level. In one possible implementation, if 3 faults are detected within 5s, a trigger signal is generated. In another possible implementation, if the abnormal importance of the detected fault exceeds a preset threshold, a trigger signal is generated. Manual triggering refers to the system administrator or operation and maintenance personnel starting the fault root cause location method 10 or step S110 through manual operation.

[0048] The architecture topology diagram in this article can indicate the association relationship and hierarchical structure between various configuration components of an application system (such as a computer application system). Configuration components refer to each specific, executable or manageable unit in the application system, which can usually include servers, network devices, databases, application modules, etc. In the architecture topology diagram, a configuration component can be represented as a node in the diagram, and a class of configuration components can be uniformly represented as nodes in the diagram. The hierarchical structure describes the hierarchical relationship of configuration components in the organization in the architecture topology diagram, and these levels usually represent different abstraction levels and functional areas. Exemplarily, the levels can be divided into business layer, application layer, service layer, logic layer and hardware layer, etc., each of which represents different configuration components and the interaction between them. The association relationship describes whether the configuration components have connection properties, which depends on whether there is a physical or logical relationship between the component instances. In the architecture topology diagram, these association relationships are usually represented by lines or arrows.

[0049] Optionally, the architecture topology diagram is generated based on the CMDB configuration management information of the application system and a preset traceability path. CMDB (Configuration Management Database) is a logical database that contains information about the entire life cycle of configuration components and the relationships between configuration components (including physical relationships, real-time communication relationships, non-real-time communication relationships, and dependencies). Exemplarily, CMDB configuration management information may include the following: configuration components; relationship information between configuration components; configuration component attribute information, such as identification, type, version, location, status, etc. The traceability path refers to the process of tracing back to other related configuration components from a certain configuration component through a series of relationship links.

[0050] Exemplarily, CMDB configuration management information and traceability paths can be pre-stored in the troubleshooting knowledge base. Optionally, in one or more embodiments according to the present application, a troubleshooting knowledge base is constructed, wherein the troubleshooting knowledge base may include one or more of the following items: CMDB configuration management information of each application system, types of various faults, abnormal indicator information of various faults, abnormal importance of various faults, importance of various sub-events, traceability paths, and processing logic for fault location. Furthermore, the troubleshooting knowledge base can also be updated based on subsequent historical fault root cause location results. In the process of fault root cause location, the troubleshooting knowledge base provides a priori knowledge that the entire process depends on, and automatically saves troubleshooting experience and iteratively updates knowledge base information after each troubleshooting, thereby feeding back to optimize the fault analysis and root cause location process.

[0051] Reference below Figure 2 , Figure 2 It is a schematic flow chart of a method 20 for generating an architecture topology diagram according to one or more embodiments of the present application.

[0052] like Figure 2 As shown, in step S210, the CMDB configuration model of the specified application system is retrieved and obtained from the troubleshooting knowledge base. In step S220, the CMDB configuration model is traversed and parsed using a breadth-first search algorithm to determine the hierarchical structure and nodes at each level of the application system. Exemplarily, the breadth-first search algorithm can be used to obtain nodes at each level from top to bottom with the application system as the center. This means starting from the top-level node, searching downward layer by layer until all reachable nodes are found. Optionally, method 20 may also include: based on the CMDB configuration model, pruning the acquired nodes at all levels and associations, and cleaning redundant data. This step is intended to remove unnecessary and redundant nodes and relationships, as well as clean data, to ensure the accuracy and consistency of information, and to prepare for the subsequent drawing of a topology map. In step S230, it is determined whether there are root nodes of other sub-levels. If there are sub-level root nodes, step S240 is executed: the hierarchical structure and nodes at each level are supplemented according to the traceability path stored in the troubleshooting knowledge base. Exemplarily, the traceability path provided by the operation and maintenance expert can be obtained from the troubleshooting knowledge base, where the traceability path includes instructions on how to trace from the bottom-level sub-nodes to the sub-level root nodes; then for each sub-level root node, the corresponding bottom-level sub-nodes are searched and supplemented according to the traceability path to ensure that all relevant nodes and relationships are fully included. If there is no sub-level root node, proceed directly to step S250. In step S250, according to the hierarchical structure, the obtained application system configuration components are used to draw the corresponding architecture topology diagram according to the association relationship between each node.

[0053] Figure 3 FIG. 1 is a schematic diagram of an architecture topology diagram drawn according to one or more embodiments of the present application. Figure 3 As shown in the figure, each node icon corresponds to a type of configuration component. The number of configuration component instances, the total number of fault alarms, and the number of critical alarms of the node can be marked below the icon. At the same time, it can also support node information drilling to display data information details. At the same time, it can be seen from the figure that the application system has a sub-level root node "cluster", such as Figure 3 According to the embodiments of the present application (eg, Figure 2 The topology diagram drawn by this top-down, root-tracing drawing scheme has obvious advantages in visual interaction and readability. Figure 3 Each type of configuration component in can contain multiple configuration component instances, and each configuration component instance has a unique configuration component identifier. Whether the various types of configuration components have a "connection" attribute depends on whether there is a physical or logical association relationship between the two configuration component instances.

[0054] Return to Figure 1 The abnormal alarm information in step S110 is key information for indicating potential or actual faults, which is crucial for quickly locating the fault, analyzing the root cause of the fault, and taking corresponding measures.

[0055] Optionally, based on the architecture topology diagram drawn in the above steps, the operating status of the configuration components in the architecture topology diagram can be monitored, and abnormal alarm information can be determined based on the monitoring data and abnormal indicator information for various types of faults. Exemplarily, reasonable abnormal indicator information can be set according to the normal operating range and historical data of the configuration components, and abnormal alarm information is generated when the monitoring data exceeds the range indicated by the abnormal indicator information. Due to the multimodal characteristics of data forms in various fields, the abnormal alarm information can be uniformly standardized for subsequent use and analysis. In a specific example, the standardized abnormal alarm information may include one or more of the following items: abnormal identification - the unique identification of this abnormal alarm information; timestamp - the specific time when the abnormality occurred; content details - alarms, events, abnormal log details and other contents related to the fault; abnormal indicator information - monitoring indicators or log templates in the fields of business, basic system, network, security, application, etc.; associated configuration instance - the configuration component identification and type of the application system associated with this fault; level - the level at which the configuration component is located, such as business layer-0, application layer-1, service layer-2, logic layer-3, hardware layer-4, etc.; abnormal importance - the severity or priority of this fault, such as general-0, important-1, critical-2, etc.

[0056] Optionally, annotation information can be added to the nodes in the architecture topology diagram based on the abnormal alarm information to form a preliminary troubleshooting diagram. The annotation information may include the number of instances of such configuration components and the number of abnormal alarms. In addition, in the preliminary troubleshooting diagram, abnormalities of different importance can be visually distinguished by different colors.

[0057] In step S120, a troubleshooting logic tree is constructed based on the architecture topology diagram and the abnormal alarm information. The troubleshooting logic tree in this article abstractly models the system architecture into a tree-shaped logic diagram, and uses the logical relationship between each node to indicate the potential fault propagation link between each configuration component in the architecture topology diagram.

[0058] Optionally, in step S120, firstly, the abnormal alarm information of each root node in the architecture topology diagram is compared, and the root node with the highest abnormal importance is determined as the root event; then the cause of the root event is decomposed into sub-events (including intermediate events that can be decomposed, and basic events that cannot be decomposed) level by level, wherein each sub-event indicates whether the operation status of the subordinate node of the root event is abnormal. Exemplarily, the sub-events and the root event can be defined as follows:

[0059] Sub-event (e.g., intermediate event, basic event): indicates whether the running status of the sub-node under the root node is abnormal:

[0060]

[0061] Root event: indicates whether the running status of the root node is abnormal. The root event is composed of all the sub-events that constitute it according to a certain logic. Therefore, the root event must be a sub-event X i function, so the root event state T is defined as:

[0062] T=T(X1,X2,...,X N )

[0063] Also define:

[0064]

[0065] In addition, according to the association relationship and hierarchical structure between the configuration components in the architecture topology diagram, the sub-events can be connected through logical AND or logical OR to transform the application system failure problem into a mathematical statistical probability problem. For example, the logical relationship mapping is as follows:

[0066] Logical AND: When all sub-events occur, the root event will occur. The structure function of the logical AND fault tree is:

[0067]

[0068] Logical OR: When at least one of the sub-events occurs, the root event will occur. The structure function of the logical OR fault tree is:

[0069]

[0070] According to the above logical relationship mapping, the troubleshooting logic tree is constructed by decomposing the root node from top to bottom and layer by layer. Figure 4 FIG. 1 is a schematic diagram of a troubleshooting logic tree constructed according to one or more embodiments of the present application. Figure 4 The structural functions of the sub-events and root events in the troubleshooting logic tree shown can be expressed as:

[0071] T=A1+A2+A3

[0072] =(C1+C2)+C n +…+C N

[0073] =(D1D2+D3D4)+D m +…+D M

[0074] =E1+E1+E k +…+E K

[0075] =F1F2+F1F2+F p +…+F P

[0076] In step S130, a qualitative analysis is performed on the troubleshooting logic tree to identify potential fault points that cause the fault.

[0077] Generally, the troubleshooting logic tree of large application systems is relatively complex, and the constructed structure function consists of multiple logical operations. The fault tree analysis method can be used to perform mathematical processing on the structure function to simplify the operation, which can reduce the complexity of fault analysis on the one hand, and provide a theoretical basis for operation and maintenance personnel to carry out qualitative fault analysis on the other hand.

[0078] Exemplarily, the Fussell algorithm can be used to simplify the troubleshooting logic tree, that is, starting from the root event of the troubleshooting logic tree, the previous layer of events are replaced with the next layer of events in order from top to bottom according to the logical relationship, until all logic gates are replaced with basic events, thereby obtaining all the minimum cut sets of the troubleshooting logic data (that is, the minimum set of basic events that cause system failures, and a minimum cut set represents a possibility of system failure). It should be noted that each minimum cut set obtained actually corresponds to a minimum troubleshooting subtree, and the root node and / or child node in each minimum troubleshooting subtree can be regarded as a potential fault point that causes the fault.

[0079] For example, Figure 4 The structure function of the troubleshooting logic tree is simplified to obtain the following minimum troubleshooting subtree structure function:

[0080] A1=C1+C2=D1D2+D3D4=E1=F1F2

[0081] Figure 5 FIG. 4 shows a schematic diagram of the minimum obstacle elimination subtree. Figure 5 As shown, the minimum troubleshooting subtree includes a root node and multiple child nodes, which can be regarded as potential failure points that cause failures.

[0082] In step S140, the troubleshooting logic tree is quantitatively analyzed to calculate the failure probability and impact degree of each potential failure point.

[0083] This step aims to analyze the probability of fault occurrence and the importance of various sub-events based on the troubleshooting logic tree, so as to quantify the impact of sub-events on the system root event and further confirm the fault propagation link and root cause. In one or more embodiments according to the present application, the quantitative analysis mainly involves key importance, structural importance and probability importance.

[0084] As described above, after performing a qualitative analysis of the fault, multiple minimum troubleshooting subtrees are obtained. The priority of the abnormality degree of each minimum troubleshooting subtree can be further determined with the help of the standardized abnormal alarm information. For example, the critical importance of the subevent is determined based on the importance of the subevent in the minimum troubleshooting subtree, the level of the configuration component associated with the subevent, and the timestamp of the subevent. i The key importance of K (i) The calculation is as follows:

[0085] I K (i) = St(a(X i ),l(X i ),t(X i ))

[0086] Among them, a(X i ) represents sub-event X i The importance of l(X i ) is the sub-event X i The level at which the associated configuration component is located, t(X i ) is the sub-event X i The timestamp generated, function St() represents the time according to a(X i )、l(X i ), t(X i ) for sub-event X i Sort and rate the importance of.

[0087] For example, the structural importance of a sub-event can be determined based on the ratio of the number of combinations of the sub-event state and the root event state to the total number of system states without considering the probability of occurrence of the sub-event. Specifically, the structural importance of a sub-event is determined by observing the results of the fault tree without considering its probability of occurrence. Since the state of a sub-event can be 1 (fault) or 0 (non-fault), when a sub-event X i When in a certain state, the number of combinations of the states of the remaining N-1 sub-events is 2 N-1 . Further, sub-event X i The structural importance of I τ (i) can be calculated as follows:

[0088]

[0089] Among them, X i _1 indicates sub-event X i The state is 1, ∑T(1,X i _1) indicates sub-event X i The number of combinations with the same status as the root event, X i _0Sub-event X i The state is 0, ∑T(1,X i _0) indicates sub-event X i The state of the root event is 0 and the state of the root event is 1, where N is the total number of sub-events.

[0090] For example, the probability importance of a sub-event may be determined based on the degree to which the probability of a root event is changed due to the change in the probability of a sub-event. i The probability importance I ρ (i) The calculation is as follows:

[0091]

[0092] Among them, P(T) is the probability of the root event occurring; P(X i ) is the sub-event X i The probability of occurrence. In general, assuming that the sub-events are independent and mutually exclusive, the following formula holds:

[0093] P(T)=P(X i )P(T,X i _1)+(1-P(X i ))P(T,X i _0)

[0094] Among them, P(T,X i _1) indicates sub-event X i The probability of the root event occurring when it occurs; P(T,Xi _0) indicates sub-event X i The probability of the root event occurring when X does not occur. Therefore, the sub-event X i The probability importance I ρ (i) can be calculated as:

[0095] I ρ (i) = P(T,X i _1)-P(T,X i _0)

[0096] Furthermore, the comprehensive importance of the sub-event can be determined by calculating the weighted sum of the key importance, structural importance, and probability importance. i The comprehensive importance of I ε (i) The calculation is as follows:

[0097] I ε (i) = kI κ (i)+tI τ (i)+pI ρ (i)

[0098] Among them, the weighted coefficients k, t, and p are obtained by learning from historical fault cases and will be iteratively optimized in actual troubleshooting.

[0099] In step S150, the root cause fault point is determined according to the results of the qualitative analysis and the quantitative analysis, and the fault propagation link of the root cause fault point is obtained.

[0100] Exemplarily, one or more sub-events with the highest comprehensive importance may be determined as the root cause fault point. After the root cause fault point is determined, a minimum troubleshooting subtree including one or more sub-events corresponding to the root cause fault point is obtained. The minimum troubleshooting subtree shows the propagation link from the sub-event to the root event, that is, the fault propagation link finally determined.

[0101] According to one or more embodiments of the present application, the fault root cause location method 10 constructs a troubleshooting logic tree and combines qualitative and quantitative analysis to achieve rapid and accurate location of the root cause of the application system fault, while obtaining the fault propagation link, effectively improving the efficiency of fault detection and resolution. Specifically, method 10 obtains the architecture topology diagram and abnormal alarm information of the application system, so as to fully and accurately understand the relationship and hierarchical structure between system components and the specific manifestation of the fault, laying a solid foundation for subsequent analysis; then, by constructing a troubleshooting logic tree, it not only intuitively shows the possible propagation path of the fault, but also utilizes the logical relationship between components, greatly improving the accuracy and efficiency of fault analysis; through qualitative analysis, method 10 can quickly narrow the scope of fault location, identify potential fault points, and reduce the time cost of blind troubleshooting; through further quantitative analysis, the failure probability and impact of each potential fault point are determined, providing a scientific basis for determining the final root cause fault point; finally, according to the comprehensive analysis results, method 10 can accurately lock the root cause fault point and clearly display its fault propagation link, which helps operation and maintenance personnel to quickly locate the problem and formulate a repair plan, thereby effectively shortening the fault recovery time and improving the stability and availability of the system.

[0102] In addition, in the fault root cause location method 10 according to one or more embodiments of the present application, a directed troubleshooting topology diagram of the application system level is constructed based on the CMDB configuration components and association relationships: based on the hierarchical management of the CMDB architecture, its configuration components are "nodes", and the association relationships are directed tree "contexts", and the full-stack "nodes" are subject to key indicators, abnormal monitoring and alarms, and log collection. Furthermore, taking the "node" abnormal events as the "crux", and the association relationships as the root cause reasoning path, the minimum troubleshooting logic tree is constructed by relying on the troubleshooting knowledge base and the fault tree analysis algorithm to qualitatively and quantitatively analyze the causes of system failures. At the same time, method 10 will accumulate troubleshooting experience, review fault cases, refine troubleshooting rules, and precipitate the troubleshooting expert knowledge base to feed back and optimize the troubleshooting process, thereby helping operation and maintenance personnel to quickly and accurately complete system-wide problem troubleshooting and improve the efficiency of fault root cause location.

[0103] Figure 6 6 is a schematic block diagram of a fault root cause location system 60 according to one or more embodiments of the present application. The fault root cause location system 60 includes a memory 610, a processor 620, and a computer program 630 stored in the memory 610 and executable on the processor 620. The execution of the computer program 630 enables the above-mentioned fault root cause location method 10 or the method 20 for generating an architecture topology diagram to be executed.

[0104] In addition, the present application can also be implemented as a computer-readable storage medium, in which a program for causing a computer to execute the following steps is stored. Figure 1 or Figure 2Here, as the computer-readable storage medium, various computer-readable storage media such as disks (e.g., magnetic disks, optical disks, etc.), cards (e.g., memory cards, optical cards, etc.), semiconductor memories (e.g., ROMs, nonvolatile memories, etc.), and tapes (e.g., magnetic tapes, cassette tapes, etc.) can be used.

[0105] The present disclosure may also be implemented as a computer program product, which includes a computer program. When the computer program is executed by a processor, the following Figure 1 or Figure 2 The steps in the method are shown in the procedure.

[0106] In the applicable situation, hardware, software or a combination of hardware and software can be used to realize the various embodiments provided by the application. Moreover, in the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be combined into a composite component comprising software, hardware and / or both. In the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be divided into subcomponents comprising software, hardware or both. In addition, in the applicable situation, it is contemplated that the software component can be implemented as a hardware component, and vice versa.

[0107] Software according to the present application (such as program code and / or data) can be stored on one or more computer-readable storage media. It is also contemplated that the software identified herein can be implemented using one or more general or special-purpose computers and / or computer systems, networked and / or otherwise. Where applicable, the order of the various steps described herein can be changed, combined into composite steps, and / or divided into sub-steps to provide the features described herein.

[0108] The embodiments and examples set forth herein are provided to best illustrate embodiments according to the present application and its specific applications, and thereby enable those skilled in the art to implement and use the present application. However, those skilled in the art will appreciate that the above description and examples are provided only for ease of illustration and example. The description set forth is not intended to cover all aspects of the present application or to limit the present application to the precise form disclosed.

Claims

1. A method for locating the root cause of a fault, characterized in that: The method comprises the following steps: Acquire an architecture topology diagram and abnormal alarm information of the application system, wherein the architecture topology diagram indicates the association relationship and hierarchical structure between the configuration components of the application system, and the abnormal alarm information includes the configuration component identifier related to the fault and the level and abnormal importance thereof; Building a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information, wherein the troubleshooting logic tree indicates potential fault propagation links between configuration components in the architecture topology diagram by using logical relationships between nodes; Performing qualitative analysis on the troubleshooting logic tree to identify potential failure points that cause the failure; Quantitatively analyzing the troubleshooting logic tree to calculate the failure probability and impact of each potential failure point; and The root cause failure point is determined according to the results of the qualitative analysis and the quantitative analysis, and the failure propagation link of the root cause failure point is obtained.

2. The method for locating the root cause of a fault according to claim 1, characterized in that: The method further comprises: The architecture topology diagram is generated based on the CMDB configuration management information of the application system and a preset traceability path.

3. The method for locating the root cause of a fault according to claim 2, characterized in that: Based on the CMDB configuration management information of the application system and the preset traceability path, generating the architecture topology diagram includes: Retrieving and acquiring the CMDB configuration model of the application system from a troubleshooting knowledge base; Using a breadth-first search algorithm to traverse and parse the CMDB configuration model to determine the hierarchical structure and nodes at each level of the application system; Supplementing the hierarchical structure and nodes at each level according to the traceability path stored in the troubleshooting knowledge base; and Based on the supplemented hierarchical structure and nodes at each level, the architecture topology diagram is generated.

4. The method for locating the root cause of a fault according to claim 2 or 3, characterized in that: The traceability path indicates how to trace back from the bottom-level child node to the sub-level root node.

5. The method for locating the root cause of a fault according to claim 1, characterized in that: The method further comprises: Monitoring the operating status of each configuration component in the architecture topology diagram; and Based on the monitoring data and abnormal indicator information for various types of faults, the abnormal alarm information is determined.

6. The method for locating the root cause of a fault according to claim 1, characterized in that: The method further comprises: According to the abnormal alarm information, annotation information is added to the nodes in the architecture topology diagram, wherein each node in the architecture topology diagram corresponds to a type of configuration component, and the annotation information includes the number of instances of the type of configuration component and the number of abnormal alarms.

7. The method for locating the root cause of a fault according to claim 1, characterized in that: The triggering conditions of the method include: manual triggering, timed polling triggering and alarm event triggering.

8. The fault root cause location method according to claim 1, characterized in that: Constructing a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information includes: Comparing the abnormal alarm information of each root node in the architecture topology diagram, and determining the root node with the highest abnormal importance as the root event; Decomposing the cause of the root event into sub-events level by level, wherein each sub-event indicates whether the running state of the subordinate node of the root event is abnormal; and According to the association relationship and hierarchical structure between the configuration components in the architecture topology diagram, the sub-events are connected through logical AND or logical OR.

9. The method for locating the root cause of a fault according to claim 1, characterized in that: Qualitative analysis of the troubleshooting logic tree to identify potential fault points that cause the fault includes: The troubleshooting logic tree is simplified by using a fault tree analysis method to identify multiple minimum cut sets that cause the fault, wherein each minimum cut set includes a minimum set of basic events that cause the fault.

10. The method for locating the root cause of a fault according to claim 9, characterized in that: Each minimum cut set corresponds to a minimum obstacle-eliminating subtree.

11. The method for locating the root cause of a fault according to claim 1, characterized in that: The quantitative analysis of the troubleshooting logic tree includes: Determining the critical importance of the sub-event based on the importance of the sub-event, the level of the configuration component associated with the sub-event, and the timestamp of the sub-event; Without considering the probability of occurrence of the sub-event, determining the structural importance of the sub-event according to the ratio of the number of combinations in which the states of the sub-event are the same as the state of the root event to the total number of system states; and The probability importance of the sub-event is determined based on the degree of change in the probability of occurrence of the root event caused by the change in the probability of occurrence of the sub-event.

12. The method for locating the root cause of a fault according to claim 11, characterized in that: Sub-event X i The key importance of K (i) The calculation is as follows: I K (i)=St(a(X i ),l(X i ),t(X i )) Among them, a(X i ) represents the sub-event X i The importance of l(X i ) is the sub-event X i The level at which the associated configuration component is located, t(X i ) is the sub-event X i The timestamp generated, function St() represents the time according to a(X i )、l(X i ), t(X i ) for sub-event X i Sort and rate the importance of.

13. The fault root cause location method according to claim 11, characterized in that: Sub-event X i The structural importance of I τ (i) The calculation is as follows: Among them, X i _1 indicates the sub-event X i The state is 1, ΣT(1,X i _1) represents the sub-event X i The number of combinations with the same state as the root event, X i _0The sub-event X i The state is 0, ΣT(1,X i _0) indicates the sub-event X i The state of the root event is 0 and the state of the root event is 1, where N is the total number of sub-events.

14. The method for locating the root cause of a fault according to claim 11, characterized in that: Sub-event X i The probability importance I ρ (i) The calculation is as follows: I ρ (i)=P(T,X i _1)-P(T,X i _0) Among them, P(T,X i _1) represents the sub-event X i The probability of the root event occurring when it occurs; P(T,X i _0) indicates the sub-event X i The probability of the root event occurring when it does not occur.

15. The method for locating the root cause of a fault according to any one of claims 11 to 14, characterized in that: The quantitative analysis of the troubleshooting logic tree also includes: A weighted sum of the key importance, the structural importance, and the probability importance is calculated to determine the comprehensive importance of the sub-event.

16. The method for locating the root cause of a fault according to claim 15, characterized in that: The root causes of failures determined based on the results of qualitative and quantitative analysis include: The one or more sub-events with the highest comprehensive importance are determined as the root cause failure point.

17. The fault root cause location method according to claim 1, characterized in that: Obtaining the fault propagation link of the root fault point includes: Obtaining a minimum troubleshooting subtree including the root cause fault point; and The fault propagation link is determined based on a minimum troubleshooting subtree.

18. The fault root cause location method according to claim 1, characterized in that: The method further comprises: Constructing a troubleshooting knowledge base, wherein the troubleshooting knowledge base includes one or more of the following: CMDB configuration management information of each application system, types of various faults, abnormal indicator information of various faults, abnormal importance of various faults, importance of various sub-events, tracing paths, and processing logic for fault location; and The troubleshooting knowledge base is updated according to historical fault root cause location results.

19. A fault root cause location system, characterized in that: The system comprises: Memory; processor; as well as A computer program stored on the memory and executable on the processor, the execution of which results in the following operations: Acquire an architecture topology diagram and abnormal alarm information of the application system, wherein the architecture topology diagram indicates the association relationship and hierarchical structure between the configuration components of the application system, and the abnormal alarm information includes the configuration component identifier related to the fault and the level and abnormal importance thereof; Building a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information, wherein the troubleshooting logic tree indicates potential fault propagation links between configuration components in the architecture topology diagram by using logical relationships between nodes; Performing qualitative analysis on the troubleshooting logic tree to identify potential failure points that cause the failure; Quantitatively analyzing the troubleshooting logic tree to calculate the failure probability and impact of each potential failure point; as well as The root cause failure point is determined according to the results of the qualitative analysis and the quantitative analysis, and the failure propagation link of the root cause failure point is obtained.

20. The fault root cause location system according to claim 19, characterized in that: The execution of the computer program also results in the following operations: The architecture topology diagram is generated based on the CMDB configuration management information of the application system and a preset traceability path.

21. The fault root cause location system according to claim 20, characterized in that: Based on the CMDB configuration management information of the application system and the preset traceability path, generating the architecture topology diagram includes: Retrieving and acquiring the CMDB configuration model of the application system from a troubleshooting knowledge base; Using a breadth-first search algorithm to traverse and parse the CMDB configuration model to determine the hierarchical structure and nodes at each level of the application system; Supplementing the hierarchical structure and nodes at each level according to the traceability path stored in the troubleshooting knowledge base; and Based on the supplemented hierarchical structure and nodes at each level, the architecture topology diagram is generated.

22. The fault root cause location system according to claim 20 or 21, characterized in that: The traceability path indicates how to trace back from the bottom-level child node to the sub-level root node.

23. The fault root cause location system according to claim 19, characterized in that: The execution of the computer program also results in the following operations: Monitoring the operating status of each configuration component in the architecture topology diagram; and Based on the monitoring data and abnormal indicator information for various types of faults, the abnormal alarm information is determined.

24. The fault root cause location system according to claim 19, characterized in that: The execution of the computer program also results in the following operations: According to the abnormal alarm information, annotation information is added to the nodes in the architecture topology diagram, wherein each node in the architecture topology diagram corresponds to a type of configuration component, and the annotation information includes the number of instances of the type of configuration component and the number of abnormal alarms.

25. The fault root cause location system according to claim 19, characterized in that: The fault root cause location system is enabled to obtain the architecture topology diagram and abnormal alarm information of the application system based on one of the following triggering conditions: manual triggering, timed polling triggering, and alarm event triggering.

26. The fault root cause location system according to claim 19, characterized in that: Constructing a troubleshooting logic tree based on the architecture topology diagram and the abnormal alarm information includes: Comparing the abnormal alarm information of each root node in the architecture topology diagram, and determining the root node with the highest abnormal importance as the root event; Decomposing the cause of the root event into sub-events level by level, wherein each sub-event indicates whether the running state of the subordinate node of the root event is abnormal; and According to the association relationship and hierarchical structure between the configuration components in the architecture topology diagram, the sub-events are connected through logical AND or logical OR.

27. The fault root cause location system according to claim 19, characterized in that: Qualitative analysis of the troubleshooting logic tree to identify potential fault points that cause the fault includes: The troubleshooting logic tree is simplified by using a fault tree analysis method to identify multiple minimum cut sets that cause the fault, wherein each minimum cut set includes a minimum set of basic events that cause the fault.

28. The fault root cause location system according to claim 19, characterized in that: Each minimum cut set corresponds to a minimum obstacle-eliminating subtree.

29. The fault root cause location system according to claim 19, characterized in that: The quantitative analysis of the troubleshooting logic tree includes: Determining the critical importance of the sub-event based on the importance of the sub-event, the level of the configuration component associated with the sub-event, and the timestamp of the sub-event; Without considering the probability of occurrence of the sub-event, determining the structural importance of the sub-event according to the ratio of the number of combinations in which the states of the sub-event are the same as the state of the root event to the total number of system states; and The probability importance of the sub-event is determined based on the degree of change in the probability of occurrence of the root event caused by the change in the probability of occurrence of the sub-event.

30. The fault root cause location system according to claim 29, characterized in that: Sub-event X i The key importance of K (i) The calculation is as follows: I K (i)=St(a(X i ),l(X i ),t(X i )) Among them, a(X i ) represents the sub-event X i The importance of l(X i ) is the sub-event X i The level at which the associated configuration component is located, t(X i ) is the sub-event X i The timestamp generated, function St() represents the time according to a(X i )、l(X i ), t(X i ) for sub-event X i Sort and rate the importance of.

31. The fault root cause location system according to claim 29, characterized in that: Sub-event X i The structural importance of I τ (i) The calculation is as follows: Among them, X i _1 indicates the sub-event X i The state is 1, ΣT(1,X i _1) represents the sub-event X i The number of combinations with the same state as the root event, X i _0The sub-event X i The state is 0, ΣT(1,X i _0) indicates the sub-event X i The state of the root event is 0 and the state of the root event is 1, where N is the total number of sub-events.

32. The fault root cause location system according to claim 29, characterized in that: Sub-event X i The probability importance I ρ (i) The calculation is as follows: I ρ (i)=P(T,X i _1)-P(T,X i _0) Among them, P(T,X i _1) represents the sub-event X i The probability of the root event occurring when it occurs; P(T,X i _0) indicates the sub-event X i The probability of the root event occurring when it does not occur.

33. The fault root cause location system according to any one of claims 29 to 32, characterized in that: The quantitative analysis of the troubleshooting logic tree also includes: A weighted sum of the key importance, the structural importance, and the probability importance is calculated to determine the comprehensive importance of the sub-event.

34. The fault root cause location system according to claim 33, characterized in that: The root causes of failures determined based on the results of qualitative and quantitative analysis include: The one or more sub-events with the highest comprehensive importance are determined as the root cause failure point.

35. The fault root cause location system according to claim 19, characterized in that: Obtaining the fault propagation link of the root fault point includes: Obtaining a minimum troubleshooting subtree including the root cause fault point; and The fault propagation link is determined based on a minimum troubleshooting subtree.

36. The fault root cause location system according to claim 19, characterized in that: The execution of the computer program also results in the following operations: Constructing a troubleshooting knowledge base, wherein the troubleshooting knowledge base includes one or more of the following items: CMDB configuration management information of each application system, types of various faults, abnormal indicator information of various faults, abnormal importance of various faults, importance of various sub-events, tracing paths, and processing logic of fault location; as well as The troubleshooting knowledge base is updated according to historical fault root cause location results.

37. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes instructions, which, when executed, execute the fault root cause locating method according to any one of claims 1-18.

38. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the fault root cause locating method according to any one of claims 1 to 18.

Citation Information

Cited By

  • Defect tracing method and system in fabric production process

    CN120912230A

  • Quality problem two-dimensional positioning method and device based on industrial internet

    CN121707432A

  • Fault root cause locating method and system, and computer storage medium and program product

    WO2026086452A1