Full-link fault root cause analysis method and device
By acquiring abnormal events from operational data of multiple departmental platforms in the banking business system, matching them with fault factors in the knowledge base to form a set of related events, and performing multi-dimensional contribution quantification and root cause analysis, the problem of low accuracy in fault root cause analysis caused by information silos is solved, and highly accurate fault root cause analysis is achieved.
Patent Information
- Application Number
- CN202511110177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-14
AI Technical Summary
The independent monitoring platforms and data storage systems of various departments in the banking business system form information silos, resulting in poor data integration and cross-layer analysis capabilities, making it difficult to conduct full-link monitoring and fault analysis, and resulting in low accuracy of root cause analysis.
By acquiring abnormal events from operational data of multiple departmental platforms and matching them with fault factors in a pre-set knowledge base, a set of related events is formed. The correlation graph is then analyzed through multi-dimensional contribution quantification and root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
It has achieved the integration of data from multiple departments, overcome the problem of data silos, provided scientific and accurate root cause analysis of failures, and improved the accuracy of root cause analysis of failures.
Smart Images

Figure CN120950290A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and apparatus for end-to-end fault root cause analysis. Background Technology
[0002] As banking operations continue to develop, their data scale and business systems are growing larger and larger. Monitoring systems used to monitor the operational status of these business systems are also facing many challenges. The reliable operation of monitoring systems is the foundation for the stable operation of banking business systems.
[0003] Currently, banking operations often involve collaboration among multiple departments (such as network, systems, and applications). However, each department often has its own independent monitoring platform and data storage system, forming information silos. This results in poor data integration and cross-layer analysis capabilities, making it difficult to conduct end-to-end monitoring and fault analysis from a global perspective, which in turn leads to low accuracy in analyzing the root causes of faults. Summary of the Invention
[0004] In view of this, this application provides a method and apparatus for end-to-end fault root cause analysis, which aims to achieve end-to-end monitoring and fault analysis, and solve the problem of low accuracy in analyzing fault root causes.
[0005] Firstly, this application provides a method for end-to-end fault root cause analysis, the method comprising:
[0006] After obtaining abnormal events from the operational data of multiple departmental platforms, they are matched with fault factors in a preset knowledge base to determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, thus forming a set of associated events.
[0007] Based on the association graph in the knowledge base and the static factors corresponding to the fault factors, the multi-dimensional contribution quantification of each fault factor in the set of associated events is performed to obtain the fault contribution degree of the fault factors.
[0008] The correlation graph is updated based on the abnormal events, the set of related events, and the fault contribution of each fault factor. The correlation graph is then analyzed by the root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
[0009] In one possible implementation, after obtaining abnormal events from the operational data of multiple departmental platforms, they are matched with fault factors in a preset knowledge base to determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, forming a set of associated events, including:
[0010] The running data is sent to processing pipelines at different levels to obtain a subset of associated events output by multiple processing pipelines. The processing pipelines are used to perform anomaly analysis on the running data from the level to which they belong, determine the abnormal events and their matching fault factors, and instantiate the fault factors to form a subset of associated events.
[0011] Construct a set of associated events from a subset of associated events containing unified target information.
[0012] In one possible implementation, sending the runtime data to processing pipelines at different levels to obtain a subset of associated events output from at least one level of the processing pipeline includes:
[0013] Each layer of the processing pipeline performs anomaly analysis on the operational data at that layer. If anomalies are found in the operational data at that layer, an anomaly event is identified. The anomaly event includes an identifier code and source data, where the source data is information in the operational data that is associated with the anomaly event.
[0014] Based on the abnormal event, it is matched with the fault factors in the knowledge base to determine the fault factor corresponding to the abnormal event;
[0015] Based on the source data of the fault factor and the abnormal event, the instantiation information of the fault factor is determined and a subset of associated events is constructed. The elements of the subset of associated events include the fault factor, the entity information corresponding to the fault factor, and the time information.
[0016] In one possible implementation, before sending the runtime data to different levels of the processing pipeline, the method further includes:
[0017] For network change events and / or suppression status alarms, the raw data obtained from the multiple departmental platforms are marked to obtain operational data;
[0018] The operational data is then transmitted to processing pipelines at different levels.
[0019] In one possible implementation, the step of performing multi-dimensional contribution quantification on each fault factor in the set of associated events based on the association graph in the knowledge base and the static factors corresponding to the fault factors to obtain the fault contribution degree of the fault factors includes:
[0020] From a pre-defined knowledge base, determine the static factors corresponding to each fault factor included in the set of associated events;
[0021] Based on the source data of the abnormal events corresponding to the fault factors, a dynamic factor reflecting the degree of abnormality of the fault factors is determined, wherein the source data is the data in the operation data associated with the abnormal events;
[0022] Assess the correlation factors between the aforementioned failure factors and overall symptoms;
[0023] Based on the association graph in the knowledge base, and the position and associated nodes of the entity information corresponding to the fault factor in the association graph, the topological influence factor of the fault factor is determined.
[0024] The fault contribution of the fault factor is determined based on the static factor, dynamic factor, correlation factor, and topological influence factor.
[0025] In one possible implementation, determining the dynamic factor reflecting the degree of abnormality of the fault factor based on the source data of the abnormal event corresponding to the fault factor includes:
[0026] Obtain multiple dynamic factor parameters from the source data; the multiple factor parameters include one or more of the following: a first parameter reflecting the severity of the abnormal event, a frequency parameter of the occurrence of the abnormal event, a duration parameter of the abnormal event, a recentity parameter of the abnormal event, and a deviation parameter of the abnormal event from the baseline;
[0027] The dynamic factors are obtained by weighted summation of the multiple dynamic factor parameters and solving the problem using a normalization function.
[0028] In one possible implementation, after determining the fault contribution of the fault factor based on the static factor, dynamic factor, correlation factor, and topological influence factor, the method further includes:
[0029] The first table is displayed on the display interface. The first table includes the fault factors corresponding to the operation data of the multiple department platforms and the fault contribution degree corresponding to each fault factor. The fault factors in the first table are sorted and displayed according to their corresponding fault contribution degrees.
[0030] In one possible implementation, updating the correlation graph based on the abnormal event, the set of related events, and the fault contribution of each fault factor, and then parsing the correlation graph through a root cause analysis engine to determine the fault evidence chain and the root cause of the fault, includes:
[0031] Based on the fault contribution of the fault factors, one of the fault factors corresponding to the operational data of the multiple departmental platforms is determined as the target fault factor.
[0032] The source data of the abnormal event, the set of related events, and the fault contribution of the fault factors included in the set of related events are loaded onto the corresponding nodes of the correlation graph to obtain the loaded correlation graph.
[0033] Using the root cause analysis engine, the correlation graph is analyzed for contextual correlation to obtain the fault evidence chain of the root cause of the target fault factor.
[0034] Based on the correlation information between different levels in the source data of the abnormal event, the correctness of the fault evidence chain and the correctness of the correlation graph are verified. If the fault evidence chain and the correlation graph are correct, the target fault factor is determined to be the root cause of the fault.
[0035] In one possible implementation, the method further includes:
[0036] The display interface shows one or more of the following: the root cause of the failure, the chain of evidence for the failure, the evidence parameters supporting the chain of evidence for the failure, the topological impact range of the root cause of the failure, and the remedial measures based on the root cause of the failure.
[0037] Secondly, this application provides a full-link fault root cause analysis device, the device comprising:
[0038] The multi-level analysis module is used to obtain abnormal events from the operation data of multiple department platforms, match them with fault factors in a preset knowledge base, determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, and form a set of related events.
[0039] The quantification module is used to perform multi-dimensional contribution quantification on each fault factor in the set of associated events based on the association graph in the knowledge base and the static factors corresponding to the fault factors, so as to obtain the fault contribution degree of the fault factors.
[0040] The root cause analysis engine module is used to update the correlation graph based on the abnormal events, the set of related events, and the fault contribution of each fault factor. The correlation graph is analyzed by the root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
[0041] This application provides a method and apparatus for end-to-end fault root cause analysis. First, abnormal events are obtained from the operational data of multiple departmental platforms and matched with fault factors in a pre-defined knowledge base to determine the fault factors corresponding to each abnormal event and their instantiation information, forming a set of associated events. Then, based on the association graph in the knowledge base and the static factors corresponding to the fault factors, the contribution of each fault factor in the set of associated events is quantified in multiple dimensions to obtain the fault contribution degree of each fault factor. Finally, the association graph is updated based on the abnormal events, the set of associated events, and the fault contribution degrees of each fault factor. The association graph is then parsed by a root cause analysis engine to determine the fault evidence chain and the fault root cause. In this way, abnormal events from the operational data of multiple departmental platforms are integrated, and fault factors can be matched to each abnormal event to determine the fault type corresponding to the abnormal event. Furthermore, the fault factor is instantiated based on the data of the abnormal event, such as the time and device location of the fault factor, to construct a set of associated events integrating information from multiple departmental platforms. Further, each fault factor included in the set of associated events is scored, i.e., its fault contribution degree. Furthermore, by updating the correlation graph with information such as the set of related events and the contribution of failures, and by parsing the contextual information through the root cause analysis engine, the chain of evidence for the failure is obtained to determine the root cause. In this way, data from multiple departments is integrated, overcoming departmental barriers and data silos, and enabling quantitative evaluation of each failure factor. This provides a scientific and accurate quantitative basis for subsequent root cause analysis, improving the accuracy of the failure root cause analysis. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A flowchart illustrating a full-link fault root cause analysis method provided in this application embodiment;
[0044] Figure 2 A flowchart illustrating a multi-dimensional contribution quantification method provided in this application embodiment;
[0045] Figure 3 A schematic diagram of a root cause analysis process provided for an embodiment of this application;
[0046] Figure 4 This is a schematic diagram of a full-link fault root cause analysis device provided in an embodiment of this application. Detailed Implementation
[0047] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0048] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0049] Unless otherwise stated, the term "multiple" means two or more. In embodiments of this disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B. The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.
[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0051] See Figure 1 , Figure 1 A flowchart illustrating a full-link fault root cause analysis method provided in this application embodiment. The full-link fault root cause analysis method includes:
[0052] S101. After obtaining abnormal events from the operational data of multiple departmental platforms, they are matched with fault factors in a preset knowledge base to determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, forming a set of associated events.
[0053] Optionally, the aforementioned multiple departmental platforms may include network, system, application, and security departments. Each department will configure its own monitoring platform and data storage system. For example, the network department mainly collects and manages network device data (such as network device performance metrics, logs, traffic, and data), the system department manages server and operating system data (such as operating system metrics and process information), and the application department manages application performance and log data. In this way, the operational data of each department's monitoring platform can be collected to obtain the operational data of multiple departmental platforms.
[0054] Optionally, the aforementioned abnormal events can refer to scattered evidence and corresponding original, objective phenomena analyzed from operational data. The aforementioned fault factors are diagnoses derived from the comprehensive evidence of abnormal events; they are labels set for the operational data corresponding to the abnormal events. These fault factors are predefined in a knowledge base and are a collection of expert experience and system knowledge. For example, these fault factors could be link packet loss, network congestion, etc. The instantiation information of the aforementioned fault factors refers to the fault factors substituted into specific transactions. For example, an abnormal event is mapped to a predefined fault factor. For instance, if multiple abnormal events (multiple phenomena, etc.) are mapped to the fault factor "link packet loss," then the instantiation specifically becomes "a link packet loss occurred on a certain port of a certain device on a certain link" (e.g., link packet loss occurred on port 12 of the switch on link 1), and this is matched and brought into the instance topology.
[0055] For example, taking network congestion as an example, the identified abnormal events are: Event 1: Abnormal indicators (phenomenon), network device ports exceeding thresholds, web server experiencing high CPU usage (evidence); Event 2: Performance degradation, application performance indicators showing a decrease in success rate; Event 3: Device logs, both network device and application system logs show packet loss; Event 4: Traffic analysis, traffic monitoring shows a surge in traffic; Event 5: Gray-scale change, a network change operation was performed at a certain moment; Event 6: User operation, user A performed a file transfer at a certain moment. The subset of associated events after these abnormal events are specifically instantiated can be {Factor type: Network congestion, Occurring entity: A device's port exceeds the threshold, Time: 2025-11-11 11:11:11}.
[0056] S102. Based on the association graph in the knowledge base and the static factors corresponding to the fault factors, perform multi-dimensional contribution quantification on each fault factor in the set of associated events to obtain the fault contribution degree of the fault factors.
[0057] The aforementioned correlation graph is a dynamic fault event correlation graph constructed by combining events, fault factors, and their relationships.
[0058] The aforementioned fault contribution can be used to quantify the fault factors determined in step S101 using information such as the association graph. The multi-dimensional contribution can include the static factors preset in the knowledge base, the dynamic factors determined by combining the source data of the abnormal events corresponding to the fault factors in step S101, the correlation factors between the fault factors and the overall symptoms (the operational faults exhibited by the overall system operation), and the topological influence factors determined by the association nodes in the association graph.
[0059] Thus, this quantification process comprehensively considers multiple dimensions such as the static importance of factors, dynamic observational impact, correlation with symptoms, and position in the system topology, calculating a quantitative score for each factor, which can be between 0 and 1. Feeding the fault contribution obtained from the multi-dimensional analysis back to the corresponding nodes in the correlation graph helps improve the accuracy of the subsequent root cause analysis in complex environments, especially for faults caused by complex event linkages and contextual relationships that are difficult to identify. By introducing multi-dimensional contribution quantification, reliable quantitative data is provided for subsequent analysis.
[0060] S103. Update the correlation graph based on the abnormal events, the set of related events, and the fault contribution of each fault factor. Analyze the correlation graph through the root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
[0061] Contextual association analysis based on the updated graph can identify the linkage patterns of events, and root cause analysis can be performed on the factors or factor sequences with the highest contribution.
[0062] As can be seen from steps S101-S103 above, the fusion of data from multiple departments overcomes departmental barriers and data silos, thereby enabling multi-dimensional quantitative evaluation of various failure factors. This provides a scientific and accurate quantitative basis for subsequent root cause analysis, improves the accuracy of root cause analysis, and effectively addresses the challenges of data integration and analysis brought about by multiple departments and platforms.
[0063] In one possible implementation, prior to step S101, the method further includes: acquiring operational data from multiple departmental platforms, specifically including:
[0064] First, it connects to data sources from multiple departments, such as network management platforms, system monitoring tools, application performance monitoring systems, security information and event management systems, log management platforms, and other heterogeneous platforms to obtain operational data from different departmental platforms.
[0065] Understandably, connectors that interface with various platforms need to support the required data acquisition protocols (such as SNMP, Syslog, Agent push, API calls, direct database connections, etc.) in order to obtain the necessary raw data from platforms of different departments and different technology stacks.
[0066] For example, the network management platform of the network department can obtain raw data of network devices such as SNMP trap / polling data (interface status, traffic, error count, CPU / memory utilization), Syslog logs, NetFlow / sFlow traffic records, and configuration change logs.
[0067] The system monitoring tools from the system department can obtain raw server data such as operating system performance metrics (CPU, memory, disk I / O, network I / O), process status, system logs, and security logs.
[0068] The application performance monitoring system of the application department can obtain raw data of application services such as application logs, transaction tracing data, API call performance, error logs, and middleware status.
[0069] Raw data from security devices, such as firewall logs, intrusion detection / prevention system (IDS / IPS) alerts, and security event logs, can be obtained from the security information and event management system of the security department.
[0070] In addition, raw data of virtual machines, such as virtual machine status, resource utilization, and migration events, can be obtained from the virtualization platform.
[0071] Then, the collected data undergoes preliminary format conversion, cleaning, and standardization for subsequent processing in step S101, thus overcoming the data silo problem caused by departments and platforms.
[0072] Furthermore, in one possible implementation, prior to step S101, the method further includes:
[0073] When collecting raw data from various departments, ongoing or planned network change events are detected in real time. Based on change event information (such as the devices involved and the time window), the relevant raw data is specially marked. Simultaneously, if a suppressed alarm is detected from a device, it is also marked. Thus, based on network change events and / or suppressed alarms, the raw data obtained from the multiple department platforms is marked; the marked raw data is used as the operational data for analysis in step S101, and the operational data is copied and distributed to different levels of processing pipelines.
[0074] The aforementioned network change events refer to any planned or unplanned modifications, maintenance, upgrades, configuration adjustments, or equipment replacements performed on infrastructure such as financial data center networks. These operations may affect the normal operation of the network and are typically accompanied by corresponding change management processes.
[0075] The aforementioned suppressed alarm refers to the temporary masking or non-displaying of alarm information from specific devices, links, or event types by the monitoring system in specific scenarios (e.g., during network change events) to avoid generating a large number of known or acceptable alarms due to anticipated operations. Although suppressed alarms are not directly presented to operations and maintenance personnel, their raw data or related information may still be collected and processed internally by the system.
[0076] Thus, by implementing multi-perspective hierarchical parallel processing of multi-source operational data through step S101, it is possible to specifically address the monitoring blind spot problem caused by suppressed state alarms during network change events, ensuring that even under alarm masking, potential anomalies can still be captured through parallel analysis of the original data at multiple levels, such as network and application.
[0077] In the embodiments of this application, the above Figure 1 There are several possible implementations of step S101, which will be described below. It should be noted that the implementations given below are merely illustrative examples and do not represent all implementations of the embodiments of this application.
[0078] In one possible implementation, step S101 above may include:
[0079] First, the running data is sent to processing pipelines at different levels to obtain a subset of associated events output by multiple processing pipelines. The processing pipelines are used to perform anomaly analysis on the running data from the level to which they belong, determine the abnormal events and their matching fault factors, and instantiate the fault factors to form a subset of associated events.
[0080] The aforementioned different levels can be divided as needed to analyze operational data at various levels of interest. For example, these different levels can be divided according to different departments, different failure factors, or other dimensions. For instance, departmental divisions could include network and system / application levels. This allows for classification, parsing, and parallel processing from different perspectives (e.g., network and system / application levels), outputting subsets of related events for each level. This ensures that different levels extract complementary information from different perspectives, rather than simply repeating the same processing. For example, the network layer focuses on link status, while the application layer focuses on service availability.
[0081] Then, the subset of associated events containing unified target information is constructed into a set of associated events.
[0082] For example, the unified target information mentioned above can be an event ID, a trace ID (in the call chain of a request, the request will always carry the trace ID to the downstream service), a timestamp, or an entity association based on system topology and dependencies, forming a set of associated time information containing multiple layers, and participating in the root cause analysis in the subsequent step S103.
[0083] In one possible implementation, the above-mentioned sending of the runtime data to processing pipelines at different levels to obtain a subset of associated events output by at least one level of processing pipeline includes:
[0084] First, the operational data is analyzed for anomalies at each level through processing pipelines at each level. If anomalies are found in the operational data at each level, an anomaly event is identified. The anomaly event includes an identifier code and source data, where the source data is information in the operational data that is associated with the anomaly event.
[0085] The aforementioned identifier can be the event ID, trace ID, or timestamp, etc.
[0086] For example, when analyzing operational data at the network level, the focus can be on in-depth analysis of operational data and regular network data at the network protocol and device status levels to identify corresponding anomalous events. For instance, even if traffic alarms for a particular interface are suppressed, this pipeline will still analyze its raw traffic data, error count, packet loss rate, etc., to detect any abnormal patterns, such as sudden drops / surges in traffic or an abnormal increase in the proportion of error packets. It will also analyze configuration change logs and compare them with the change plan to identify potential configuration errors. When analyzing operational data at the system / application level, the focus is on in-depth analysis of operational data and regular system / application data at the operating system and application service levels to identify corresponding anomalous events. For example, analyzing the system logs of servers related to the changed equipment can help identify signs of abnormal process startup / exit, resource exhaustion, etc.; analyzing application logs and transaction tracing data can detect application-level anomalies such as service response latency, increased error rates, and connection interruptions.
[0087] Then, based on the abnormal event, it is matched with the fault factors in the knowledge base to determine the fault factor corresponding to the abnormal event.
[0088] Finally, based on the source data of the fault factor and the abnormal event, the instantiation information of the fault factor is determined and a subset of associated events is constructed. The elements of the subset of associated events include the fault factor, the entity information corresponding to the fault factor, and the time information.
[0089] For example, the subset of associated events after these abnormal events are specifically instantiated can be {factor type: network congestion, occurrence entity: a certain device has a port exceeding the threshold, time: 2025-11-11 11:11:11}.
[0090] In the embodiments of this application, the above Figure 1 There are several possible implementations of step S102, which will be described below. It should be noted that the implementations given below are merely illustrative examples and do not represent all implementations of the embodiments of this application.
[0091] See Figure 2 The diagram illustrates a multi-dimensional contribution quantification method. Based on the association graph in the knowledge base and the static factors corresponding to the fault factors, multi-dimensional contribution quantification is performed on each fault factor in the set of associated events to obtain the fault contribution degree of the fault factor, including:
[0092] S201. Determine the static factors corresponding to each fault factor included in the set of associated events from the preset knowledge base.
[0093] The knowledge base pre-defines multiple potential fault factors, such as "high device CPU usage," "packet loss," "service response timeout," and "configuration error." Each fault factor is assigned a static weight Wi∈[0,1], reflecting its inherent importance and potential impact range within the data center environment. For example, the weight of fault factors corresponding to core components is higher than that of fault factors corresponding to edge components.
[0094] S202. Based on the source data of the abnormal events corresponding to the fault factor, determine a dynamic factor that reflects the degree of abnormality of the fault factor, wherein the source data is the data in the operating data that is associated with the abnormal event.
[0095] In one example, the specific calculation method of the dynamic factor can be as follows: obtain multiple dynamic factor parameters from the source data and perform over-normalization processing; the multiple factor parameters include one or more of the following: a first parameter Si' reflecting the severity of the abnormal event, a frequency parameter Fi' of the abnormal event, a duration parameter Di' of the abnormal event, a recentity parameter Ri' of the abnormal event, and a deviation parameter Bi' of the abnormal event from the baseline; the multiple dynamic factor parameters are weighted and summed, and solved by a normalization function to obtain the dynamic factor. Specifically, the dynamic factor Ii = Normalize(α'×Si'+β'×Fi'+γ'×Di'+δ'×Ri'+ε'×Bi'), where α', β', γ', δ', and ε' are configurable weight coefficients that reflect the relative importance of each parameter in evaluating the dynamic factor.
[0096] S203. Evaluate the correlation factors between the failure factors and the overall symptoms.
[0097] The overall symptom refers to the operational failure exhibited by the system as a whole; for example, if an overall symptom occurs in a banking system, one or more departments within the system may experience corresponding abnormal events. Therefore, there may be a certain correlation between the failure factors identified based on the abnormal events and the overall symptom. In one example, the correlation factor Ci∈[0,1] between the abnormal state of the failure factor and the currently observed overall symptom is used to evaluate the correlation. Ci reflects the probability or logical correlation between the co-occurrence of the factor abnormality and the overall symptom. It can be calculated using methods such as real-time co-occurrence frequency, matching degree with historical failure patterns, and judgment confidence based on expert rules.
[0098] S204. Based on the association graph in the knowledge base and the position and associated nodes of the entity information corresponding to the fault factor in the association graph, determine the topological influence factor of the fault factor.
[0099] In one instance, based on the relational graph in the database (including physical connections, logical links, service calls, application dependencies, etc.), the topological impact factor Ti∈[0,1] of the failure factor is calculated. This reflects the potential impact path and scope of the failure factor on affected services or users within the data system. The calculation can be performed based on data such as the node position of the failure factor in the graph (e.g., critical path), the number or weight of its downstream affected nodes, and the topological distance to the node containing the overall symptom.
[0100] S205. Determine the fault contribution of the fault factor based on the static factor, dynamic factor, correlation factor, and topological influence factor.
[0101] Optionally, the fault contribution of the fault factor can be calculated by combining the static factor Wi, the dynamic factor Ii, the correlation factor Ci, and the topological influence factor Ti, which can be expressed as MCPi = Sigmoid(ωW×Wi+ωI×Ii+ωC×Ci+ωT×Ti); or, after weighted averaging and normalization, MCPi = (ωW×Wi+ωI×Ii+ωC×Ci+ωT×Ti) / (ωW+ωI+ωC+ωT), where ωW, ωI, ωC, and ωT are configurable weight coefficients used to adjust the relative importance of each factor in the final fault contribution calculation. The above Sigmoid function ensures that MCPi ultimately remains between 0 and 1.
[0102] Optionally, a first table is displayed on the display interface. The first table includes fault factors corresponding to the operation data of the multiple department platforms and the fault contribution degree corresponding to each fault factor. The fault factors in the first table are sorted and displayed according to their corresponding fault contribution degrees.
[0103] Thus, a list of potential failure factors (Table 1) is output, with each failure factor accompanied by its calculated MCPi score (failure contribution, ranging from 0 to 1), sorted from highest to lowest score for easy viewing by the user. Subsequently, this failure contribution is also attached as a failure factor attribute to the corresponding node in the anomaly event association graph.
[0104] In the embodiments of this application, the above Figure 1 There are several possible implementations of step S103, which will be described below. It should be noted that the implementations given below are merely illustrative examples and do not represent all implementations of the embodiments of this application.
[0105] In one possible implementation, see Figure 3 The diagram shown illustrates a root cause analysis process. Step S103 may include:
[0106] S301. Based on the fault contribution of the fault factors, determine one of the fault factors corresponding to the operation data of the multiple department platforms as the target fault factor.
[0107] For example, two fault factors are identified: Fault Factor A: Network congestion, high CPU usage on the database server; and Fault Factor B: Configuration error, network configuration error on the web server. Through calculation in step S102, the fault contribution of fault factor A is determined to be 0.92, and the fault contribution of fault factor B is determined to be 0.65. Therefore, network congestion is initially selected as the target fault factor. Subsequently, the root cause analysis engine can prioritize root cause analysis targeting the network congestion problem.
[0108] S302. Load the source data of the abnormal event, the set of related events, and the fault contribution of the fault factors included in the set of related events onto the node corresponding to the correlation graph to obtain the loaded correlation graph.
[0109] The information in the aforementioned relationship graph can include: node type, edge type, and attributes. Node types can include devices (servers, switches, application hosts, etc.), services (database services, application services, middleware services, etc.), processes, network interfaces, IP addresses, users, network change events, suppression status alarms, raw monitoring events (log entries, performance metric anomalies, traffic events, etc.), and potential fault factors. Edge types (edges in the graph can represent various relationships between nodes) can include physical connections (device-device, device-interface), logical connections (IP-IP), service dependencies (application A depends on service B), process calls, event occurrence (event-device / service), event association (events associated by time, ID, or attributes), causal relationships (derived through inference), temporal relationships (event B occurs after event A), belonging to the same change (event-change event), and impact (factor-symptom, change event-device). Attributes (nodes and edges can have attributes) can include event occurrence time, severity, performance metric values, MCP scores, change type, and change status.
[0110] The core of the root cause analysis engine is a graph database used to store and manage entities, events, and their complex relationships within the system. Specifically, the root cause analysis engine generates dynamic information based on real-time acquired runtime data, such as the fault factors obtained in step S101, the source data of the abnormal events corresponding to the fault factors, the instantiation information of the fault factors, and the fault contribution of the fault factors. This information is combined with a pre-set association graph (with static topology and dependency information) in the knowledge base, and the engine updates and transforms these information into nodes and edges in the association graph in real time, thus constructing a dynamic association graph of abnormal events.
[0111] S303. Using the root cause analysis engine, perform contextual root cause analysis on the correlation graph for the target failure factor to obtain the failure evidence chain of the root cause of the target failure factor.
[0112] When the root cause analysis engine traverses and analyzes the correlation graph to determine the fault evidence chain, it can not only focus on nodes with high fault contribution, but also combine the timing of events, topological dependencies, and other contextual information to identify paths or event linkage patterns related to the target fault factor. Specifically, one or more of the following methods can be used to analyze correlation paths or events:
[0113] Method 1: Path / pattern exploration based on fault contribution (MCP score). This method uses the MCP score as the weight of nodes or edges to find the path or subgraph with the highest weight in the graph. For example, it searches for a path starting from the overall symptom node and passing through the node corresponding to the target fault factor.
[0114] Method 2: Through time series graph analysis, analyze the edges between event nodes in the graph based on time order to identify abnormal time series patterns or event sequences. For example, find a series of events that occur consecutively within a short period of time and are related to a specific entity (e.g., the entity corresponding to the target failure factor).
[0115] Method 3: Through topological dependency path tracing, starting from the affected service or device node on the association graph, trace upwards along the dependency edge, and combine the events on the path and the MCP score to locate the potential root cause upstream.
[0116] Method 4: Graph pattern matching is used to match subgraph patterns in the graph with a pre-defined library of known fault patterns in the knowledge base. These fault patterns are defined as combinations of specific node types and edge relationships, which can represent complex fault scenarios. For example, the "specific configuration change event" node is connected to the "device restart" node through the "affect" edge, and then connected to the "service unavailable" node through the "cause" edge.
[0117] Method 5: Analyze the context information of events in the graph through context sensitivity analysis, such as system load, network traffic, and related process status when the event occurs. This information is stored in the graph as attributes of nodes or edges and can be considered during analysis.
[0118] Method Six: Analysis of event linkages involves designing graph analysis algorithms to identify linkage patterns formed by multiple independent event nodes connected by temporal, relational, or causal edges. For example, a seemingly normal process initiation event, if immediately followed by an abnormal network connection attempt and accompanied by a security alert within a short period, might indicate a potential attack, which a single event alone would not reveal. These linkage patterns are identified by searching for specific edge sequences or subgraph structures within the graph. The graph analysis algorithms mentioned above can include MCP-based graph traversal algorithms (such as weighted depth-first search and weighted breadth-first search), temporal pathfinding algorithms, graph pattern matching algorithms, community detection algorithms (for identifying clusters of highly correlated events), and machine learning models based on graph neural networks (GNNs).
[0119] Of course, it is not limited to the above six methods; other methods that can determine the chain of evidence and root cause of the failure can also be chosen.
[0120] S304. Based on the correlation information between different levels in the source data of the abnormal event, verify the correctness of the fault evidence chain and the correctness of the correlation graph. If the fault evidence chain and the correlation graph are correct, then determine the target fault factor as the root cause of the fault.
[0121] Optionally, if the fault evidence chain is incorrect but the correlation graph is correct, then based on the fault contribution ranking of the fault factors, the fault factor with the second highest fault contribution ranking is determined as the target fault factor, and step S302 and subsequent steps are executed until the root cause of the fault is determined. If the correlation graph is incorrect, then the information of erroneous nodes or edges in the correlation graph is adjusted based on expert experience, and steps S101-S103 are re-executed based on the adjusted correlation graph.
[0122] In step S101, multi-layered analysis from different perspectives is performed, and associations are formed based on identifiers such as event IDs, forming a set containing cross-layer data. This data is represented as attributes of nodes or edges when constructing the association graph, or through specific edge types (such as "network event associated with system event"). The graph is used to determine the fault evidence chain and the root cause of the fault. Similarly, if the fault root cause engine presents a complete fault evidence chain that is caused by "network congestion and high CPU usage of the database server", this complete evidence chain and graph are then verified using the running data from step S101 to ensure the accuracy of the associations between nodes. If there are no errors, the correctness of the fault evidence chain and the correctness of the association graph are verified.
[0123] Furthermore, the association information between nodes in the association graph updated in steps S101-S103 can be iteratively updated to the association graph in the knowledge base, so that the optimized association graph can be used again in the next call. In addition, the initial association graph can be constructed based on human experience.
[0124] Thus, based on the steps S301-S304 above, by using the graph data of the correlation map and contextual correlation analysis, it is possible to identify those event linkage patterns with diagnostic significance formed by a series of seemingly normal or unrelated events under specific time sequence and correlation relationships, thereby discovering the root cause hidden under the complex appearance, which is difficult to achieve with existing isolated event analysis methods.
[0125] Furthermore, as mentioned above, this application marks network change times. Step S101 implements multi-perspective, layered, and parallel processing of multi-source operational data. Specifically addressing the monitoring blind spot issue caused by suppressed state alarms during network change events, it ensures that even under alarm masking, potential anomalies can still be captured through parallel analysis of the raw data at multiple levels, including network and application, avoiding redundant work. The multi-source data and corresponding analysis in step S101, combined with the quantitative evaluation in S102, are dynamically loaded into the correlation graph. Through the root cause analysis engine in subsequent step S103, the paths and patterns in the graph are traversed and analyzed to identify event linkage patterns, thereby locating the root causes of complex faults caused by multiple events under specific time sequences and correlations. Furthermore, it can consider the correlation of network change event nodes, distinguish between anomalies caused by the change itself and independent faults occurring during the change period, effectively analyzing the true faults.
[0126] In one possible implementation, the display interface shows one or more of the following: the root cause of the failure, the chain of evidence for the failure, evidence parameters supporting the chain of evidence for the failure, the topological impact range of the root cause of the failure, and the remedial measures based on the root cause of the failure.
[0127] The above describes some specific implementations of a full-link fault root cause analysis method provided in this application. Based on this, this application also provides a corresponding apparatus. The apparatus provided in this application will be described below from the perspective of functional modularity.
[0128] See Figure 4 The diagram shows a structural schematic of a full-link fault root cause analysis device. The full-link fault root cause analysis device includes:
[0129] The multi-level analysis module 401 is used to obtain abnormal events from the operation data of multiple department platforms, match them with fault factors in a preset knowledge base, determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, and form a set of related events.
[0130] The quantification module 402 is used to perform multi-dimensional contribution quantification on each fault factor in the set of associated events based on the association graph in the knowledge base and the static factor corresponding to the fault factor, so as to obtain the fault contribution degree of the fault factor.
[0131] The root cause analysis engine module 403 is used to update the correlation graph based on the abnormal event, the set of related events, and the fault contribution of each fault factor, and to parse the correlation graph through the root cause analysis engine to determine the fault evidence chain and the fault root cause.
[0132] According to the aforementioned apparatus, this application integrates data from multiple departments through a multi-level analysis module 401, overcoming departmental barriers and data silos. Then, a quantification module 402 quantitatively evaluates each failure factor, providing a scientific and accurate quantitative basis for the subsequent root cause analysis engine module 403. In this way, the deep integration of multi-source heterogeneous data, failure contribution quantification, and correlation maps ensures the comprehensiveness and accuracy of the analysis, helps break down data barriers between departments, and improves the accuracy of root cause analysis.
[0133] In one possible implementation, the multi-level analysis module 401 is specifically used to send the running data to processing pipelines at different levels to obtain a subset of related events output by multiple processing pipelines. The processing pipeline is used to perform anomaly analysis on the running data from the level to which the processing pipeline belongs, determine the abnormal events and their matching fault factors, and instantiate the fault factors to form a subset of related events; the subset of related events containing unified target information is constructed into a set of related events.
[0134] In one possible implementation, the multi-layer analysis module 401 is specifically used for the processing pipelines at each layer to perform anomaly analysis on the operational data at that layer. If anomalies exist in the operational data at that layer, an anomaly event is identified. The anomaly event includes an identifier code and source data, where the source data is information in the operational data associated with the anomaly event. Based on the anomaly event, it is matched with fault factors in the knowledge base to determine the fault factor corresponding to the anomaly event. Based on the fault factor and the source data of the anomaly event, the instantiation information of the fault factor is determined, and a subset of associated events is constructed. The elements of the subset of associated events include the fault factor, the entity information corresponding to the fault factor, and time information.
[0135] In one possible implementation, the device further includes a data acquisition and processing unit, which is used to mark raw data obtained from the multiple departmental platforms to obtain operational data in response to network change events and / or suppression status alarms; and to transmit the operational data to processing pipelines at different levels.
[0136] In one possible implementation, the quantification module 402 is specifically used to: determine the static factors corresponding to each fault factor included in the set of associated events from a preset knowledge base; determine the dynamic factors reflecting the degree of abnormality of the fault factor based on the source data of the abnormal events corresponding to the fault factor, wherein the source data is the data associated with the abnormal events in the operational data; evaluate the correlation factors between the fault factor and the overall symptoms, wherein the overall symptoms are the operational faults occurring in the multiple departmental platforms; determine the topological influence factor of the fault factor based on the correlation graph in the knowledge base and the position and associated nodes of the entity information corresponding to the fault factor in the correlation graph; and determine the fault contribution of the fault factor based on the static factors, dynamic factors, correlation factors, and topological influence factors.
[0137] In one possible implementation, the quantization module 402 is specifically used to acquire multiple dynamic factor parameters in the source data; the multiple factor parameters include one or more of the following: a first parameter reflecting the severity of the abnormal event, a frequency parameter of the occurrence of the abnormal event, a duration parameter of the abnormal event, a recentity parameter of the abnormal event, and a deviation parameter of the abnormal event from the baseline; the multiple dynamic factor parameters are weighted and summed, and the dynamic factors are obtained by solving through a normalization function.
[0138] In one possible implementation, the device further includes a display module, which is used to display a first table on a display interface. The first table includes fault factors corresponding to the operating data of the multiple department platforms and the fault contribution degree corresponding to each fault factor. The fault factors in the first table are sorted and displayed according to their corresponding fault contribution degrees.
[0139] In one possible implementation, the root cause analysis engine module 403 is specifically used to determine one of the fault factors corresponding to the operational data of the multiple departmental platforms as the target fault factor based on the fault contribution of the fault factors; load the source data of the abnormal event, the set of related events, and the fault contribution of the fault factors included in the set of related events onto the nodes corresponding to the association graph to obtain the loaded association graph; through the root cause analysis engine, perform contextual root cause analysis on the association graph for the target fault factor to obtain the fault evidence chain of the root cause of the target fault factor; verify the correctness of the fault evidence chain and the correctness of the association graph based on the association information between different levels in the source data of the abnormal event; if the fault evidence chain and the association graph are correct, then the target fault factor is determined to be the root cause of the fault.
[0140] In one possible implementation, the display module is further configured to display on a display interface one or more of the following: the root cause of the fault, the chain of evidence for the fault, evidence parameters supporting the chain of evidence for the fault, the topological impact range of the root cause of the fault, and the remedial measures based on the root cause of the fault.
[0141] This application also provides corresponding devices and computer storage media for implementing the solutions provided in this application.
[0142] The device includes a memory and a processor. The memory stores instructions or code, and the processor executes the instructions or code to cause the device to perform the method described in any embodiment of this application.
[0143] The computer storage medium stores code, and when the code is run, the device running the code implements the method described in any embodiment of this application.
[0144] In the embodiments of this application, the terms "first" and "second" (if they exist) are used only as name identifiers and do not represent the order of first and second.
[0145] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus a general-purpose hardware platform. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0146] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0147] The above description is merely an exemplary implementation of this application and is not intended to limit the scope of protection of this application.
Claims
1. A method for end-to-end fault root cause analysis, characterized in that, The method includes: After obtaining abnormal events from the operational data of multiple departmental platforms, they are matched with fault factors in a preset knowledge base to determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, thus forming a set of associated events. Based on the association graph in the knowledge base and the static factors corresponding to the fault factors, the multi-dimensional contribution quantification of each fault factor in the set of associated events is performed to obtain the fault contribution degree of the fault factors. The correlation graph is updated based on the abnormal events, the set of related events, and the fault contribution of each fault factor. The correlation graph is then analyzed by the root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
2. The method according to claim 1, characterized in that, After obtaining abnormal events from the operational data of multiple departmental platforms, they are matched with fault factors in a preset knowledge base to determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, forming a set of associated events, including: The running data is sent to processing pipelines at different levels to obtain a subset of associated events output by multiple processing pipelines. The processing pipelines are used to perform anomaly analysis on the running data from the level to which they belong, determine the abnormal events and their matching fault factors, and instantiate the fault factors to form a subset of associated events. Construct a set of associated events from a subset of associated events containing unified target information.
3. The method according to claim 2, characterized in that, The step of sending the runtime data to processing pipelines at different levels to obtain a subset of associated events output from at least one level of the processing pipeline includes: Each layer of the processing pipeline performs anomaly analysis on the operational data at that layer. If anomalies are found in the operational data at that layer, an anomaly event is identified. The anomaly event includes an identifier code and source data, where the source data is information in the operational data that is associated with the anomaly event. Based on the abnormal event, it is matched with the fault factors in the knowledge base to determine the fault factor corresponding to the abnormal event; Based on the source data of the fault factor and the abnormal event, the instantiation information of the fault factor is determined and a subset of associated events is constructed. The elements of the subset of associated events include the fault factor, the entity information corresponding to the fault factor, and the time information.
4. The method according to claim 2, characterized in that, Before sending the runtime data to different levels of the processing pipeline, the method further includes: Based on network change events and / or suppression status alarms, the raw data obtained from the multiple departmental platforms is tagged to obtain operational data; The operational data is sent to processing pipelines at different levels.
5. The method according to any one of claims 1-4, characterized in that, The step of performing multi-dimensional contribution quantification on each fault factor in the set of associated events based on the association graph in the knowledge base and the static factors corresponding to the fault factors to obtain the fault contribution degree of the fault factors includes: From a pre-defined knowledge base, determine the static factors corresponding to each fault factor included in the set of associated events; Based on the source data of the abnormal events corresponding to the fault factors, a dynamic factor reflecting the degree of abnormality of the fault factors is determined, wherein the source data is the data in the operation data associated with the abnormal events; Assess the correlation factors between the aforementioned failure factors and overall symptoms; Based on the association graph in the knowledge base, and the position and associated nodes of the entity information corresponding to the fault factor in the association graph, the topological influence factor of the fault factor is determined. The fault contribution of the fault factor is determined based on the static factor, dynamic factor, correlation factor, and topological influence factor.
6. The method according to claim 5, characterized in that, The step of determining the dynamic factor reflecting the degree of abnormality of the fault factor based on the source data of the abnormal events corresponding to the fault factor includes: Obtain multiple dynamic factor parameters from the source data; the multiple factor parameters include one or more of the following: a first parameter reflecting the severity of the abnormal event, a frequency parameter of the occurrence of the abnormal event, a duration parameter of the abnormal event, a recentity parameter of the abnormal event, and a deviation parameter of the abnormal event from the baseline; The dynamic factors are obtained by weighted summation of the multiple dynamic factor parameters and solving the problem using a normalization function.
7. The method according to claim 5, characterized in that, After determining the fault contribution of the fault factor based on the static factor, dynamic factor, correlation factor, and topological influence factor, the method further includes: The first table is displayed on the display interface. The first table includes the fault factors corresponding to the operation data of the multiple department platforms and the fault contribution degree corresponding to each fault factor. The fault factors in the first table are sorted and displayed according to their corresponding fault contribution degrees.
8. The method according to claim 6 or 7, characterized in that, The process of updating the correlation graph based on the abnormal events, the set of related events, and the fault contribution of each fault factor, and then analyzing the correlation graph through a root cause analysis engine to determine the fault evidence chain and the root cause of the fault, includes: Based on the fault contribution of the fault factors, one of the fault factors corresponding to the operational data of the multiple departmental platforms is determined as the target fault factor. The source data of the abnormal event, the set of related events, and the fault contribution of the fault factors included in the set of related events are loaded onto the corresponding nodes of the correlation graph to obtain the loaded correlation graph. Using the root cause analysis engine, the correlation graph is analyzed for contextual correlation to obtain the fault evidence chain of the root cause of the target fault factor. Based on the correlation information between different levels in the source data of the abnormal event, the correctness of the fault evidence chain and the correctness of the correlation graph are verified. If the fault evidence chain and the correlation graph are correct, the target fault factor is determined to be the root cause of the fault.
9. The method according to claim 8, characterized in that, The method further includes: The display interface shows one or more of the following: the root cause of the failure, the chain of evidence for the failure, the evidence parameters supporting the chain of evidence for the failure, the topological impact range of the root cause of the failure, and the remedial measures based on the root cause of the failure.
10. A full-link fault root cause analysis device, characterized in that, The device includes: The multi-level analysis module is used to obtain abnormal events from the operation data of multiple department platforms, match them with fault factors in a preset knowledge base, determine the fault factors corresponding to each abnormal event and the instantiation information of the fault factors, and form a set of related events. The quantification module is used to perform multi-dimensional contribution quantification on each fault factor in the set of associated events based on the association graph in the knowledge base and the static factors corresponding to the fault factors, so as to obtain the fault contribution degree of the fault factors. The root cause analysis engine module is used to update the correlation graph based on the abnormal events, the set of related events, and the fault contribution of each fault factor. The correlation graph is analyzed by the root cause analysis engine to determine the fault evidence chain and the root cause of the fault.
Citation Information
Cited By
Fault diagnosis and early warning method and system for mobile energy storage equipment
CN121412598A
Service fault detection method and system
CN121478664A
Internet of Things connection fault root cause positioning method, device, equipment and program product
CN121509212A