Fault root cause positioning method and device, electronic equipment and medium

By using knowledge graphs and causal inference models to analyze alarm information in a cloud environment, combined with the fault root cause prediction model, the problem of fault location difficulties caused by multi-layer monitoring alarms is solved, and more efficient and accurate fault location is achieved.

CN120216236APending Publication Date: 2025-06-27THE FOURTH PARADIGM BEIJING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279325.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In a cloud environment, multi-layer monitoring and alarms lead to difficulties in fault location, and operation and maintenance personnel find it difficult to quickly and accurately locate the root cause of the fault, affecting the efficiency of fault processing.

Method used

By receiving at least two alarm information, the association relationship between the pre-determined multiple alarm types, a pre-constructed knowledge graph and a pre-trained causal inference model are determined, and the failure root cause probability is determined by combining the association relationship and the failure root cause probability.

Benefits of technology

It improves the accuracy and efficiency of fault location, reduces the time for operation and maintenance personnel to check when facing multiple alarms, and improves the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216236A_ABST
    Figure CN120216236A_ABST
Patent Text Reader

Abstract

The invention provides a fault root cause positioning method and device, electronic equipment and a medium, and belongs to the technical field of information. The method comprises the following steps: under the condition that at least two pieces of alarm information are received, determining an association relationship between the at least two pieces of alarm information according to a predetermined association degree among a plurality of alarm types, a pre-constructed knowledge graph and / or a pre-trained causal inference model; determining the fault root cause probability of each component corresponding to the at least two pieces of alarm information by using a pre-trained fault root cause prediction model; and determining a target component corresponding to target alarm information in the at least two pieces of alarm information as a fault root cause according to an association relationship between the at least two pieces of alarm information and the fault root cause probability of each component. Therefore, the fault root cause is finally positioned by performing deep analysis on the multiple pieces of alarm data, identifying the incidence relation between the alarms and introducing the machine learning model to perform fault root cause prediction, so that the accuracy and efficiency of fault positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of information technology, and particularly relates to a method, apparatus, electronic device, and medium for fault root cause location. Background Art

[0002] In a cloud environment, an application usually adopts a microservices architecture, involving multiple components and services, such as gateways, k8s svc, k8s nodes, pods, MySQL, Alibaba Cloud translation service, Microsoft Cloud translation service, etc. To ensure the stability of the system, it is necessary to monitor each component. However, when a component fails, it will trigger alarms of multiple monitoring modules. Currently, fault location mainly relies on the experience of operation and maintenance personnel. When faced with multiple alarms triggered simultaneously, it is difficult for operation and maintenance personnel to quickly and accurately locate the root cause of the fault, and it may take a lot of time to check one by one, thus affecting the efficiency of fault handling. Summary of the Invention

[0003] The purpose of the embodiments of the present disclosure is to provide a method, apparatus, electronic device, and medium for fault root cause location, which can solve the problem of low fault handling efficiency in existing fault location methods.

[0004] In a first aspect, the embodiments of the present disclosure provide a method for fault root cause location, including:

[0005] When receiving at least two alarm messages, determine the association relationship between the at least two alarm messages according to target information, where the target information includes at least one of the following: the association degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, and a pre-trained causal inference model. The knowledge graph includes multiple nodes and multiple edges. The nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component;

[0006] Use a pre-trained fault root cause prediction model to determine the fault root cause probability of each component corresponding to the at least two alarm messages, where the fault root cause prediction model is trained based on historical alarm data;

[0007] Determine the fault root cause from the components corresponding to the at least two alarm messages according to the association relationship between the at least two alarm messages and the fault root cause probability of each component.

[0008] Optionally, the determining the association relationship between the at least two alarm messages according to the target information includes:

[0009] Determine the association degree between the at least two alarm messages according to the association degree between multiple pre-determined alarm types and the alarm types corresponding to each alarm message in the at least two alarm messages;

[0010] Determine the causal relationship between the at least two pieces of alarm information according to a pre-constructed knowledge graph and / or a pre-trained causal inference model;

[0011] Among them, the association relationship includes the degree of association and the causal relationship.

[0012] Optionally, the determining the causal relationship between the at least two pieces of alarm information according to a pre-constructed knowledge graph and / or a pre-trained causal inference model includes:

[0013] Infer a first causal relationship between the at least two pieces of alarm information according to the relationships between the components in the pre-constructed knowledge graph;

[0014] And / or, arrange the at least two pieces of alarm information in the order of alarm time to obtain an alarm event sequence; input the alarm event sequence into the pre-trained causal inference model to obtain a second causal relationship between the at least two pieces of alarm information output by the causal inference model;

[0015] Among them, the causal relationship between the at least two pieces of alarm information includes the first causal relationship and / or the second causal relationship.

[0016] Optionally, the determining the root cause of the fault from the components corresponding to the at least two pieces of alarm information according to the association relationship between the at least two pieces of alarm information and the probability of the root cause of the fault of each component includes:

[0017] Determine a first confidence level of each component as the root cause of the fault according to the probability of the root cause of the fault of each component;

[0018] Adjust the first confidence level according to the association relationship between the at least two pieces of alarm information to obtain a target confidence level;

[0019] Sort the components according to the target confidence level to generate a list of root causes of the fault, and determine the target component with the highest target confidence level as the root cause of the fault.

[0020] Optionally, the method further includes at least one of the following:

[0021] Search for the corresponding handling suggestions when the target component fails in a pre-established fault handling suggestion library, and output the handling suggestions, where the target component is the component determined as the root cause of the fault from the components corresponding to the at least two pieces of alarm information;

[0022] Graphically display the fault location result, where the fault location result includes at least one of the following: the association relationship between the at least two pieces of alarm information, the knowledge graph, and the location process of the root cause of the fault.

[0023] Optionally, the method further includes:

[0024] Determine the correlation degree between every two of the multiple alarm types according to the co-occurrence times of every two of the multiple alarm types in the historical alarm information.

[0025] Optionally, the method further includes:

[0026] Receive feedback information from the user on the positioning result of the root cause of the fault for the target component, where the feedback information includes at least one of confirmation, modification, and negation;

[0027] Adjust the root cause prediction model of the fault according to the feedback information.

[0028] Optionally, the root cause prediction model of the fault is trained in the following manner:

[0029] Obtain historical alarm data, extract alarm features from the historical alarm data, and determine the component information involved in the historical alarm data; use the alarm features and the component information as model input features, input them into an initial root cause prediction model of the fault, and adjust the model parameters of the root cause prediction model of the fault according to the deviation between the model output result and the true root cause of the fault marked in the historical alarm data, so as to obtain the trained root cause prediction model of the fault;

[0030] Wherein, the alarm features include at least one of the following: alarm time, alarm type, alarm frequency, alarm duration, and alarm level.

[0031] In a second aspect, an embodiment of the present disclosure provides a root cause localization device for a fault, including:

[0032] A first determination module, configured to determine the correlation relationship between at least two alarm messages according to target information when receiving at least two alarm messages, where the target information includes at least one of the following: the correlation degree between multiple predetermined alarm types, a pre-constructed knowledge graph, and a pre-trained causal inference model, the knowledge graph includes multiple nodes and multiple edges, the nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component;

[0033] A second determination module, configured to use a pre-trained root cause prediction model of the fault to determine the root cause probability of each component corresponding to the at least two alarm messages, where the root cause prediction model of the fault is trained based on historical alarm data;

[0034] A third determination module, configured to determine a root cause of a fault from the components corresponding to the at least two pieces of alarm information according to the association relationship between the at least two pieces of alarm information and the probability of the root cause of the fault of each component.

[0035] Optionally, the first determination module includes:

[0036] A first determination sub-module, configured to determine the association degree between the at least two pieces of alarm information according to the association degree between a plurality of pre-determined alarm types and the alarm types corresponding to each piece of alarm information in the at least two pieces of alarm information;

[0037] A second determination sub-module, configured to determine the causal relationship between the at least two pieces of alarm information according to a pre-constructed knowledge graph and / or a pre-trained causal inference model;

[0038] Wherein, the association relationship includes the association degree and the causal relationship.

[0039] Optionally, the second determination sub-module includes:

[0040] A first inference unit, configured to infer a first causal relationship between the at least two pieces of alarm information according to the relationship between the components in a pre-constructed knowledge graph;

[0041] And / or, a second inference unit, configured to arrange the at least two pieces of alarm information in the order of alarm time to obtain an alarm event sequence; input the alarm event sequence into a pre-trained causal inference model to obtain a second causal relationship between the at least two pieces of alarm information output by the causal inference model;

[0042] Wherein, the causal relationship between the at least two pieces of alarm information includes the first causal relationship and / or the second causal relationship.

[0043] Optionally, the third determination module includes:

[0044] A first determination unit, configured to determine a first confidence level of each component as a root cause of a fault according to the probability of the root cause of the fault of each component;

[0045] An adjustment unit, configured to adjust the first confidence level according to the association relationship between the at least two pieces of alarm information to obtain a target confidence level;

[0046] A second determination unit, configured to sort the components according to the target confidence level, generate a list of root causes of faults, and determine the target component with the highest target confidence level as the root cause of the fault.

[0047] Optionally, the fault root cause location device further includes at least one of the following:

[0048] An output module, configured to look up a processing suggestion corresponding to a target component when a failure occurs from a pre-established failure processing suggestion library, and output the processing suggestion, where the target component is a component determined as the root cause of the failure from the components corresponding to the at least two alarm messages;

[0049] A display module, configured to graphically display the failure location result, where the failure location result includes at least one of the following: the association relationship between the at least two alarm messages, the knowledge graph, and the location process of the root cause of the failure.

[0050] Optionally, the root cause of failure location device further includes:

[0051] A fourth determination module, configured to determine the association degree between every two alarm types among the multiple alarm types according to the co-occurrence times of every two alarm types in the historical alarm messages.

[0052] Optionally, the root cause of failure location device further includes:

[0053] A receiving module, configured to receive feedback information of the user on the location result of the target component as the root cause of the failure, where the feedback information includes at least one of confirmation, modification, and negation;

[0054] An adjustment module, configured to adjust the root cause of failure prediction model according to the feedback information.

[0055] Optionally, the root cause of failure prediction model is trained in the following manner:

[0056] Obtain historical alarm data, extract alarm features from the historical alarm data, and determine the component information involved in the historical alarm data; use the alarm features and the component information as model input features, input them into an initial root cause of failure prediction model, and adjust the model parameters of the root cause of failure prediction model according to the deviation between the model output result and the true root cause of the failure marked in the historical alarm data, so as to obtain the trained root cause of failure prediction model;

[0057] Wherein, the alarm features include at least one of the following: alarm time, alarm type, alarm frequency, alarm duration, and alarm level.

[0058] In a third aspect, an embodiment of the present disclosure provides an electronic device, which includes a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0059] Fourthly, embodiments of the present disclosure provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0060] Fifthly, embodiments of the present disclosure provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0061] Sixthly, embodiments of the present disclosure provide a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0062] In the embodiments of the present disclosure, when at least two alarm messages are received, according to the target information, the association relationship between the at least two alarm messages is determined, where the target information includes at least one of the following: the association degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, and a pre-trained causal inference model. The knowledge graph includes multiple nodes and multiple edges. The nodes represent components in the system, and the edges represent the relationships between the nodes. One alarm message corresponds to one component. A pre-trained fault root cause prediction model is used to determine the fault root cause probabilities of the components corresponding to the at least two alarm messages, where the fault root cause prediction model is trained based on historical alarm data. According to the association relationship between the at least two alarm messages and the fault root cause probabilities of the components, the fault root cause is determined from the components corresponding to the at least two alarm messages. In this way, by deeply analyzing multiple alarm data, the association relationship between alarms is identified, and a machine learning model is introduced to predict the fault root cause of multiple alarms. Finally, combining the association relationship between alarms and the fault root cause prediction result, the fault root cause is located. This method can improve the accuracy and efficiency of fault location. Description of the Drawings

[0063] Figure 1 is a flowchart of the fault root cause location method provided by the embodiments of the present disclosure;

[0064] Figure 2 is a structural diagram of the fault root cause location device provided by the embodiments of the present disclosure;

[0065] Figure 3 is a structural diagram of the electronic device provided by the embodiments of the present disclosure. Detailed Embodiments

[0066] Next, the technical solutions in the embodiments of the present disclosure will be clearly described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present disclosure.

[0067] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the associated objects before and after are in an "or" relationship.

[0068] To make the embodiments of the present disclosure clearer, the following introduces the relevant background technical knowledge involved in the embodiments of the present disclosure:

[0069] The following technical defects exist in the prior art:

[0070] 1) Multi-layer monitoring and alarming make it difficult to locate faults

[0071] In the cloud environment, applications usually adopt a microservices architecture, involving multiple components and services, such as gateways, container cluster management systems (kubernetes, k8s) services (service, svc), k8s nodes (node), containers (pod), My Structured Query Language (MySQL), cloud translation services, etc. To ensure the stability of the system, it is necessary to monitor each component. However, when a component fails, it will trigger alarms from multiple monitoring modules.

[0072] The normal data flow of a translation service is: gateway -> k8s svc -> node -> pod. In addition, the pod needs to depend on third parties such as MySQL, cloud translation service 1, and cloud translation service 2.

[0073] As a translation service with a relatively comprehensive coverage, its monitoring and inspection should cover the following aspects:

[0074] Monitoring 1: Monitor whether the functions of the entire process are normal by injecting traffic through the gateway;

[0075] Monitoring 2: Determine whether k8s is normal by injecting traffic through k8s svc;

[0076] Monitoring 3: Monitor k8s nodes;

[0077] Monitoring 4: Monitor pod status and pod logs;

[0078] Monitoring 5: Monitor the third-party dependency MySQL;

[0079] Monitoring 6: Monitor Cloud Translation Service 1;

[0080] Monitoring 7: Conduct patrol monitoring on Cloud Translation Service 2.

[0081] Example 1: When a network paralysis occurs in a k8s node, Monitoring 1 (Gateway Injected Traffic Monitoring), Monitoring 2 (k8s svc Traffic Monitoring), Monitoring 3 (k8s Node Monitoring), and Monitoring 4 (Pod Status Monitoring) will issue alarms simultaneously.

[0082] Example 2: When an exception occurs in MySQL, Monitoring 1, Monitoring 2, Monitoring 4, and Monitoring 5 (MySQL Monitoring) will issue alarms simultaneously.

[0083] Facing multiple alarms triggered simultaneously, it is difficult for operation and maintenance personnel to accurately locate the root cause of the failure in a timely manner. They may need to spend a lot of time troubleshooting one by one, which affects the efficiency of failure handling.

[0084] 2) Existing methods lack intelligent alarm correlation analysis

[0085] Currently, failure location mainly relies on the experience of operation and maintenance personnel or conducts alarm correlation analysis based on preset rules. These methods have the following problems:

[0086] Lack of flexibility: Preset rules cannot cover all possible failure scenarios. Facing a complex and changeable cloud environment, the cost of maintaining and updating the rules is high.

[0087] Unable to deeply analyze causal relationships: Existing methods are difficult to automatically analyze the deep causal relationships between alarms and cannot quickly lock the root cause of the failure.

[0088] Facing the above defects, the present disclosure proposes a method for quickly locating the root cause of a failure based on AI. By collecting alarm information from each monitoring module and using machine learning and knowledge graph technologies, it analyzes the correlation and causal relationships between alarms, quickly locates the root cause of the failure, improves the processing efficiency of operation and maintenance personnel, and reduces the impact of failures.

[0089] The following will, in conjunction with the accompanying drawings, explain in detail the failure root cause location method provided by the embodiments of the present disclosure through specific embodiments and their application scenarios.

[0090] Please refer to Figure 1 , Figure 1The flowchart of the fault root cause location method provided by the embodiments of the present disclosure is as follows Figure 1 As shown, the method includes the following steps:

[0091] Step 101, when receiving at least two alarm messages, determine the association relationship between the at least two alarm messages according to the target information, where the target information includes at least one of the following: the association degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, a pre-trained causal inference model, the knowledge graph includes multiple nodes and multiple edges, the nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component.

[0092] When a component in the system fails, multiple monitored alarms will be triggered, so the system will receive alarms from multiple monitoring modules, that is, receive at least two alarm messages. Each alarm message corresponds to an alarm from a monitoring module, and each monitoring module monitors one component, that is, one alarm message corresponds to an alarm of one component.

[0093] Specifically, the following data can be collected from each monitoring module in real time:

[0094] Alarm information: alarm time, alarm type, alarm level, involved components, alarm description, etc.

[0095] Log data: system logs, application logs, error logs, etc.

[0096] Performance metrics: Central Processing Unit (CPU) usage rate, memory usage rate, network traffic, response time, etc.

[0097] In some embodiments, user experience data can also be collected, such as collecting user feedback, application performance metrics, etc., to comprehensively evaluate the scope of the fault impact.

[0098] Then, the above data is preprocessed, including:

[0099] Data cleaning: Remove duplicate, incorrect, and missing data to ensure data quality.

[0100] Data formatting: Unify the formats of data from different sources to facilitate subsequent analysis.

[0101] Feature extraction: Extract key features, such as alarm frequency, alarm duration, context information of involved components, etc.

[0102] In some embodiments, anomaly detection algorithms can also be introduced, such as Isolation Forest, One-Class SVM, etc., to identify abnormal alarm patterns.

[0103] Subsequently, subsequent analysis is performed based on the extracted feature data.

[0104] In the embodiments of the present disclosure, when at least two alarm messages are received, correlation analysis can be performed on the at least two alarm messages, such as analyzing the correlation degree between the at least two alarm messages and inferring the causal relationship between the at least two alarm messages. Specifically, it can be combined with the correlation degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, a pre-trained causal inference model, etc. for analysis. For example, the alarm type to which each alarm message in the at least two alarm messages belongs can be determined, and then the correlation degree between the at least two alarm messages can be determined according to the correlation degree between multiple pre-determined alarm types; it is also possible to determine the components corresponding to each alarm message in the at least two alarm messages, and then infer the possible causal relationship between the at least two alarm messages according to the dependency relationship and call relationship defined between the components in the pre-constructed knowledge graph. Additionally, a pre-trained causal inference model can be used to infer the causal relationship between the at least two alarm messages.

[0105] Among them, in the knowledge graph, each component in the system can be represented by a node, and the components with a dependency relationship or an interaction relationship are connected by an edge. That is to say, the elements in the knowledge graph can be defined to include nodes (entities) and edges (relationships). For example:

[0106] Nodes (entities): Components and services in the system, including gateways, k8s svc, k8s nodes, pods, MySQL, cloud translation services, etc.

[0107] Edges (relationships): Dependency relationships and interaction relationships between components, services, and between components and services in the system, such as call relationships, deployment relationships, network connections, etc.

[0108] The knowledge graph can be constructed in the following manner:

[0109] Manual construction: Based on the system architecture document, initially establish the relationship between components.

[0110] Automatic update: Use the configuration management database and automatic discovery tools to dynamically update the knowledge graph and maintain an accurate description of the current state of the system.

[0111] An example of the constructed knowledge graph is as follows:

[0112] Gateway --> k8s svc --> k8s node --> pod --> MySQL / Cloud translation service

[0113] Optionally, determining the correlation relationship between the at least two alarm messages according to the target information includes:

[0114] Determine the correlation degree between the at least two pieces of alarm information according to the correlation degrees between multiple pre-determined alarm types and the alarm types corresponding to each piece of alarm information in the at least two pieces of alarm information;

[0115] Determine the causal relationship between the at least two pieces of alarm information according to a pre-constructed knowledge graph and / or a pre-trained causal inference model;

[0116] Wherein, the association relationship includes the correlation degree and the causal relationship.

[0117] That is, in some embodiments, the alarm types corresponding to each piece of alarm information in the at least two pieces of alarm information can be determined, and then, in combination with the correlation degrees between multiple pre-determined alarm types, the correlation degrees between the alarm types corresponding to the at least two pieces of alarm information can be determined, and further the correlation degrees between the at least two pieces of alarm information can be determined. For example, the correlation degree between alarm type 1 and alarm type 2 is 80, the correlation degree between alarm type 1 and alarm type 3 is 90, alarm information A in the at least two pieces of alarm information corresponds to alarm type 1, alarm information B corresponds to alarm type 3, and alarm information C corresponds to alarm type 2. Therefore, it can be determined that the correlation degree between alarm information A and alarm information B is 90, and the correlation degree between alarm information A and alarm information C is 80.

[0118] The causal relationship between the at least two pieces of alarm information can also be inferred according to a pre-constructed knowledge graph and / or a pre-trained causal inference model. Specifically, according to the pre-constructed knowledge graph, the relationships between the components corresponding to the at least two pieces of alarm information, such as the dependency relationship and the call relationship, can be found, and the causal relationship between the at least two pieces of alarm information can be inferred based on this relationship. For example, if component A depends on component B, then the failure of component B may cause component A to alarm. Therefore, it can be inferred that the root cause of component A's alarm is component B's alarm. A causal inference model can also be pre-trained and used to infer the causal relationship between the at least two pieces of alarm information.

[0119] In some embodiments, a graph neural network (GNN) can also be used to perform a more in-depth analysis of the knowledge graph to capture complex alarm association patterns.

[0120] Wherein, the association relationship includes the correlation degree and the causal relationship. In this way, the root cause of the failure can be comprehensively determined by combining the correlation degree, the causal relationship between the at least two pieces of alarm information, and the probability of the root cause of the failure of each predicted component, ensuring the accuracy of fault location.

[0121] Optionally, the method further includes:

[0122] Determine the correlation degree between every two of the multiple alarm types according to the co-occurrence times of every two alarm types described in the historical alarm information.

[0123] In some embodiments, the historical alarm information can be analyzed to calculate the co-occurrence times of different alarm types. For example, when alarm types A and B trigger alarms simultaneously, it is recorded as 1 co-occurrence. If alarm types A and B trigger alarms simultaneously 5 times within the past 7 days, the co-occurrence times of alarm types A and B can be determined to be 5. Then, according to the co-occurrence times between different alarm types, statistical methods such as the Pearson correlation coefficient can be used to measure the correlation degree between different alarm types. Further, a correlation matrix between different alarm types can be constructed to represent the correlation strength between different alarm types.

[0124] In this way, through this implementation method, it can be ensured that the correlation degree between each alarm type is measured more accurately, and when multiple alarms are triggered simultaneously subsequently, according to the pre-constructed correlation degree of each alarm type, the correlation degree between each alarm event can be analyzed to help accurately locate the root cause of the fault.

[0125] Optionally, the determining of the causal relationship between the at least two alarm messages according to the pre-constructed knowledge graph and / or the pre-trained causal inference model includes:

[0126] Infer the first causal relationship between the at least two alarm messages according to the relationships between the components in the pre-constructed knowledge graph;

[0127] and / or, arrange the at least two alarm messages in the order of alarm time to obtain an alarm event sequence; input the alarm event sequence into the pre-trained causal inference model to obtain the second causal relationship between the at least two alarm messages output by the causal inference model;

[0128] wherein, the causal relationship between the at least two alarm messages includes the first causal relationship and / or the second causal relationship.

[0129] In some embodiments, a knowledge graph or a causal inference model can be used to perform causal inference, or a knowledge graph and a causal inference model can be combined to perform causal inference to ensure the reliability of the inference results. Specifically, the relationships defined between the components in the knowledge graph can be used to infer the possible causal relationships between the at least two pieces of alarm information, and this inference result is the above-mentioned first causal relationship. A causal inference model can also be used to infer the causal relationships between the at least two pieces of alarm information. Specifically, the at least two pieces of alarm information can be arranged in the order of alarm time to form an alarm event sequence, and then the alarm event sequence is input into a pre-trained causal inference model. This model can adopt Granger causality test, structural equation model, etc. By analyzing the alarm event sequence through the causal inference model, the causal relationships between the alarms are determined, and this inference result is the above-mentioned second causal relationship.

[0130] Among them, the causal relationship includes the first causal relationship and / or the second causal relationship, that is, the causal relationship between the at least two pieces of alarm information can be determined according to the inference result of the first causal relationship or the inference result of the second causal relationship, or the inference results of these two causal relationships can be combined to comprehensively determine the causal relationship between the at least two pieces of alarm information. For example, first, the possible causal relationship between the at least two pieces of alarm information is preliminarily determined according to the first causal relationship, and then the final causal relationship between the at least two pieces of alarm information is determined according to the second causal relationship. This implementation manner can ensure the accuracy of causal relationship inference.

[0131] In the embodiments of the present application, artificial intelligence (AI) technology is introduced to automatically perform alarm correlation analysis and causal analysis. Machine learning and knowledge graph technologies are used to deeply analyze multi-source monitoring data to automatically identify the correlation and causal relationships between alarms, so as to quickly locate the root cause of the fault. A dependency relationship model between the components of the system is established and visually represented in the form of a knowledge graph to assist the AI model in understanding the system structure and improving the accuracy and efficiency of fault location.

[0132] Step 102: Use a pre-trained fault root cause prediction model to determine the fault root cause probabilities of the components corresponding to the at least two pieces of alarm information, where the fault root cause prediction model is trained based on historical alarm data.

[0133] In the embodiments of the present disclosure, a fault root cause prediction model can also be pre-trained to predict the probability that each component corresponding to each piece of alarm information is the fault root cause. Specifically, historical alarm data and corresponding fault root causes can be collected, and a suitable model such as a deep learning model is selected as the fault root cause prediction model, and then the collected data is used to train the fault root cause prediction model.

[0134] Optionally, the fault root cause prediction model is trained in the following manner:

[0135] Obtain historical alarm data, extract alarm features from the historical alarm data, and determine the component information involved in the historical alarm data; use the alarm features and the component information as model input features, input them into an initial fault root cause prediction model, and adjust the model parameters of the fault root cause prediction model according to the deviation between the model output result and the true fault root cause marked in the historical alarm data, so as to obtain the trained fault root cause prediction model;

[0136] Wherein, the alarm features include at least one of the following: alarm time, alarm type, alarm frequency, alarm duration, alarm level.

[0137] That is, in some embodiments, a supervised learning algorithm such as random forest, Extreme Gradient Boosting (XGBoost), etc., or a deep learning algorithm such as Recurrent Neural Network (RNN), Transformer, etc., can be used to train the fault prediction model, extract alarm features such as alarm time, alarm type, alarm frequency, alarm duration, alarm level, etc., and component information from the collected historical alarm data as model input features, train the fault prediction model, and specifically adjust the model parameters of the fault root cause prediction model according to the deviation between the model output result and the true fault root cause marked in the historical alarm data, that is, the model loss. Repeat the above training process to minimize the model training loss, and finally obtain the trained fault root cause prediction model.

[0138] In this way, by training the fault root cause prediction model in the above training manner, a relatively high model accuracy can be ensured, and further the reliability of the model prediction result can be ensured.

[0139] In the embodiments of the present application, based on historical faults and alarm data, a fault root cause prediction model is obtained by training a machine learning model, learning alarm patterns and fault characteristics, and realizing intelligent fault root cause prediction.

[0140] Step 103: Determine the fault root cause from the components corresponding to the at least two alarm messages according to the association relationship between the at least two alarm messages and the fault root cause probabilities of the respective components.

[0141] In this step, the analysis results of the correlation relationships and the predicted results of the root causes of the at least two alarm messages can be fused to locate the final root cause of the fault. For example, the component that has a strong correlation with other components among the at least two alarm messages and has the highest probability of the root cause of the fault can be determined as the final root cause of the fault. That is to say, the root cause of the fault can be verified from multiple dimensions to ensure the accuracy and reliability of the location of the root cause of the fault.

[0142] Optionally, step 103 includes:

[0143] Determine the first confidence level of each component as the root cause of the fault according to the probability of the root cause of the fault of each component;

[0144] Adjust the first confidence level according to the correlation relationship between the at least two alarm messages to obtain the target confidence level;

[0145] Sort the components according to the target confidence level to generate a list of root causes of the fault, and determine the target component with the highest target confidence level as the root cause of the fault.

[0146] In some embodiments, according to the predicted probability of the root cause of the fault of each component, the first confidence level of each component as the root cause of the fault can be determined. For example, the probability of the root cause of the fault of the component can be directly converted into a confidence level. If the predicted probability of the root cause of the fault of a certain component is 0.8, then the first confidence level of this component as the root cause of the fault can be directly determined to be 80% (the highest confidence level is 100%); then, according to the correlation relationship between the at least two alarm messages, the first confidence level can be adjusted to obtain the final confidence level, that is, the target confidence level. For example, it is determined that there is a strong correlation between alarm message 1 and alarm message 2, and the alarm of component 2 will be caused by the fault of component 1 corresponding to alarm message 1, that is, alarm message 2 is generated. The confidence level of component 1 as the root cause of the fault is predicted to be 80%, and the confidence level of component 2 as the root cause of the fault is 70%. Then, the confidence level of component 1 as the root cause of the fault can be adjusted to 90%, and the confidence level of component 2 as the root cause of the fault can be adjusted to 60%.

[0147] It should be noted that in some embodiments, a confidence level adjustment model can also be trained according to relevant historical data, so as to use this model to perform confidence level adjustment to ensure the adjustment efficiency and adjustment accuracy. Among them, the relevant historical data can include, for example, historical alarm data, corresponding root causes of faults, historical adjustment data of the confidence level of each predicted component as the root cause of the fault, etc.

[0148] After determining the target confidence levels of each component as the root cause of the fault, the components can also be sorted in descending order of the target confidence levels to generate a list of root causes of the fault, which can be output for the user to view, facilitating the user to understand the possible situations of component failures currently. The target component with the highest target confidence level can also be determined as the root cause of the fault, facilitating the user to quickly troubleshoot the fault of the target component.

[0149] In this way, by fusing the correlation degree, the results of causal inference, and the prediction results of the machine learning model, the accuracy of fault location can be effectively improved.

[0150] Optionally, the method further includes at least one of the following:

[0151] Search for the corresponding handling suggestions when the target component fails from a pre-established fault handling suggestion library, and output the handling suggestions, where the target component is the component determined as the root cause of the fault from the components corresponding to the at least two alarm messages;

[0152] Graphically display the fault location results, where the fault location results include at least one of the following: the correlation relationship between the at least two alarm messages, the knowledge graph, and the location process of the root cause of the fault.

[0153] In some embodiments, services such as providing fault handling suggestions and graphical display of fault location results can also be provided. For example, a fault handling suggestion library for common faults can be established in advance, and then after the root cause of the fault is located, the corresponding handling suggestions can be retrieved from the fault handling suggestion library and output for the user to view, helping the operation and maintenance personnel to quickly solve the fault problem. The correlation relationship between each analyzed alarm message, the knowledge graph, the entire fault location process, etc. can also be graphically displayed, facilitating the user to intuitively understand the location process of the root cause of the fault.

[0154] Optionally, the method further includes:

[0155] Receive feedback information from the user on the location result of the target component as the root cause of the fault, where the feedback information includes at least one of confirmation, modification, and negation;

[0156] Adjust the fault root cause prediction model according to the feedback information.

[0157] In some embodiments, a feedback mechanism can be provided to support the user to provide feedback on the fault root cause location result. For example, the operation and maintenance personnel can confirm, modify, or negate the fault root cause location result, and the fault root cause prediction model can also be optimized and adjusted based on the user's feedback. For example, when the user negates the target component as the root cause of the fault, the participation of the fault root cause prediction model can be adjusted to reduce the probability that the model outputs the target component as the root cause of the fault in this case.

[0158] Through this implementation, the root cause prediction model of faults can be continuously optimized based on user feedback, and the accuracy of root cause prediction of faults can be continuously improved.

[0159] The following are several examples to illustrate the implementation of the disclosed embodiments of the book:

[0160] The normal data flow of a translation service is: gateway -> k8s svc -> k8s node -> pod. In addition, the pod needs to depend on third parties such as MySQL, Cloud Translation Service 1, and Cloud Translation Service 2. The following monitoring modules are included:

[0161] Monitoring 1: Monitor whether the functions of the entire process are normal by injecting traffic through the gateway.

[0162] Monitoring 2: Determine whether k8s is normal by injecting traffic through k8s svc.

[0163] Monitoring 3: Monitor the k8s nodes.

[0164] Monitoring 4: Monitor the pod status and pod logs;

[0165] Monitoring 5: Monitor the third-party dependency MySQL;

[0166] Monitoring 6: Monitor Cloud Translation Service 1;

[0167] Monitoring 7: Conduct patrol monitoring on Cloud Translation Service 2.

[0168] Scenario 1: Node network paralysis

[0169] Alarm trigger: Monitoring 1, Monitoring 2, Monitoring 3, and Monitoring 4 alarm simultaneously.

[0170] The analysis process is as follows:

[0171] The knowledge graph shows that problems with k8s nodes will affect the normal operation of pods, k8s svc, and the gateway.

[0172] Alarm correlation analysis: The alarm of Monitoring 3 (k8s node monitoring) has the highest correlation with other alarms.

[0173] Causal inference: A network failure of the k8s node may cause the pod to fail to operate normally, thereby affecting the traffic of k8s svc and the gateway.

[0174] Machine learning prediction: Based on historical data, the probability of k8s node network paralysis is the highest under such an alarm combination.

[0175] Root cause location: Determine that the k8s node network paralysis is the root cause of the fault.

[0176] Processing suggestion: Check and restore the network connection of the k8s node.

[0177] Scenario 2: MySQL anomaly

[0178] Alarm trigger: Monitoring 1, Monitoring 2, Monitoring 4, and Monitoring 5 alarm simultaneously.

[0179] The analysis process is as follows:

[0180] The knowledge graph shows that the pod depends on MySQL, and the MySQL anomaly will affect the normal operation of the pod, and further affect the k8s svc and the gateway.

[0181] Alarm correlation analysis: The alarm of Monitoring 5 (MySQL monitoring) has a high correlation with other alarms.

[0182] Causal inference: The MySQL failure may cause the pod to be unable to connect to the database, affecting the normal provision of services.

[0183] Machine learning prediction: Based on historical data, under this combination of alarms, the possibility of a MySQL anomaly is high.

[0184] Root cause location: Determine that the MySQL anomaly is the root cause of the failure.

[0185] Processing suggestion: Check the status of the MySQL service and restart or repair the database service.

[0186] According to the above introduction, the system of the embodiments of the present disclosure may mainly include the following modules:

[0187] Data collection module: Collect alarm information, log data, performance metrics, etc. of all monitoring modules in the system.

[0188] Data preprocessing module: Clean, format, extract features from the collected data, and convert it into structured data required for analysis.

[0189] Knowledge graph construction module: Construct the knowledge graph of the system according to the system architecture and component dependencies.

[0190] Fault analysis and prediction module: Use machine learning algorithms and knowledge graphs to perform correlation analysis and causal relationship inference on multi-source alarm data, and predict the root cause of the failure.

[0191] Root cause location and suggestion module: Based on the analysis results, output the root cause of the failure and processing suggestions to assist the operation and maintenance personnel in quickly solving the problem.

[0192] Human-computer interaction and feedback module: Provide a visual interface to display the analysis process and results, accept the feedback from the operation and maintenance personnel, and continuously optimize the model.

[0193] The embodiments of the present disclosure can also be extended to the following scenarios:

[0194] 1) Fault location across cloud environments

[0195] For multi-cloud and hybrid cloud environments, integrate the monitoring data of different cloud platforms for unified fault analysis and location.

[0196] 2) DevOps process optimization

[0197] Integrate the root cause analysis of faults into the DevOps (Development & Operations) process, automatically trigger corresponding repair strategies, and achieve automated operation and maintenance.

[0198] 3) Intelligent operation and maintenance platform

[0199] Integrate the fault location method of the embodiments of the present disclosure into a comprehensive intelligent operation and maintenance platform to provide operation and maintenance support throughout the life cycle.

[0200] The embodiments of the present disclosure solve the problems of excessive multi-layer monitoring alarms and difficult fault location for cloud applications. By introducing AI technology, rapid location of the root cause of faults is achieved, the fault handling time is reduced, and the reliability and stability of the system are improved. This method is universal and applicable to the monitoring and operation and maintenance of various cloud applications, and has important commercial value and application prospects.

[0201] In the root cause location method of the embodiments of the present disclosure, when receiving at least two alarm messages, according to the target information, determine the correlation relationship between the at least two alarm messages, where the target information includes at least one of the following: the correlation degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, and a pre-trained causal inference model. The knowledge graph includes multiple nodes and multiple edges, the nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component; use a pre-trained root cause prediction model of faults to determine the probability of the root cause of faults for each component corresponding to the at least two alarm messages, where the root cause prediction model of faults is trained based on historical alarm data; according to the correlation relationship between the at least two alarm messages and the probability of the root cause of faults for each component, determine the root cause of faults from the components corresponding to the at least two alarm messages. In this way, by deeply analyzing multiple alarm data, identifying the correlation relationship between alarms, and introducing a machine learning model to predict the root cause of faults for multiple alarms, and finally combining the correlation relationship between alarms and the root cause prediction result, the root cause of faults is located, and this method can improve the accuracy and efficiency of fault location.

[0202] The fault root cause location method provided by the embodiments of the present disclosure may be executed by a fault root cause location device. In the embodiments of the present disclosure, taking the fault root cause location device executing the fault root cause location method as an example, the fault root cause location device provided by the embodiments of the present disclosure is described.

[0203] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the fault root cause location device provided by the embodiments of the present disclosure. As Figure 2 shown, the fault root cause location device 200 includes:

[0204] A first determination module 201, configured to determine the association relationship between the at least two alarm messages according to the target information when receiving the at least two alarm messages, where the target information includes at least one of the following: the association degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, and a pre-trained causal inference model. The knowledge graph includes multiple nodes and multiple edges, the nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component;

[0205] A second determination module 202, configured to determine the fault root cause probability of each component corresponding to the at least two alarm messages by using a pre-trained fault root cause prediction model, where the fault root cause prediction model is trained based on historical alarm data;

[0206] A third determination module 203, configured to determine the fault root cause from the components corresponding to the at least two alarm messages according to the association relationship between the at least two alarm messages and the fault root cause probability of each component.

[0207] Optionally, the first determination module 201 includes:

[0208] A first determination sub-module, configured to determine the association degree between the at least two alarm messages according to the association degree between multiple pre-determined alarm types and the alarm types corresponding to each alarm message in the at least two alarm messages;

[0209] A second determination sub-module, configured to determine the causal relationship between the at least two alarm messages according to the pre-constructed knowledge graph and / or the pre-trained causal inference model;

[0210] where the association relationship includes the association degree and the causal relationship.

[0211] Optionally, the second determination sub-module includes:

[0212] A first inference unit, configured to infer the first causal relationship between the at least two alarm messages according to the relationships between the components in the pre-constructed knowledge graph;

[0213] and / or, a second inference unit, configured to arrange the at least two pieces of alarm information in the order of alarm time to obtain an alarm event sequence; input the alarm event sequence into a pre-trained causal inference model to obtain a second causal relationship between the at least two pieces of alarm information output by the causal inference model;

[0214] Wherein, the causal relationship between the at least two pieces of alarm information includes the first causal relationship and / or the second causal relationship.

[0215] Optionally, the third determination module 203 includes:

[0216] A first determination unit, configured to determine a first confidence level of each component as a root cause of the failure according to the probability of the root cause of the failure of each component;

[0217] An adjustment unit, configured to adjust the first confidence level according to the association relationship between the at least two pieces of alarm information to obtain a target confidence level;

[0218] A second determination unit, configured to sort the components according to the target confidence level to generate a list of root causes of the failure, and determine the target component with the highest target confidence level as the root cause of the failure.

[0219] Optionally, the root cause of failure location device 200 further includes at least one of the following:

[0220] An output module, configured to search for a corresponding handling suggestion when the target component fails from a pre-established failure handling suggestion library, and output the handling suggestion, where the target component is the component determined as the root cause of the failure from the components corresponding to the at least two pieces of alarm information;

[0221] A display module, configured to graphically display the failure location result, where the failure location result includes at least one of the following: the association relationship between the at least two pieces of alarm information, the knowledge graph, and the location process of the root cause of the failure.

[0222] Optionally, the root cause of failure location device 200 further includes:

[0223] A fourth determination module, configured to determine the association degree between every two of the multiple alarm types according to the number of co-occurrences of every two of the multiple alarm types in the historical alarm information.

[0224] Optionally, the root cause of failure location device 200 further includes:

[0225] A receiving module, configured to receive feedback information of the user on the location result of the target component as the root cause of the failure, where the feedback information includes at least one of confirmation, modification, and negation;

[0226] An adjustment module for adjusting the fault root cause prediction model according to the feedback information.

[0227] Optionally, the fault root cause prediction model is trained in the following manner:

[0228] Obtain historical alarm data, extract alarm features from the historical alarm data, and determine the component information involved in the historical alarm data; use the alarm features and the component information as model input features, input them into an initial fault root cause prediction model, and adjust the model parameters of the fault root cause prediction model according to the deviation between the model output result and the true fault root cause marked in the historical alarm data, so as to obtain the trained fault root cause prediction model;

[0229] Wherein, the alarm features include at least one of the following: alarm time, alarm type, alarm frequency, alarm duration, alarm level.

[0230] In the fault root cause location device in the embodiments of the present disclosure, when at least two alarm messages are received, according to the target information, determine the association relationship between the at least two alarm messages, where the target information includes at least one of the following: the association degree between multiple pre-determined alarm types, a pre-constructed knowledge graph, a pre-trained causal inference model, the knowledge graph includes multiple nodes and multiple edges, the nodes represent components in the system, the edges represent the relationships between the nodes, and one alarm message corresponds to one component; use the pre-trained fault root cause prediction model to determine the fault root cause probabilities of the components corresponding to the at least two alarm messages, where the fault root cause prediction model is trained based on historical alarm data; according to the association relationship between the at least two alarm messages and the fault root cause probabilities of the components, determine the fault root cause from the components corresponding to the at least two alarm messages. In this way, by deeply analyzing multiple alarm data, identifying the association relationship between alarms, and introducing a machine learning model to predict the fault root cause of multiple alarms, and finally combining the association relationship between alarms and the fault root cause prediction result, the fault root cause is located. This method can improve the accuracy and efficiency of fault location.

[0231] The fault root cause location device in the embodiments of the present disclosure can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present disclosure do not make specific limitations.

[0232] The fault root cause location device in the embodiments of the present disclosure can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present disclosure do not make specific limitations.

[0233] The fault root cause location device provided by the embodiments of the present disclosure can implement Figure 1 each process implemented by the method embodiments and can achieve the same technical effects. To avoid repetition, details are not described here again.

[0234] Optionally, as Figure 3 shown, the embodiments of the present disclosure further provide an electronic device 300, including a processor 301 and a memory 302. A program or instruction that can run on the processor 301 is stored on the memory 302. When the program or instruction is executed by the processor 301, it implements each step of the above-mentioned fault root cause location method embodiments and can achieve the same technical effects. To avoid repetition, details are not described here again.

[0235] It should be noted that the electronic devices in the embodiments of the present disclosure include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0236] The embodiments of the present disclosure further provide a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements each process of the above-mentioned fault root cause location method embodiments and can achieve the same technical effects. To avoid repetition, details are not described here again.

[0237] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.

[0238] Another embodiment of the present disclosure provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-mentioned embodiment of the root cause localization method for faults, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0239] It should be understood that the chip mentioned in the embodiments of the present disclosure may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0240] The embodiments of the present disclosure provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the root cause localization method for faults, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0241] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present disclosure is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0242] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present disclosure, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present disclosure.

[0243] The embodiments of the present disclosure have been described above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present disclosure, those of ordinary skill in the art can also make many forms without departing from the purpose of the present disclosure and the scope protected by the claims, and all of them belong to the protection scope of the present disclosure.

Claims

1. A method for locating the root cause of a fault, characterized in that: include: In the case of receiving at least two pieces of alarm information, determining the association relationship between the at least two pieces of alarm information according to target information, wherein the target information includes at least one of the following: a degree of association between a plurality of predetermined alarm types, a pre-built knowledge graph, and a pre-trained causal inference model, wherein the knowledge graph includes a plurality of nodes and a plurality of edges, wherein the nodes represent components in the system, the edges represent the relationship between the nodes, and one piece of alarm information corresponds to one component; Determine the probability of a root cause of failure of each component corresponding to the at least two pieces of alarm information by using a pre-trained root cause prediction model, wherein the root cause prediction model is trained based on historical alarm data; According to the association relationship between the at least two pieces of alarm information and the probability of a root cause of failure of each component, a root cause of failure is determined from the components corresponding to the at least two pieces of alarm information.

2. The method according to claim 1, characterized in that The determining, according to the target information, the association relationship between the at least two pieces of warning information includes: Determining the correlation between the at least two pieces of alarm information according to the correlation between the predetermined plurality of alarm types and the alarm type corresponding to each piece of alarm information in the at least two pieces of alarm information; Determine the causal relationship between the at least two pieces of warning information according to a pre-built knowledge graph and / or a pre-trained causal inference model; The correlation relationship includes the correlation degree and the causal relationship.

3. The method according to claim 2, characterized in that The determining the causal relationship between the at least two pieces of warning information according to the pre-built knowledge graph and / or the pre-trained causal inference model includes: Inferring a first causal relationship between the at least two pieces of warning information based on the relationship between the components in the pre-constructed knowledge graph; And / or, arranging the at least two pieces of alarm information in order of alarm time to obtain an alarm event sequence; inputting the alarm event sequence into a pre-trained causal inference model to obtain a second causal relationship between the at least two pieces of alarm information output by the causal inference model; The causal relationship between the at least two pieces of warning information includes the first causal relationship and / or the second causal relationship.

4. The method according to claim 1, characterized in that The determining the root cause of the fault from the components corresponding to the at least two pieces of alarm information according to the association relationship between the at least two pieces of alarm information and the probability of the root cause of the fault of each component includes: Determining, according to the probability of a root cause of failure of each component, a first confidence level of each component as a root cause of failure; According to the correlation relationship between the at least two pieces of warning information, the first confidence level is adjusted to obtain a target confidence level; The components are sorted according to the target confidence, a fault root cause list is generated, and the target component with the highest target confidence is determined as the fault root cause.

5. The method according to claim 1, characterized in that The method further comprises at least one of the following: Searching for a processing suggestion corresponding to a target component failure from a pre-established fault processing suggestion library, and outputting the processing suggestion, wherein the target component is a component determined as a root cause of the failure from the components corresponding to the at least two pieces of alarm information; The fault location result is displayed in a graphical manner, and the fault location result includes at least one of the following: the correlation relationship between the at least two alarm information, the knowledge graph and the location process of the root cause of the fault.

6. The method according to claim 1, characterized in that The method further comprises: The correlation between each two of the multiple alarm types is determined according to the number of common occurrences of each two of the multiple alarm types in the historical alarm information.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: receiving feedback information from a user regarding a positioning result in which a target component is a root cause of a fault, wherein the feedback information includes at least one of confirmation, modification, and negation, wherein the target component is a component determined as the root cause of the fault from the components corresponding to the at least two pieces of alarm information; The fault root cause prediction model is adjusted according to the feedback information.

8. The method according to any one of claims 1 to 6, characterized in that The fault root cause prediction model is trained in the following way: Acquire historical alarm data, extract alarm features from the historical alarm data, and determine component information involved in the historical alarm data; The alarm feature and the component information are used as model input features, input into an initial fault root cause prediction model, and according to the deviation between the model output result and the actual fault root cause marked in the historical alarm data, the model parameters of the fault root cause prediction model are adjusted to obtain the trained fault root cause prediction model; The alarm characteristics include at least one of the following: alarm time, alarm type, alarm frequency, alarm duration, and alarm level.

9. A fault root cause location device, characterized in that: include: A first determination module is used to determine, when receiving at least two pieces of alarm information, a correlation relationship between the at least two pieces of alarm information according to target information, wherein the target information includes at least one of the following: a correlation degree between a plurality of predetermined alarm types, a pre-built knowledge graph, and a pre-trained causal inference model, wherein the knowledge graph includes a plurality of nodes and a plurality of edges, wherein the nodes represent components in the system, the edges represent the relationship between the nodes, and one piece of alarm information corresponds to one component; A second determination module is used to determine the probability of a root cause of failure of each component corresponding to the at least two pieces of alarm information by using a pre-trained root cause prediction model, wherein the root cause prediction model is trained based on historical alarm data; The third determination module is used to determine the root cause of the fault from the components corresponding to the at least two pieces of alarm information according to the association relationship between the at least two pieces of alarm information and the root cause probability of the fault of each component.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the fault root cause locating method according to any one of claims 1 to 8 are implemented.