Alarm fault delimiting method, device and equipment based on probabilistic graph model and medium

Through the method based on the probability graph model, a probability graph model of system failure is constructed, and the failure probability and posterior probability of each node are calculated, which solves the problem of difficulty in delimiting faults of large online service systems, and achieves rapid and effective delimiting faults.

CN120075029APending Publication Date: 2025-05-30SHANGHAI QINGCHUANG INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510302355.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Due to the numerous components and complex relationships of large online service systems, it is difficult to delineate faults, resulting in alarm storms that bring huge pressure to operation and maintenance personnel and troubleshooting difficulties.

Method used

Using a method based on the probability graph model, by determining the initial fault node and the fault reachable node, a probability graph model is constructed, and the failure probability and posterior probability of each node are calculated, thereby delimiting the root cause node of the fault.

Benefits of technology

Quickly and effectively delimit the root cause nodes of the fault, reduce the processing pressure of operation and maintenance personnel, and save troubleshooting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075029A_ABST
    Figure CN120075029A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an alarm fault delimiting method and device based on a probabilistic graph model, equipment and a medium. The method comprises the following steps: determining an initial fault node and a fault reachable node according to fault alarm data; wherein the fault reachable node refers to a node associated with an initial fault node; determining a fault probability, a transition probability and an observation probability of each node, and constructing a probability graph model; determining alarm evidence according to the fault alarm data, and determining a posterior probability that each node is a fault root cause according to the alarm evidence and a probability graph model; and according to the posterior probability that each node is the fault root cause, determining the root cause node with the fault from the initial fault node and the fault reachable node. By adopting the technical scheme of the embodiment of the invention, the root cause node which is most likely to have a fault is inferred based on the probabilistic graph model, and precious troubleshooting time is won for operation and maintenance personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of communication technologies, and in particular, to an alarm fault delimitation method, device, equipment and medium based on a probabilistic graph model. Background Art

[0002] With the development of information technology, the architecture of online service systems has gradually changed from the original monolithic architecture to a distributed architecture. A large online service system usually includes several or even hundreds or thousands of components, which cooperate with each other to provide a unified service to users. To ensure the high stability and high availability of the online service system and minimize the economic losses caused by service anomalies as much as possible, operation and maintenance personnel will define various alarm rules to monitor the operation of each service component.

[0003] Due to the large number of components in a large online service system and the very complex association and dependency relationships among components, a fault in one component will cause its related components to also report alarms, resulting in a large number of alarms being reported simultaneously in a short period of time, forming an alarm storm, bringing huge processing pressure to operation and maintenance personnel and being extremely unfavorable for troubleshooting the root cause of the fault. To quickly delimit the fault and find the root cause, operation and maintenance personnel often define many rules to handle various alarms, but these rules are difficult to reuse and have a strong dependence on experience; there are also some root cause location based on algorithms, but they have high requirements for the quality of alarms. In fact, faulty components often do not send alarms, and existing algorithms will fail.

[0004] Therefore, how to quickly and effectively delimit the fault is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] The embodiments of the present invention provide an alarm fault delimitation method, device, equipment and medium based on a probabilistic graph model to infer the root cause node most likely to have a fault based on the probabilistic graph model, and strive for valuable troubleshooting time for operation and maintenance personnel.

[0006] In a first aspect, the embodiments of the present invention provide an alarm fault delimitation method based on a probabilistic graph model, including:

[0007] Determine an initial fault node and fault reachable nodes according to fault alarm data; wherein, the fault reachable nodes refer to the nodes associated with the initial fault node;

[0008] Determine the fault probability, transition probability and observation probability of each node, and construct a probabilistic graph model;

[0009] Determine alarm evidence according to fault alarm data, and determine the posterior probability of each node being the root cause of the fault according to the alarm evidence and the probabilistic graph model;

[0010] Determine the root cause node of the failure from the initial failure node and the failure-reachable nodes according to the posterior probability of each node being the root cause of the failure.

[0011] In a second aspect, an embodiment of the present invention further provides an alarm fault delimitation device based on a probabilistic graph model, including:

[0012] A node determination module, configured to determine an initial failure node and failure-reachable nodes according to failure alarm data; wherein, the failure-reachable nodes refer to nodes associated with the initial failure node;

[0013] A probabilistic graph model construction module, configured to determine the failure probability, transition probability, and observation probability of each node, and construct a probabilistic graph model;

[0014] A node root cause probability prediction module, configured to determine alarm evidence according to failure alarm data, and determine the posterior probability of each node being the root cause of the failure according to the alarm evidence and the probabilistic graph model;

[0015] A root cause node determination module, configured to determine the root cause node of the failure from the initial failure node and the failure-reachable nodes according to the posterior probability of each node being the root cause of the failure.

[0016] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes:

[0017] One or more processors;

[0018] A storage device, configured to store one or more programs;

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the alarm fault delimitation method based on a probabilistic graph model according to any embodiment of the present invention.

[0020] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the alarm fault delimitation method based on a probabilistic graph model according to any embodiment of the present invention.

[0021] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the alarm fault delimitation method based on a probabilistic graph model according to any embodiment of the present invention.

[0022] An embodiment of the present invention provides a method, device, electronic device, and storage medium for alarm fault delimitation based on a probabilistic graphical model. The method includes determining an initial fault node and fault reachable nodes based on fault alarm data; determining the fault probability, transition probability, and observation probability of each node, and constructing a probabilistic graphical model; determining alarm evidence based on the fault alarm data, and determining the posterior probability of each node being the root cause of the fault based on the alarm evidence and the probabilistic graphical model; and determining the root cause node of the fault from the initial fault node and the fault reachable nodes based on the posterior probability of each node being the root cause of the fault. By adopting the technical solution of the embodiment of the present invention, the root cause node most likely to have a fault is inferred based on the probabilistic graphical model, saving valuable troubleshooting time for maintenance personnel. Description of the Drawings

[0023] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0024] Figure 1 is a flowchart of a method for alarm fault delimitation based on a probabilistic graphical model provided in an embodiment of the present invention;

[0025] Figure 2 is a flowchart of another method for alarm fault delimitation based on a probabilistic graphical model provided in an embodiment of the present invention;

[0026] Figure 3 is a schematic diagram of a fault topology provided in an embodiment of the present invention;

[0027] Figure 4 is a flowchart of yet another method for alarm fault delimitation based on a probabilistic graphical model provided in an embodiment of the present invention;

[0028] Figure 5 is a schematic structural diagram of an alarm fault delimitation device based on a probabilistic graphical model provided in an embodiment of the present invention;

[0029] Figure 6 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Embodiments

[0030] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all structures.

[0031] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations (or steps) can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.

[0032] Among them, the acquisition, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, or models may be mentioned. They should be considered exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.

[0034] Embodiment 1

[0035] Figure 1 It is a flowchart of an alarm fault localization method based on a probabilistic graph model provided in an embodiment of the present invention. This embodiment is applicable to the situation of alarm fault localization based on a probabilistic graph model. The method of this embodiment can be executed by an alarm fault localization device based on a probabilistic graph model, and this device can be implemented in a hardware and / or software manner. This device can be configured in a server for alarm fault localization based on a probabilistic graph model. The method specifically includes the following steps:

[0036] S110. Determine the initial fault node and the fault reachable nodes according to the fault alarm data.

[0037] Among them, the fault alarm data can refer to the alarm data in the most recent period of time when the fault occurs as the fault alarm data after the fault occurs.

[0038] The initial failure node may refer to the failure node associated with the failure alarm data. The failure alarm data includes the attribute information of the monitored node, and the initial failure node is determined based on the attribute information; wherein, the attribute information includes the node name and the node Internet protocol address; the node may refer to a component or a device, including but not limited to a physical machine and a switch. Taking the failure of an online service system as an example, the failure alarm data at the time of failure is obtained, and the initial failure node, that is, the initial failure component, is determined based on the node name and the node Internet protocol address (node IP address) associated with the failure alarm data.

[0039] The failure reachable node may refer to the node directly or indirectly associated with the initial failure node. For example, if node A is associated with node B, node B is associated with node C, and node A is the initial failure node, then node B and node C are failure reachable nodes. Or, if node A is associated with node B, node A is associated with node C, and node A is the initial failure node, then node B and node C are failure reachable nodes.

[0040] S120. Determine the failure probability, transition probability, and observation probability of each node, and construct a probabilistic graph model.

[0041] The failure probability may refer to the probability that a node is the root cause of the failure. The transition probability may refer to the probability that a node failure causes other associated nodes to fail. The observation probability may refer to the probability that a node issues an alarm in the failure or normal state.

[0042] The probabilistic graph model may refer to a theory that represents the probabilistic dependence relationship of variables using a graph. The probabilistic graph model includes but is not limited to a Bayesian network model, a Markov network model, and a hidden Markov network model.

[0043] In the embodiment of the present invention, a probabilistic graph model is constructed based on the failure probability, transition probability, and observation probability of the initial failure node and the failure reachable nodes.

[0044] S130. Determine the alarm evidence based on the failure alarm data, and determine the posterior probability that each node is the root cause of the failure based on the alarm evidence and the probabilistic graph model.

[0045] Among them, the alarm evidence may refer to the node observation result, that is, the alarm situation of each node. For example, if nodes n1 and n2 issue alarms and node n3 does not issue an alarm, taking values [0, 1], the corresponding alarm evidence is n1 = 1, n2 = 1, n3 = 0.

[0046] The posterior probability can refer to the probability that is re - revised after obtaining the information of the "result". For example, in Bayes' formula, it is the "cause" in the problem of "seeking the cause from the result". For example, given that node n1 fails, the probability that node n1 and its associated nodes are the root causes of the failure is inferred through a probabilistic graphical model to determine the root cause of the failure of node n1.

[0047] S140. Determine the root cause node of the failure from the initial failure node and the failure - reachable nodes according to the posterior probability of each node being the root cause of the failure.

[0048] Among them, after determining the posterior probability of each node, select the node with the largest posterior probability value as the root cause node of the failure. For example, if the posterior probability of node n1 is A, the posterior probability of node n2 is B, and the posterior probability of node n3 is C, and A is the largest, then determine that node n1 is the root cause node of the failure.

[0049] An alarm fault location method based on a probabilistic graphical model is provided in an embodiment of the present invention. By determining an initial failure node and failure - reachable nodes according to fault alarm data; where the failure - reachable nodes refer to the nodes associated with the initial failure node; determining the fault probability, transition probability, and observation probability of each node, and constructing a probabilistic graphical model; determining alarm evidence according to the fault alarm data, and determining the posterior probability of each node being the root cause of the failure according to the alarm evidence and the probabilistic graphical model; determining the root cause node of the failure from the initial failure node and the failure - reachable nodes according to the posterior probability of each node being the root cause of the failure. By adopting the technical solution of the embodiment of the present invention, the root cause node most likely to fail is inferred based on the probabilistic graphical model, saving valuable troubleshooting time for operation and maintenance personnel.

[0050] Embodiment Two

[0051] Figure 2 It is a flowchart of an alarm fault location method based on a probabilistic graphical model provided in an embodiment of the present invention. The embodiment of the present invention further optimizes the foregoing embodiment on the basis of the above - mentioned embodiment, and the embodiment of the present invention can be combined with each optional solution in one or more of the above - mentioned embodiments. As Figure 2 shown, the alarm fault location method based on a probabilistic graphical model provided in an embodiment of the present invention may include the following steps:

[0052] S210. Determine an initial failure node and failure - reachable nodes according to the fault alarm data.

[0053] Among them, obtain the alarm data in the most recent period after the failure occurs, and determine the initial failure node and failure - reachable nodes according to the alarm data.

[0054] As an optional but non-limiting implementation, determining the initial fault node and the fault reachable nodes based on the fault alarm data includes, but is not limited to, steps A1 - A2:

[0055] Step A1: Obtain the fault alarm data, and determine the initial fault node based on the fault alarm data.

[0056] Step A2: Map the fault alarm data to the fault propagation topology graph based on the Configuration Management Database, and determine the fault reachable nodes according to the fault propagation topology graph; wherein, the Configuration Management Database records the association relationships between nodes.

[0057] Among them, the Configuration Management Database (CMDB) stores and manages various configuration information of devices in the enterprise IT architecture. It is closely associated with all service support and service delivery processes, supports the operation of these processes, gives play to the value of the configuration information, and at the same time depends on the relevant processes to ensure the accuracy of the data. For example, virtual machines run on physical machines, and physical machines are connected to switches. A fault propagation topology graph is constructed based on the relationships between virtual machines, physical machines, and switches.

[0058] After determining the fault propagation topology graph based on the Configuration Management Database, determine the initial fault node according to the fault alarm data, and determine the corresponding fault reachable nodes according to the fault propagation topology graph. For example, there is the fault alarm data shown in Table 1. According to the fault alarm data, the initial fault nodes can be determined as n1 and n2. After determining the initial fault nodes, a fault propagation topology graph based on the Configuration Management Database can be constructed according to the fault alarm data, and the fault propagation topology graph about nodes n1 and n2 as shown in Figure 3 (a) can be obtained. According to Table 1, the initial fault nodes are determined as n1 and n2. According to the fault propagation topology graph shown in Figure 3 (a), the fault reachable node can be determined as n0.

[0059] Table 1 Fault Alarm Data

[0060] Time Content Node 2025-01-01 00:00 Host n1 is unreachable n1 2025-01-01 00:00 Host n2 is unreachable n2

[0061] S220. Construct a fault subgraph based on the initial fault node and the fault reachable nodes, and determine the node information included in the fault subgraph.

[0062] Among them, the fault subgraph may refer to a topology graph including the node information of each node. The node information included in the fault subgraph includes the fault status, node alarm situation, and node root cause situation of each node. Taking the construction of a Bayesian network fault subgraph as an example, see Figure 3(b), the N-type variables (n0, n1, n2) indicate whether the corresponding nodes are faulty, the A-type variables (a0, a1, a2) represent the alarm conditions of the corresponding nodes, and the R-type variables (r0, r1, r2) represent whether the corresponding nodes are the root causes.

[0063] S230. Determine the failure probability, transition probability, and observation probability of each node according to the node information included in the faulty subgraph.

[0064] Among them, taking the probability graph model as the Bayesian network model as an example, each node derives three variables ri, ni, and ai in the Bayesian network.

[0065] Among them, ri is the initial variable, taking values 0 or 1, representing the probability that the corresponding node is the root cause, that is, the failure probability; for example, the example value corresponding to r1 is [0.9, 0.1], indicating that the probability that node r1 is the root cause is 0.1, and the probability that node r1 is not the root cause is 0.9.

[0066] ni represents the failure state of the corresponding node, taking values 0, 1; ni depends on ri and other N-type variables, representing the probability that the associated nodes fail after the corresponding node fails. Taking the determination of the transition probability of n2 as an example, the failure probability depends on r2 and n0; construct a CPD table for n2 to determine the corresponding transition probability. The CPD table (Conditional Probability Distribution) can refer to a table used to describe the conditional probability distribution of each node in the Bayesian network under different value combinations of all its parent nodes. For example, referring to Table 2, construct a CPD table for n2, and the example values of the corresponding transition probability are [[0.9, 0.1, 0.9, 0.2], [0.1, 0.9, 0.1, 0.8]].

[0067] Table 2 CPD Table

[0068] n2 n0 r2 Transition probability 0 0 0 0.9 0 0 1 0.1 0 1 0 0.9 0 1 1 0.2 1 0 0 0.1 1 0 1 0.9 1 1 0 0.1 1 1 1 0.8

[0069] ai represents the alarm condition of the corresponding node, that is, the observation result, taking values 0, 1, and depends on ni. Determine the corresponding observation probability according to the alarm condition of ai and the failure state of ni.

[0070] S240. Assign weights to the failure probability, transition probability, and observation probability of each node, and construct a probability graph model according to the failure probability, transition probability, observation probability of each node, and the corresponding weights.

[0071] Among them, referring to Figure 4, taking the construction of a Bayesian network model as an example, after determining the failure probabilities, transition probabilities, and observation probabilities of each node, weights are assigned to the failure probabilities, transition probabilities, and observation probabilities of each node, and a Bayesian network model is established based on the pgmpy package.

[0072] S250. Determine alarm evidence based on the fault alarm data, and determine the posterior probability of each node being the root cause of the fault based on the alarm evidence and the probabilistic graphical model.

[0073] Among them, determine alarm evidence based on the fault alarm data, and determine the posterior probability of each node being the root cause of the fault based on the constructed probabilistic graphical model.

[0074] As an optional but non-limiting implementation manner, the determining alarm evidence based on the fault alarm data and determining the posterior probability of each node being the root cause of the fault based on the alarm evidence and the probabilistic graphical model includes, but is not limited to, steps B1 - B2:

[0075] Step B1: Determine whether each node has an alarm based on the fault alarm data, and generate alarm evidence.

[0076] Step B2: Input the alarm evidence into the probabilistic graphical model to infer the posterior probability of each node being the root cause of the fault.

[0077] Among them, it can be observed from Table 1 that n1 and n2 have alarms, that is, the corresponding alarm evidence is a1 = 1 and n2 = 1. Input the alarm evidence as evidence into the probabilistic graphical model to infer the posterior probability of r0, r1, and r2 being the root cause of the fault.

[0078] S260. Determine the root cause node of the fault from the initial fault node and the fault-reachable nodes based on the posterior probability of each node being the root cause of the fault.

[0079] Among them, based on the alarm evidence observed in Table 1, the posterior probabilities of r0, r1, and r2 obtained by inference are shown in Table 3. Based on this, determine the root cause node of the fault from the initial fault node and the fault-reachable nodes. As shown in Table 3, take the values of r0, r1, and r2 corresponding to the maximum posterior probability value, and thus the root cause node of the fault is obtained as r0.

[0080] Table 3 Posterior Probability Table of Each Node

[0081] r0 r1 r2 Posterior probability 0 0 0 0.26 0 0 1 0.06 0 1 0 0.06 0 1 1 0.01 1 0 0 0.4 1 0 1 0.09 1 1 0 0.09 1 1 1 0.02

[0082] As an optional but non-limiting implementation manner, the determining the root cause node of the fault from the initial fault node and the fault-reachable nodes based on the posterior probability of each node being the root cause of the fault includes, but is not limited to, steps C1 - C2:

[0083] Step C1: Sort the posterior probabilities of each node being the root cause of the fault in a preset sorting manner, and determine the target posterior probability ranked first; wherein, the preset sorting manner is to sort the posterior probabilities from large to small.

[0084] Step C2: Determine the root cause node of the fault from the initial fault nodes and the fault reachable nodes according to the target posterior probability.

[0085] Optionally, after determining the posterior probabilities of each node being the root cause of the fault, the posterior probabilities can be arranged in descending order to determine the maximum posterior probability, so as to determine the root cause node of the fault.

[0086] The embodiment of the present invention provides an alarm fault boundary determination method based on a probabilistic graph model. By determining the initial fault nodes and the fault reachable nodes according to the fault alarm data; constructing a fault subgraph based on the initial fault nodes and the fault reachable nodes, and determining the node information included in the fault subgraph; determining the fault probability, transition probability and observation probability of each node according to the node information included in the fault subgraph; assigning weights to the fault probability, transition probability and observation probability of each node, and constructing a probabilistic graph model; determining the alarm evidence according to the fault alarm data, and inputting the alarm evidence into the probabilistic graph model to determine the posterior probability of each node being the root cause of the fault, so as to determine the root cause node of the fault. The relatively general and efficient fault boundary determination method proposed in the embodiment of the present invention uses the configuration management database to construct a fault propagation topology graph and a fault subgraph, and infers the root cause node most likely to have a fault based on the probabilistic graph model, saving valuable troubleshooting time for maintenance personnel.

[0087] Embodiment III

[0088] Figure 5 It is a schematic structural diagram of an alarm fault boundary determination device based on a probabilistic graph model provided in the embodiment of the present invention. The technical solution of this embodiment can be applied to the situation of alarm fault boundary determination based on a probabilistic graph model. The device can be implemented by software and / or hardware, and is generally integrated in any electronic device with network communication functions. The electronic device includes but is not limited to: devices such as servers, computers, and personal digital assistants. As Figure 5 shown, the alarm fault boundary determination device based on a probabilistic graph model provided in this embodiment may include: a node determination module 510, a probabilistic graph model construction module 520, a node root cause probability prediction module 530, and a root cause node determination module 540; wherein,

[0089] The node determination module 510 is used to determine the initial fault nodes and the fault reachable nodes according to the fault alarm data; wherein, the fault reachable nodes refer to the nodes associated with the initial fault nodes;

[0090] The probability graph model construction module 520 is used to determine the failure probability, transition probability, and observation probability of each node, and construct a probability graph model;

[0091] The node root cause probability prediction module 530 is used to determine alarm evidence based on the failure alarm data, and determine the posterior probability of each node being the root cause of the failure according to the alarm evidence and the probability graph model;

[0092] The root cause node determination module 540 is used to determine the root cause node of the failure from the initial failure node and the failure reachable nodes according to the posterior probability of each node being the root cause of the failure.

[0093] Based on the above embodiments, optionally, the node determination module is specifically used for:

[0094] Obtain the failure alarm data, and determine the initial failure node according to the failure alarm data;

[0095] Map the failure alarm data to the failure propagation topology graph based on the configuration management database, and determine the failure reachable nodes according to the failure propagation topology graph; wherein, the configuration management database records the association relationship between nodes.

[0096] Based on the above embodiments, optionally, the failure alarm data includes the attribute information of the monitored node, and the initial failure node is determined according to the attribute information; wherein, the attribute information includes the node name and the node Internet protocol address.

[0097] Based on the above embodiments, optionally, the probability graph model construction module is specifically used for:

[0098] Construct a failure subgraph according to the initial failure node and the failure reachable nodes, and determine the node information included in the failure subgraph; wherein, the node information included in the failure subgraph includes the failure state, node alarm situation, and node root cause situation of each node;

[0099] Determine the failure probability, transition probability, and observation probability of each node according to the node information included in the failure subgraph;

[0100] Construct a probability graph model according to the failure probability, transition probability, and observation probability of each node.

[0101] Based on the above embodiments, optionally, the node root cause probability prediction module is specifically used for:

[0102] Determine whether each node has an alarm according to the failure alarm data, and generate alarm evidence;

[0103] Input the alarm evidence into a probabilistic graphical model to infer the posterior probability of each node being the root cause of the fault.

[0104] Based on the above embodiments, optionally, the root cause node determination module is specifically configured to:

[0105] Sort the posterior probabilities of each node being the root cause of the fault in a preset sorting manner, and determine the target posterior probability ranked first; wherein, the preset sorting manner is to sort the posterior probabilities from large to small.

[0106] Determine the root cause node of the fault from the initial fault nodes and the fault-reachable nodes according to the target posterior probability.

[0107] The alarm fault boundary determination device based on a probabilistic graphical model provided in the embodiments of the present invention can execute the alarm fault boundary determination method based on a probabilistic graphical model provided in any of the above embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the alarm fault boundary determination method based on a probabilistic graphical model. For the detailed process, refer to the relevant operations of the alarm fault boundary determination method based on a probabilistic graphical model in the foregoing embodiments.

[0108] Embodiment 4

[0109] Figure 6 It is a schematic structural diagram of an electronic device provided in the embodiments of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0110] As Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0111] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0112] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the alarm fault delimitation method based on a probabilistic graphical model.

[0113] In some embodiments, the alarm fault delimitation method based on a probabilistic graphical model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the alarm fault delimitation method based on a probabilistic graphical model described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the alarm fault delimitation method based on a probabilistic graphical model by any other appropriate means (such as by means of firmware).

[0114] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0115] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0116] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0118] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0119] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0120] Example Five

[0121] An embodiment of the present invention further provides a computer program product, including a computer program which, when executed by a processor, implements the method for alarm fault delimitation based on a probabilistic graphical model provided in any embodiment of the present application.

[0122] In the process of implementing the computer program product, computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0123] It should be understood that the various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0124] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for delimiting alarm faults based on a probabilistic graphical model, characterized in that: The method comprises: Determine the initial fault node and the fault reachable node according to the fault alarm data; wherein the fault reachable node refers to the node associated with the initial fault node; Determine the failure probability, transition probability, and observation probability of each node, and build a probabilistic graph model; Determine alarm evidence based on the fault alarm data, and determine the posterior probability that each node is the root cause of the fault based on the alarm evidence and the probability graph model; According to the posterior probability that each node is the root cause of the fault, the root cause node of the fault is determined from the initial fault node and the fault reachable nodes.

2. The method according to claim 1, characterized in that The determining of the initial fault node and the fault reachable node according to the fault alarm data includes: Acquire fault alarm data, and determine the initial fault node according to the fault alarm data; The fault alarm data is mapped to a fault propagation topology map based on a configuration management database, and the fault-reachable nodes are determined according to the fault propagation topology map; wherein the configuration management database records the association relationship between nodes.

3. The method according to claim 1, characterized in that The fault alarm data includes attribute information of the monitored node, and the initial fault node is determined based on the attribute information; wherein the attribute information includes the node name and the node Internet Protocol address.

4. The method according to claim 1, characterized in that: Determining the failure probability, transition probability and observation probability of each node and constructing a probability graph model includes: Construct a fault subgraph based on the initial fault node and the fault reachable nodes, and determine the node information contained in the fault subgraph; wherein the node information contained in the fault subgraph includes the fault status of each node, the node alarm status and the node root cause status; Determine the failure probability, transition probability and observation probability of each node according to the node information contained in the fault subgraph; Weights are assigned to the failure probability, transition probability and observation probability of each node, and a probabilistic graph model is constructed based on the failure probability, transition probability, observation probability of each node and the corresponding weights.

5. The method according to claim 1, characterized in that Determining the alarm evidence based on the fault alarm data, and determining the posterior probability that each node is the root cause of the fault based on the alarm evidence and the probability graph model, includes: Determine whether each node has an alarm based on the fault alarm data and generate alarm evidence; The alarm evidence is input into a probabilistic graphical model to infer the posterior probability of each node being the root cause of the fault.

6. The method according to claim 1, characterized in that The step of determining the root cause node of the fault from the initial fault node and the fault reachable nodes based on the posterior probability that each node is the root cause of the fault includes: The posterior probabilities of each node being the root cause of the fault are sorted in a preset sorting manner, and the target posterior probability ranked first is determined; wherein the preset sorting manner is to sort the posterior probabilities from large to small; According to the target a posteriori probability, the root cause node of the fault is determined from the initial fault node and the fault reachable nodes.

7. An alarm fault demarcation device based on a probabilistic graphical model, characterized in that: The device comprises: A node determination module, used to determine an initial fault node and a fault-reachable node according to the fault alarm data; wherein the fault-reachable node refers to a node associated with the initial fault node; The probability graph model building module is used to determine the failure probability, transfer probability and observation probability of each node and build a probability graph model; A node root cause probability prediction module is used to determine the alarm evidence based on the fault alarm data, and determine the posterior probability of each node being the root cause of the fault based on the alarm evidence and the probability graph model; The root cause node determination module is used to determine the root cause node of the fault from the initial fault node and the fault reachable nodes according to the posterior probability that each node is the root cause of the fault.

8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the alarm fault demarcation method based on the probabilistic graphical model as described in any one of claims 1-6.

9. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the alarm fault demarcation method based on the probabilistic graphical model as described in any one of claims 1 to 6 when executed by a computer processor.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the alarm fault demarcation method based on a probabilistic graphical model according to any one of claims 1 to 6.

Citation Information

Cited By

  • Alarm event root cause analysis method, device and equipment based on large language model and service topology

    CN120994446A