Fault root cause positioning method and device and related products
By acquiring network alarm data and topology resource data, and using a fault causal reasoning model for slicing and filtering, the problems of inaccurate and time-consuming fault root cause localization in existing technologies are solved, achieving fast and accurate fault root cause localization and improving processing efficiency.
Patent Information
- Application Number
- CN202411460477.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-10-18
AI Technical Summary
Existing fault location methods based on expert experience cannot be applied to complex fault scenarios, resulting in inaccurate root cause location, long processing time, and low efficiency.
By acquiring network alarm data and topology resource data, and using a fault causal reasoning model to perform time and space dimension slicing, filtering, and reasoning, the root cause nodes and causal relationships of the faults are determined.
It enables rapid and accurate fault root cause location, improving fault handling efficiency.
Smart Images

Figure CN119520247B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network security, and in particular to a fault root cause positioning method and device and related products. BACKGROUND
[0002] With the increase of network communication services, the pressure on network element devices of operators is increasing, the number of network faults is rapidly growing, and network fault root cause positioning has become an indispensable link. At present, the existing technology usually adopts a fault positioning method based on expert experience to position the network fault root cause. However, the fault positioning method based on expert experience cannot be applied to complex fault scenarios, resulting in inaccurate and time-consuming fault root cause positioning and low efficiency. SUMMARY
[0003] The embodiments of the present application aim to provide a fault root cause positioning method, device and related products to solve the problem of being unable to quickly and accurately locate the fault root cause in the related art.
[0004] To solve the above technical problems, the embodiments of the present application are implemented as follows:
[0005] In a first aspect, the embodiments of the present application provide a fault root cause positioning method, comprising:
[0006] obtaining network alarm data and corresponding network topology resource data;
[0007] performing slice processing on the network alarm data and the corresponding network topology resource data in the time dimension and the space dimension through a fault cause and effect reasoning model to obtain an initial alarm correlation data group;
[0008] performing screening and reasoning processing on the initial alarm correlation data group through the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data.
[0009] In a second aspect, the embodiments of the present application provide a fault root cause positioning device, comprising:
[0010] an acquisition module configured to obtain network alarm data and corresponding network topology resource data;
[0011] a processing module configured to perform slice processing on the network alarm data and the corresponding network topology resource data in the time dimension and the space dimension through a fault cause and effect reasoning model to obtain an initial alarm correlation data group;
[0012] a reasoning module configured to perform screening and reasoning processing on the initial alarm correlation data group through the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data.
[0013] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory complete mutual communication through the bus; the memory is used for storing a computer program; the processor is used for executing the program stored on the memory to realize the steps of the fault root cause positioning method according to the first aspect.
[0014] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps of the fault root cause positioning method according to the first aspect.
[0015] In a fifth aspect, a computer program product is provided, comprising a computer program, and the computer program is executed by a processor to realize the steps of the fault root cause positioning method according to the first aspect.
[0016] As can be seen from the technical solutions provided by the embodiments of the present application, first, the network alarm data and the corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are sliced in the time dimension and the space dimension by the fault cause and effect reasoning model to obtain an initial alarm correlation data set; and then, the initial alarm correlation data set is screened and reasoned by the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data. As can be seen, by the embodiments of the present application, the network alarm data and the corresponding network topology resource data acquired can be sliced in the time dimension and the space dimension by the fault cause and effect reasoning model to obtain an initial alarm correlation data set with time information and space information, and the corresponding alarm propagation path can be obtained by screening and reasoning the initial alarm correlation data set with time information and space information, so as to determine the corresponding fault cause and effect relationship, locate the fault root cause node, realize fast and accurate positioning of the fault root cause, and thus improve the efficiency of fault processing. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0018] Figure 1 A flowchart of a fault root cause positioning method provided by the embodiments of the present application is shown in the figure.
[0019] Figure 2An alarm space-time correlation relationship schematic diagram provided for an embodiment of the present application;
[0020] Figure 3 A module composition schematic diagram of a fault root cause positioning device provided for an embodiment of the present application;
[0021] Figure 4 A structure schematic diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0022] An abnormal request monitoring method and device are provided in an embodiment of the present application.
[0023] In order to enable personnel in the technical field to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work should belong to the protection scope of the present application.
[0024] Figure 1 A flowchart of a fault root cause positioning method provided for an embodiment of the present application, an execution subject of the method can be a server, wherein the server can be an independent server, can be a server cluster composed of multiple servers, can also be a cloud server for cloud computing processing, and the server can be a server capable of network operation processing. As shown in the figure, the method can specifically include the following steps: Figure 1
[0025] Step S102, network alarm data and corresponding network topology resource data are acquired;
[0026] Step S104, the network alarm data and the corresponding network topology resource data are subjected to slicing processing in time dimension and space dimension through a fault cause and effect reasoning model, to obtain an initial alarm correlation data group;
[0027] Step S106, the initial alarm correlation data group is subjected to screening and reasoning processing through the fault cause and effect reasoning model, to obtain a fault root cause node and corresponding fault cause and effect relationship data.
[0028] From the technical solutions provided by the embodiments of the application, it can be seen that, in the embodiments of the application, firstly, network alarm data and corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are processed in time dimension and space dimension by a fault causal reasoning model to obtain an initial alarm correlation data set; and then, the initial alarm correlation data set is processed by screening and reasoning by the fault causal reasoning model to obtain a fault root cause node and corresponding fault causal relationship data. It can be seen that, by the embodiments of the application, the acquired network alarm data and corresponding network topology resource data can be processed in time dimension and space dimension by a fault causal reasoning model to obtain an initial alarm correlation data set with time information and space information, and the initial alarm correlation data set with time information and space information can be processed by screening and reasoning to obtain a corresponding alarm propagation path, to determine the corresponding fault causal relationship and locate the fault root cause node, so that the fault root cause can be quickly and accurately located, and the efficiency of fault processing is improved.
[0029] In the step S102, the network alarm data and the corresponding network topology resource data are acquired.
[0030] The network alarm data refers to the data generated by various devices and systems in the network management system, which is used to indicate the existence of abnormalities or potential problems in the network. These data are crucial for network administrators, as they can help quickly locate problems, assess the scope of impact, and take appropriate measures to restore normal operation of the network. Alarm data, for example, includes basic information of alarms, alarm types and description information, and related system and device information. Among them, the basic information includes: alarm ID, alarm time and source address / target address. The alarm ID is a unique identifier for each alarm, used to distinguish different alarm events. The alarm time is the timestamp of the alarm occurrence, usually accurate to seconds or milliseconds, which helps to quickly locate the time point of the problem occurrence. The source address / target address is the source IP address and target IP address that triggers the alarm, which helps to identify the source of the attack or the affected system.
[0031] In the embodiments of the present application, the network alarm data can be alarm data in a 5G (5th Generation Mobile Communication Technology) core network, and the network topology resource data can be network topology resource data corresponding to a cross-domain 5G core network. Network topology is an abstract representation method for describing the physical or logical connection manner and layout between various devices (such as computers, routers, switches, etc.) in a network, and is an application of graph theory, in which different network devices are modeled as nodes, and the connections between devices are modeled as links or lines between nodes. Network topology describes the physical or logical connection relationship between various nodes (such as computers, routers, switches, etc.) in the network, and the flow of data in the network.
[0032] In the embodiments of the present application, the manner of obtaining network alarm data can be obtained through a security monitoring system and tool, obtained through a threat intelligence sharing platform, obtained through an API (Application Programming Interface) interface, or obtained through data sharing and cooperation.
[0033] In the above step S104, the network alarm data and the corresponding network topology resource data are sliced in the time dimension and the space dimension by the fault causal reasoning model to obtain an initial alarm correlation data set.
[0034] In the embodiments, the network data and the related network topology resources are sliced according to the preset time step to obtain the initial alarm correlation data set with time information and space information by setting the preset time step.
[0035] For specific training methods of the fault causal reasoning model in the embodiments of the present application, refer to the related steps or embodiments below, which will not be repeated here.
[0036] In one embodiment, the network alarm data and the corresponding network topology resource data are sliced in the time dimension and the space dimension by the fault causal reasoning model to obtain an initial alarm correlation data set, including:
[0037] According to the preset time step, the fault causal reasoning model obtains a sub-topology alarm set in each preset time step; the sub-topology is a region formed by network nodes with alarms; and the sub-topology alarm set is a set formed by all alarms existing in the sub-topology within the preset time step;
[0038] Based on each sub-topology alarm set, the fault causal reasoning model generates an initial alarm correlation data set.
[0039] In the embodiments of the present application, first, the obtained network alarm data and corresponding network topology resource data are input into the trained fault causal reasoning model, and through the fault causal reasoning model, the maximum relevant area formed by the network nodes with alarms is obtained according to the preset time step, and all alarms existing in the maximum relevant area are obtained, and the set formed by all alarms in the maximum relevant area is used to obtain the sub-topology alarm set in each preset time step. Then, through the fault causal reasoning model, based on each sub-topology alarm set, an initial alarm correlation data group is generated.
[0040] By setting the preset time step, the network alarm data and the related network topology resource data are sliced in the time dimension and the space dimension, the maximum relevant area formed by the network nodes with alarms in the preset time step can be determined, and all alarms existing in the maximum relevant area can be determined. The preset time step can be the time length from the arrival of the first alarm.
[0041] In the network topology, if an alarm occurs on a network node, the network node is said to be activated or lit up. The maximum relevant area formed by the lit-up node is a sub-topology. It can be understood that the sub-topology will continuously expand or remain unchanged in the space dimension as time goes on. An alarm on a network node can affect all other network nodes related to the network node. In the preset time step, the related network nodes may or may not be affected by the alarm. If the related network nodes are affected by the alarm and also generate alarms, the sub-topology will expand in the space dimension.
[0042] In one example, assuming that the preset time step is 15 minutes, node 1 is detected to have an alarm at time 0, node 2 is detected to have an alarm 10 minutes after time 0, node 2 is closely related to node 1 in performance, and no other alarms are found after 15 minutes of time 0. Then, in the 15 minutes, the sub-topology formed by node 1 expands to include node 2 in the space dimension, and the set formed by the above two alarms is the sub-topology alarm set.
[0043] The expansion of the sub-topology in the time dimension and the space dimension can also form an alarm space-time correlation relationship graph. In the alarm space-time correlation relationship graph, each alarm data is marked with time and location (i.e., network device or node), forming a multi-dimensional alarm data set. Based on this, through the fault causal reasoning model, based on each sub-topology alarm set, an initial alarm correlation data group is generated. The initial alarm correlation data group can be in the form of an alarm space-time correlation relationship graph, and the initial alarm correlation data group can include alarm time, alarm location, alarm cause, and alarm impact.
[0044] In one example, Figure 2 An alarm space-time correlation relationship diagram provided by an embodiment of the present application is shown, as shown in Figure 2 As shown, the sub-topology includes a network element layer, a network cloud, and a data communication layer, and nodes 1 to 13 exist between the network element layer, the network cloud, and the data communication layer. An alarm of PDU session establishment success rate degradation and UPF node unreachability is generated in node 2 of the network element layer, then an alarm of host port in error packet rate being too high is generated in related node 9 in the network cloud, then an interface terminal alarm is generated in related node 12 of the data communication layer, and alarms of message forwarding exception and interface interruption are generated in node 13. The multiple nodes of the network cloud and the data communication layer are closely related to the nodes in the network element layer, when an alarm is generated in a node in the network element layer, an alarm is also generated in the related nodes in the network cloud and the data communication layer, the alarms generated between the network element layer, the network cloud, and the data communication layer have a certain causal relationship, forming alarm causal relationship data, and within a certain time range, the alarm causal relationship data has space-time correlation.
[0045] In the step S106, the initial alarm correlation data set is filtered and processed by the fault causal reasoning model to obtain the fault root cause node and the corresponding fault causal relationship data.
[0046] After obtaining the initial alarm correlation data set, the initial alarm correlation data set can be filtered and processed by the fault causal reasoning model to obtain the fault root cause node and the corresponding fault causal relationship data.
[0047] In one embodiment, the initial alarm correlation data set is filtered and processed by the fault causal reasoning model to obtain the fault root cause node and the corresponding fault causal relationship data, including:
[0048] The initial alarm correlation data set is filtered and processed by the fault causal reasoning model to obtain an alarm propagation directed graph;
[0049] The fault causal reasoning model is used to determine the alarm root cause node and the corresponding alarm causal relationship data based on the alarm propagation directed graph;
[0050] The fault causal reasoning model is used to determine the fault root cause node and the corresponding fault causal relationship data based on the alarm root cause node and the corresponding alarm causal relationship data, and a knowledge document is used; the knowledge document includes at least one of the following: a device alarm description document, a historical fault case document, and an event discovery rule defined based on expert experience.
[0051] In the embodiments of the present application, the initial alarm causal relationship data is filtered and processed by a fault causal reasoning model to obtain an alarm propagation directed graph. Then, alarm root cause nodes and corresponding alarm causal relationship data can be determined in the alarm propagation directed graph. Next, device alarm description documents, historical fault case documents, and event discovery rules defined based on expert experience, and other knowledge documents are used to determine the faults caused by the alarms to the devices or systems, determine fault root cause nodes, and determine corresponding fault causal relationship data.
[0052] In one embodiment, first, the initial alarm correlation data set can be filtered and processed by a fault causal reasoning model. Redundant or repeated alarm data and unreal alarm data in the initial alarm correlation data set can be filtered based on the time correlation and spatial correlation between the alarm data to obtain an alarm propagation directed graph. Then, alarm root cause nodes and alarm causal relationship data can be determined in the alarm propagation directed graph by the fault causal reasoning model. The alarm root cause node can be a network node with an in-degree of 0 in the alarm causal relationship graph. The alarm causal relationship data includes alarm causes, impacts caused by the alarms, and the like. Next, device alarm description documents, historical fault case documents, and event discovery rules defined based on expert experience, and other knowledge documents can be used to determine the corresponding system or device faults caused by the alarms, determine fault root cause nodes, and determine corresponding fault causal relationship data.
[0053] In another embodiment, after the alarm spatio-temporal correlation graph formed by the sub-topology expansion continues to expand to the boundary, that is, the alarm data has fully unfolded in time and space and reached a certain stable state, the initial alarm correlation data is filtered and processed by the fault causal reasoning model to obtain fault root cause nodes and corresponding fault causal relationship data. The fault spatio-temporal candidate correlation group can also be solidified based on the initial alarm causal relationship data, and the fault spatio-temporal candidate correlation group can be filtered to obtain a fault propagation directed graph. Based on the fault propagation directed graph, fault root cause nodes and fault causal relationship data can be determined.
[0054] The solidification process of the fault spatio-temporal candidate correlation group mainly includes the following steps: correlation rule extraction: effective correlation rules are extracted from the alarm spatio-temporal correlation graph. These rules describe the time correlation and spatial correlation between the alarm data. Candidate correlation group generation: according to the extracted correlation rules, alarm data with similar characteristics are divided into different candidate correlation groups. Each candidate correlation group represents a possible fault mode or scenario. Verification and confirmation: the candidate correlation groups are verified and confirmed to exclude false positives and false negatives. This usually needs to be achieved by comparing and analyzing with actual fault records. Solidification storage: the verified and confirmed candidate correlation groups are solidified as fault spatio-temporal candidate correlation groups and stored in the fault management system for subsequent use.
[0055] In one embodiment, based on the alarm propagation directed graph, alarm root cause nodes and corresponding alarm causal relationship data are determined by the fault causal reasoning model, comprising:
[0056] At least one initial alarm root cause node is determined in the alarm propagation directed graph by the fault causal reasoning model; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm propagation directed graph;
[0057] For each initial alarm root cause node, the influence range of the initial alarm root cause node is determined by the fault causal reasoning model using a breadth-first search algorithm;
[0058] Alarm root cause nodes and corresponding alarm causal relationship data are determined by the fault causal reasoning model according to the influence range of each initial alarm root cause node.
[0059] In one implementation, first, each network node with an in-degree of 0 in the alarm propagation directed graph is taken as an initial alarm root cause node by the fault causal reasoning model, and the in-degree of a node refers to the number of edges pointing to the node. If the in-degree of a node is 0, it means that there is no edge directly pointing to the node from other nodes in the graph, and therefore, the network node with an in-degree of 0 in the alarm propagation directed graph can be taken as an initial alarm root cause node. Then, for each initial alarm root cause node, the influence range of the initial alarm root cause node is evaluated by the fault causal reasoning model using a breadth-first search algorithm. Next, the initial alarm node with the largest influence range is determined as an alarm root cause node by the fault causal reasoning model, and alarm causal relationship data are determined based on the alarm propagation directed graph.
[0060] In another embodiment, the fault spatiotemporal candidate association groups are screened to obtain a fault propagation directed graph, and based on the fault propagation directed graph, the fault root cause node and the fault causal relationship data are determined. First, according to the system monitoring data, the fault alarm information and the preliminary diagnostic analysis, a plurality of groups of possible candidate faults are determined. These candidate faults can involve a plurality of system components or subsystems. For each candidate fault, information of system components or subsystems related to the candidate fault is collected, and these components or subsystems will be regarded as nodes in the fault propagation directed graph. Based on this, each system component or subsystem is taken as a node in the directed graph, and according to the fault propagation relationship between the components or subsystems, the related nodes are connected by directed edges to form a fault propagation directed graph. The directed edge represents the direction of fault propagation from one node (component or subsystem) to another node. In the fault propagation directed graph, each node with an in-degree of 0 is taken as an initial fault root cause node, and the initial fault root cause nodes are traversed by using a breadth-first search algorithm to evaluate the influence range of each initial fault root cause node, the initial fault root cause node with the largest influence range is selected as the fault root cause node, and the fault causal relationship data is determined based on the fault propagation directed graph.
[0061] In one embodiment, according to the influence range of each initial alarm root cause node, the alarm root cause node is determined by the fault causal reasoning model, including:
[0062] For each initial alarm root cause node, the number of descendant nodes affected by the initial alarm root cause node is determined according to the alarm propagation path by the fault causal reasoning model;
[0063] According to a preset sorting rule, the number of descendant nodes affected by each initial alarm node is sorted, and the alarm root cause node is determined according to the sorting result.
[0064] In one embodiment, for each initial alarm root cause node, the influence range of the initial alarm root cause node is evaluated by using a breadth-first search algorithm to traverse the initial alarm root cause node by the fault causal reasoning model, and the number of descendant nodes (including child nodes or grandchild nodes) affected by the initial alarm root cause node is calculated according to the alarm propagation path. Then, the number of descendant nodes affected by each initial alarm node is sorted by the fault causal reasoning model, and the initial alarm root cause node with the largest influence range, i.e. the largest number of affected descendant nodes, is determined as the alarm root cause node. After the alarm root cause node is determined, the alarm causal relationship data is determined based on the alarm causal relationship graph.
[0065] In another implementation, for the initial fault root cause node in the fault propagation directed graph, the initial fault root cause node is traversed and evaluated using a breadth-first search algorithm to evaluate the impact range of each initial fault root cause node, the number of descendant nodes (including child nodes or grandchild nodes) affected by the alarm propagation path is calculated, the number of descendant nodes is sorted from many to few, and the initial fault root cause node with the largest impact range, i.e., the largest number of affected descendant nodes, is selected as the fault root cause node. After determining the fault root cause node, the fault causal relationship data is determined based on the fault propagation directed graph.
[0066] In one embodiment, further comprising:
[0067] Obtaining network alarm sample data and corresponding network topology resource sample data;
[0068] Determining alarm causal relationship sample data corresponding to the network alarm sample data using the causal strength function, and determining fault causal relationship label data based on the alarm causal relationship sample data;
[0069] Determining fault causal relationship prediction data by a pre-built neural network based on the network alarm sample data and the corresponding network topology resource sample data, and the fault causal relationship label data;
[0070] Iteratively training the neural network according to the fault causal relationship label data, the fault causal relationship prediction data, and a target loss function, and determining the trained neural network as a fault causal reasoning model.
[0071] In one implementation, the training process of the fault causal reasoning model is as follows: obtaining network alarm sample data and corresponding network topology resource sample data, the network alarm sample data including a plurality of alarm data that have occurred in the system. The network topology resource sample data is the network topology resource corresponding to the network alarm sample data. The network sample alarm data can be obtained through a security monitoring system and tool, a threat intelligence sharing platform, an API interface, or data sharing and collaboration.
[0072] After obtaining the network alarm sample data and the corresponding network topology resource sample data, the causal trigger strength between the network alarm sample data can be mined based on the Hawkes process algorithm. When a network failure occurs, a series of alarms or abnormalities will be generated over time. The alarms can be arranged in chronological order on a time axis, and the alarms can be regarded as a process of a series of discrete events. The occurrence of each event will increase or stimulate the possibility of the occurrence of the next event, and the effect of the stimulation will decay over time. Based on this, the above stimulation process can be modeled based on the Hawkes process algorithm. Based on the one / multi-dimensional Hawkes process strength function, it is assumed that the causal trigger strength between sub-topology nodes depends on the topological distance and is inversely proportional to the topological distance. Based on this, the Hawkes process causal strength function of various types of alarms combined with topology can be obtained. In this way, the causal strength between different alarms can be calculated based on the causal strength function, so as to obtain the alarm causal relationship sample data corresponding to the network alarm sample data. Through the annotation of the alarm causal relationship sample data by the knowledge document and the like, the accurate fault causal relationship label data can be obtained.
[0073] After obtaining the alarm causal relationship sample data and the fault causal relationship label data, the alarm causal relationship sample data and the fault causal relationship label data, and the corresponding network topology resource sample can be input into the neural network built in advance. Through the neural network built in advance, the alarm causal relationship prediction data is obtained. The alarm causal relationship prediction data is filtered and reasoned by the neural network built in advance, and the predicted fault root cause node and the fault causal relationship prediction data are output. Then, according to the fault causal relationship label data, the fault causal relationship prediction data and the target loss function, the neural network is iteratively trained, and the trained neural network is determined as the fault causal reasoning model.
[0074] In another embodiment, when obtaining the alarm causal relationship sample data, the obtained network alarm sample data and the corresponding network topology resource sample data can also be processed in the time dimension and the space dimension by the neural network built in advance to obtain an initial alarm correlation sample data group. Then, the initial alarm correlation sample data group is filtered by the neural network built in advance to obtain an alarm sample propagation directed graph. Based on the alarm sample propagation directed graph, the alarm sample root cause node and the corresponding alarm causal relationship sample data are determined.
[0075] In an implementation, when labeling the alarm causal relationship sample data, the alarm causal relationship sample data can be labeled by using knowledge documents. The knowledge documents include device alarm description documents, historical fault case documents, and event discovery rules defined by relying on expert experience. For the device alarm description documents, since the number of documents is large and the information in the documents is scattered, it is very inconvenient to directly label the alarm causal relationship sample data by using the device alarm description documents. Therefore, when directly labeling the alarm causal relationship sample data by using the device alarm description documents, the following steps a and b can be performed on the device alarm description documents first, and then the alarm causal relationship sample data is labeled. In step a, inter-sentence or intra-sentence fault causal relationship groups are extracted from the device alarm description documents. The step a specifically includes: 1) extracting the “possible cause”, “alarm”, and “impact on the system” and other related information in different inter-sentences in the device alarm description documents, combining “possible cause”->“alarm” and “alarm”->“impact on the system” to form a fault causal relationship pair. 2) For the fault causal relationship that may exist in the same inter-sentence in the device alarm document, the connection words representing causality, such as cause, cause, and trigger, are used to extract the fault causal relationship pair in detail. Since the fault causal relationship pair cannot be directly applied to production, in step b, the obtained fault causal relationship pair is matched for consistency with the alarm causal relationship data description, and the hidden alarm causal relationship pair is further extracted from the fault causal relationship pair.
[0076] When training the fault causal reasoning model, the fault causal relationship prediction data can also be compared with the fault causal relationship label data to confirm whether the fault causal relationship prediction data is accurate and consistent with the actual business logic. The fault causal reasoning model can be adjusted in parameters according to the fault causal relationship label data and the actual business logic.
[0077] In addition, the embodiments of the present application also provide a fault root cause positioning system, which supports inputting alarm causal relationship data with spatiotemporal correlation, performing a fault causal reasoning process, checking whether the generated fault causal relationship data is consistent with the actual business logic, adjusting the fault root cause reasoning model according to whether the generated fault causal relationship data is consistent with the actual business logic, supporting offline verification of the adjustment result, supporting offline import of network topology resource data, network alarm data, and the fault root cause reasoning model, ensuring that the adjustment can achieve the expected effect, and updating the parameters of the fault root cause reasoning model after verification succeeds, to better support business scenarios.
[0078] From the technical solutions provided by the embodiments of the present application, it can be seen that, in the embodiments of the present application, firstly, network alarm data and corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are processed in time dimension and space dimension by a fault cause and effect reasoning model to obtain an initial alarm correlation data group; and then, the initial alarm correlation data group is filtered and reasoned by the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data. It can be seen that, by the embodiments of the present application, the acquired network alarm data and corresponding network topology resource data can be processed in time dimension and space dimension by the fault cause and effect reasoning model to obtain an initial alarm correlation data group with time information and space information, and the initial alarm correlation data group with time information and space information can be filtered and reasoned to obtain a corresponding alarm propagation path, determine the corresponding fault cause and effect relationship, locate the fault root cause node, and realize rapid and accurate positioning of the fault root cause, thereby improving the efficiency of fault processing.
[0079] Corresponding to the fault root cause positioning method provided by the above embodiments, based on the same technical concept, the embodiments of the present application further provide a fault root cause positioning device, Figure 3 A module composition schematic diagram of a fault root cause positioning device provided by the embodiments of the present application, the fault root cause positioning device is used for executing Figures 1 to 2 The fault root cause positioning method described above, as shown in Figure 3 The fault root cause positioning device includes:
[0080] The acquisition module 31 is configured to acquire network alarm data and corresponding network topology resource data.
[0081] The processing module 32 is configured to process the network alarm data and the corresponding network topology resource data in time dimension and space dimension by a fault cause and effect reasoning model to obtain an initial alarm correlation data group.
[0082] The reasoning module 33 is configured to filter and reason the initial alarm correlation data group by the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data.
[0083] Optionally, the processing module 32 is specifically configured to
[0084] According to the fault cause and effect reasoning model, a sub-topology alarm set in each preset time step is acquired according to a preset time step; the sub-topology is a region formed by network nodes with alarms; and the sub-topology alarm set is a set formed by all alarms existing in the sub-topology in the preset time step.
[0085] An initial alarm correlation data set is generated based on each sub-topology alarm set through a fault causal reasoning model.
[0086] Optionally, the reasoning module 33 is specifically configured to
[0087] The initial alarm correlation data set is filtered through the fault causal reasoning model to obtain an alarm causal relationship graph.
[0088] The alarm root cause node and corresponding alarm causal relationship data are determined based on the alarm causal relationship graph through the fault causal reasoning model.
[0089] The fault root cause node and corresponding fault causal relationship data are determined based on the alarm root cause node and corresponding alarm causal relationship data through the fault causal reasoning model by using a knowledge document; the knowledge document includes at least one of the following: a device alarm description document, a historical fault case document, and an event discovery rule defined by relying on expert experience.
[0090] Optionally, the reasoning module 33 is specifically configured to
[0091] At least one initial alarm root cause node is determined in the alarm causal relationship graph through the fault causal reasoning model; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm causal relationship graph.
[0092] The influence range of the initial alarm root cause node is determined through the fault causal reasoning model by using a breadth-first search algorithm for each initial alarm root cause node.
[0093] The alarm root cause node and corresponding alarm causal relationship data are determined according to the influence range of each initial alarm root cause node through the fault causal reasoning model.
[0094] Optionally, the reasoning module 33 is specifically configured to
[0095] The number of descendant nodes affected by the initial alarm root cause node is determined according to an alarm propagation path through the fault causal reasoning model for each initial alarm root cause node.
[0096] The number of descendant nodes affected by each initial alarm node is sorted according to a preset sorting rule, and the alarm root cause node is determined according to a sorting result.
[0097] Optionally, the apparatus further includes a training module configured to acquire network alarm sample data and corresponding network topology resource sample data.
[0098] The alarm causal relationship sample data corresponding to the network alarm sample data is determined by using a causal strength function, and fault causal relationship label data is determined based on the alarm causal relationship sample data.
[0099] determining fault causal relationship prediction data according to the network alarm sample data, the corresponding network topology resource sample data, and the fault causal relationship label data through the pre-built neural network;
[0100] According to the fault causal relationship label data, the fault causal relationship prediction data, and a target loss function, the neural network is iteratively trained, and the trained neural network is determined as a fault causal reasoning model.
[0101] As can be seen from the technical solutions provided by the above embodiments of the present application, in the embodiments of the present application, first, network alarm data and corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are sliced in the time dimension and the space dimension through a fault causal reasoning model to obtain an initial alarm correlation data set; and then, the initial alarm correlation data set is screened and reasoned through the fault causal reasoning model to obtain a fault root cause node and corresponding fault causal relationship data. As can be seen, through the embodiments of the present application, the acquired network alarm data and corresponding network topology resource data can be sliced in the time dimension and the space dimension through the fault causal reasoning model to obtain an initial alarm correlation data set with time information and space information, and the corresponding alarm propagation path can be obtained by screening and reasoning the initial alarm correlation data set with time information and space information, so as to determine the corresponding fault causal relationship and locate the fault root cause node, thereby realizing rapid and accurate positioning of the fault root cause and improving the efficiency of fault processing.
[0102] The fault root cause positioning apparatus provided by the embodiments of the present application can realize each process in the embodiments corresponding to the fault root cause positioning method, and thus details are not repeated here.
[0103] It should be noted that the fault root cause positioning apparatus provided by the embodiments of the present application is based on the same inventive concept as the fault root cause positioning method provided by the embodiments of the present application, and thus the specific implementation of this embodiment can be referred to the implementation of the foregoing fault root cause positioning method, and details are not repeated.
[0104] Corresponding to the fault root cause positioning method provided by the above embodiments and based on the same technical concept, the embodiments of the present application further provide an electronic device for executing the foregoing fault root cause positioning method, Figure 4 A structural schematic diagram of an electronic device provided by the embodiments of the present application is shown in Figure 4As shown, the electronic device can have a large difference due to different configurations or performance, and can include one or more processors 401 and memories 402, and one or more storage applications or data can be stored in the memories 402. Among them, the memory 402 can be temporary storage or persistent storage. The application stored in the memory 402 can include one or more modules (not shown in the figure), each of which can include a series of computer executable instructions in the electronic device. Further, the processor 401 can be configured to communicate with the memory 402 and execute a series of computer executable instructions in the memory 402 on the electronic device. The electronic device can also include one or more power supplies 403, one or more wired or wireless network interfaces 404, one or more input / output interfaces 405, and one or more keyboards 406.
[0105] In particular, in the present embodiment, the electronic device includes a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory complete the communication among each other through the bus; the memory is used to store computer programs; the processor is used to execute the programs stored on the memory to realize the following method steps:
[0106] Obtain network alarm data and corresponding network topology resource data;
[0107] Through the fault causal reasoning model, the network alarm data and the corresponding network topology resource data are sliced in time dimension and space dimension to obtain an initial alarm correlation data set;
[0108] Through the fault causal reasoning model, the initial alarm correlation data set is screened and reasoned to obtain a fault root cause node and corresponding fault causal relationship data.
[0109] From the technical solutions provided by the above embodiments of the present application, it can be seen that, in the embodiments of the present application, firstly, network alarm data and corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are processed in time dimension and space dimension by a fault causal reasoning model to obtain an initial alarm correlation data group; and then, the initial alarm correlation data group is processed by the fault causal reasoning model to obtain a fault root cause node and corresponding fault causal relationship data. It can be seen that, by the embodiments of the present application, the acquired network alarm data and corresponding network topology resource data can be processed in time dimension and space dimension by the fault causal reasoning model to obtain an initial alarm correlation data group with time information and space information, and the initial alarm correlation data group with time information and space information can be processed by screening and reasoning to obtain a corresponding alarm propagation path, thereby determining the corresponding fault causal relationship and locating the fault root cause node, and the fault root cause is quickly and accurately located, thereby improving the efficiency of fault processing.
[0110] The embodiments of the present application also provide a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the following method steps:
[0111] acquiring network alarm data and corresponding network topology resource data;
[0112] processing the network alarm data and the corresponding network topology resource data in time dimension and space dimension by a fault causal reasoning model to obtain an initial alarm correlation data group;
[0113] processing the initial alarm correlation data group by the fault causal reasoning model to obtain a fault root cause node and corresponding fault causal relationship data.
[0114] From the technical solutions provided by the embodiments of the present application, it can be seen that, in the embodiments of the present application, firstly, network alarm data and corresponding network topology resource data are acquired; then, the network alarm data and the corresponding network topology resource data are processed in time dimension and space dimension by a fault cause and effect reasoning model to obtain an initial alarm correlation data set; and then, the initial alarm correlation data set is filtered and reasoned by the fault cause and effect reasoning model to obtain a fault root cause node and corresponding fault cause and effect relationship data. It can be seen that, by the embodiments of the present application, the acquired network alarm data and corresponding network topology resource data can be processed in time dimension and space dimension by the fault cause and effect reasoning model to obtain an initial alarm correlation data set with time information and space information, and the initial alarm correlation data set with time information and space information can be filtered and reasoned to obtain a corresponding alarm propagation path, to determine the corresponding fault cause and effect relationship and locate the fault root cause node, so that the fault root cause can be quickly and accurately located, and the efficiency of fault processing is improved.
[0115] The embodiments of the present application also provide a computer program product, which, when executed by a processor, realizes each process of the above fault root cause positioning method embodiments and can achieve the same technical effects. To avoid repetition, details are not described herein.
[0116] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product in the form of being implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.
[0117] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks.
[0118] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0120] In a typical configuration, an electronic device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0121] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) and / or non-volatile memory, e.g., read only memory (ROM) or flash memory. The memory is an example of computer readable media.
[0122] Computer readable media includes permanent and non-permanent, moveable and non- moveable media that can be implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that is accessible to a computing device. According to the definition provided herein, computer readable media excludes transitory media, such as modulated data signals and carrier waves.
[0123] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0124] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, an apparatus or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.
[0125] The embodiments of the present application described above are intended to be merely exemplary and those skilled in the art shall understand that modifications and variations can be made to the embodiments without departing from the spirit and scope of the application. Accordingly, the application is not limited to the embodiments described above.
Claims
1. A method of fault root cause localization, the method comprising: The method comprises: obtaining network alarm data and corresponding network topology resource data; performing slice processing on the network alarm data and the corresponding network topology resource data in the time dimension and the space dimension through a fault causal reasoning model to obtain an initial alarm correlation data group; performing screening processing on the initial alarm correlation data group through the fault causal reasoning model to obtain an alarm propagation directed graph; determining at least one initial alarm root cause node in the alarm propagation directed graph; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm propagation directed graph; for each initial alarm root cause node, determining the influence range of the initial alarm root cause node by using a breadth-first search algorithm through the fault causal reasoning model; determining alarm root cause nodes and corresponding alarm causal relationship data according to the influence range of each initial alarm root cause node through the fault causal reasoning model; determining fault root cause nodes and corresponding fault causal relationship data based on the alarm root cause nodes and the corresponding alarm causal relationship data by using a knowledge document; the knowledge document comprises at least one of the following: a device alarm description document, a historical fault case document, and an event discovery rule defined based on expert experience; wherein, determining alarm root cause nodes according to the influence range of each initial alarm root cause node through the fault causal reasoning model comprises: for each initial alarm root cause node, determining the number of descendant nodes affected by the initial alarm root cause node according to an alarm propagation path through the fault causal reasoning model; sorting the number of descendant nodes affected by each initial alarm node according to a preset sorting rule, and determining alarm root cause nodes according to the sorting result.
2. The method of claim 1, wherein, The method comprises: obtaining network alarm data and corresponding network topology resource data; performing slice processing on the network alarm data and the corresponding network topology resource data in the time dimension and the space dimension through a fault causal reasoning model to obtain an initial alarm correlation data group; 3. The method of claim 1, wherein, performing screening processing on the initial alarm correlation data group through the fault causal reasoning model to obtain an alarm propagation directed graph; determining at least one initial alarm root cause node in the alarm propagation directed graph; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm propagation directed graph; for each initial alarm root cause node, determining the influence range of the initial alarm root cause node by using a breadth-first search algorithm through the fault causal reasoning model; determining alarm root cause nodes and corresponding alarm causal relationship data according to the influence range of each initial alarm root cause node through the fault causal reasoning model; determining fault root cause nodes and corresponding fault causal relationship data based on the alarm root cause nodes and the corresponding alarm causal relationship data by using a knowledge document; the knowledge document comprises at least one of the following: a device alarm description document, a historical fault case document, and an event discovery rule defined based on expert experience; wherein, determining alarm root cause nodes according to the influence range of each initial alarm root cause node through the fault causal reasoning model comprises: for each initial alarm root cause node, determining the number of descendant nodes affected by the initial alarm root cause node according to an alarm propagation path through the fault causal reasoning model; sorting the number of descendant nodes affected by each initial alarm node according to a preset sorting rule, and determining alarm root cause nodes according to the sorting result. The method comprises: obtaining network alarm data and corresponding network topology resource data; performing slice processing on the network alarm data and the corresponding network topology resource data in the time dimension and the space dimension through a fault causal reasoning model to obtain an initial alarm correlation data group; performing screening processing on the initial alarm correlation data group through the fault causal reasoning model to obtain an alarm propagation directed graph; determining at least one initial alarm root cause node in the alarm propagation directed graph; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm propagation directed graph; for each initial alarm root cause node, determining the influence range of the initial alarm root cause node by using a breadth-first search algorithm through the fault causal reasoning model; determining alarm root cause nodes and corresponding alarm causal relationship data according to the influence range of each initial alarm root cause node through the fault causal reasoning model; determining fault root cause nodes and corresponding fault causal relationship data based on the alarm root cause nodes and the corresponding alarm causal relationship data by using a knowledge document; the knowledge document comprises at least one of the following: a device alarm description document, a historical fault case document, and an event discovery rule defined based on expert experience; wherein, determining alarm root cause nodes according to the influence range of each initial alarm root cause node through the fault causal reasoning model comprises: for each initial alarm root cause node, determining the number of descendant nodes affected by the initial alarm root cause node according to an alarm propagation path through the fault causal reasoning model; sorting the number of descendant nodes affected by each initial alarm node according to a preset sorting rule, and determining alarm root cause nodes according to the sorting result. According to the fault causality label data, the fault causality prediction data and a target loss function, the neural network is iteratively trained, and the trained neural network is determined as the fault causality reasoning model.
4. A fault root cause localization apparatus characterized by, The apparatus comprises: An acquisition module is configured to acquire network alarm data and corresponding network topology resource data; A processing module is configured to perform slicing processing on the network alarm data and the corresponding network topology resource data in a time dimension and a space dimension by using a fault causality reasoning model to obtain an initial alarm correlation data set; A reasoning module is configured to perform screening processing on the initial alarm correlation data set by using the fault causality reasoning model to obtain an alarm propagation directed graph, and determine at least one initial alarm root cause node in the alarm propagation directed graph; the initial alarm root cause node is a network node with an in-degree of 0 in the alarm propagation directed graph; For each initial alarm root cause node, the influence range of the initial alarm root cause node is determined by using a breadth-first search algorithm based on the fault causality reasoning model; An alarm root cause node and corresponding alarm causality data are determined according to the influence range of each initial alarm root cause node based on the fault causality reasoning model; a fault root cause node and corresponding fault causality data are determined based on the alarm root cause node and the corresponding alarm causality data by using a knowledge document; the knowledge document comprises at least one of the following: a device alarm description document, a historical fault case document and an event discovery rule defined based on expert experience; The determination of the alarm root cause node based on the fault causality reasoning model according to the influence range of each initial alarm root cause node comprises: For each initial alarm root cause node, the number of descendant nodes affected by the initial alarm root cause node is determined according to an alarm propagation path based on the fault causality reasoning model; The number of descendant nodes affected by each initial alarm root cause node is sorted according to a preset sorting rule, and an alarm root cause node is determined according to a sorting result.
5. An electronic device, comprising: The apparatus comprises a processor, a communication interface, a memory and a communication bus; the processor, the communication interface and the memory communicate with each other through the bus; the memory is configured to store a computer program; the processor is configured to execute the program stored in the memory to implement the fault root cause positioning method steps of any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the fault root cause positioning method steps of any one of claims 1-3.
7. A computer program product, characterised in that, The computer program is executed by the processor to implement the fault root cause positioning method steps of any one of claims 1-3.
Citation Information
Patent Citations
Transmission line receiving direction interruption fault positioning method
CN117014067A