Fault root cause location method and device, and computer storage medium
By generating the knowledge graph of the target network, the problem of low efficiency in obtaining network information in the existing technology is solved, and intuitive visualization and efficient query of network information are realized.
Patent Information
- Application Number
- CN201980099130.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2019-11-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-11-11
AI Technical Summary
In the prior art, the efficiency of obtaining network information is low, and it is necessary to manually query the command line, which is inefficient.
Through the management device, the network data of the target network is obtained, the knowledge graph triplets are extracted, the knowledge graph of the target network is generated, and the network data is visualized. Users can obtain network information by viewing the knowledge graph.
It improves the efficiency of users to obtain network information, can intuitively reflect network entities and their relationships, and simplifies the network information query process.
Smart Images

Figure CN114208128B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network technology, and in particular to a method and device for locating a root cause of a fault, and a computer storage medium. Background Art
[0002] As the scale of the network continues to expand, the amount of network data in the communication network is increasing. How to obtain valuable network information from the massive amount of network data has become a problem that people in all fields need to face.
[0003] Currently, technicians can obtain network information in a communication network by manually querying command lines. For example, when a technician needs to query which routing protocol a certain interface of a network device carries, he can query the command line corresponding to the interface configuration of the network device. However, the efficiency of obtaining network information by querying command lines is low. Summary of the invention
[0004] The present application provides a data processing method and device, and a computer storage medium, which can solve the current problem of low efficiency in obtaining network information.
[0005] In a first aspect, a data processing method is provided. A management device obtains network data of a target network, the network data including the network topology of the target network and device information of multiple network devices in the target network, the device information including one or more of interface configuration information, protocol configuration information, and service configuration information. The management device extracts multiple knowledge graph triples from the network data, each knowledge graph triple including two network entities and a relationship between the two network entities, and the type of the network entity is a network device, an interface, a protocol, or a service. The management device generates a knowledge graph of the target network based on the multiple knowledge graph triples.
[0006] In this application, after obtaining the network data of the target network, the management device processes the network data to generate a knowledge graph of the target network. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationship between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain network information in the target network, he can view the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of users in obtaining network information.
[0007] Optionally, the device information also includes routing table entries. That is, the management device can extract the knowledge graph triples from the network configuration information of the network device, or can obtain the knowledge graph triples based on the routing table entries of the network device.
[0008] Optionally, the relationship between the two network entities is a dependency relationship, a subordinate relationship or a peer relationship.
[0009] Optionally, after generating the knowledge graph of the target network, the management device can also obtain abnormal events generated in the target network when a failure occurs in the target network; and identify abnormal network entities that generate abnormal events in the target network on the knowledge graph of the target network.
[0010] In this application, after the management device identifies the abnormal network entity on the knowledge graph of the target network, it can send the knowledge graph of the target network to the OSS or other terminal devices connected to the management device for display by the OSS or terminal devices, so that the operation and maintenance personnel can view the abnormal network entities in the target network. The management device can also send the abnormal events corresponding to the abnormal network entities to the OSS or other terminal devices connected to the management device for display by the OSS or terminal devices, so that the operation and maintenance personnel can obtain the abnormal type of the abnormal network entity.
[0011] Optionally, the abnormal event includes one or more of an alarm log, a state change log, and an abnormal key performance indicator.
[0012] Optionally, after the management device identifies the abnormal network entity that generates the abnormal event in the target network on the knowledge graph of the target network, the management device may also determine one or more root cause fault network entities among all the abnormal network entities on the knowledge graph of the target network based on the fault propagation relationship between the network entities; and identify one or more root cause fault network entities on the knowledge graph of the target network.
[0013] Optionally, the management device can obtain multiple knowledge graph samples, each knowledge graph sample is marked with all abnormal network entities and root cause fault network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a fault occurs in the network to which the knowledge graph sample belongs; the management device determines the fault propagation relationship based on the multiple knowledge graph samples.
[0014] In an embodiment of the present application, the management device can use multiple knowledge graph samples to learn the fault propagation relationship between network entities, and based on the fault propagation relationship, determine the root cause fault network entity among the abnormal network entities on the knowledge graph of the target network, thereby realizing automatic reasoning and positioning of the root cause of the network fault.
[0015] Optionally, the network to which the knowledge graph sample belongs is the target network, or the network to which the knowledge graph sample belongs is another network of the same network type as the target network.
[0016] Optionally, the process of extracting a plurality of knowledge graph triples from network data by the management device includes:
[0017] The management device extracts multiple knowledge graph triples from the network data based on the abstract business model corresponding to the network type of the target network. The abstract business model is used to reflect the relationship between different network entities.
[0018] In a second aspect, a data processing device is provided. The device includes multiple functional modules, and the multiple functional modules interact with each other to implement the method in the first aspect and its respective embodiments. The multiple functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the multiple functional modules can be arbitrarily combined or divided based on specific implementations.
[0019] In a third aspect, a data processing device is provided, comprising: a processor and a memory;
[0020] The memory is used to store a computer program, wherein the computer program includes program instructions;
[0021] The processor is used to call the computer program to implement the data processing method as described in any one of the first aspects.
[0022] According to a fourth aspect, a computer storage medium is provided, on which instructions are stored. When the instructions are executed by a processor, the data processing method as described in any one of the first aspects is implemented.
[0023] In a fifth aspect, a chip is provided, the chip including a programmable logic circuit and / or program instructions, and when the chip is running, it implements the data processing method as described in any one of the first aspects.
[0024] The beneficial effects of the technical solution provided by this application include at least:
[0025] After acquiring the network data of the target network, the management device processes the network data to generate a knowledge graph of the target network. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationship between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain network information in the target network, he can view the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of the user in obtaining network information. In addition, after a failure occurs in the target network, the management device can also identify abnormal network entities on the knowledge graph of the target network, intuitively reflecting the abnormal location in the target network. Furthermore, the management device can also determine the root cause fault network entity in the abnormal network entity on the knowledge graph of the target network based on the fault propagation relationship between network entities, thereby realizing automatic reasoning and positioning of the root cause of the network failure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1It is a schematic diagram of an application scenario involved in the data processing method provided in an embodiment of the present application;
[0027] Figure 2 is a flow chart of a data processing method provided in an embodiment of the present application;
[0028] Figure 3 It is a schematic diagram of a knowledge graph triple provided in an embodiment of the present application;
[0029] Figure 4 It is a schematic diagram of a knowledge graph provided in an embodiment of the present application;
[0030] Figure 5 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;
[0031] Figure 6 is a structural schematic diagram of another data processing device provided in an embodiment of the present application;
[0032] Figure 7 is a structural diagram of another data processing device provided in an embodiment of the present application;
[0033] Figure 8 is a structural diagram of another data processing device provided in an embodiment of the present application;
[0034] Fig. 9 It is a block diagram of a data processing device provided in an embodiment of the present application.
[0035] Fig.10 This is a schematic diagram of a fault root cause location device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0037] Figure 1 Schematic diagram of application scenarios involved in the data processing method provided in the embodiment of the present application. Figure 1 As shown, the application scenario includes a management device 101 and network devices 102a-102c (collectively referred to as network devices 102) in a communication network. Figure 1The number of management devices and network devices is only used as an illustration and is not intended to limit the application scenarios involved in the data processing method provided in the embodiments of the present application. The communication network may be a data center network (DCN), a metropolitan area network, a wide area network, a campus network, a virtual local area network (VLAN) or a virtual extensible local area network (VXLAN), etc. The embodiments of the present application do not limit the type of the communication network.
[0038] Optionally, the management device 101 may be a server, or a server cluster consisting of several servers, or a cloud computing service center. The network device 102 may be a switch or a router, etc. Optionally, please continue to refer to Figure 1 , the application scenario may also include a control device 103. The control device 103 is used to manage and control the network device 102 in the communication network. The management device 101 and the control device 103 are connected via a wired network or a wireless network, and the control device 103 and the network device 102 are connected via a wired network or a wireless network. The control device 103 may be a network controller, a network management device, a gateway or other device with control capabilities. The control device 103 may be one or more devices.
[0039] Among them, the control device 103 may store the networking topology of the communication network managed by the control device 103. The control device 103 is also used to collect the device information of the network device 102 in the communication network and the abnormal events generated in the communication network, and provide the management device 101 with the networking topology of the communication network, the device information of the network device 102 and the abnormal events generated in the communication network. The device information of the network device includes the network configuration information and / or routing table items of the network device. The network configuration information generally includes interface configuration information, protocol configuration information and service configuration information. Optionally, the control device 103 may periodically collect the device information of the network device 102 and the abnormal events generated in the communication network. Alternatively, when the device information of the network device 102 changes, the network device 102 actively reports the changed device information to the control device 103; when a failure occurs in the communication network, the network device 102 actively reports the abnormal events generated to the control device 103. Of course, in some application scenarios, the management device may also be directly connected to the network device in the communication network, that is, the application scenario may not include the control device, and the embodiments of the present application do not limit this.
[0040] Figure 2 is a flow chart of a data processing method provided in an embodiment of the present application. Figure 1 The management device 101 in the application scenario shown. Figure 2 As shown, the method includes:
[0041] Step 201: Acquire network data of the target network.
[0042] The network data includes the networking topology of the target network and the device information of multiple network devices in the target network. The device information of the network device includes the network configuration information of the network device, specifically including one or more of the interface configuration information, the protocol configuration information and the service configuration information. The device information may also include routing table items, etc. Optionally, the interface configuration information of the network device includes the Internet Protocol (IP) address of the interface, the protocol type supported by the interface, and the service type supported by the interface, etc.; the protocol configuration information of the network device includes the identifier of the protocol, which is used to uniquely identify the protocol, and the identifier of the protocol can be represented by characters, letters and / or numbers, etc.; the service configuration information of the network device includes the services used by the network device, such as virtual private network (VPN) services and / or dynamic host configuration protocol (DHCP) services, etc.
[0043] Optionally, the implementation process of step 201 includes: the management device receives network data of the target network sent by the control device of the target network.
[0044] Step 202: extract multiple knowledge graph triples from the network data of the target network.
[0045] Each knowledge graph triple includes two network entities and the relationship between the two network entities. For example, the relationship between the two network entities may be a dependency relationship, a subordinate relationship, or a peer relationship.
[0046] In an embodiment of the present application, the type of network entity may be a network device, an interface, a protocol, or a service. For example, when two network entities in a certain knowledge graph triple are a network device and an interface, respectively, the relationship between the two network entities is a subordinate relationship, that is, the interface belongs to the network device. For another example, when two network entities in a certain knowledge graph triple are two interfaces for establishing a communication connection, the relationship between the two network entities is a peer-to-peer relationship.
[0047] Optionally, a network entity of type network device can be represented by the name of the network device, media access control (MAC) address, hardware address or other identifier that can uniquely identify the network device. A network entity of type interface can be represented by the name of the interface. A network entity of type protocol can be represented by the identifier of the protocol. The knowledge graph triple is represented in the form of a graph, and the knowledge graph triple is composed of two basic elements, points and edges. The points represent network entities, and the edges represent the relationship between two network entities. Among them, the edges in the knowledge graph triple can be directed or undirected. The edges in the knowledge graph triple can also be used to represent the specific relationship between two network entities, such as a dependency, a subordinate relationship, or a peer relationship. For example, when the two network entities are in a peer relationship, the two network entities can be connected by an undirected edge; when the two network entities are in a dependency or a subordinate relationship, the two network entities can be connected by a directed edge (such as an arrow), and the direction of the edge is from the dependent network entity to the dependent network entity, or the direction of the edge is from the attached network entity to the attached network entity.
[0048] For example, assuming that the network data acquired by the management device contains a piece of data "network device 1 has an interface named "10GE1 / 0 / 6"", the management device can extract the following information: Figure 3 The knowledge graph triples shown. Figure 3 In this knowledge graph triplet, the types of the two network entities are network devices and interfaces, respectively. The arrow points from the interface to the network device, indicating that the interface named "10GE1 / 0 / 6" belongs to network device 1, that is, this knowledge graph triplet can reflect that network device 1 has an interface named "10GE1 / 0 / 6".
[0049] Optionally, the implementation process of step 202 includes: the management device extracts multiple knowledge graph triples from the network data based on the abstract business model corresponding to the network type of the target network, and the abstract business model is used to reflect the relationship between different network entities. The abstract business models corresponding to different network types may be different. The abstract business model is essentially a data object used to define the dependency relationship between different network entities. For example, the abstract business model can be defined as follows: each network device has one or more interfaces, that is, the interface belongs to the network device; the interface can carry forwarding services, for example, the interface can carry three-layer IP forwarding services, that is, the interface supports the use of interior gateway protocol (IGP) to forward packets, that is, the three-layer IP forwarding service or IGP depends on the interface; the three-layer IP forwarding service can carry VXLAN tunnels, traffic engineering (TE) tunnels and border gateway protocols (BorderGateway Protocol, BGP), that is, VXLAN tunnels, TE tunnels and BGP depend on the three-layer IP forwarding service; VPN services can be carried on TE tunnels, that is, VPN services depend on TE tunnels; and so on. Among them, the three-layer IP forwarding service can carry VXLAN tunnels, which means that the interface carrying the three-layer IP forwarding service can be used as the endpoint of the VXLAN tunnel; the three-layer IP forwarding service can carry TE tunnels, which means that the interface carrying the three-layer IP forwarding service can be used as the endpoint of the TE tunnel; the three-layer IP forwarding service can carry BGP, which means that the interface carrying the three-layer IP forwarding service can send and receive BGP-based protocol messages; the TE tunnel can carry VPN services, which means that the interface carrying the TE tunnel can support VPN services.
[0050] Optionally, the management device may extract the knowledge graph triples from the network configuration information of the network device, or may obtain the knowledge graph triples based on the routing table entries of the network device.
[0051] For example, assume that the content of the knowledge graph triple extracted by the management device based on the network configuration information of the network device includes: Network device 1 has an interface named "10GE1 / 0 / 6"; Interface "10GE1 / 0 / 6" supports Layer 3 IP forwarding services, and the IP address of interface "10GE1 / 0 / 6" is 192.168.1.1; Network device 1 has a VXLAN tunnel, the source IP address of the VXLAN tunnel is 192.168.1.1, and the destination IP address is 192.168.10.1. In order to determine the relationship between the VXLAN tunnel and the interface, the management device can use 192.168.10.1 as the destination IP address to search the routing table entry. If it is determined based on the routing table entry that the outgoing interface to the destination IP address is interface "10GE1 / 0 / 6", then the content of the knowledge graph triple obtained by the management device based on the routing table entry is: One endpoint of the VXLAN tunnel is located on interface "10GE1 / 0 / 6" of network device 1.
[0052] Step 203: Generate a knowledge graph of the target network based on multiple knowledge graph triples.
[0053] For example, assume that the target network includes two network devices, namely network device A and network device B. Network device A has three interfaces, and the names of the three interfaces are 10GE1 / 0 / 1, 10GE1 / 0 / 2, and 10GE1 / 0 / 3. Network device B has four interfaces, and the names of the four interfaces are 10GE3 / 0 / 1, 10GE3 / 0 / 2, 10GE3 / 0 / 3, and 10GE3 / 0 / 4. Network device A and network device B both support the open shortest path first (OSPF) protocol, which is an IGP. The identifier of the OSPF protocol in network device A is represented by 10.89.46.25, including three routing IPs, namely 11.11.11.11, 11.11.11.12, and 11.11.11.13. The identifier of the OSPF protocol in network device B is 10.89.49.37, which includes four routing IPs, namely 11.12.11.11, 11.12.11.12, 11.12.11.13 and 11.12.11.14. The interface "10GE1 / 0 / 2" of network device A is connected to the interface "10GE3 / 0 / 2" of network device B, and the two interfaces communicate using the OSPF protocol. The routing IP used by the interface "10GE1 / 0 / 2" of network device A is 11.11.11.11, and the routing IP used by the interface "10GE3 / 0 / 2" of network device B is 11.12.11.14. Based on the above network data, we can get the following: Figure 4 The knowledge graph shown. Figure 4 , the device information of network device A and network device B in the target network are expressed in the form of graphs on the knowledge graph.
[0054] Optionally, after generating the knowledge graph of the target network, the management device can send the knowledge graph of the target network to an operations support system (OSS) or other terminal devices connected to the management device for display by the OSS or terminal devices. Of course, if the management device itself has a display function, the management device can also directly display the knowledge graph of the target network on its own display interface.
[0055] In the embodiment of the present application, after acquiring the network data of the target network, the management device processes the network data to generate a knowledge graph of the target network. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationships between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain network information in the target network, it can be done by viewing the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of the user in obtaining network information.
[0056] Optionally, after generating the knowledge graph of the target network, the management device may also store the knowledge graph of the target network in the management device or in a storage device connected to the management device for subsequent use. For example, the knowledge graph of the target network can be used as a basis for determining the fault propagation relationship between network entities, and / or as a basis for fault root cause reasoning, etc.
[0057] Step 204: When a target network failure occurs, abnormal events generated in the target network are obtained.
[0058] A target network failure refers to a failure of a network device in the target network. The types of network device failures include interface failures, protocol failures (including failures to send and receive protocol messages normally), and service failures. Optionally, abnormal events include one or more of alarm logs, state change logs, and abnormal key performance indicators. The alarm log includes the identifier of the abnormal network entity in the network device and the alarm type. The state change log includes configuration file change information and / or routing table entry change information. For example, the state change log may include information such as "access sub-interface deletion" and "destination IP host route deletion". Abnormal key performance indicators are used to describe the abnormality of a certain indicator of a network entity.
[0059] Step 205: Identify abnormal network entities that generate abnormal events in the target network on the knowledge graph of the target network.
[0060] Optionally, the management device may determine abnormal network entities based on abnormal events in the acquired target network, and identify all abnormal network entities on the knowledge graph of the target network. Abnormal network entities on the knowledge graph may be distinguished from normal network entities by using different patterns and / or different colors, or words such as "abnormal" may be marked near the abnormal network entities. The embodiment of the present application does not limit the method of identifying abnormal network entities on the knowledge graph.
[0061] Optionally, after the management device identifies the abnormal network entity on the knowledge graph of the target network, it can send the knowledge graph of the target network to the OSS or other terminal devices connected to the management device for display by the OSS or terminal devices, so that the operation and maintenance personnel can view the abnormal network entities in the target network. The management device can also send the abnormal events corresponding to the abnormal network entities to the OSS or other terminal devices connected to the management device for display by the OSS or terminal devices, so that the operation and maintenance personnel can obtain the abnormal type of the abnormal network entity.
[0062] The knowledge graph generated by step 203 may also be referred to as an initial knowledge graph. The knowledge graph in which abnormal network entities are identified by step 205 may also be referred to as a first knowledge graph.
[0063] Step 206: Based on the fault propagation relationship between network entities, one or more root cause fault network entities are determined from all abnormal network entities on the knowledge graph of the target network.
[0064] The root cause fault network entity refers to an abnormal network entity that is the root cause of the fault.
[0065] Optionally, the process of the management device obtaining the fault propagation relationship includes: the management device obtains multiple knowledge graph samples, each of which is marked with all abnormal network entities and root cause fault network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a fault occurs in the network to which the knowledge graph sample belongs. The management device determines the fault propagation relationship based on the multiple knowledge graph samples. Each knowledge graph sample is a fault case, and the abnormal network entities and root cause fault network entities in the knowledge graph sample can be manually determined. Optionally, the management device can use a graph embedding algorithm to learn the fault propagation relationship in the multiple knowledge graph samples. Alternatively, when the probability of two network entities in the same knowledge graph triplet simultaneously having abnormalities is greater than a certain threshold, the management device can determine that fault propagation will occur between the two network entities.
[0066] For example, in Figure 4In the knowledge graph shown, when the interface "10GE1 / 0 / 2" of network device A fails, the interface will not be able to communicate normally, which will cause the routing IP "11.12.11.14" used by the interface to be unavailable. Therefore, the management device can obtain a set of fault propagation relationships: interface failure will cause the routing IP used by the interface to be unavailable. When the management device obtains a first abnormal event indicating an interface failure and a second abnormal event indicating that the routing IP used by the interface is unavailable, the management device determines that the first abnormal event is a root cause abnormal event, and determines that the interface is a root cause fault network entity.
[0067] Optionally, the network to which the above-mentioned knowledge graph sample belongs is the target network, or the network to which the above-mentioned knowledge graph sample belongs is another network of the same network type as the target network. The above-mentioned multiple knowledge graph samples may belong to the same network or to multiple networks, which is not limited in the embodiments of the present application.
[0068] In an embodiment of the present application, the management device can use multiple knowledge graph samples to learn the fault propagation relationship between network entities, and based on the fault propagation relationship, determine the root cause fault network entity among the abnormal network entities on the knowledge graph of the target network, thereby realizing automatic reasoning and positioning of the root cause of the network fault.
[0069] Optionally, the fault propagation relationship between network entities can also be determined by other devices and sent to the management device. The way in which other devices determine the fault propagation relationship between network entities can refer to the way in which the above-mentioned management device determines the fault propagation relationship between network entities, and the embodiments of the present application will not be elaborated here.
[0070] Step 207: Identify one or more root cause failure network entities on the knowledge graph of the target network.
[0071] Optionally, after the management device identifies the root cause fault network entity on the knowledge graph of the target network, the knowledge graph of the target network can be sent to the OSS or other terminal devices connected to the management device for display by the OSS or terminal devices, so that operation and maintenance personnel can view the root cause fault network entity in the target network and quickly locate the fault, thereby improving the efficiency of fault repair, that is, it can shorten the time taken for the network device to switch from a faulty state to a working state. The time taken for the network device to switch from a faulty state to a working state can also be called the mean time to recovery (MTTR).
[0072] Optionally, in an embodiment of the present application, the management device may include one device or multiple devices. When the management device includes one device, the above steps 201 to 207 are all performed by the device. Alternatively, when the management device includes multiple devices, such as a first device and a second device, the above steps 201 to 205 and step 207 can be performed by the first device, and the above step 206 can be performed by the second device, that is, the first device generates a knowledge graph of the target network and identifies abnormal network entities on the knowledge graph of the target network; the first device sends the knowledge graph with abnormal network entities to the second device, and the second device determines the root cause fault network entity in the abnormal network entity based on the fault propagation relationship between network entities; the second device sends the identification of the root cause fault network entity to the first device; the first device identifies the root cause fault network entity on the knowledge graph of the target network based on the received identification of the root cause fault network entity.
[0073] The order of steps in the data processing method provided in the embodiment of the present application can be adjusted appropriately, and the steps can also be increased or decreased accordingly according to the situation. For example, step 207 can also be not executed. After the management device determines the root cause fault network entity on the knowledge graph of the target network, it directly outputs the root cause abnormal event corresponding to the root cause fault network entity. Any technician familiar with the technical field can easily think of a change within the technical scope disclosed in this application, which should be covered within the protection scope of this application, so it will not be repeated.
[0074] In summary, in the data processing method provided by the embodiment of the present application, after obtaining the network data of the target network, the management device processes the network data to generate a knowledge graph of the target network. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationship between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain network information in the target network, it can be done by viewing the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of the user in obtaining network information. In addition, the management device can also identify abnormal network entities on the knowledge graph of the target network after a failure occurs in the target network, and intuitively reflect the abnormal location in the target network. Furthermore, the management device can also determine the root cause fault network entity in the abnormal network entity on the knowledge graph of the target network based on the fault propagation relationship between network entities, thereby realizing automatic reasoning and positioning of the root cause of the network failure.
[0075] It should be noted that the management device may include multiple devices, each of which is used to execute different steps. If two adjacent steps are executed by different devices, the device that executes the previous step sends the execution result of the previous step to the device that executes the next step. Assume that the management device includes device A and device B. In one implementation, steps 201-203 can be executed by device A, and steps 204-207 are executed by device B, wherein device A sends the initial knowledge graph generated by steps 201-203 to device B, and device B executes steps 204-207 based on the initial knowledge graph received from device A. In another implementation, steps 201-205 can be executed by device A, and steps 206-207 are executed by device B, wherein device A sends the first knowledge graph generated by steps 201-205 to device B, and device B executes steps 206-207 based on the first knowledge graph received from device A.
[0076] Figure 5 Schematic diagram of a data processing device provided in an embodiment of the present application. The data processing device can be applied to Figure 1 The management device 101 in the application scenario shown. Figure 5 As shown, the device 50 includes:
[0077] The first acquisition module 501 is used to acquire network data of the target network, where the network data includes the network topology of the target network and device information of multiple network devices in the target network, where the device information includes one or more of interface configuration information, protocol configuration information and service configuration information.
[0078] The extraction module 502 is used to extract multiple knowledge graph triples from the network data, each knowledge graph triple includes two network entities and the relationship between the two network entities, and the type of the network entity is a network device, an interface, a protocol or a service.
[0079] The generation module 503 is used to generate a knowledge graph of the target network based on multiple knowledge graph triples.
[0080] In summary, the data processing device provided in the embodiment of the present application, after the management device obtains the network data of the target network through the first acquisition module, processes the network data and generates the knowledge graph of the target network through the generation module. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationship between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain network information in the target network, it can be done by viewing the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of the user in obtaining network information.
[0081] Optionally, the device information also includes routing table entries.
[0082] Optionally, the relationship between the two network entities is a dependency relationship, a subordinate relationship, or a peer relationship.
[0083] Alternatively, if Figure 6 As shown, the device 50 also includes:
[0084] The second acquisition module 504 is used to acquire abnormal events generated in the target network when a fault occurs in the target network;
[0085] The identification module 505 is used to identify abnormal network entities that generate abnormal events in the target network on the knowledge graph of the target network.
[0086] Optionally, the abnormal event includes one or more of an alarm log, a state change log, and an abnormal key performance indicator.
[0087] Alternatively, if Figure 7 As shown, the device 50 also includes:
[0088] A first determination module 506 is used to determine one or more root cause fault network entities from all abnormal network entities on the knowledge graph of the target network based on the fault propagation relationship between network entities;
[0089] The identification module 505 is further used to identify one or more root cause failure network entities on the knowledge graph of the target network.
[0090] Alternatively, if Figure 8 As shown, the device 50 also includes:
[0091] The third acquisition module 507 is used to acquire multiple knowledge graph samples, each of which is marked with all abnormal network entities and root cause failure network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a failure occurs in the network to which the knowledge graph sample belongs;
[0092] The second determination module 508 is used to determine the fault propagation relationship based on multiple knowledge graph samples.
[0093] Optionally, the network to which the knowledge graph sample belongs is the target network, or the network to which the knowledge graph sample belongs is another network of the same network type as the target network.
[0094] Optionally, an extraction module is used to:
[0095] Based on the abstract business model corresponding to the network type of the target network, multiple knowledge graph triples are extracted from the network data. The abstract business model is used to reflect the relationship between different network entities.
[0096] In summary, the data processing device provided by the embodiment of the present application, after the management device obtains the network data of the target network through the first acquisition module, processes the network data and generates the knowledge graph of the target network through the generation module. The knowledge graph of the target network can intuitively reflect the network entities in the target network and the relationship between different network entities, that is, it can realize the visualization of network data. When the user needs to obtain the network information in the target network, it can be done by viewing the knowledge graph of the target network without manually querying the command line, thereby improving the efficiency of the user in obtaining network information. In addition, the management device can also identify the abnormal network entity on the knowledge graph of the target network through the identification module after the target network fails, and intuitively reflect the abnormal location in the target network. Furthermore, the management device can also determine the root cause fault network entity in the abnormal network entity on the knowledge graph of the target network through the first determination module based on the fault propagation relationship between network entities, thereby realizing the automatic reasoning and positioning of the root cause of the network fault.
[0097] Fig. 9 is a block diagram of a data processing device provided in an embodiment of the present application. The data processing device may be as follows Figure 1 The management device in the application scenario shown in FIG. Fig. 9 As shown, the management device 90 includes: a processor 901 and a memory 902 .
[0098] A memory 902, used to store a computer program, wherein the computer program includes program instructions;
[0099] Processor 901 is used to call a computer program to implement the following Figure 2 The data processing method shown.
[0100] Optionally, the management device 90 further includes a communication bus 903 and a communication interface 904 .
[0101] The processor 901 includes one or more processing cores, and the processor 901 executes various functional applications and data processing by running computer programs.
[0102] The memory 902 may be used to store computer programs. Optionally, the memory may store an operating system and an application unit required for at least one function. The operating system may be a real-time operating system (RTX), LINUX, UNIX, WINDOWS, or OS X.
[0103] There may be multiple communication interfaces 904, and the communication interfaces 904 are used to communicate with other devices, such as a control device or a network device.
[0104] The memory 902 and the communication interface 904 are connected to the processor 901 via a communication bus 903 , respectively.
[0105] Fig.10 Schematic diagram of a fault root cause location device provided in an embodiment of the present application. The fault root cause location device can be applied to Figure 1 The management device 101 in the application scenario shown. Fig.10 As shown, the device 100 includes:
[0106] Acquisition module 1010 is used to acquire a first knowledge graph of a target network in which a fault occurs, wherein the first knowledge graph identifies abnormal network entities in the target network that generate abnormal events, and the types of network entities in the first knowledge graph are network devices, interfaces, protocols, or services. Specifically, the first knowledge graph may be received from other devices, or the first knowledge graph may be generated by executing steps 203-205. In the process of generating the first knowledge graph, an initial knowledge graph may be generated or received from other devices, and the abnormal network entities in the target network that generate the abnormal events may be identified on the initial knowledge graph to obtain the first knowledge graph.
[0107] The determination module 1020 is used to determine one or more root cause fault network entities among all the abnormal network entities on the first knowledge graph of the target network based on the fault propagation relationship between the network entities.
[0108] The apparatus 100 may further include an identification module 1030, which is used to identify the one or more root cause failure network entities on the knowledge graph of the target network.
[0109] Furthermore, the acquisition module 1010 is also used to acquire multiple knowledge graph samples, each of which is respectively marked with all abnormal network entities and root cause fault network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a fault occurs in the network to which the knowledge graph sample belongs; accordingly, the determination module 1020 is also used to determine the fault propagation relationship based on the multiple knowledge graph samples. The network to which the knowledge graph sample belongs is the target network, or the network to which the knowledge graph sample belongs is another network of the same network type as the target network.
[0110] The embodiment of the present application further provides a computer storage medium, wherein instructions are stored on the computer storage medium, and when the instructions are executed by a processor, the following is implemented: Figure 2 The data processing method shown.
[0111] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0112] In the embodiments of the present application, the terms "first", "second" and "third" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. The term "at least one" means one or more, and the term "plurality" means two or more, unless otherwise expressly defined.
[0113] The term "and / or" in this application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0114] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the concept and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for locating the root cause of a fault. It is characterized in that The method comprises: Obtain a first knowledge graph of a target network where a fault occurs, the first knowledge graph comprising a plurality of knowledge graph triples, each of the knowledge graph triples comprising two network entities and a relationship between the two network entities, the relationship between the two network entities being a dependency relationship, a subordinate relationship or a peer relationship, the plurality of knowledge graph triples being extracted from network data of the target network, the network data comprising a networking topology of the target network and device information of a plurality of network devices in the target network, the device information comprising one or more of interface configuration information, protocol configuration information and service configuration information, an abnormal network entity generating an abnormal event in the target network being identified on the first knowledge graph, and the type of the network entity on the first knowledge graph being a network device, an interface, a protocol or a service; Based on the fault propagation relationship between network entities, one or more root cause fault network entities are determined from all the abnormal network entities on the first knowledge graph of the target network.
2. The method according to claim 1, It is characterized in that The obtaining of the first knowledge graph of the target network having a fault includes: When a fault occurs in the target network, obtaining an abnormal event generated in the target network; An abnormal network entity that generates the abnormal event in the target network is identified on an initial knowledge graph of the target network to obtain the first knowledge graph, wherein the initial knowledge graph is generated based on network data of the target network.
3. The method according to claim 1, It is characterized in that The device information also includes routing table entries.
4. The method according to claim 2, It is characterized in that Before identifying the abnormal network entity generating the abnormal event in the target network on the initial knowledge graph of the target network, the method further includes: Acquire network data of the target network; Extracting the plurality of knowledge graph triples from the network data; An initial knowledge graph of the target network is generated according to the multiple knowledge graph triples.
5. The method according to claim 4, It is characterized in that The step of extracting a plurality of knowledge graph triples from the network data comprises: Based on an abstract business model corresponding to the network type of the target network, the multiple knowledge graph triples are extracted from the network data, and the abstract business model is used to reflect the relationship between different network entities.
6. The method according to any one of claims 1 to 5, It is characterized in that After determining one or more root-cause fault network entities among all the abnormal network entities on the first knowledge graph of the target network, the method further includes: The one or more root cause failure network entities are identified on a first knowledge graph of the target network.
7. The method according to any one of claims 1 to 5, It is characterized in that The abnormal event includes one or more of an alarm log, a state change log, and an abnormal key performance indicator.
8. The method according to any one of claims 1 to 5, It is characterized in that The method further comprises: Acquire multiple knowledge graph samples, each of which is marked with all abnormal network entities and root cause fault network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a fault occurs in the network to which the knowledge graph sample belongs; Based on the multiple knowledge graph samples, the fault propagation relationship is determined.
9. The method according to claim 8, It is characterized in that The network to which the knowledge graph sample belongs is the target network, or the network to which the knowledge graph sample belongs is another network of the same network type as the target network.
10. A fault root cause location device, It is characterized in that The device comprises: A first acquisition module is used to acquire a first knowledge graph of a target network where a fault occurs, wherein the first knowledge graph includes a plurality of knowledge graph triples, each of which includes two network entities and a relationship between the two network entities, wherein the relationship between the two network entities is a dependency relationship, a subordinate relationship, or a peer relationship, wherein the plurality of knowledge graph triples are extracted from network data of the target network, wherein the network data includes a networking topology of the target network and device information of a plurality of network devices in the target network, wherein the device information includes one or more of interface configuration information, protocol configuration information, and service configuration information, wherein an abnormal network entity generating an abnormal event in the target network is identified on the first knowledge graph, and the type of the network entity on the first knowledge graph is a network device, an interface, a protocol, or a service; The first determination module is used to determine one or more root cause fault network entities among all the abnormal network entities on the first knowledge graph of the target network based on the fault propagation relationship between the network entities.
11. The device according to claim 10, It is characterized in that The first acquisition module is used to acquire abnormal events generated in the target network when a fault occurs in the target network, and to identify abnormal network entities that generate the abnormal events in the target network on an initial knowledge graph of the target network to obtain the first knowledge graph, wherein the initial knowledge graph is generated based on network data of the target network.
12. The device according to claim 10, It is characterized in that The device information also includes routing table entries.
13. The device according to claim 11, It is characterized in that The device also includes: A second acquisition module is used to acquire network data of the target network before identifying an abnormal network entity that generates the abnormal event in the target network on an initial knowledge graph of the target network; An extraction module, used for extracting the plurality of knowledge graph triples from the network data; A generation module is used to generate an initial knowledge graph of the target network based on the multiple knowledge graph triples.
14. The device according to claim 13, It is characterized in that The extraction module is used to: Based on an abstract business model corresponding to the network type of the target network, the multiple knowledge graph triples are extracted from the network data, and the abstract business model is used to reflect the relationship between different network entities.
15. The device according to any one of claims 10 to 14, It is characterized in that The device also includes: An identification module is used to identify one or more root cause fault network entities on the first knowledge graph of the target network after determining the one or more root cause fault network entities among all the abnormal network entities on the first knowledge graph of the target network.
16. The device according to any one of claims 10 to 14, It is characterized in that The abnormal event includes one or more of an alarm log, a state change log, and an abnormal key performance indicator.
17. The device according to any one of claims 10 to 14, It is characterized in that The device also includes: A third acquisition module is used to acquire multiple knowledge graph samples, each of which is marked with all abnormal network entities and root cause failure network entities that generate abnormal events in the network to which the knowledge graph sample belongs when a failure occurs in the network to which the knowledge graph sample belongs; The second determination module is used to determine the fault propagation relationship based on the multiple knowledge graph samples.
18. The device according to claim 17, It is characterized in that The network to which the knowledge graph sample belongs is the target network, or the network to which the knowledge graph sample belongs is another network of the same network type as the target network.
19. A fault root cause location device, It is characterized in that include: Processor and memory; The memory is used to store a computer program, wherein the computer program includes program instructions; The processor is used to call the computer program to implement the fault root cause locating method according to any one of claims 1 to 9.
20. A computer storage medium, It is characterized in that The computer storage medium stores instructions, and when the instructions are executed by the processor, the fault root cause locating method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Network topology layout method and apparatus
CN107749803A
Method and apparatus for efficient problem resolution via incrementally constructed causality model based on history data
US20090055684A1