Failure estimation device, failure estimation method, and program

The fault estimation device improves fault location accuracy by integrating physical and logical connection relationships in network management, addressing the limitations of existing technologies in handling ripple effects from logical connections, thus enhancing fault identification in diverse networks.

WO2025163825A1PCT designated stage Publication Date: 2025-08-07NT T INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/003134
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing fault detection technologies in networks fail to accurately identify fault locations due to their inability to consider ripple effects from logical connection relationships, leading to incomplete extraction of necessary alarms and reduced accuracy in fault identification, especially in networks with diverse protocols and services.

Method used

A fault estimation device that utilizes a topology database managing physical and logical connection relationships, enabling the search for a fault-related alarm occurrence range by considering both physical and logical connections, thereby expanding the range of alarms considered for fault identification.

Benefits of technology

This approach enhances fault location estimation accuracy without increasing upfront operational work, by incorporating logical connections and reducing the need to define alarm ranges individually for various network configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024003134_07082025_PF_FP_ABST
    Figure JP2024003134_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A failure estimation device according to an embodiment of the present invention is provided with a topology database and a failure-related alarm occurrence range search function. The topology database is a database that manages information specifying the connection relationships between a plurality of nodes that constitute a target network, on the basis of a resource management model defined by a unified standard that includes physical and logical layers. On the basis of a physical resource to be faulted, the failure-related alarm occurrence range search function refers to the topology database and searches for a failure-related alarm occurrence range caused by the physical resource.
Need to check novelty before this filing date? Find Prior Art

Description

Fault estimation device, fault estimation method, and program

[0001] One aspect of the present invention relates to a fault estimation device, a fault estimation method, and a program.

[0002] When a fault or failure occurs in a network, some kind of alarm is generated. There are technologies that capture and analyze messages such as alarms to estimate the cause of the fault, the location where it occurred, etc. A typical example is a technology that uses If-then rules. For example, for each fault case, If-then rules are defined in advance, with the If conditions being the "type of alarm" and the "positional relationship between the fault location and the alarm occurrence location," and analysis is performed based on these rules. Creating If-then rules requires human experience and skill, so a technology that automates the creation of rules is known (see Patent Document 1).

[0003] Patent No. 6637854

[0004] As mentioned above, there is a rule-based fault detection technology that uses the alarm type and the relative position from the fault location. However, existing technology only takes into account the physical connection relationships, and does not consider the range of alarms that have been generated due to ripple effects based on logical connection relationships. As a result, it is not possible to extract all alarms necessary for fault identification as rule conditions, and there is room for improvement in the accuracy of fault identification.

[0005] Since the configuration of real networks varies widely, it is not realistic to individually organize the range of alarms that should be adopted in rules according to the wide variety of protocols and services.There is a demand for technology that can reduce the work required to define the range of alarms that should be generated.

[0006] Therefore, an object of the present invention is to provide a technique that can estimate the location of a failure with high accuracy without increasing the amount of prior operation.

[0007] According to an embodiment, a fault estimation device includes a topology database and a fault-related alarm occurrence range search function. The topology database is a database that manages information that defines the connection relationships between multiple nodes that make up a target network, based on a resource management model defined by a unified standard including physical layers and logical layers. The fault-related alarm occurrence range search function refers to the topology database based on a physical resource that is a target of a fault, and searches for a fault-related alarm occurrence range caused by the physical resource.

[0008] According to one aspect of the present invention, it is possible to provide a technique that can estimate a fault location with high accuracy without increasing prior operation.

[0009] FIG. 1 is a block diagram showing an example of a fault estimation system and a network according to an embodiment. FIG. 2 is a block diagram showing an example of the hardware configuration of a fault estimation device according to an embodiment. FIG. 3 is a functional block diagram showing an example of the software configuration of a fault estimation device according to an embodiment. FIG. 4 is a functional block diagram showing details of a fault-related alarm occurrence range search function 24. FIG. 5 is a diagram showing a model definition related to a resource management model according to an embodiment. FIG. 6 is a diagram for explaining a parent-child relationship in a network configuration. FIG. 7 is a diagram showing an example of a resource management model. FIG. 8 is a diagram showing a legend in FIG. 7. FIG. 9 is a flowchart showing an example of a processing procedure of a fault estimation device according to an embodiment. FIG. 10 is a diagram showing an initial state of fault location search in the first embodiment. FIG. 11 is a diagram showing resources for which an alarm has occurred that have been searched for in step S1 from the state of FIG. 10. FIG. 12 is a diagram showing resources for which an alarm has occurred that have been searched for in step S2 from the state of FIG. 11. FIG. 13 is a diagram showing resources for which an alarm has occurred that have been searched for in step S3 from the state of FIG. 12. FIG. 14 is a diagram showing resources for which an alarm has occurred that have been searched for in step S4 from the state of FIG. 13. FIG. 15 is a diagram showing resources causing an alarm to be searched for by the procedure of step S4 from the state of FIG. 14. FIG. 16 is a diagram showing resources causing an alarm to be searched for by the procedure of step S5 from the state of FIG. 15. FIG. 17 is a diagram showing resources causing an alarm to be searched for by the procedure of step S6 from the state of FIG. 16. FIG. 18 is a diagram showing resources causing an alarm to be searched for by the procedure of step S7 from the state of FIG. 17. FIG. 19 is a diagram showing the initial state of failure location search in the second embodiment. FIG. 20 is a diagram showing resources causing an alarm to be searched for by the procedure of step S1 from the state of FIG. 19. FIG. 21 is a diagram showing resources causing an alarm to be searched for by the procedure of step S2 from the state of FIG. 20. FIG. 22 is a diagram showing resources causing an alarm to be searched for by the procedure of step S3 from the state of FIG. 21. FIG. 23 is a diagram showing resources causing an alarm to be searched for by the procedure of step S4 from the state of FIG. 22.Fig. 24 is a diagram showing resources for which an alarm has occurred that have been searched for by the procedure of step S4 from the state of Fig. 23. Fig. 25 is a diagram showing resources for which an alarm has occurred that have been searched for by the procedure of step S5 from the state of Fig. 24. Fig. 26 is a diagram showing resources for which an alarm has occurred that have been searched for by the procedure of step S6 from the state of Fig. 25. Fig. 27 is a diagram showing resources for which an alarm has occurred that have been searched for by the procedure of step S7 from the state of Fig. 26. Fig. 28 is a functional block diagram showing another example of the software configuration of a fault estimation device according to an embodiment.

[0010] 1 is a block diagram showing an example of a fault estimation system and a network according to an embodiment. The fault estimation system 1 shown in FIG. 1 estimates the location of a fault that has occurred in a network NW, for example. The fault estimation system 1 includes a monitoring device 2 and a fault estimation device 3.

[0011] The network NW is an example of a topology to be monitored by the fault estimation system 1. The network NW is composed of multiple communication devices PD having connection relationships. The communication devices PD are entities such as switches that constitute the network. The communication devices PD are also called physical devices. The example of FIG. 4 shows a case where the network NW is composed of eight communication devices PD1, PD2, PD3, PD4, PD5, PD6, PD7, and PD8.

[0012] The monitoring device 2 is, for example, a computer such as a server. The monitoring device 2 monitors the state of the network NW. When a failure occurs in the network NW, the monitoring device 2 collects events generated in the network NW. The events include, for example, the type of alarm (event type) issued from the communication device PD and information indicating the communication device PD (event-generating device) that issued the alarm. The monitoring device 2 inputs the collected group of events to the failure estimation device 3.

[0013] The failure estimation device 3 is, for example, a computer such as a server, etc. Based on the input event group, the failure estimation device 3 estimates what kind of failure has occurred in which communication device PD within the network NW.

[0014] 2 is a block diagram showing an example of the hardware configuration of the fault estimation device 3 according to the embodiment. As shown in FIG. 2, the fault estimation device 3 includes a control circuit 11, a communication module 12, a user interface 13, a storage 14, a drive 15, and a storage medium 16.

[0015] The control circuit 11 is a circuit that controls the overall components of the fault estimation device 3. The control circuit 11 includes a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), etc. The ROM of the control circuit 11 stores programs and the like used in various processes in the fault estimation device 3. The CPU of the control circuit 11 controls the entire fault estimation device 3 in accordance with the programs stored in the ROM of the control circuit 11. The RAM of the control circuit 11 is used as a work area for the CPU of the control circuit 11.

[0016] The communication module 12 is, for example, a circuit used for communicating data with the monitoring device 2. The user interface 13 is an interface that manages communication between the user and the control circuit 11. The user interface 13 includes input devices and output devices. The input devices include, for example, a keyboard, a touch panel, and operation buttons. The output devices include, for example, an LCD (Liquid Crystal Display) or an EL (Electroluminescence) display. The user interface 13 converts input from the user into an electrical signal and then transmits it to the control circuit 11. The user interface 13 outputs the execution result based on the user input to the user.

[0017] The storage 14 includes, for example, a hard disk drive (HDD) or a solid state drive (SSD). The storage 14 stores information used in various processes in the failure estimation device 3.

[0018] The drive 15 is a device for reading software stored in the storage medium 16. The drive 15 includes, for example, a CD (Compact Disk) drive and a DVD (Digital Versatile Disk) drive.

[0019] The storage medium 16 is a medium that stores software electrically, magnetically, optically, mechanically, or chemically. The storage medium 16 may store programs for executing various processes in the fault estimation device 3.

[0020] 3 is a functional block diagram showing an example of the software configuration of the fault estimation device according to the embodiment. A case will be described with reference to FIG. 3 in which information on the fault-related alarm occurrence range is created in advance for all nodes.

[0021] 3, the storage 14 (FIG. 2) of the fault estimation device 3 stores the topology database (DB) 22, the rule database 25, and the fault-related alarm occurrence range database 27 of FIG.

[0022] The topology database 22 is a database of information (network topology) that defines the connection relationships between multiple communication devices SW (nodes) that make up the network NW. In the topology database 22, the resources of the target network NW are managed based on a resource management model that is defined by a unified standard that includes a physical layer and a logical layer.

[0023] The rule database 25 is a database of information in which if-then rules are predefined for each failure case, with the if conditions being the "type of alarm" and the "positional relationship between the failure location and the alarm occurrence location."

[0024] The fault-related alarm occurrence range database 27 is a database that registers information used when generating fault determination rules and estimating the cause of a fault.

[0025] Furthermore, the CPU of the control circuit 11 of the fault estimation device 3 loads a program stored in the ROM of the control circuit 11 or the storage medium 16 into the RAM of the control circuit 11. The CPU of the control circuit 11 then interprets and executes the program loaded into the RAM of the control circuit 11. This allows the fault estimation device 3 to function as a computer equipped with a fault-related alarm occurrence range searching function 24 and a fault cause estimation function 26.

[0026] 3, (a) a maintenance person pre-registers fault information for rule learning (fault resource, cause, time of fault occurrence, etc.) in the rule generation function 28. (b) The topology database 22 works in conjunction with other external databases to hold the latest network topology.

[0027] The fault-related alarm occurrence range search function 24 acquires topology information from the topology database 22 (c), and registers the fault-related alarm occurrence range in the fault-related alarm occurrence range database 27 (d).

[0028] The rule generation function 28 acquires the fault-related alarm occurrence range from the fault-related alarm occurrence range database 27 based on the fault information registered by the maintenance person or information acquired from the external monitoring system 31, etc. (e) and generates a rule for identifying the fault (f). The generated rule is registered in the rule database 25.

[0029] The failure cause estimation function 26 acquires topology information from the topology database 22 (h) in response to an event group (g) provided, for example, from a monitoring system, acquires a failure-related alarm occurrence range from the failure-related alarm occurrence range database 27 (j), and acquires rules from the rule database 25 (i). Based on this information, the failure cause estimation function 26 then estimates the cause of the failure that occurred in the network NW (i.e., the name of the failure and the faulty device), and presents the estimation result to a maintenance person (k). The maintenance person checks the estimation result and determines whether the estimation result is correct.

[0030] 4 is a functional block diagram showing details of the fault-related alarm occurrence range searching function 24. The fault-related alarm occurrence range searching function 24 includes a fault-related alarm occurrence range searching unit 24a, a fault-related alarm occurrence range registration unit 24b, and a resource information reference unit 24c.

[0031] The fault-related alarm occurrence range searching unit 24a searches for a fault-related alarm occurrence range based on a specified physical resource, and defines a fault-related alarm occurrence range corresponding to this physical resource.

[0032] The fault-related alarm occurrence range registering unit 24 b registers the fault-related alarm occurrence range defined by the fault-related alarm occurrence range searching unit 24 a in the fault-related alarm occurrence range database 27 .

[0033] The resource information reference unit 24c references the resource information registered in the topology database 22 and notifies the fault-related alarm occurrence range search unit 24a of the connection relationships of the physical layer and logical layer required for searching for a path.

[0034] 5 is a diagram showing a model definition related to a resource management model according to an embodiment. In FIG. 5, entities such as PPort (Physical Port), PLink (Physical Link), EQP (Equipment), EH (Equipment Holder), and PD (Physical Device) are defined in the physical layer.

[0035] A PPort is a physical port of a network device. There are also virtual PPorts.

[0036] A PLink represents a physical line connecting network devices, such as an optical fiber, a UTP cable, or a wireless connection.

[0037] EQP represents a network device package, card, optical module (SFP, etc.), etc. A PPort is connected to the EQP.

[0038] An EH is a container for equipment, and represents a rack, chassis, slot, etc. A PD represents the actual device to be managed. The physical configuration is represented by the EH / EQP connected from the PD.

[0039] In the logical layer, entities called TPE (Termination Point Encapsulation), FRE (Forwarding Relationship Encapsulation), NFD (Network Forwarding Domain), and TL (Topological Link) are defined.

[0040] TPE represents the end point of a logical layer. TPE refers to a TPE in a lower logical layer or a PPort in a physical layer, and represents the hierarchical relationship of layers.

[0041] The TPE includes a TCP (Termination Connection Point) and a CP (Connection Point).

[0042] TCP is a termination point: TCP is responsible for terminating connectivity within the layer.

[0043] A CP is a relay point, which has the role of relaying connectivity within a layer.

[0044] FRE represents the flow of information between endpoints.

[0045] FRE includes NC (Network Connection), LC (Link Connection), and XC (Cross Connection).

[0046] NC indicates the connectivity from a TCP in a layer to a reachable TCP. The structure of the NW within a layer is represented by the subordinate LC and XC.

[0047] LC represents connectivity between devices based on the lower layer NC.

[0048] XC represents connectivity within a device based on the NFD of the lower layer, and corresponds to the routing function of a router or the switch function of a switch.

[0049] The NFD represents the transfer capability of a device, which is the basis for creating an XC. The NFD is basically used only in the Logical Device (LD) layer.

[0050] The TL represents logical connectivity between devices in the LD layer, and is located at the lowest level of the logical layer.

[0051] 6 is a diagram for explaining parent-child relationships in a network configuration. The parent-child relationships in physical resources are shown as a prerequisite. It is also assumed that the components of each device in the physical resources form a tree structure. That is, the components of each device in the physical resources form a tree structure, and the parent-child relationships are determined in this tree structure. For example, from the perspective of a communication port, the package to which the communication port belongs and the chassis to which the package belongs are the direct parents from the perspective of the communication port.

[0052] In Figure 6, the left side of the arrow (→) is the parent node and the right side is the child node. In physical resources, a child node cannot have multiple parent nodes. Note that the symbols "<" and ">" may be used instead of arrows (as will be explained later).

[0053] For example, in a box-type device, PD is the parent node of PPort.

[0054] When managing a chassis / card, the PD is the parent of the chassis, the chassis is the parent of the card, and the card is the parent of the PPort.

[0055] In a large-scale device such as a transmission device, the PD is the parent of the Rack, the Rack is the parent of the Shelf, the Shelf is the parent of the Card, the Card is the parent of the Module, and the Module is the parent of the PPort.

[0056] By utilizing the layer-by-layer relationships of logical resources, the connectivity between the termination points of the upper layer is composed of multiple combinations of the connectivity between the termination points of information transfer in the lower layer.

[0057] FIG. 7 is a diagram showing an example of a resource management model according to an embodiment. For example, a network model can be considered in which multiple logical layers are layered on top of a physical layer. The logical layer further comprises multiple sublayers, with logical layer 0 (LD), logical layer 1 (transmission), logical layer 2 (ETH), and logical layer 3 (IP) hierarchically organized in descending order of proximity to the physical layer. Each entity is distinguished by line type or hatching, and a legend is shown in FIG. 8. The following description will follow the legend in FIG. 8.

[0058] 9 is a flowchart showing an example of a processing procedure of the fault estimation device according to the embodiment. Prior to explaining the details, the definition of the fault-related alarm generation range according to the embodiment will be summarized as follows.

[0059] <1> The network is managed using a unified model (resource management model) that includes the physical layer and the logical layer.

[0060] <2> Acquire the physical resource to be faulted.

[0061] <3> The alarm occurrence range (candidate occurrence locations) for the acquired physical resource (fault location) is searched for according to the following steps S1 to S7, and for each end resource (reached resource), route information expressed by the entity type with the faulty resource as the start resource is recorded. The steps <1> to <3> are explained in detail below.

[0062] (Step S1) The fault estimation device 3 searches for physical resources in the device to which the failed physical resource belongs.

[0063] (Step S2) The fault estimation device 3 searches for physical resources in adjacent devices of the failed physical resource or its direct descendants (see FIG. 6).

[0064] (Step S3) The fault estimation device 3 searches for a communication termination point (TCP) of the upper logical layer (LD layer) corresponding to the physical port (pport) from among the physical resources searched for in step S1, a connection (TL) between two points including that communication termination point (TCP), and the other communication termination point (TCP) corresponding to that TL.

[0065] (Step S4) The fault estimation device 3 searches for a connection (NC) between two points in a higher layer that includes the connection (TL or NC) between two points in the logical layer searched for in step S3 or this step S4 (step S4 is repeated recursively until the top layer is reached), and for its communication termination point (TCP).

[0066] (Step S5) The fault estimation device 3 searches for a communication termination point (TCP) in a lower logical layer that corresponds to the communication termination point (TCP) in the logical layer searched for in step S4.

[0067] (Step S6) The fault estimation device 3 searches for a physical port (pport) corresponding to the communication termination point (TCP) of the logical layer searched for in step S5, and for a physical resource that is a direct parent of that physical port.

[0068] (Step S7) The fault estimation device 3 searches for a communication termination point (TCP) in the upper layer corresponding to the physical port from among the physical resources searched for in step S2.

[0069] The fault estimation device 3 registers route information from the end point resource and the start point resource that fall within the alarm occurrence range searched for in steps S1 to S7 in the fault-related alarm occurrence range database 27. If there are multiple routes that reach the end point from the start point, all of these may be registered.

[0070] First Embodiment In the first embodiment, it is assumed that a failure occurs in (PD) in a failure node (physical layer), that is, the failed resource is (PD).

[0071] 10 is a diagram showing the initial state of fault location search in the first embodiment. It is assumed that the leftmost starting node (router) has failed, and that a failure has occurred in the PD of this failed node.

[0072] Fig. 11 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S1 from the state of Fig. 10. (1) in Fig. 11 shows a path represented by a parent-child relationship of PD > EQP > PP. (2) shows a path represented by a parent-child relationship of PD > EQP.

[0073] Here, the symbol "<" indicates that the left node is a child node of the right node. The symbol ">" indicates that the left node is a parent node of the right node. The symbol "=" is used when there is no parent-child relationship between the left node and the right node. Hereinafter, a representation format of a path in which multiple nodes are connected with the symbols "<", ">", and "=" is also referred to as a "path representation". The leftmost node and the rightmost node in a path represented by a path representation are also referred to as the starting node and the ending node, respectively.

[0074] Fig. 12 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S2 from the state of Fig. 11. (3) in Fig. 12 shows a path expressed by a parent-child relationship of PD>EQP>PP=PL=PP<EQP<PD>EQP>PP. (4) shows a path expressed by a parent-child relationship of PD>EQP>PP=PL=PP<EQP<PD.

[0075] Fig. 13 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S3 from the state of Fig. 12. (5) in Fig. 13 shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP. (6) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP=NC. (7) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP=NC=TCP.

[0076] Fig. 14 is a diagram showing the alarm generating resources found by the procedure of step S4 from the state of Fig. 13. (8) in Fig. 14 shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP. (9) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP=NC. (10) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP=NC=TCP.

[0077] Figure 15 shows the alarm-generating resources searched for by the procedure of step S4 from the state of Figure 14. In Figure 14, the search did not reach the top layer (logical layer 3 (IP)), so step S4 is further executed recursively. As a result, the search of step S4 reaches the top layer in Figure 15.

[0078] 15 (11) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP<CP=XC=TCP, or PD>PP<TCP<TCP. (12) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP<CP=XC=TCP=NC, or PD>PP<TCP<TCP=NC. (13) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP<CP=XC=TCP=NC=TCP, or PD>PP<TCP<TCP=NC=TCP.

[0079] Fig. 16 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S5 from the state of Fig. 15. (14) in Fig. 16 shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP=NC=TCP>TCP. (15) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP<CP=XC=TCP=NC=TCP>TCP, or PD>PP<TCP<TCP=NC=TCP>TCP.

[0080] Fig. 17 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S6 from the state of Fig. 16. (16) in Fig. 17 shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP=NC=TCP>TCP>PP. (17) shows a route expressed by a parent-child relationship of PD>EQP>PP<TCP<TCP<CP=XC=TCP=NC=TCP>TCP>PP, or PD>PP<TCP<TCP=NC=TCP>TCP>PP. (18) shows a route expressed by a parent-child relationship: PD>EQP>PP<TCP<TCP=NC=TCP>TCP>PP<EQP<PD, or PD>EQP>PP<TCP<TCP<CP=XC=TCP=NC=TCP>TCP>PP<PD, or PD>PP<TCP<TCP=NC=TCP>TCP>PP<PD.

[0081] Figure 18 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S7 from the state of Figure 17. (19) in Figure 18 shows a path expressed by a parent-child relationship of PD>EQP>PP=PL=PP<EQP<PD>EQP>PP<TCP<TCP. (20) shows a path expressed by a parent-child relationship of PD>EQP>PP=PL=PP<EQP<PD>EQP>PP<TCP.

[0082] Second Embodiment In a second embodiment, it is assumed that a failure occurs in (EQP) in a failure node (physical layer), that is, the failure resource is (EQP).

[0083] 19 is a diagram showing the initial state of fault location search in the second embodiment. It is assumed that a node (L2 switch) adjacent to the leftmost starting node (router) has failed, and that a failure has occurred in the EQP of this failed node.

[0084] Fig. 20 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S1 from the state of Fig. 19. (21) in Fig. 20 shows a path expressed by a parent-child relationship of EQP>PP. (22) shows a path expressed by a parent-child relationship of EQP<PD>EQP>PP.

[0085] Fig. 21 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S2 from the state of Fig. 20. (23) in Fig. 21 shows a path expressed by a parent-child relationship of EQP>PP=PL=PP<EQP<PD>PP. (24) shows a path expressed by a parent-child relationship of EQP>PP=PL=PP.

[0086] Fig. 22 is a diagram showing the alarm generating resource searched for by the procedure of step S3 from the state of Fig. 21. (25) in Fig. 22 shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP. (26) shows a route expressed by a parent-child relationship of EQP>PP<TCP. (27) shows a route expressed by a parent-child relationship of EQP<PD>EQP>PP<TCP.

[0087] Fig. 23 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S4 from the state of Fig. 22. (28) in Fig. 23 shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP. (29) shows a route expressed by a parent-child relationship of EQP<PD>EQP>PP<TCP<TCP=NC.

[0088] Figure 24 shows the alarm-generating resources searched for by the procedure of step S4 from the state of Figure 23. In Figure 23, the search did not reach the top layer (logical layer 3 (IP)), so step S4 is further executed recursively. As a result, the search of step S4 reaches the top layer in Figure 24.

[0089] 24 (30) shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP<CP<XC<TCP. (31) shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP<CP<XC<TCP=NC.

[0090] Figure 25 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S5 from the state of Figure 24. (32) in Figure 25 shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP=NC=TCP>TCP. (33) shows a route expressed by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP<CP<XC<TCP=NC=TCP>TCP.

[0091] Figure 26 is a diagram showing the resources for which an alarm has occurred, searched for by the procedure of step S6 from the state of Figure 25. (34) in Figure 26 shows a path represented by a parent-child relationship of EQP<PD>EQP>PP<TCP=NC=TCP>PP<EQP<PD. (35) shows a path represented by a parent-child relationship of EQP<PD>EQP>PP<TCP<TCP=NC=TCP>TCP>PP<EQP<PD. (36) shows a path represented by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP=NC=TCP>TCP>PP<EQP. (37) shows a path represented by a parent-child relationship of EQP>PP<TCP=TL=TCP<TCP<CP<XC<TCP=NC=TCP>TCP>PP<PD.

[0092] Fig. 27 is a diagram showing the resource for which an alarm has occurred, searched for by the procedure of step S7 from the state of Fig. 26. (38) in Fig. 27 shows a path expressed by a parent-child relationship of EQP>PP=PL=PP<EQP<PD>PP<TCP.

[0093] The following items <1> to <4> show a generalized search procedure in the first and second embodiments: <1> Management is performed using a unified network management model that includes physical connections and logical connections.

[0094] <2> The alarm occurrence range is the connection destination that is directly (physically) connected to the faulty location.

[0095] <3> The range of the communication section whose logical connection is affected by the fault location and its termination point is determined as the alarm generation range by utilizing the relationship between the communication sections of the upper and lower layers.

[0096] <4> If there are physical ports in the physical layer and communication endpoints in the logical layer that are targets of the alarm generation range, the corresponding communication endpoints in the logical layer and communication ports in the physical layer and their direct parents are set as the alarm generation range. In other words, physical ports and logical communication endpoints are treated the same.

[0097] As described above, in the embodiment, a network is managed by setting a unified network management model that includes physical connections and logical connections. When a failure occurs, the destinations directly (physically) connected to the failure point are defined as the alarm generation range. Furthermore, in the embodiment, in addition to the destinations directly connected to the failure point, the alarm generation range is defined to include the communication section whose logical connection is affected by the failure point and the endpoint of that communication (= if the failure point is used due to the rules of the communication protocol, the communication section and the endpoint of that communication).

[0098] By searching the alarm occurrence range in this way, it becomes possible to handle alarms that lead to the identification of faults occurring in devices that are not directly connected physically in the fault detection rules, which improves the accuracy of fault estimation rules and reduces the work required to define the alarm occurrence range to be adopted in each rule in a wide variety of networks.

[0099] Existing fault detection rules for estimating the cause of network faults have the problem that they cannot handle alarms that lead to the identification of faults that occur from devices that are not directly physically connected, due to logical connection relationships at the protocol or service level.

[0100] It is considered that an alarm due to a failure may occur within the range of the communication section affected by the failure. Therefore, in the embodiment, the network to be managed is managed using a unified model, and elements corresponding to the range of the impact of the failure are defined as the alarm occurrence range based on the modeled elements and the relationships between layers. This makes it possible to extract and adopt the alarms necessary to identify the failure location.

[0101] With existing technology, the alarms used in the If section of a failure detection rule are limited to alarms of adjacent devices directly connected to the failed device by cable, etc., as the range in which related alarms are generated due to the impact of the failure. In other words, the search range is limited to the physical layer. Since it is believed that alarms effective for failure detection will be generated within the range affected by the failure, it is necessary to consider these as candidates for alarm generation locations to be used in the rule.

[0102] On the other hand, in large-scale networks, alarms that lead to the identification of faults may be generated not only by the physical connections between devices but also by logical connections between protocols and services, even from devices that are not directly physically connected.In a wide variety of networks, it is not realistic to organize the range of alarms that should be adopted in rules according to the individually different protocols and services.

[0103] Therefore, in the embodiment, the following viewpoints (1) and (2) are introduced when searching for the occurrence range of a fault-related alarm.

[0104] <1> The alarm occurrence range is defined to include not only physical connections but also logical connections.

[0105] (2) Generalize the definition of the alarm occurrence range so that it is independent of the communication protocols and services of a wide variety of networks.

[0106] The viewpoint of (1) is that the range in which an alarm that should be adopted in a rule occurs is considered to be the range in which the impact of the failure occurs, so in addition to "connection destinations that are directly (physically) connected to the failure point," the search range should also include "communication sections whose logical connections are affected by the failure point and the end points of that communication (= when the failure point is used in relation to the rules of the communication protocol, that communication section and the end points of that communication)."

[0107] The point of view of (2) is to realize a definition method that utilizes a unified network management model that includes logical connections.

[0108] By implementing a definition method that takes into account the points (1) and (2) regarding the definition of the range in which an alarm to be adopted in a rule will occur (the range in which an alarm may occur due to the effects of a fault), it becomes possible to adopt an appropriate alarm as a rule condition, thereby improving the accuracy of the fault estimation rule.

[0109] The above configuration makes it possible to extract alarms that should be considered when creating fault detection rules, thereby improving the accuracy of fault identification. Furthermore, by generalizing the definition of the alarm occurrence range, it is possible to reduce the work required to define the alarm occurrence range that should be adopted in each rule in a wide variety of networks. As a result, according to the embodiment, it is possible to estimate the fault location with high accuracy without increasing the amount of work required upfront.

[0110] The present invention is not limited to the above-described embodiment.

[0111] 28 is a functional block diagram showing another example of the software configuration of the fault estimation device according to the embodiment. A case will be described with reference to FIG. 28 where the fault estimation device 3 searches for information on the fault-related alarm occurrence range for a target node each time learning is performed.

[0112] 28, a maintenance person pre-registers failure information for rule learning (failure resource, cause, time of failure, etc.) in a rule generation function 28 (a). A topology database 22 holds the latest network topology. When a failure-related alarm occurrence range search function 24 receives a resource specification from a failure cause estimation function 26 or a rule generation function 28 ((o) or (r)), it refers to the topology database 22 (l) and registers the failure-related alarm occurrence range in a failure-related alarm occurrence range database 27 (m).

[0113] The rule generation function 28 acquires (p) the fault-related alarm occurrence range from the fault-related alarm occurrence range database 27 based on the fault information (a) registered by the maintenance person or the information (n) acquired from the external monitoring system 31, etc., and generates (f) a rule for identifying the fault. The generated rule is registered in the rule database 25.

[0114] The failure cause estimation function 26 acquires (q) targets of the failure-related alarm occurrence range from the failure-related alarm occurrence range database 27 in response to an event group (g) provided, for example, from a monitoring system, and acquires rules from the rule database 25. Then, based on this information, the failure cause estimation function 26 estimates the cause of the failure that occurred in the network NW (i.e., the name of the failure and the faulty device), and presents the estimation result to the maintenance person (k). The maintenance person checks the estimation result and determines whether the estimation result is correct.

[0115] Furthermore, the present invention can be variously modified in the implementation stage without departing from the spirit of the invention. Furthermore, each embodiment may be implemented in appropriate combination, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.

[0116] Furthermore, in the implementation stage, the components of this invention can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

[0117] 1...Fault estimation system 2...Monitoring device 3...Fault estimation device 11...Control circuit 12...Communication module 13...User interface 14...Storage 15...Drive 16...Storage medium 22...Topology database 24...Fault-related alarm occurrence range search function 24a...Fault-related alarm occurrence range search unit 24b...Fault-related alarm occurrence range registration unit 24c...Resource information reference unit 25...Rule database 26...Fault cause estimation function 27...Fault-related alarm occurrence range database 28...Rule generation function 30...Topology database 31...Monitoring system, etc.

Claims

1. A fault estimation device comprising: a topology database that manages information specifying the connection relationships between multiple nodes that make up a target network, based on a resource management model defined by a unified standard including the physical layer and the logical layer; and a fault-related alarm occurrence range search function that references the topology database based on the physical resource that is the target of a fault, and searches for the range of fault-related alarms caused by the physical resource.

2. The fault estimation device according to claim 1, wherein the fault-related alarm occurrence range search function comprises: a resource information reference unit that references resource information registered in the topology database; a fault-related alarm occurrence range search unit that searches for a fault-related alarm occurrence range based on the physical resource and defines a fault-related alarm occurrence range corresponding to the physical resource; and a fault-related alarm occurrence range registration unit that registers the defined fault-related alarm occurrence range in a fault-related alarm occurrence range database.

3. The fault estimation device according to claim 2, wherein the fault-related alarm occurrence range search unit: searches for physical resources within the node to which the physical resource belongs; searches for physical resources in adjacent nodes of the physical resource or direct descendants of the physical resource; searches for a communication termination point in a higher logical layer corresponding to a physical port, a connection between two points including that communication termination point, and the other communication termination point of a connection including that communication termination point; searches for a connection between two points in a higher layer including the connection between two points in the searched logical layer and its communication termination point; searches for a communication termination point in a lower logical layer corresponding to the communication termination point in the searched logical layer; searches for a physical port corresponding to the communication termination point in the searched logical layer and a physical resource that is a direct parent of that physical port; and searches for a communication termination point in a higher layer corresponding to a physical port among the searched physical resources.

4. The fault estimation device according to claim 3, wherein the fault-related alarm occurrence range search unit recursively searches for a connection between two points in an upper layer that includes the connection between the two points in the searched logical layer, and for that connection's communication termination point.

5. A fault estimation method comprising: managing, in a topology database, information defining the connection relationships of multiple nodes constituting a target network based on a resource management model defined by a unified standard including the physical layer and the logical layer; and searching, based on the physical resource to be faulted, for the range of occurrence of a fault-related alarm caused by the physical resource by referring to the topology database.

6. A program for causing a computer to execute the functions of the fault estimation device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Network management device, method and program

    WO2023233635A1