Failure estimation device and failure estimation method

The fault estimation device addresses inefficiencies in multi-layer networks by using unified network configuration and adaptable rule learning to efficiently estimate fault locations, reducing operator workload and processing load.

WO2025177563A1PCT designated stage Publication Date: 2025-08-28NT T INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/006628
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing rule-learning fault location estimation technologies struggle with increased workload and processing load in large-scale, multi-layer networks due to differences in relationships between resources, leading to cumbersome rule learning and inefficient fault location estimation.

Method used

A fault estimation device that stores network configuration information using a unified standard, learns rules based on resource management models, and estimates fault locations by defining relationships between alarms and faults, while adapting rule definitions to account for the presence or absence of logical layer resources.

Benefits of technology

Reduces the workload and processing load required for fault location estimation by enabling flexible rule application across varying network configurations, improving estimation efficiency and reducing the number of rules needed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024006628_28082025_PF_FP_ABST
    Figure JP2024006628_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A failure estimation device according to one embodiment of the present invention comprises a network configuration information storage unit, a rule learning control unit, and a failure location estimation function unit. The network configuration information storage unit stores network configuration information stipulating a connection relationship between a plurality of nodes constituting a target network, on the basis of a resource management model defined by a unified standard including a physical layer and a logical layer. The rule learning control unit learns, on the basis of the resource management model and the network configuration information, a rule indicating, by a route, a relationship between the location of occurrence of an alarm based on a failure occurring in the network and the location of occurrence of the failure, and registers the rule in a rule database. When a failure occurs in the network, the failure location estimation function unit estimates the location of occurrence of the failure, on the basis of an alarm group generated due to the failure and the rule stored in the rule database. The rule learning control unit changes the presence or absence of logical layer resources in a rule definition in accordance with a condition determined depending on the presence or absence of a return route in the logical layer, and registers the change in the rule database.
Need to check novelty before this filing date? Find Prior Art

Description

Fault estimation device and fault estimation method

[0001] One aspect of the present invention relates to a fault estimation device and a fault estimation method.

[0002] When a network failure occurs, maintenance personnel identify the location and cause of the failure through alarm analysis, normal / abnormal system isolation, etc. In order to minimize human intervention in network maintenance and operation work (zero-touch operation), technologies have been developed to autonomously estimate the location of the failure. For example, rule-learning fault location estimation technology is known.

[0003] The rule-learning fault location estimation technology is a technology for estimating the fault location based on rules that represent the relationship between alarms that characterize the fault and the fault location and cause. The rules are defined and generated in advance, for example, by the technology described in Patent Document 1.

[0004] Patent No. 6637854

[0005] The scale of communication networks is growing year by year. Large-scale networks, in particular, are often layered, combining a wide variety of protocols from the physical layer to the logical layer. When applying rule-learning fault location estimation technology to such multi-layer networks, differences in the relationships between resources described in the rule definitions can result in essentially identical faults being classified as separate faults. This increases the number of rules, and the operator's work (operation) required to learn new rules becomes cumbersome and time-consuming. Furthermore, the processing load for the estimation process when a fault occurs increases.

[0006] Therefore, an object of the present invention is to provide a technology that can reduce the workload involved in estimating a fault location and reduce the load.

[0007] A fault estimation device according to one aspect of the present invention includes a network configuration information storage unit, a rule learning control unit, and a fault location estimation function unit. The network configuration information storage unit stores network configuration information that defines the connection relationships between multiple nodes that constitute a target network, based on a resource management model defined by a unified standard including a physical layer and a logical layer. The rule learning control unit learns rules that indicate, using paths, the relationship between the location of an alarm generated based on a fault that has occurred in the network and the location of the fault, based on the resource management model and the network configuration information, and registers the rules in a rule database. When a fault occurs in the network, the fault location estimation function unit estimates the location of the fault based on the alarms generated due to the fault and the rules stored in the rule database. The rule learning control unit changes the presence or absence of a logical layer resource in the rule definition in accordance with a condition determined by the presence or absence of a return path in the logical layer, and registers the rule in the rule database.

[0008] According to one aspect of the present invention, it is possible to provide a technology that can reduce the workload involved in estimating a fault location and reduce the load.

[0009] FIG. 1 is a block diagram showing an example of a fault estimation system and a monitored network according to an embodiment. FIG. 2 is a block diagram showing an example of the hardware configuration of a fault estimation device according to an embodiment. FIG. 3 is a functional block diagram showing an example of the software configuration of a fault estimation device according to an embodiment. FIG. 4 is a diagram showing a model definition related to a resource management model according to an embodiment. FIG. 5 is a schematic diagram showing an example of a fault that occurs in a network. FIG. 6 is a diagram showing events and entities corresponding to FIG. 5. FIG. 7 is a diagram showing rules defined corresponding to FIG. 6. FIG. 8 is a schematic diagram showing another example of a fault that occurs in a network. FIG. 9 is a diagram showing a first rule that can be defined corresponding to FIG. 8. FIG. 10 is a diagram showing events and entities corresponding to FIG. 9. FIG. 11 is a diagram showing a second rule that can be defined corresponding to FIG. 8. FIG. 12 is a diagram showing events and entities corresponding to FIG. 11. FIG. 13 is a schematic diagram showing another example of a fault that occurs in a network. FIG. 14 is a diagram showing an example of an abbreviated notation of logical layers in a rule definition. FIG. 15 is a diagram showing events and entities corresponding to FIG. 13. FIG. 16 is a diagram for explaining the difference between rule notation in existing technology and the embodiment. FIG. 17 is a diagram for explaining the difference between rule notation in the existing technology and the embodiment. FIG. 18 is a diagram for explaining the difference between rule notation in the existing technology and the embodiment. FIG. 19 is a diagram for explaining the difference between rule notation in the existing technology and the embodiment. FIG. 20 is a diagram for further explaining rule notation in the embodiment. FIG. 21 is a diagram showing an example of a path expression of physics (PPort) -> physics (PPort). FIG. 22 is a diagram showing a specific example of the rule notation in FIG. 21. FIG. 23 is a diagram showing an example of a path expression of logic (TPE) -> logic (TPE). FIG. 24 is a diagram showing a specific example of the rule notation in FIG. 23.

[0010] 1 is a block diagram showing an example of a fault estimation system according to an embodiment and a monitored network. A monitoring system 1 and a fault estimation system 2 are provided at a base for monitoring a target network (NW) 100, for example, via a hub device 5. The fault estimation system 2 estimates the location of a fault that has occurred in the network 100, and includes a fault location estimation device 3 and a GUI display terminal 4 that can access the fault location estimation device 3.

[0011] 1, a monitoring system 1 receives alarms issued from a network 100 and periodically transmits the received alarms to a fault estimation system 2. The fault estimation system 2 generates rules and estimates the fault location based on the received alarms. A GUI display terminal 4 displays the fault location estimation results and countermeasures in a GUI environment, allowing an operator to take appropriate action.

[0012] 2 is a block diagram showing an example of the hardware configuration of the fault estimation device 3 according to the embodiment. As shown in FIG. 2, the fault estimation device 3 includes a control circuit 11, a communication module 12, a user interface 13, a storage 14, a drive 15, and a storage medium 16.

[0013] The control circuit 11 is a circuit that controls the overall components of the fault estimation device 3. The control circuit 11 includes a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), etc. The ROM of the control circuit 11 stores programs and the like used in various processes in the fault estimation device 3. The CPU of the control circuit 11 controls the entire fault estimation device 3 in accordance with the programs stored in the ROM of the control circuit 11. The RAM of the control circuit 11 is used as a work area for the CPU of the control circuit 11.

[0014] The communication module 12 is, for example, a circuit used for communication via the hub device 5. The user interface 13 is an interface that manages communication between the user and the control circuit 11. The user interface 13 includes input devices and output devices. The input devices include, for example, a keyboard, a touch panel, and operation buttons. The output devices include, for example, an LCD (Liquid Crystal Display) or an EL (Electroluminescence) display. The user interface 13 converts input from the user into an electrical signal and then transmits it to the control circuit 11. The user interface 13 outputs the execution result based on the user input to the user.

[0015] The storage 14 includes, for example, a hard disk drive (HDD) or a solid state drive (SSD). The storage 14 stores information used in various processes in the failure estimation device 3.

[0016] The drive 15 is a device for reading software stored in the storage medium 16. The drive 15 includes, for example, a CD (Compact Disk) drive and a DVD (Digital Versatile Disk) drive.

[0017] The storage medium 16 is a medium that stores software electrically, magnetically, optically, mechanically, or chemically. The storage medium 16 may store programs for executing various processes in the fault estimation device 3.

[0018] 3 is a functional block diagram showing an example of the software configuration of the fault estimation device according to the embodiment. The fault estimation device 3 includes functional blocks of a data acquisition unit 31 that acquires alarm groups and network configuration information, a rule learning control unit 32, a fault location estimation function unit 33, a countermeasure management function unit 34, an alarm information storage unit 36, a network configuration information storage unit 37, a rule database (DB) 38, a fault history / countermeasure history storage unit 39, a GUI unit 35, and an API unit 40.

[0019] The network configuration information storage unit 37 stores network configuration information that defines the connection relationships of multiple nodes that make up the target network, based on a resource management model defined by a unified standard including the physical layer and the logical layer.

[0020] When a failure occurs (for the first time), the rule learning control unit 32 acquires a group of alarms that characterize the failure, and generates rules from the relationships between network resources, and records the rules in the rule database 38. That is, the rule learning control unit 32 learns rules that indicate, by a path, the relationship between the location of an alarm based on a failure that occurred in the network 100 and the location of the failure itself, based on the resource management model and network configuration information.

[0021] When a failure occurs (for the second time or later), the failure location estimation function unit 33 executes failure estimation processing using rule definitions that include the generated alarm. That is, when a failure occurs in the network 100, the failure location estimation function unit 33 estimates the location of the failure based on the alarms that are generated due to the failure and the rules stored in the rule database 38.

[0022] In the embodiment, the rule learning control unit 32 changes the presence or absence of resources in the logical layer in the rule definition according to the condition determined by the presence or absence of a return path in the logical layer, and registers the result in the rule database 38.

[0023] 4 is a diagram showing a model definition related to a resource management model according to an embodiment, which shows an example of a resource management model in which a network is defined based on a unified standard including a physical layer and a logical layer.

[0024] In FIG. 4, entities such as PPort (Physical Port), PLink (Physical Link), EQP (Equipment), EH (Equipment Holder), and PD (Physical Device) are defined in the physical layer.

[0025] A PPort is a physical port of a network device. There are also virtual PPorts.

[0026] A PLink represents a physical line connecting network devices, such as an optical fiber, a UTP cable, or a wireless connection.

[0027] EQP represents a network device package, card, optical module (SFP, etc.), etc. A PPort is connected to the EQP.

[0028] An EH is a container for equipment, and represents a rack, chassis, slot, etc. A PD represents the actual device to be managed. The physical configuration is represented by the EH / EQP connected from the PD.

[0029] In the logical layer, entities called TPE (Termination Point Encapsulation), FRE (Forwarding Relationship Encapsulation), NFD (Network Forwarding Domain), and TL (Topological Link) are defined.

[0030] TPE represents the end point of a logical layer. TPE refers to a TPE in a lower logical layer or a PPort in a physical layer, and represents the hierarchical relationship of layers.

[0031] The TPE includes a TCP (Termination Connection Point) and a CP (Connection Point).

[0032] TCP is a termination point: TCP is responsible for terminating connectivity within the layer.

[0033] A CP is a relay point, which has the role of relaying connectivity within a layer.

[0034] FRE represents the flow of information between endpoints.

[0035] FRE includes NC (Network Connection), LC (Link Connection), and XC (Cross Connection).

[0036] NC indicates the connectivity from a TCP in a layer to a reachable TCP. The structure of the NW within a layer is represented by the subordinate LC and XC.

[0037] LC represents connectivity between devices based on the lower layer NC.

[0038] XC represents connectivity within a device based on the NFD of the lower layer, and corresponds to the routing function of a router or the switch function of a switch.

[0039] The NFD represents the transfer capability of a device, which is the basis for creating an XC. The NFD is basically used only in the Logical Device (LD) layer.

[0040] The TL represents logical connectivity between devices in the LD layer, and is located at the lowest level of the logical layer.

[0041] Fig. 5 is a schematic diagram showing an example of a fault that occurs in a network. In the rule-learning fault location estimation technology, alarms and fault locations are associated with information on each network resource in a network model that is managed in advance. In the example of Fig. 5, rules for fault location estimation, such as those shown in Figs. 6 and 7, are defined using the relationships (routes) between network resources in the network model.

[0042] For example, the rule in FIG. 7 is defined based on the series of events shown in FIG. 6: [Occurrence of a port failure in device B (1)], [Occurrence of a "HARDFAIL" alarm from device B (2)], and [Occurrence of a "LINKDOWN" alarm from device A (3)].

[0043] In Figure 7, the path PD > EH > EQP > PPort corresponding to "HARDFAIL" indicates the dotted line path of device B in Figure 6, and the path PPort > PLink < PPort corresponding to "LINKDOWN" indicates the dashed line path from PPort of device A to PPort of device B in Figure 6.

[0044] FIG. 8 is a schematic diagram showing another example of a network failure. FIG. 8 illustrates a case in which a logical alarm (CC FAIL) is issued as an IP communication interruption error when a port failure occurs on device A. In this case, two rules can be defined depending on whether a VLAN (Virtual LAN) is configured between devices A and B. The first rule is shown in FIG. 9 (when a VLAN is present), and the second rule is shown in FIG. 11 (when a VLAN is not present). FIG. 10 illustrates events and entities corresponding to FIG. 9, and FIG. 12 illustrates events and entities corresponding to FIG. 11. While TPE and FRE appear in the logical layer (VLAN layer) in FIG. 10, they do not appear in FIG. 12. Reflecting this, different rules are defined for each. In other words, the presence or absence of a VLAN layer results in different network management models, and different rule definitions result depending on the "relationship (path) between the alarm occurrence location and the failure location" in the rule definition. It should be noted that, when a port fails, other alarms such as LINKDOWN are also issued, but for simplicity of explanation, they will be omitted.

[0045] FIG. 13 is a schematic diagram showing another example of a fault occurring in a network. FIG. 13 illustrates the occurrence of a fault in the logical layer (IP layer), such as clearing IP settings. In this embodiment, the "relationship (route) between the alarm occurrence location and the fault location" in the rule definition is used in a rule in which the logical layer through which the fault occurs is omitted. That is, rules generated based on conditions (1), (2), and (3) are registered in the rule database 38 depending on whether or not there is a return route in the logical layer. (1) For a route without a return route in the logical layer, logical resources other than the start point and end point are omitted. (2) For a route with a return route in the logical layer, logical resources other than the start point, return point, and end point are omitted. (3) To distinguish between return routes, a description distinguishing the layer is written at the beginning of each logical resource.

[0046] By doing so, it becomes possible to estimate the location of a fault without the influence of the logical layer, which is essentially unaffected.

[0047] Fig. 14 is a diagram showing an example of an abbreviated notation of logical layers in a rule definition. Fig. 15 is a diagram showing events and entities corresponding to Fig. 14. Fig. 15 shows a case where there is no entity related to the logical layer, but even if an entity such as a VLAN exists in the logical layer, in this embodiment, a rule omitting the logical layer to be passed through is registered in the rule database 38.

[0048] 16 and 17 are diagrams for explaining the difference between the rule notation in the existing technology and the embodiment. Both show the case of a route without turnarounds in the logical layer. In this case, logical resources other than the start point and end point are omitted from the notation.

[0049] 16, data registered in the rule database 38 has conventionally been expressed as PPort>Plink<PPort<TPE<TPE<TPE, whereas in the embodiment, it is expressed as PPort>Plink<PPort<<TPE. Also, in FIG. 17, data registered in the rule database 38 has conventionally been TPE>TPE>TPE, whereas in the embodiment, it is TPE>>TPE.

[0050] 18 and 19 are diagrams for explaining the difference between the rule notation in the existing technology and the embodiment. Both illustrate the case of a route with a turnaround in the logical layer. In this case, logical resources other than the start point, turnaround point, and end point are omitted from the notation.

[0051] 18, data registered in the rule database 38 has conventionally been expressed as PPort<TPE<TPE<FRE>TPE>TPE>PPort, whereas in the embodiment it is expressed as PPort<<FRE>>PPort. Also, in FIG. 19, data registered in the rule database 38 has conventionally been TPE<TPE<FRE>TPE>TPE, whereas in the embodiment it is expressed as TPE<<FRE>>TPE.

[0052] 20 is a diagram for further explaining the rule notation in the embodiment. As a consideration for the abbreviated notation, in a route with a loopback in a logical layer, in order to distinguish the loopback layer, for example, a description (layer identifier) ​​that distinguishes the layer is written at the beginning of each logical resource. For example, in the case of a route via FRE of layer 4, an identifier combining the ID of layer 4 (4) with an underscore (_) is added to the FRE of the logical resource, and the notation is PPort>>4_FRE>>PPort. Furthermore, in the case of a route via FRE of layer 3, an identifier combining the ID of layer 3 (3) with an underscore is added to the FRE, and the notation is PPort>>3_FRE>>PPort.

[0053] FIG. 21 is a diagram showing an example of a path expression from physical (P Port) to physical (P Port). In FIG. 21, the paths from P Port to P Port include (1) a route via a physical link and (2) a route via a logical resource. Furthermore, in (2), different routes are used when there are multiple layers. Each of these routes must be expressed as being different in the definition of the relationship between alarms and faults. Therefore, in the example of FIG. 21, logical layers are not uniformly omitted, but are abbreviated in a way that distinguishes between physical and logical, and, if logical, which layer.

[0054] Fig. 22 is a diagram showing a specific example of the rule notation of Fig. 21. In the case of (1) via PL, PLink is written to indicate that it is via a physical link, and it is expressed as PPort > PLink > PPort. In the case of (2) via a logical resource, it is PPort > 3_FRE > PPort in layer 3, and PPort > 3_FRE > PPort in layer 4.

[0055] FIG. 23 is a diagram showing an example of a path expression from logic (TPE) to logic (TPE). In FIG. 23, the paths from 2_TPE to 2_TPE include (1) a route via a Topological Link and (2) a route via a Network Connection. In addition, in (2), different routes are used when there are multiple layers. Each of these routes must be expressed as a different route in the definition of the relationship between alarms and faults. Therefore, in the example of FIG. 23, instead of omitting all logical layers, an abbreviated expression is used that allows distinction between physical and logical, and if logical, which layer.

[0056] Fig. 24 is a diagram showing a specific example of the rule notation of Fig. 23. When going via TL, 2_TPE>TLink>2_TPE. When going via NC, 2_TPE>3_FRE>2_TPE in layer 3, and 2_TPE>4_FRE>2_TPE in layer 4.

[0057] In this embodiment, in the rule definition "relationship (route) between alarm occurrence location and fault location," the logical layers passed through are used in the rule with an abbreviated notation as follows. By doing so, it becomes possible to estimate the fault location without the influence of logical layers that have no effect essentially. (1) In the case of a route without a turnaround in the logical layer, logical resources other than the start point and end point are omitted. (2) In the case of a route with a turnaround in the logical layer, logical resources other than the start point, turnaround point, and end point are omitted. (3) To distinguish turnaround layers, a description that distinguishes the layer is written at the beginning of each logical resource.

[0058] In other words, according to the embodiment, it is possible to create rules that can be applied more flexibly to differences in the configuration of a multi-layered network than with existing technologies, and therefore it is possible to reduce the amount of work required by operators to learn new rules. Furthermore, as the number of rules to be defined is reduced, it is possible to reduce the processing load during fault location estimation processing and shorten the estimation processing time.

[0059] With existing rule-learning-based fault location estimation technology, it is sometimes not possible to estimate that essentially identical faults are the same because the "path from the alarm occurrence point to the fault point" varies in the rule definitions generated in networks that are realized by combining a variety of protocols, such as those found in large-scale networks.

[0060] For example, if a physical layer port failure causes an IP layer alarm from a remote device, the presence or absence of an intermediate logical layer (VLAN) is thought to have no essential effect on fault location estimation. However, existing technologies define the "path from the alarm occurrence point to the fault point" according to a pre-managed network model, so the presence or absence of a logical layer (VLAN), which is essentially unaffected, affects differences in rules, resulting in a lack of flexibility for different configurations in multi-layered networks. This also creates issues with the reusability of existing rules for fault estimation and the degradation of estimation processing performance due to the increase in rules generated. According to the embodiments, these issues are resolved, reducing the workload required for fault location estimation in a network.

[0061] The present invention is not limited to the above-described embodiments. Furthermore, the present invention can be embodied by modifying the components within the scope of the gist of the invention when implemented. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

[0062] 1...Monitoring system 2...Fault estimation system 3...Fault estimation device 4...GUI display terminal 5...Hub device 11...Control circuit 12...Communication module 13...User interface 14...Storage 15...Drive 16...Storage medium 31...Data acquisition unit 32...Rule learning control unit 33...Fault location estimation function unit 34...Countermeasure management function unit 35...GUI unit 36...Alarm information storage unit 37...Network configuration information storage unit 38...Rule database 39...Fault history and countermeasure history storage unit 40...API unit 100...Network.

Claims

1. A fault estimation device comprising: a network configuration information storage unit that stores network configuration information that specifies the connection relationships of multiple nodes that make up a target network, based on a resource management model defined by a unified standard including a physical layer and a logical layer; a rule learning control unit that learns rules that indicate the relationship between the location of an alarm based on a fault that has occurred in the network and the location of the fault using a path, based on the resource management model and the network configuration information; a rule database that registers the learned rules; and a fault location estimation function unit that, when a fault occurs in the network, estimates the location of the fault based on a group of alarms that are generated due to the fault and the rules stored in the rule database, wherein the rule learning control unit changes the presence or absence of logical layer resources in the rule definitions in accordance with conditions determined by the presence or absence of a return path in the logical layer, and registers the result in the rule database.

2. The fault estimation device of claim 1, wherein the rule learning control unit: if the route does not include a turnaround in the logical layer, registers in the rule database a rule that omits logical resources other than the start point and end point; if the route includes a turnaround in the logical layer, registers in the rule database a rule that omits logical resources other than the start point, turnaround point, and end point.

3. The fault estimation device according to claim 2, wherein the rule learning control unit registers a layer identifier for each logical resource in a rule when the route includes a turnaround in a logical layer.

4. A fault estimation method comprising: a computer storing network configuration information that specifies the connection relationships of multiple nodes that make up a target network, based on a resource management model defined by a unified standard including a physical layer and a logical layer; a computer learning rules that indicate, by a route, the relationship between the location of an alarm based on a fault that has occurred in the network and the location of the fault, based on the resource management model and the network configuration information; a computer registering the learned rules in a rule database; a computer changing the presence or absence of logical layer resources in the rule definition in accordance with a condition determined by the presence or absence of a return route in the logical layer, and registering the changed rules in the rule database; and a computer estimating, when a fault occurs in the network, the location of the fault based on a group of alarms that are generated due to the fault and the rules stored in the rule database.

Citation Information

Patent Citations

  • Fault analysis method and device

    CN115396287A

  • Method and device for locating a failed link, and method, device and system for analyzing alarm root cause

    US20120093005A1

  • Rule generation device, method, and program

    WO2021079521A1

  • Network management device, method and program

    WO2021131002A1