Information processing device and method
The information processing device and method enhance fault identification by extracting and restricting candidate events based on predefined rules, ensuring accurate and efficient fault location determination in complex network failure scenarios.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2026-03-12
AI Technical Summary
Existing fault estimation methods struggle to accurately identify the cause and location of network failures due to the complexity of alarms generated, leading to difficulties in interpreting evaluation values and missing true fault locations when multiple faults occur simultaneously.
An information processing device and method that extracts candidate events using predefined rules, restricts these candidates to combinations that cover all events, and presents the corresponding relationships between these candidates and the events, minimizing the number of candidates for operators to check.
This approach improves the recognition of fault locations by presenting a minimal set of candidates that can explain all alarms, reducing the number of potential fault locations that need verification, even in cases of multiple simultaneous failures.
Smart Images

Figure JP2024032261_12032026_PF_FP_ABST
Abstract
Description
Information processing device and method
[0001] FIELD Embodiments of the present invention relate to an information processing apparatus and method.
[0002] With the provision of various IT services, the maintenance and management of networks by telecommunications carriers has become important. When the servers or transmission equipment on the network required to provide the services become unavailable due to a malfunction or other reason, alarms indicating an abnormality are issued from various points, such as the network equipment or the servers providing the services.
[0003] Alarms that occur are collected by a monitoring system. Maintenance personnel must use the collected alarms and network configuration information to analyze where and what type of failure has occurred and then take action to restore the system.
[0004] However, various alarms are generated depending on the cause of the failure and the type of failed device, and the number of devices generating alarms and the type of alarm also vary depending on the network configuration. Understanding the characteristics of such alarms and identifying the location and cause of the failure requires a vast amount of know-how, and when many alarms are generated, even an experienced technician finds it difficult to instantly identify the cause of the failure. Therefore, it is effective to create rules for alarms generated by failures and to instantly identify the cause and location of the failure when an alarm is generated.
[0005] There is a technology that estimates the cause and location of a failure from messages such as alarms that are generated when a failure occurs on a network (see, for example, Patent Document 1 and Non-Patent Document 1). For example, if-then rules are defined in advance for each failure case, with the if conditions being the "type of alarm" and the "positional relationship between the failure location and the alarm occurrence location," and the compliance rate with these rules is calculated as an evaluation value, and the suspected failure location and cause are estimated from the evaluation value.
[0006] Japanese Patent No. 6637854
[0007] Fumika Asai et al.: Study on Autonomous Network Fault Estimation and Applicability Evaluation (IEICE Technical Report, vol. 122, no. 442, ICM2022-59, pp. 95-100, 2022)
[0008] However, when using known fault estimation methods, it can be difficult to interpret the results of suspected fault locations based on the above-mentioned evaluation values. For example, when all suspected fault candidates are presented to the operator, or when suspected fault candidates based on evaluation values above a threshold are presented to the operator, it is up to the operator to determine the true fault location from the multiple suspected fault candidates presented. On the other hand, when the presentation is based on the evaluation value with the highest rule conformance, some true fault locations may be missed when multiple faults occur simultaneously, making it difficult to properly identify the fault locations.
[0009] The present invention has been made in light of the above circumstances, and its purpose is to provide an information processing device and method that can appropriately identify the phenomenon that triggers the occurrence of an event.
[0010] An information processing device according to one aspect of the present invention includes a candidate extraction unit that extracts candidates for underlying events that trigger the occurrence of an event to be evaluated, based on rules that define the relationship between the event and the event that triggers the occurrence of the event to be evaluated, a candidate restriction unit that restricts the candidates extracted by the candidate extraction unit to combinations of underlying events that trigger the occurrence of all events to be evaluated, based on the event that occurs as a trigger for the occurrence of the candidate, for each candidate, and a presentation unit that presents the correspondence between the candidates for underlying events restricted by the candidate restriction unit and the events that occur as a trigger for the occurrence of the candidate.
[0011] An information processing method according to one aspect of the present invention is a method performed by an information processing device, comprising: a candidate extraction unit of the information processing device extracts candidates for underlying events that trigger the occurrence of an event to be evaluated, based on rules that define the relationship between the event and the event that triggers the occurrence of the event to be evaluated; a candidate restriction unit of the information processing device restricts the candidates extracted by the candidate extraction unit to combinations of underlying events that trigger the occurrence of all events to be evaluated, based on the event that triggers the occurrence of the candidate for each candidate; and a presentation unit of the information processing device presents the correspondence between the candidates for underlying events limited by the candidate restriction unit and the events that occur as triggers of the candidate.
[0012] According to the present invention, it is possible to appropriately identify the phenomenon that triggers the occurrence of an event.
[0013] FIG. 1 is a diagram illustrating an application example of an information processing device according to an embodiment of the present invention. FIG. 2 is a diagram illustrating an example of a network configuration according to the first embodiment. FIG. 3 is a diagram illustrating an example of a calculation result of an evaluation value according to the first embodiment. FIG. 4 is a diagram illustrating an example of an alarm caused by a suspected fault candidate according to the first embodiment. FIG. 5 is a diagram illustrating an example of a suspected fault candidate corresponding to a decision variable that minimizes an objective function according to the first embodiment. FIG. 6 is a diagram illustrating an example of a fault determination result and an associated alarm according to the first embodiment. FIG. 7 is a diagram illustrating an example of a fault determination result and an associated alarm according to the first embodiment. FIG. 8 is a diagram illustrating an example of a network configuration according to the second embodiment. FIG. 9 is a diagram illustrating an example of a calculation result of an evaluation value according to the second embodiment. FIG. 10 is a diagram illustrating an example of an alarm caused by the occurrence of a suspected fault candidate according to the second embodiment. FIG. 11 is a diagram illustrating an example of an alarm coverage rate, an objective function, and an alarm overlap rate for each combination of decision variables according to the second embodiment. FIG. 12 is a diagram illustrating an example of a fault determination result and an associated alarm according to the second embodiment. FIG. 13 is a diagram illustrating an example of a fault determination result and an associated alarm according to the second embodiment. FIG. 14 is a diagram showing an example of a fault determination result and associated alarms according to the second embodiment. FIG. 15 is a diagram showing an example of a network configuration according to the third embodiment. FIG. 16 is a diagram showing an example of an alarm caused by a suspected fault candidate according to the third embodiment. FIG. 17 is a diagram showing an example of an alarm coverage rate, an objective function, and an alarm duplication rate for each combination of decision variables according to the third embodiment. FIG. 18 is a diagram showing an example of a fault determination result and associated alarms according to the third embodiment. FIG. 19 is a diagram showing an example of a fault determination result and associated alarms according to the third embodiment. FIG. 20 is a diagram showing an example of a fault determination result and associated alarms according to the third embodiment. FIG. 21 is a block diagram showing an example of the hardware configuration of an information processing device according to an embodiment of the present invention.
[0014] An embodiment of the present invention will now be described. Fig. 1 is a diagram showing an application example of an information processing device according to an embodiment of the present invention. As shown in Fig. 1, an information processing device 10 according to an embodiment of the present invention includes a root cause candidate extraction unit 11, a candidate limiting unit 12, and a causal relationship presentation unit 13.
[0015] When an event occurs as a result of a certain event, and the root cause candidate extraction unit 11 evaluates the possibility of the occurrence of the underlying event that triggered the event, the root cause candidate extraction unit 11 references an external event information DB (database) 20 that stores information on events that occur as a result of a certain event, and an external knowledge DB 30 that stores rules that define the relationship between the event and the event that triggered the event, and evaluates the possibility of the underlying event occurring numerically depending on whether or not there is an event that is expected to occur based on the rules.
[0016] The candidate limiting unit 12 compares a set of events caused by each of one or more combinations of root events with the events to be evaluated, i.e., the events that are occurring, for each root event whose occurrence is suggested by the root cause candidate extraction unit 11, and limits the combinations to the minimum combination of root events that can cover all of the events to be evaluated. The causal relationship presenting unit 13 presents the correspondence between the candidates represented by the combinations of root events limited by the candidate limiting unit 12 and the events that occur as a result of the occurrence of these candidates.
[0017] In this embodiment, the following preconditions (1-1) and (1-2) are set. (1-1) An event caused by a certain event does not necessarily occur, but when an event occurs, the underlying event is always occurring. (1-2) The possibility of multiple events occurring simultaneously is low, but if the probability of occurrence is high, such as when the evaluation value indicating the possibility of the underlying event occurring is 100%, the occurrence of multiple events is allowed.
[0018] Next, as a specific case, for example, when an alarm is generated due to the root cause of a failure, the location of the failure cause is identified from the generated alarm based on an If-then rule, etc. In this case, the above-mentioned preconditions can be replaced with the following (2-1) and (2-2).
[0019] (2-1) An alarm caused by a fault does not necessarily occur, but when the target alarm occurs, a fault is sure to have occurred. (2-2) The probability of faults occurring at multiple locations at the same time is low, but when the probability of a fault occurring according to the rule is 100%, let M be a set of I alarms suggesting the fault that is the evaluation target and is adopted as the occurring fault, and let N be a set of J suspected fault candidates that are nominated by the fault detection rule.
[0020] A specific example of the process will be described below. (3-1) The root cause candidate extraction unit 11 extracts candidates for suspected fault locations in the fault detection rules stored in the knowledge DB 30 from the event information stored in the event information DB (database) 20. At this time, for each suspected fault candidate j, the probability of occurrence of the candidate is calculated as an evaluation value v j j is a suspected fault candidate ID, which is set for each location and fault type, and is in the range of "1≦j≦J".
[0021] In addition, the relationship between alarm i and suspected fault candidate j is expressed as w ij i is an alarm ID, which is set for each occurrence date and time, occurrence location, and alarm type, and is in the range of "1≦i≦I".
[0022] w ij indicates whether an alarm with alarm ID "i" occurs due to the occurrence of a suspected fault candidate j, and w ij If is "1", an alarm will occur, and ij "0" indicates that no alarm will occur.
[0023] (3-2) The candidate limiting unit 12 determines whether or not a suspected fault candidate j occurs by using a decision variable x j Introduce x jIf x = 1, it means that the suspected fault candidate has actually occurred. j If it is 0, it means that the suspected fault candidate has not actually occurred.
[0024] The candidate limiting unit 12 calculates the alarm coverage Cov shown in the following (1) for the combination of decision variables for each suspected fault candidate, and calculates a set of suspected fault candidates that minimizes the objective function shown in the following (2) under the condition that "alarm coverage Cov = 1".
[0025]
[0026] However, the evaluation value v j When is extremely high, for example 100%, x j When multiple combinations of suspected fault candidates that minimize the objective function are found, a relatively small alarm overlap rate or evaluation value v j The plurality of combinations of suspected fault candidates may be further limited to combinations where the average value of is relatively high, for example, by using other indicators.
[0027]
[0028] (3-3) The causal relationship presenting unit 13 presents a set of suspected fault candidates limited by the candidate limiting unit 12 and an alarm that occurs as a result of the occurrence of each suspected fault candidate, in this case, for each suspected fault candidate j, a “w ij . . . 1”.
[0029] This makes it possible to present a minimum number of suspected fault candidates that can explain all alarms that suggest a fault has occurred, thereby improving the recognizability of rule-based fault evaluation.
[0030] In this embodiment, it is possible to evaluate a combination of events that can cover all events that occur as a result of a root event. Therefore, when multiple root events occur simultaneously, which is not possible to determine based on a numerical evaluation of the possibility of the root event occurring, it is possible to indicate the possibility of an event occurring due to this event, and it becomes possible to reduce the number of candidate root events that the operator needs to check.
[0031] (First embodiment) Fig. 2 is a diagram showing an example of a network configuration according to a first embodiment. In the network shown in Fig. 2, devices a, b, c, d, and e are physically connected in a row at the physical layer. Devices A and E are IP devices, and devices B, C, and D are transmission devices.
[0032] In the network shown in Figure 2, communication services between two IP devices are provided via three transmission devices. The transmission layer indicates the connection section where optical communication is performed between the transmission devices, and the Ether layer indicates the section where communication is performed between the two devices via an Ethernet (registered trademark) cable. Furthermore, the service layer connection indicates service communication performed between the IP devices, and is shown to be composed of two Ether layer connections and one transmission layer connection.
[0033] In "Case 1," an example in this embodiment, four alarms to be evaluated occur following a failure in device B. The type, alarm ID (i = 1 to 4), and occurrence location of each alarm are as follows: Alarm a (i = 1): Device A adjacent to device B Alarm a (i = 2): Device C adjacent to device B Alarm b (i = 3): Communication endpoint of the service layer supported by device B Alarm b (i = 4): Communication endpoint of the service layer supported by device B
[0034] (4-1) Fig. 3 is a diagram showing an example of the calculation result of the evaluation value according to the first embodiment. When the root cause candidate extraction unit 11 extracts suspected fault candidates based on the alarm to be evaluated using the following "fault detection rule 1," the five suspected fault candidates shown in Fig. 3 are found to be applicable. The evaluation value is evaluated based on the conformance rate to the rule conditions.
[0035] (Failure detection rule 1) "If condition (alarm occurrence location, alarm type) if1 "adjacent device of the faulty device", "alarm a" if2 "communication end point of the service layer supported by the faulty device", "alarm b" Then (cause, failure location) "failure cause 1", "device""
[0036] 4 is a diagram showing an example of an alarm that occurs due to a suspected fault candidate according to the first embodiment. Also, w indicates whether or not an alarm with alarm ID "i" occurs due to the occurrence of a suspected fault candidate j. ij The values are shown in FIG.
[0037] (4-2) The candidate limiting unit 12 determines whether or not any suspected fault candidate j occurs by using a decision variable x j is introduced, and a set of suspected fault candidates that minimizes the objective function shown in (2) above is obtained.
[0038] 5 is a diagram illustrating an example of a fault suspect candidate corresponding to a decision variable that minimizes an objective function according to the first embodiment. In the example shown in FIG. 5, j When Cov = (1,0,0,0,0), the alarm coverage rate Cov = 1 and the objective function is minimized. In this embodiment, since there are no other suspected fault candidates that minimize the objective function, no other indicators are used, but the alarm duplication rate Dup at this time is 1.0.
[0039] (4-3) The causal relationship presenting unit 13 presents a set of suspected fault candidates limited by the candidate limiting unit 12 and an alarm caused by the failure of each suspected fault candidate, in this case, for each candidate j, a “w ij . . . 1”.
[0040] 6 and 7 are diagrams showing examples of fault determination results and associated alarms according to the first embodiment. In this embodiment, information associating device B (j=1), which is a suspected fault candidate limited by the candidate limiting unit 12, with the alarm caused by the fault of device B is presented in the form of a table shown in Fig. 6 or in the form of an image of a network configuration shown in Fig. 7.
[0041] Here, it is shown that a failure occurs in device B, and that this failure causes alarm a to be generated in devices A and C, and alarm b to be generated at the service layer communication endpoint on the device A side and the service layer communication endpoint on the device E side.
[0042] Second Embodiment Fig. 8 is a diagram showing an example of a network configuration according to a second embodiment. In the network shown in Fig. 8, the relationships between each device, which is an IP device, and its adjacent devices are as follows: Device K: Devices L, M Device L: Devices K, O Device M: Devices K, N, P, Q Device N: Devices M, O, Q Device O: Devices L, N, Q Device P: Devices M, Q, R Device Q: Devices M, N, O, P, R, T Device R: Devices P, Q, S Device S: Devices R, T Device T: Devices Q, S
[0043] In "Case 2" shown in Figure 8, devices O and P are faulty, and as a result of this fault, an alarm is issued from each device. The relationship between the device where the alarm was issued, the type of alarm issued from that device, and the alarm ID is as follows: Device L: Alarm x (i = 1) Device M: Alarm y (i = 2) Device N: Alarm x (i = 3), y (i = 4) Device Q: Alarm x (i = 5) Device R: Alarm y (i = 6)
[0044] 8, the six alarms to be evaluated have occurred. The difference between "Case 2" and "Case 3" described later is that in "Case 2," only one alarm x has occurred in device Q.
[0045] (5-1) FIG. 9 is a diagram showing an example of the calculation result of the evaluation value according to the second embodiment. When the root cause candidate extraction unit 11 extracts suspected fault candidates based on the alarm to be evaluated using the following "fault detection rule 2," the nine suspected fault candidates shown in FIG. 9 apply. The evaluation value is evaluated based on the conformance rate to the rule conditions. (Fault detection rule 2) "If condition (alarm occurrence location, alarm type) if1 "adjacent device of faulty device," "alarm x" if2 "adjacent device of faulty device," "alarm y" Then (cause, fault location) "fault cause 2," "device""
[0046] 10 is a diagram showing an example of an alarm that occurs due to the occurrence of a suspected fault candidate according to the second embodiment. Also, w indicates whether or not an alarm with alarm ID “i” occurs due to the occurrence of a suspected fault candidate j. ij The values are shown in FIG.
[0047] (5-2) The candidate limiting unit 12 determines whether or not any suspected fault candidate j occurs by using a decision variable x j is introduced, and a set of suspected fault candidates that minimizes the objective function shown in (2) above is obtained.
[0048] In "Case 2" in this embodiment, it is not possible to explain all alarm occurrences when the objective function shown in (2) above is set to 1. For this reason, in this embodiment, in order to minimize the objective function, it is necessary to be able to explain all alarm occurrences using two failure occurrence candidates.
[0049] Fig. 11 is a diagram showing an example of the alarm coverage rate, objective function, and alarm duplication rate for each combination of decision variables according to the second embodiment. For the occurrence of the alarm (i = 1) shown in Fig. 11, the suspected fault candidate (j = 1 or 2) shown in Fig. 11 must be selected. Therefore, as shown in Fig. 11, there are 15 possible combinations of decision variables in which two suspected fault candidates including this suspected fault candidate (j = 1 or 2) are selected.
[0050] Of the patterns of combinations of decision variables shown in FIG. 11 , i.e., patterns “1” to “15,” there are two patterns that have an alarm coverage rate Cov=1, i.e., can explain all alarm occurrences: pattern “2,” i.e., a suspected fault candidate (j=1) and a suspected fault candidate (j=3), and pattern “4,” i.e., a suspected fault candidate (j=1) and a suspected fault candidate (j=5) (symbols a and b in FIG. 11 ).
[0051] The candidates for the suspected fault location to be limited may be limited to the above-mentioned patterns "2" and "4", but the candidates may also be further limited by using, for example, the alarm duplication rate or the average value of the above-mentioned evaluation values.
[0052] For example, as shown in FIG. 11, the alarm duplication rate for pattern "2" is 1.17, and the alarm duplication rate for pattern "4" is 1.33, so pattern "2", which has a relatively low alarm duplication rate, may be selected as a further limited candidate for the suspected fault location.
[0053] Furthermore, as shown in FIG. 11, the average value of the evaluation values of each suspected fault location in pattern "2" is 0.59, and the average value of the evaluation values of each suspected fault location in pattern "4" is 0.5, so pattern "2" with this relatively high average value may be selected as a further limited candidate for the suspected fault location.
[0054] (5-3) The causal relationship presenting unit 13 presents a set of suspected fault candidates limited by the candidate limiting unit 12 and an alarm caused by the failure of each suspected fault candidate, in this case, for each candidate j, a “w ij . . . 1”.
[0055] 12, 13, and 14 are diagrams showing examples of fault determination results and associated alarms according to the second embodiment. In this embodiment, the suspected fault candidates limited by the candidate limiting unit 12 are device O (j=1) and device P (j=3).
[0056] Therefore, information associating each suspected failure candidate with the alarms that have occurred as a result of the failure of this candidate is presented in the form of a table as shown in Fig. 12 or in the form of network configuration images as shown in Fig. 13 and Fig. 14. Fig. 13 shows alarms that have occurred as a result of the failure of device O, which is the first limited suspected failure candidate, and Fig. 14 shows alarms that have occurred as a result of the failure of device P, which is the second limited suspected failure candidate.
[0057] (Third embodiment) Fig. 15 is a diagram showing an example of a network configuration according to a third embodiment. In this embodiment, the relationships between each device, which is an IP device, and the devices adjacent to those devices are the same as the configuration shown in Fig. 8 described in the second embodiment. In addition, in "Case 3" in this embodiment, the failure locations are devices O and P, the same as in the second embodiment. On the other hand, the relationship between the device in which an alarm occurred, and the type and alarm ID of the alarm generated from that device is as follows, and compared to the second embodiment, it differs in that multiple alarms x are generated from device Q.
[0058] Device L: Alarm x (i = 1) Device M: Alarm y (i = 2) Device N: Alarm x (i = 3), y (i = 4) Device Q: Alarm x (i = 5), alarm x (i = 6) Device R: Alarm y (i = 7) In "Case 3" in this embodiment, the seven alarms mentioned above that are the subject of evaluation have occurred.
[0059] (6-1) When the root cause candidate extraction unit 11 extracts suspected fault candidates based on the alarm to be evaluated using the "fault detection rule 2" described in the second embodiment, the nine cases shown in Fig. 9 described in the second embodiment are applicable. The evaluation value is evaluated based on the conformance rate to the rule conditions.
[0060] 16 is a diagram showing an example of an alarm caused by a suspected fault candidate according to the third embodiment. Also, w indicates whether or not an alarm with alarm ID “i” is caused by the occurrence of a suspected fault candidate j. ij The values are shown in FIG.
[0061] In this example, the number of occurrences of an alarm caused by the failure of a suspected failure candidate is set to 1. Here, alarm (i=5) and alarm x (i=6) are alarms generated by device Q, and if the alarm caused by the failure of a certain suspected failure candidate is alarm x from device Q, the occurrence of only one of alarm x (i=5) and alarm x (i=6) caused by the failure is indicated as "1" as shown in FIG. *In other words, in the example shown in FIG. 16, when an alarm occurs due to a failure of suspected failure candidate "1", only one of alarm x (i=5) or alarm x (i=6) occurs as a result of the failure.
[0062] (6-2) The candidate limiting unit 12 determines whether or not any suspected fault candidate j occurs by using a decision variable x j is introduced, and a set of suspected fault candidates that minimizes the objective function shown in (2) above is obtained.
[0063] In "Case 3" of this embodiment, similar to "Case 2" of the second embodiment, not all alarm occurrences can be explained when the objective function shown in (2) above is set to 1. For this reason, in this embodiment, in order to minimize the objective function, it is necessary to make it possible to explain all alarm occurrences using two failure occurrence candidates.
[0064] Fig. 17 is a diagram showing an example of the alarm coverage rate, objective function, and alarm duplication rate for each combination of decision variables according to the third embodiment. For the occurrence of the alarm (i = 1) shown in Fig. 17, the suspected fault candidate (j = 1 or 2) shown in Fig. 17 must be selected. Therefore, as shown in Fig. 17, there are 15 combinations of decision variables in which two suspected fault candidates including this suspected fault candidate (j = 1 or 2) are selected.
[0065] 17, that is, patterns "1" to "15," the only pattern that has an alarm coverage rate Cov=1, that is, that can explain all alarm occurrences, is pattern "2," that is, a pattern consisting of a suspected fault candidate (j=1) and a suspected fault candidate (j=3) (symbol a in FIG. 17). Note that, as shown in FIG. 17, the alarm duplication rate Dup for this pattern "2" is 1.
[0066] From the above, the combination of decision variables is "x j = (1,0,1,0,0,0,0,0,0)" the alarm coverage rate Cov = 1 and the objective function is minimized. Note that in this example, the alarm duplication rate index is not used.
[0067] Also, as mentioned above, in this example, the number of alarms caused by the failure of a suspected failure candidate is set to 1. Therefore, alarm (i=5) and alarm (i=6) shown in Fig. 15 are alarm x generated in device Q, but as shown in Fig. 17, when a suspected failure candidate (j=1) occurs, only one of alarm (i=5) or alarm (i=6) is generated as a result of the failure.
[0068] In other words, since the evaluation of the alarm coverage rate or the alarm duplication rate also follows the above idea, when pattern "1" of the combination of decision variables shown in FIG. 17 is selected, that is, when suspected fault candidates "1" and "2" are selected, the only alarms that can be explained out of the total seven alarms are alarms "i=1," "i=2," "i=3," "i=4," and "i=5," or alarms "i=1," "i=2," "i=3," "i=4," and "i=6," so the coverage rate Cov is 0.71 (= 5 / 7) and the alarm duplication rate Dup is 0.86 (= 6 / 7).
[0069] Furthermore, when pattern "2" of the combination of decision variables shown in FIG. 17 is selected (i.e., when suspected fault candidates "1" and "3" are selected), the alarms that can be explained are alarms "1" to "7" out of the total seven alarms, so the coverage rate Cov is 1.00 (= 7 / 7) and the alarm duplication rate Dup is also 1.00 (= 7 / 7).
[0070] (6-3) The causal relationship presenting unit 13 presents a set of suspected fault candidates limited by the candidate limiting unit 12 and an alarm caused by the failure of each suspected fault candidate, in this case, for each candidate j, a “w ij . . . 1”.
[0071] 18, 19, and 20 are diagrams showing examples of fault determination results and associated alarms according to the third embodiment. In this embodiment, the suspected fault candidates limited by the candidate limiting unit 12 are device O (j=1) and device P (j=3).
[0072] Therefore, information associating each suspected failure candidate with the alarms triggered by the failure of this candidate is presented in the form of a table as shown in Fig. 18 or in the form of network configuration images as shown in Fig. 19 and Fig. 20. Fig. 19 shows alarms triggered by the failure of device O, which is the first limited suspected failure candidate, and Fig. 20 shows alarms triggered by the failure of device P, which is the second limited suspected failure candidate.
[0073] 21 is a block diagram showing an example of the hardware configuration of an information processing device according to an embodiment of the present invention. In the example shown in FIG. 21, the information processing device 10 according to the embodiment is configured, for example, as a server computer or a personal computer, and has a hardware processor 111A such as a CPU. A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115.
[0074] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from a communication network. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.
[0075] An input device 200 and an output device 300 attached to the information processing device 10 and used by a user or the like are connected to the input / output interface 113. The input / output interface 113 receives operation data input by a user or the like through the input device 200 such as a keyboard, a touch panel, a touchpad, or a mouse, and outputs output data to an output device 300 including a display device using a liquid crystal or an organic electroluminescence (EL) display, for display. The input device 200 and the output device 300 may be devices built into the information processing device 10, or may be input devices and output devices of other information terminals that can communicate with the information processing device 10 via a network (NW).
[0076] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), and a non-volatile memory such as a read only memory (ROM), and stores programs necessary to execute various control processes, etc., according to one embodiment.
[0077] The data memory 112 is a tangible storage medium that is a combination of, for example, the above-mentioned non-volatile memory and a volatile memory such as RAM (Random Access Memory), and is used to store various data acquired and created in the course of various processes performed by the information processing device 10.
[0078] An information processing device 10 according to an embodiment of the present invention can be configured as an information processing device having a processing function unit implemented by software.
[0079] The storage areas used as work memories or the like by the various components of the information processing device 10 can be configured using the data memory 112 shown in Fig. 21. However, these configured storage areas are not essential components within the information processing device 10, and may be areas provided in, for example, an external storage medium such as a USB (Universal Serial Bus) memory, or a storage device such as a database server located in the cloud.
[0080] The processing function unit can be realized by having the hardware processor 111A read and execute a program stored in the program memory 111B, but the processing function unit may also be realized in various other forms, including an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0081] The methods described in each embodiment can be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (floppy disk, hard disk, etc.), optical disk (CD-ROM, DVD, MO, etc.), or semiconductor memory (ROM, RAM, flash memory, etc.), and can also be distributed by transmitting it via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only execution programs but also tables or data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-mentioned processing by controlling the operation of this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.
[0082] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.
[0083] REFERENCE SIGNS LIST 10... Information processing device 11... Root cause candidate extraction unit 12... Candidate limitation unit 13... Causal relationship presentation unit 20... Event information DB 30... Knowledge DB
Claims
1. An information processing device comprising: a candidate extraction unit that extracts candidates for underlying events that trigger the occurrence of an event to be evaluated, based on rules that define the relationship between the event and the event that triggers the occurrence of the event to be evaluated; a candidate restriction unit that restricts the candidates extracted by the candidate extraction unit to combinations of underlying events that trigger the occurrence of all events to be evaluated, based on the event that occurs as a trigger for the occurrence of the candidate, for each candidate; and a presentation unit that presents the correspondence between the candidates for underlying events and the events that occur as a trigger for the occurrence of the candidates, restricted by the candidate restriction unit.
2. The information processing device according to claim 1, wherein the candidate limiting unit determines combinations of candidates that include events that are triggers for the occurrence of any of the events from among the candidates extracted by the candidate extraction unit and candidates indicated in relation to events that occur as a result of the occurrence of the candidates, and limits the candidates extracted by the candidate extraction unit to the smallest combination of fundamental events that can cover the occurrence of all events from among the combinations determined.
3. The information processing device of claim 1, wherein the candidate extraction unit calculates an evaluation value for each candidate to evaluate the likelihood that the candidate will trigger the occurrence of the event, and the candidate limitation unit further limits the number of candidates for the limited underlying event to a smaller number based on the calculated evaluation value for each of the limited candidates.
4. A method performed by an information processing device, comprising: a candidate extraction unit of the information processing device extracts candidates for underlying events that trigger the occurrence of an event to be evaluated, based on rules that define the relationship between the event and the event that triggers the occurrence of the event to be evaluated; a candidate restriction unit of the information processing device restricts the candidates extracted by the candidate extraction unit to combinations of underlying events that trigger the occurrence of all events to be evaluated, based on the event that triggers the occurrence of the candidate, for each candidate; and a presentation unit of the information processing device presents the correspondence between the candidates for underlying events limited by the candidate restriction unit and the events that trigger the occurrence of the candidates.
Citation Information
Patent Citations
Performance monitoring device, performance monitoring method and program
JP2005327261A
Network failure diagnostic device, network failure diagnostic method and network failure diagnostic program
JP2007096796A
Fault separation method and administrative server for performing fault separation
JP2017069895A