Failure handling device, system, method, and program

JPWO2024135322A5Pending Publication Date: 2025-08-19
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024565751
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-06-11
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing systems face challenges in effectively handling failures in information systems due to the complexity of event occurrence patterns, making it difficult to identify the appropriate recovery methods.

Method used

A failure handling device and system that identifies the degree of association between events and failure causes, calculates a certainty factor based on elapsed time, and selects the most reliable cause to execute countermeasures.

Benefits of technology

Enables appropriate measures to be taken in response to failures by accurately identifying the cause of events and executing relevant countermeasures, improving system recovery efficiency.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This failure handling device comprises: an identifying unit that identifies, from definition information defining relevance between events that could occur in an information system and failure causes, relevance between each of a plurality of occurred events that have occurred in a predetermined period and each failure cause; a calculation unit that, on the basis of the time that has elapsed from the occurrence of each of the plurality of occurred events, uses the identified relevance to calculate a confidence level for each failure cause in a set of a plurality of occurred events; a selection unit that selects one of a plurality of failure causes that has the highest calculated confidence level as a failure cause candidate; and an execution control unit that causes the information system to execute the contents of handling associated with the failure cause candidate. In this way, appropriate handling is implemented in accordance with a group of events that have occurred due to failure of an information system.
Need to check novelty before this filing date? Find Prior Art

Description

Fault handling device, system, method, and program

[0001] The present disclosure relates to a failure handling device, system, method, and program.

[0002] When a system failure occurs during operation of an information system, an operator identifies the cause of the failure, considers and decides on an appropriate countermeasure (such as a recovery process) depending on the cause of the failure, and implements the decided countermeasure to restore the information system. Patent Document 1 discloses a technology related to a computer system for identifying an appropriate recovery method from a failure. The computer system disclosed in Patent Document 1 includes meta-rules that define a failure and a group of events that are expected to occur as a result of the failure, and displays a failure recovery method corresponding to the meta-rule depending on the event that has occurred.

[0003] International Publication No. 2011 / 007394

[0004] There are many types of events that can occur due to a fault in an information system, and in particular, the combinations and sequences of events that can occur vary widely. Therefore, there is a problem in that it is difficult to deal with unknown event occurrence patterns simply by creating definition information such as the meta-rules in Patent Document 1.

[0005] In view of the above-mentioned problems, the object of the present disclosure is to provide a failure handling device, system, method, and program for taking appropriate measures in response to a group of events that occur due to a failure in an information system.

[0006] The fault handling device according to the present disclosure comprises: an identification means for identifying the degree of association between each fault cause and each of a plurality of occurrence events that occur within a predetermined period of time, based on definition information that defines the degree of association between an event that can occur in an information system and a fault cause; a calculation means for calculating a certainty factor for each fault cause in a set of the plurality of occurrence events, using the identified degree of association, based on the elapsed time since the occurrence of each of the plurality of occurrence events; a selection means for selecting, from the plurality of fault causes, the one with the highest calculated certainty factor as a candidate fault cause; and an execution control means for causing the information system to execute the countermeasure content associated with the candidate fault cause.

[0007] The fault handling system according to the present disclosure comprises: a storage device that stores definition information that defines the degree of association between an event that may occur in an information system and a fault cause, and history information of a plurality of occurrence events that have occurred in the information system over a predetermined period of time; and a fault handling device connected to the storage device, wherein the fault handling device comprises: an identification means that refers to the storage device and identifies the degree of association between each of the plurality of occurrence events and each fault cause from the definition information; a calculation means that calculates a certainty for each fault cause in the set of the plurality of occurrence events using the identified degree of association based on the elapsed time since the occurrence of each of the plurality of occurrence events included in the history information; a selection means that selects the fault cause with the highest calculated certainty from the plurality of fault causes as a candidate fault cause; and an execution control means that causes the information system to execute the response content associated with the candidate fault cause.

[0008] The failure handling method disclosed herein involves a computer identifying the degree of association between each failure cause and each of a plurality of events that occurred during a specified period of time, based on definition information that defines the degree of association between events that can occur in an information system and the failure cause; calculating a certainty factor for each failure cause in the set of the plurality of events using the identified degree of association based on the elapsed time since the occurrence of each of the plurality of events; selecting the failure cause with the highest calculated certainty factor from the plurality of failure causes as a candidate failure cause; and executing the handling content associated with the candidate failure cause on the information system.

[0009] The failure handling program disclosed herein causes a computer to execute the following steps: a specification process that specifies the degree of association between each of a plurality of occurrence events that occurred within a specified period and each failure cause, based on definition information that defines the degree of association between events that can occur in an information system and the failure cause; a calculation process that calculates the certainty of each failure cause in a set of the plurality of occurrence events using the specified degree of association based on the elapsed time since the occurrence of each of the plurality of occurrence events; a selection process that selects the failure cause with the highest calculated certainty from the plurality of failure causes as a candidate failure cause; and an execution control process that causes the information system to execute the response content associated with the candidate failure cause.

[0010] The present disclosure makes it possible to provide a failure handling device, system, method, and program for taking appropriate measures in response to a group of events that occur due to a failure in an information system.

[0011] 1 is a block diagram showing the configuration of a failure handling device. FIG. 2 is a flowchart showing the flow of a failure handling method. FIG. 3 is a block diagram showing the overall configuration including a failure handling system. FIG. 4 is a block diagram showing the configuration of a failure handling device. FIG. 5 is a block diagram showing an example configuration of a failure handling rule. FIG. 6 is a diagram showing an example of relevance definition information. FIG. 7 is a block diagram showing an example configuration of an event occurrence history. FIG. 8 is a sequence diagram showing a flow including failure handling processing. FIG. 9 is a flowchart showing the flow of failure handling processing in a failure handling device. FIG. 10 is a diagram showing an example of the relationship between the elapsed time from the occurrence time of each event in an event occurrence pattern and the importance level. FIG. 11 is a diagram showing an example of an identified relevance level. FIG. 12 is a diagram showing an example of an update result of the relevance level. FIG. 13 is a diagram showing an example where a failure cause (hypothesis) is added to the relevance definition information. FIG. 14 is a diagram showing an example where an event is added to the relevance definition information.

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and for clarity of explanation, duplicate explanations will be omitted as necessary.

[0013] 1 is a block diagram showing the configuration of a failure handling device 100. The failure handling device 100 is an information processing device or system that determines appropriate countermeasures for a set of events that occur due to failures in an information system (not shown) to be operated, and executes the countermeasures to recover the information system and continue operation.

[0014] Here, "event" refers to events or fault information that occur in an information system, or information (such as monitoring alerts) that is detected as meeting predetermined conditions through monitoring of resources, logs, etc. of the information system. The "cause of a fault" in an information system can be identified not only by a single event, but also by a combination (set) of multiple events or the order of occurrence (occurrence pattern). Furthermore, "countermeasures" refer to countermeasures for restoring the information system to a normal state after a fault has occurred. Countermeasures include, for example, predetermined commands or scripts for the information system, or parameters for generating these. Note that countermeasures do not have to be final commands or scripts. For example, countermeasures include identification information for the countermeasure, the system (server) to which the countermeasure is to be applied, parameters used in the final command, etc.

[0015] Here, the failure handling device 100 is capable of referencing definition information that defines the degree of association between events that may occur in the information system and the cause of the failure. The failure handling device 100 is also capable of referencing information (e.g., failure handling rules) that pre-associates the above-mentioned failure causes with the details of how to deal with them. The definition information of the degree of association and the failure handling rules are information that are pre-defined by the operator of the information system, etc. The definition information of the degree of association and the failure handling rules are stored in a storage device (not shown) inside or outside the failure handling device 100.

[0016] The fault handling device 100 includes an identification unit 11, a calculation unit 12, a selection unit 13, and an execution control unit 14. The identification unit 11 identifies the degree of relevance between each of a plurality of occurrence events that occurred during a predetermined period and each of the failure causes based on the relevance definition information. An "occurrence event" is one of a plurality of events that may occur in an information system and is defined in the definition information. The calculation unit 12 calculates a certainty factor for each failure cause in a set of the occurrence events using the relevance factor identified by the identification unit 11 based on the elapsed time since the occurrence of each of the occurrence events. In other words, the calculation unit 12 calculates a plurality of certainty factors corresponding to each of a plurality of failure causes in various events and failures that may occur in an information system. Here, the "certainty factor" is a value indicating the likelihood of the assumption that the corresponding failure cause is the cause of the set of occurrence events. Therefore, the higher the certainty factor, the more likely the failure cause corresponding to the certainty factor is the cause of the set of occurrence events. The selection unit 13 selects, from among the plurality of failure causes, the one with the highest certainty factor calculated by the calculation unit 12 as a candidate for the failure cause. The execution control unit 14 causes the information system to execute a countermeasure associated with the candidate cause of the failure. For example, the execution control unit 14 causes the information system to execute a command corresponding to the countermeasure. Alternatively, the execution control unit 14 may separately input the countermeasure into a command execution tool and cause the command corresponding to the countermeasure to be executed with the information system as the destination.

[0017] 2 is a flowchart showing the flow of the failure handling method. First, the identification unit 11 identifies the degree of association between each of multiple occurrence events that occurred during a predetermined period and each failure cause based on definition information that defines the degree of association between events that can occur in the information system and the failure cause (S11). Next, the calculation unit 12 calculates the certainty factor for each failure cause in the set of multiple occurrence events using the degree of association identified in step S11 based on the elapsed time since each of the occurrence events occurred (S12). Then, the selection unit 13 selects the failure cause with the highest certainty factor calculated in step S12 from the multiple failure causes as a candidate failure cause (S13). Then, the execution control unit 14 causes the information system to execute the response content associated with the candidate failure cause (S14).

[0018] In this way, the failure response device 100 according to this embodiment calculates the certainty of the failure cause in a set of occurrence events using the individual associations between the events and the failure cause. At this time, the failure response device 100 dynamically calculates the certainty based on the elapsed time since the occurrence of each occurrence event included in the set. In other words, even if the set of occurrence events is the same, the certainty of each failure cause may be calculated as a different value depending on the elapsed time. For example, even if the set of occurrence events is the same, the certainty of each failure cause may be calculated as a different value depending on the order in which the events occur or the interval between events. Then, based on the certainty, hypotheses (failure causes) can be narrowed down and appropriate countermeasures can be implemented. Therefore, appropriate countermeasures can be implemented according to the group of events that occurred due to the failure of the information system.

[0019] The failure handling device 100 includes a processor, memory, and storage device (not shown). The storage device stores a computer program that implements the processing of the failure handling method according to this embodiment. The processor then loads the computer program from the storage device into the memory and executes the computer program. This allows the processor to implement the functions of an identification unit 11, a calculation unit 12, a selection unit 13, and an execution control unit 14.

[0020] Alternatively, each component of the failure handling device 100 may be realized by dedicated hardware. Furthermore, some or all of the components of each device may be realized by general-purpose or dedicated circuits, processors, etc., or a combination of these. These may be configured by a single chip, or by multiple chips connected via a bus. Some or all of the components of each device may be realized by a combination of the above-mentioned circuits, etc., and a program. Furthermore, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field-Programmable Gate Array), a quantum processor (quantum computer control chip), etc., may be used as the processor.

[0021] Furthermore, when some or all of the components of the failure handling device 100 are realized by multiple information processing devices, circuits, etc., the multiple information processing devices, circuits, etc. may be centrally or decentralized. For example, the information processing devices, circuits, etc. may be realized as a client-server system, a cloud computing system, or the like, in a form in which each is connected via a communication network. Furthermore, the functions of the failure handling device 100 may be provided in a SaaS (Software as a Service) format.

[0022] Second Embodiment FIG. 3 is a block diagram showing an overall configuration including a fault handling system 1000. The fault handling system 1000 is an information system that monitors an information system 1 and includes a monitoring tool 2, a fault handling device 3, and a countermeasure execution tool 4. The information system 1 is a monitored information system and is composed of one or more computers. The monitoring tool 2 monitors messages output from the information system 1, and if a specific fault message is detected, notifies the fault handling device 3 of the detected fault message. The fault handling device 3 is an example of the fault handling device 100 described above. The fault handling device 3 records information extracted from the fault messages obtained from the monitoring tool 2 as "occurring events" in an event occurrence history (database) described below. The fault handling device 3 then selects a fault cause and corresponding countermeasure according to the order of occurrence of multiple events that occurred within a predetermined period, and outputs an execution instruction to the countermeasure execution tool 4 to execute the selected countermeasure on the information system 1, etc. The countermeasure execution tool 4 (command execution tool) executes the action command included in the execution instruction input from the fault handling device 3 to the destination specified in the execution instruction. For example, if the destination is the information system 1, the countermeasure execution tool 4 executes the action command in the information system 1. If the destination is the mail server 5, the countermeasure execution tool 4 executes the action command and outputs the outgoing mail to the operation terminal 6. The monitoring tool 2 and the countermeasure execution tool 4 are realized by a computer program executed on the same computer as the failure handling device 3 or on a different computer.

[0023] The failure handling device 3 is also connected to an operation terminal 6. The operation terminal 6 is a computer operated by an operator to operate the information system 1. If the information system 1 is not restored by executing the countermeasures, the failure handling device 3 notifies the operation terminal 6 of this fact. In response to the notification from the failure handling device 3, the operation terminal 6 displays a screen, outputs a sound, etc.

[0024] FIG. 4 is a block diagram showing the configuration of the failure handling device 3. The failure handling device 3 may be configured redundantly with multiple servers, and each functional block may be implemented by multiple computers. The failure handling device 3 includes a storage unit 310, a memory 320, a communication unit 330, and a control unit 340. The storage unit 310 is an example of a non-volatile storage device such as a hard disk or flash memory. The storage unit 310 stores a program 311, a failure handling rule 312, relevance definition information 313, and an event occurrence history 314. Some or all of the information stored in the storage unit 310 may be stored in an external storage device. The program 311 is a computer program (failure handling program) that implements the failure handling process and the like according to the second embodiment. The storage unit 310 may also store an OS (Operating System), which is not shown in the figure.

[0025] The fault handling rules 312 are information that associates in advance the order of occurrence (occurrence pattern) of multiple events that may occur in the information system 1, the cause of the failure, and the details of the response. Fig. 5 is a block diagram showing an example of the configuration of the fault handling rules 312. The fault handling rules 312 include rule information 3121, ..., rule information 312m (m is a natural number). The rule information 3121 is information that associates an event occurrence pattern 351, a cause of the failure 352, and details of the response 353. Note that the rule information 3122 (not shown) to 312m have the same configuration as the rule information 3121, and therefore detailed description thereof will be omitted.

[0026] The event occurrence pattern 351 is information that defines the occurrence order for each of event IDs 361, 362, ..., 36n (n is a natural number equal to or greater than 2). The event ID is identification information for the type of event. In this example, the event ID 361 is associated with the occurrence order number "1", the event ID 362 is associated with the occurrence order number "2", ..., and the event ID 36n is associated with the occurrence order number "n". The event occurrence pattern 351 may use information indicating the type of event or the content of the event instead of the event ID. The event occurrence pattern 351 is a specific example of information that indicates the occurrence order of multiple events that can occur in the information system 1, and the occurrence order is not limited to this.

[0027] The failure cause 352 is information indicating the cause of the failure that is considered to be the cause of the occurrence of the event occurrence pattern 351, such as identification information of the failure, text information indicating the cause of the failure (the name of the failure), etc. The countermeasure content 353 is information indicating the method and content for recovering the information system 1 from the failure. Examples of the countermeasure content 353 include, but are not limited to, identification information of the countermeasure content, a destination server, an action command, command parameters, etc.

[0028] The relevance definition information 313 defines the relevance between events that may occur in the information system 1 and failure causes. FIG. 6 is a diagram showing an example of the relevance definition information 313. The relevance definition information 313 indicates the individual relevance between events P1 to P5 and failure causes (hypothesis) A1 to A5 in one-to-one combinations. For example, the relevance between event P1 and failure cause A1 is defined as "1," the relevance between event P3 and failure cause A1 is defined as "3," the relevance between event P2 and failure cause A2 is defined as "0," and the relevance between event P1 and failure cause A4 is defined as "2." Therefore, in this example, the relevance between failure cause A1 and event P3 is higher than that between event P1. Furthermore, the relevance between event P1 and failure cause A4 is higher than that between event P1 and failure cause A1. Note that the relevance is merely an example and is not limited to these. Events P1 to P5 are information equivalent to the event IDs described above. Examples of information indicated by events P1 to P5 include the occurrence of an unknown error in an application, the occurrence of a URL (Uniform Resource Locator) monitoring error, the occurrence of a NW (Network) communication error, the occurrence of a high CPU load, and the occurrence of a disk write error. However, examples of events are not limited to these. Furthermore, failure causes (assumptions) A1 to A5 are information equivalent to the identification information of the failure causes described above. Examples of information indicated by failure causes (assumptions) A1 to A5 include failure of external API (Application Programming Interface) integration due to a communication failure, hang-up due to a bug, disk failure, high load due to excessive access, and service failure due to incorrect web server settings. However, examples of failure causes (assumptions) are not limited to these.

[0029] The event occurrence history 314 is a database of history information that records events that have occurred in the information system 1. Fig. 7 is a block diagram showing an example of the configuration of the event occurrence history 314. The event occurrence history 314 includes history records 3141, ..., 314j (j is a natural number). The history record 3141 is information that associates a history ID 371, an occurrence time 372, a detected device 373, an event ID 374, and an event content 375. Note that the history records 3142 (not shown) to 314j have the same configuration as the history record 3141, and therefore detailed description thereof will be omitted.

[0030] The history ID 371 is identification information for the history record 3141. The occurrence time 372 is the time when the event occurred. The occurrence time 372 may be the time when the monitoring tool 2 detected a fault message in the information system 1 or the time when the fault handling device 3 received a fault message from the monitoring tool 2. The detection device 373 is information indicating the device in which the event was detected among multiple devices (servers, network devices, etc.) that make up the information system 1. The event ID 374 is identification information for the type of event. The event content 375 may include, but is not limited to, text information of the fault message, a code extracted from the fault message, a parameter value, a fault level, etc.

[0031] The memory 320 is a volatile storage device such as RAM (Random Access Memory), and is a storage area for temporarily storing information while the control unit 340 is operating. The communication unit 330 is a communication interface between the failure handling device 3 and the outside (for example, the monitoring tool 2, the response execution tool 4, and the operation terminal 6). For example, the communication unit 330 receives a failure message from the monitoring tool 2 and outputs the received failure message to the control unit 340. The communication unit 330 also receives an execution instruction including an action command from the control unit 340 and outputs it to the response execution tool 4. The communication unit 330 also receives a notification message from the control unit 340 and outputs it to the operation terminal 6.

[0032] The control unit 340 is a processor, i.e., a control device, that controls each component of the failure handling device 3. The control unit 340 loads the OS and program 311 from the storage unit 310 into the memory 320 and executes the OS and program 311. In this way, the control unit 340 realizes the functions of an occurrence event recording unit 341, an identification unit 342, a determination unit 343, a calculation unit 344, a selection unit 345, an execution control unit 346, and an update unit 347.

[0033] The event recording unit 341 acquires a fault message notified from the monitoring tool 2, generates a history record using information extracted from the fault message, and records the history record in the event occurrence history 314. For example, the event recording unit 341 analyzes the acquired fault message to extract the time of event occurrence, event content, etc., and identifies the detected device, event ID, etc. Note that various fault message analysis logics can be used. The event recording unit 341 then issues a history ID and registers a history record including the extracted and identified information in the event occurrence history 314.

[0034] The identifying unit 342 is a specific example of the identifying unit 11 described above. The identifying unit 342 identifies a set of occurrence events that occurred within a predetermined period from the event occurrence history 314. The "predetermined period" may be a time period previously set for the normal operation of the information system 1, or a time period extending a predetermined time back from a reference time. In other words, the "predetermined period" is a period from the determination start time to the determination end time, and can be considered a time period in which the interval between the determination start time and the determination end time is a "predetermined time." Here, the "predetermined time" is a predetermined time, such as five minutes, but is not limited thereto. Furthermore, the "reference time" may be the current time or the occurrence time of the most recent event, which can be considered the "determination end time." Therefore, the "determination start time" can be considered to be the time extending a "predetermined time" back from the "determination end time." In other words, the "determination end time" can be considered to be the time when a "predetermined time" has elapsed since the "determination start time."

[0035] In particular, the identification unit 342 identifies the order of occurrence of events that occurred within a predetermined period (event occurrence pattern). Furthermore, the identification unit 342 identifies the degree of association between each occurrence event included in the identified set (or event occurrence pattern) and each failure cause, based on the association definition information 313.

[0036] The determining unit 343 determines whether the failure handling rule 312 includes the order of occurrence of the plurality of events identified by the identifying unit 342 .

[0037] The calculation unit 344 is a specific example of the calculation unit 12 described above. The calculation unit 344 calculates the importance of each occurring event according to the elapsed time since the occurrence. Specifically, the calculation unit 344 calculates the importance of each occurring event identified by the identification unit 342 so that the shorter the elapsed time from the occurrence of the event to the reference time, the higher the value of the importance. In other words, the identification unit 342 calculates the importance of each occurring event so that the more recently the event occurred, the higher the value of the importance.

[0038] The calculation unit 344 then calculates the certainty factor for each fault cause in the event occurrence pattern using the calculated importance and the identified relevance. Furthermore, the calculation unit 344 calculates the (importance and) certainty factor when the determination result by the determination unit 343 indicates that the event occurrence pattern is not included in the fault handling rules 312.

[0039] The selection unit 345 is a specific example of the selection unit 13 described above. The selection unit 345 selects, as a candidate for the fault cause, the fault cause with the highest certainty calculated by the calculation unit 344 from among the multiple fault causes defined in the fault handling rules 312 or the relevance definition information 313. The selection unit 345 also selects a countermeasure from the fault handling rules 312 based on the determination result by the determination unit 343. Specifically, when the determination result by the determination unit 343 indicates that the fault handling rules 312 do not include an event occurrence pattern, the selection unit 345 selects, from the fault handling rules 312, a countermeasure associated with the candidate for the fault cause selected based on the certainty. When the determination result by the determination unit 343 indicates that the fault handling rules 312 include an event occurrence pattern, the selection unit 345 also selects, from the fault handling rules 312, a countermeasure associated with the event occurrence pattern.

[0040] If the highest certainty calculated by the calculation unit 344 is equal to or less than 0, the selection unit 345 does not select a candidate cause of the failure, and may notify the operation terminal 6 of the information system 1 that the cause of the failure in the set of multiple occurring events cannot be identified. This allows the operation staff to consider and implement countermeasures for an unknown event occurrence pattern.

[0041] Furthermore, the selection unit 345 may select a candidate cause of failure from among a plurality of causes of failure remaining after excluding causes of failure whose certainty is calculated to be equal to or less than 0. This narrows down the selection targets, thereby reducing the load of the selection process.

[0042] The execution control unit 346 is a specific example of the above-mentioned execution control unit 14. The execution control unit 346 outputs an execution instruction to the response execution tool 4 to cause the information system 1 or the like to execute the response content selected by the selection unit 345. Specifically, the execution control unit 346 generates an execution instruction including an action command, a destination server, command parameters, etc., and transmits the execution instruction to the response execution tool 4.

[0043] The update unit 347 updates the degree of association between the candidate cause of the failure and each of the multiple occurring events according to the execution result of the countermeasure content. Specifically, if the execution result indicates a normal termination or the information system 1 has been recovered, the update unit 347 adds a predetermined value to the degree of association between the selected candidate cause of the failure and each of the multiple occurring events, and updates the degree of association definition information 313. Furthermore, if the execution result indicates an abnormal termination or the information system 1 has not been recovered, the update unit 347 subtracts a predetermined value from the degree of association between the selected candidate cause of the failure and each of the multiple occurring events, and updates the degree of association definition information 313.

[0044] 8 is a sequence diagram showing a flow including a fault handling process. First, the monitoring tool 2 periodically monitors the information system 1 and detects fault messages (S21-1). Then, the monitoring tool 2 notifies the fault handling device 3 of the detected fault messages (S22-1).

[0045] Next, the event recording unit 341 of the failure handling device 3 analyzes the failure message obtained from the monitoring tool 2 and extracts various information. For example, the event recording unit 341 extracts from the failure message the message ID, message body, service type, failure name, failure level, server temperature, CPU usage rate, occurrence date and time (hour and minute), memory usage rate, etc. Note that the extracted information is not limited to these. The event recording unit 341 then generates a history record using the extracted information and records it in the event occurrence history 314 (S23-1). For example, the event recording unit 341 records in the event occurrence history 314 that event P1 has occurred.

[0046] The monitoring tool 2 periodically monitors the information system 1 after step S21-1. During some monitoring, the monitoring tool 2 detects a fault message (S21-2). The monitoring tool 2 then notifies the fault handling device 3 of the detected fault message (S22-2). The occurring event recording unit 341 then records the event in the event occurrence history 314 in the same manner as above (S23-2). For example, the occurring event recording unit 341 records the occurrence of event P2 in the event occurrence history 314.

[0047] Thereafter, the monitoring tool 2 continues to monitor the information system 1 periodically in the same manner, and as soon as it detects a fault message, it notifies the fault handling device 3, and the fault handling device 3 records the event that has occurred in the event recording unit 341 in response to receiving the fault message.

[0048] Thereafter, the failure handling device 3 executes the failure handling process periodically or at an arbitrary timing (S24). For example, the failure handling device 3 may execute step S24 at intervals longer than the monitoring interval of the monitoring tool 2.

[0049] FIG. 9 is a flowchart showing the flow of the fault handling process in the fault handling device 3. First, the identification unit 342 identifies an event occurrence pattern within a predetermined period (S301). For example, the identification unit 342 periodically starts execution and references the event occurrence history 314 to identify the occurrence time t3 of event P3, which occurred immediately before the current time. The identification unit 342 then sets the occurrence time t3 as the determination end time te and sets the time preceding the determination end time te by a predetermined time (e.g., 5 minutes) as the determination start time ts. The identification unit 342 then identifies a set of events that occurred during the predetermined period from the event occurrence history 314, with the time period from the determination start time ts to the determination end time te as the predetermined period. The identification unit 342 then sorts the occurrence times of the identified events in ascending order and identifies the order of occurrence of the events as the event occurrence pattern. For the sake of explanation, the event occurrence pattern will be described below as being in the order of P1, P2, and P3.

[0050] Then, the determination unit 343 determines whether the identified event occurrence pattern matches an existing event occurrence pattern (S302). Specifically, the determination unit 343 determines whether an event occurrence pattern in the order of P1, P2, P3 exists in the fault handling rules 312. Note that in this case, the determination unit 343 determines whether an event occurrence pattern that completely matches the identified event occurrence pattern exists in the fault handling rules 312.

[0051] If the determination result in step S302 indicates that an existing event occurrence pattern that completely matches the existing event occurrence pattern exists, the selection unit 345 selects a cause of the failure and a countermeasure corresponding to the event occurrence pattern identified in step S301 from the failure handling rules 312 (S303). Then, the execution control unit 346 instructs the countermeasure execution tool 4 to execute the selected countermeasure (S304).

[0052] On the other hand, if the determination result of step S302 indicates that there is no existing event occurrence pattern that completely matches, the calculation unit 344 calculates the importance level based on the elapsed time since the occurrence of each event (S305). Note that the "elapsed time since the occurrence of each event" can be interpreted as "the time elapsed from the occurrence time of each event to the reference time (the determination end time)" or "the retroactive time from the reference time to the occurrence time of each event." As described above, the reference time may be the current time or the occurrence time of the most recent event. The calculation unit 344 may calculate the importance level using a calculation logic in which, for example, the shorter the elapsed time, the higher the value (the longer the elapsed time, the lower the value).

[0053] Specifically, the calculation unit 344 calculates the occurrence time t p The importance of the occurrence event P w t (P) may be calculated, where t s is the judgment start time, t e is the end time of the judgment.

[0054] FIG. 10 is a diagram showing an example of the relationship between the elapsed time from the occurrence time of each event in an event occurrence pattern and the importance level. Here, the occurrence order (occurrence event pattern) of three occurrence events is P1, P2, and P3, with the occurrence time of occurrence event P1 being t1, the occurrence time of occurrence event P2 being t2, and the occurrence time of occurrence event P3 being t3. The occurrence time t3 of occurrence event P3, which is the most recent of the three occurrence events, is defined as the judgment end time te, and the time ts, a predetermined time before the judgment end time te, is defined as the judgment start time. The occurrence times t1 and t2 are assumed to be times after the judgment start time ts. In the example of FIG. 10, the importance level wt(P1) of occurrence event P1 is calculated as 0.2, the importance level wt(P2) of occurrence event P2 is 0.6, and the importance level wt(P3) of occurrence event P3 is 1.0. In other words, the more recent an event is with respect to the reference time (judgment end time te), the higher the calculated importance level. Here, the time elapsed from the occurrence time t2 to the reference time is shorter than the time elapsed from the occurrence time t1 to the reference time. Therefore, the importance of the occurring event P2, "0.6," is calculated to be higher than the importance of the occurring event P1, "0.2." Since the freshness of information can be said to decrease over time, a weight according to the elapsed time from the occurrence time of the event is reflected in subsequent certainty calculations. This allows for increased accuracy of the certainty.

[0055] After step S305, the identification unit 342 identifies the degree of association between each occurrence event and each failure cause (S306). Fig. 11 is a diagram showing an example of the identified degrees of association. Here, the degrees of association (each value surrounded by a dashed line in Fig. 11) are individually defined for the combination of each occurrence event P1 to P3 and each of all failure causes A1 to A5. Note that step S306 may be executed before step S305 or in parallel with step S305.

[0056] Thereafter, the calculation unit 344 calculates the certainty of each fault cause (hypothesis) using the importance calculated in step S305 and the relevance identified in step S306 (S307). Specifically, the calculation unit 344 calculates the event occurrence pattern (set of occurring events) P a From P k Each cause of failure (assumption) A1 From A N Here, R(P, A) is the degree of association between the occurrence event P and the hypothesis A.

[0057] The determination unit 343 then determines whether any of the calculated certainties is equal to or greater than 0 (S308). If any of the certainties is equal to or greater than 0, the selection unit 345 selects the hypothesis with the highest certainty as a candidate for the cause of the failure (S309). The selection unit 345 then selects a countermeasure corresponding to the selected candidate for the cause of the failure from the fault handling rules 312 (S310). The execution control unit 346 then issues an execution instruction to the countermeasure execution tool 4 for the selected countermeasure (S304). Specifically, the execution control unit 346 generates an execution instruction including an action command, a destination server, command parameters, etc., and transmits the execution instruction to the countermeasure execution tool 4 (S25 in FIG. 8 ).

[0058] In response to this, the countermeasure execution tool 4 executes an action command with set parameters, etc., on the destination server (information system 1) included in the received execution instruction (S26). For example, the countermeasure execution tool 4 remotely turns the power OFF / ON on a specific server of the information system 1. Alternatively, the countermeasure execution tool 4 executes a recovery script on the specific server. Note that the action commands and execution instructions are not limited to these. The countermeasure execution tool 4 then receives the execution result from the information system 1 (S27) and transmits the execution result to the fault handling device 3 (S28). Note that if the countermeasure execution tool 4 is unable to receive the execution result after a certain period of time has elapsed, it transmits to the fault handling device 3 an execution result indicating that the execution has failed or recovery has not been completed.

[0059] Thereafter, the update unit 347 of the failure handling device 3 receives the execution result from the response execution tool 4 and determines whether the information system 1 has been restored based on the execution result (S311 in FIG. 9 ). If it is determined in step S311 that the information system 1 has been restored, the update unit 347 adds a predetermined value to the degree of association between each event in the event occurrence pattern and the selected cause of the failure (S312). On the other hand, if it is determined in step S311 that the information system 1 has not been restored, the update unit 347 subtracts a predetermined value from the degree of association between each event in the event occurrence pattern and the selected cause of the failure (S313).

[0060] If the certainty level is not 0 or higher in step S308, or after step S313, the failure handling device 3 (e.g., the selection unit 345) notifies the operation terminal 6 that the information system 1 has not been restored (S314). For example, the failure handling device 3 may notify the details of the event occurrence pattern identified in step S301 (the contents of the history record of each event), the importance of each event, the degree of association between each event and each failure cause, the certainty level of each failure cause, and the fact that the failure cause in the event occurrence pattern cannot be identified. The failure handling device 3 may also notify the execution result. The failure handling device 3 may also output a warning email directly to the mail server 5.

[0061] Next, specific examples of importance and certainty calculated by this embodiment, examples of update results of relevance, and examples of adding definitions to relevance definition information will be described.

[0062] (Example 1) Here, it is assumed that the event occurrence pattern is the order of occurrence of P1, P2, and P3. In this case, the failure handling device 3 calculates the importance of each occurring event by the following formula (3).

[0063] Then, the fault handling device 3 identifies the degree of association between each of the occurring events P1 to P3 and each of the fault causes (assumptions) A1 to A5, and calculates the confidence C(A) for each assumption A using the following formula (4).

[0064] As a result, among the confidence levels of all the calculated assumptions, confidence level C (A 1 ) is 3.8, which is the highest. Therefore, the selection unit 345 selects 1) is selected as a candidate for the cause of the fault.

[0065] (Example 2) Here, it is assumed that the event occurrence pattern is the order of occurrence of P1, P3, and P2. In other words, in Example 2, the order of occurrence of the events P2 and P3 is reversed from that in Example 1. In this case, the fault handling device 3 calculates the importance of each occurrence event using the following formula (5).

[0066] Then, the failure handling device 3 identifies the degree of association in the same manner as in the first embodiment, and calculates the confidence factor C(A) for each assumption A using the following formula (6).

[0067] As a result, of the certainty factors of all the calculated hypotheses, certainty factor C(A5) is 3.4, which is the highest. Therefore, the selection unit 345 selects hypothesis A5 with certainty factor C(A5) as a candidate for the cause of the failure. Therefore, although the relevance is the same in Examples 1 and 2, different candidates for the cause of the failure are selected because the importance differs depending on the order in which the events occurred. In particular, a hypothesis with a higher relevance to the most recently occurred event is selected as it is more likely to have a significant impact as a cause of the failure. This allows more appropriate countermeasures to be implemented.

[0068] Third Embodiment In the first embodiment described above, when a countermeasure associated with the failure cause A1 is executed and the execution result indicates recovery, the update unit 347 adds 1 to each of the degrees of association between the failure cause A1 and each of the event occurrence patterns P1, P2, and P3. FIG. 12 is a diagram showing an example of the update results of the degrees of association. Here, the update unit 347 adds 1 to the degree of association R(P1, A1) "1" to update it to "2," adds 1 to the degree of association R(P2, A1) "1" to update it to "2," and adds 1 to the degree of association R(P3, A1) "3" to update it to "4."

[0069] In the third embodiment, it is assumed that events subsequently occur again in the order of event occurrence patterns P1, P2, and P3. In this case, the fault handling device 3 calculates the importance of each occurring event in the same manner as in equation (3) above. Then, the fault handling device 3 identifies the degrees of association between each of the occurring events P1 to P3 and the fault cause A1 as "2," "2," and "4" from the association definition information 313 in FIG. 12 above. Note that the other degrees of association are identified in the same manner as in the first embodiment. Then, the fault handling device 3 calculates the confidence C(A) for each assumption A using equation (7) below.

[0070] As a result, among the confidence levels of all the calculated assumptions, confidence level C (A 1 ) is 5.6, which is the highest. Therefore, the selection unit 345 selects the confidence C(A 1 ) is selected as a candidate for the cause of the fault.

[0071] 12, events occur in the order of event occurrence patterns P1, P3, and P2, similarly to the above-described Example 2. In this case, the fault handling device 3 calculates the importance of each occurring event in the same manner as in the above-described formula (5).

[0072] Then, the failure handling device 3 identifies the degree of association in the same manner as in the third embodiment, and calculates the confidence factor C(A) for each assumption A using the following formula (8).

[0073] As a result, among the confidence levels of all the calculated assumptions, confidence level C (A 1 ) is 4.8, which is the highest. Therefore, the selection unit 345 1 ) is selected as a candidate for the cause of the failure. In this way, in Example 4, even if the event occurrence patterns are the same as those of Example 2, P1, P3, and P2, a different cause of the failure than in Example 2 is selected because the execution results of the countermeasures are fed back to the relevance. In other words, by feeding back the execution results of the countermeasures to the relevance, it is possible to influence the likelihood of a hypothesis being selected.

[0074] In a fifth embodiment, it is assumed that events occur in the order of event occurrence patterns P1 and P2 after the relevance definition information 313 is updated to that shown in Fig. 12. In this case, the fault handling device 3 calculates the importance of each occurring event using the following formula (9).

[0075] 12, the failure handling device 3 identifies the degrees of association between each of the occurring events P1 and P2 and the failure cause A1 as "2" and "2." The failure handling device 3 identifies the degrees of association between each of the occurring events P1 and P2 and the other failure causes A2 to A5 in the same manner as in Example 1. The failure handling device 3 then calculates the confidence C(A) for each assumption A using the following formula (10):

[0076] As a result, among the confidence levels of all the calculated assumptions, confidence level C (A 1 ) is 3.0, which is the highest. Therefore, the selection unit 345 selects the confidence C(A 1 ) is selected as a candidate for the cause of the fault.

[0077] If the execution result of the countermeasure indicates that the system is not restored, the update unit 347 subtracts 1 from each of the degrees of association between the event occurrence patterns P1 and P2 and the failure cause A1. FIG. 13 is a diagram showing an example of the update results of the degrees of association. Here, the update unit 347 subtracts 1 from the degree of association R(P1, A1) of "2" to update it to "1," and subtracts 1 from the degree of association R(P2, A1) of "2" to update it to "1." Therefore, among the degrees of association with the failure cause A1, the degree of association with the event P3, which has a greater impact, increases, and the degrees of association with the events P1 and P2, which have a smaller impact, decrease.

[0078] After that, it is assumed that the events occur again in the order of event occurrence patterns P1 and P2. In this case, the fault handling device 3 calculates the importance of each occurring event in the same manner as in equation (9) above. Then, the fault handling device 3 specifies the degree of association between each of the occurring events P1 and P2 and the fault cause A1 as "1" and "1" from the association definition information 313 in FIG. 13 above. Note that the fault handling device 3 also specifies the degree of association between each of the occurring events P1 and P2 and the other fault causes A2 to A5 in the same manner as above. Then, the fault handling device 3 calculates the confidence C(A) for each assumption A using equation (11) below.

[0079] As a result, of the certainty factors of all the calculated hypotheses, certainty factor C(A5) is 2.5, which is the highest. Therefore, the selector 345 selects hypothesis A5 with certainty factor C(A5) as a candidate for the fault cause. Therefore, if the association between fault cause A1 and events P1 and P2 decreases and a set of events P1 and P2 occurs, the certainty factor of A5 will be higher than that of fault cause A1, and A5 will be selected. In other words, feedback on the association factor allows appropriate countermeasures to be implemented.

[0080] Sixth Embodiment When a definition of a new failure cause (hypothesis) is added to the failure handling rules 312, it is also necessary to add the relevance between the corresponding failure cause and each event to the relevance definition information 313. In this case, a user such as an operations staff member specifies, to the failure handling device 3, an event in the relevance definition information 313 that is considered to be relevant (provides evidence) to the failure cause to be added. In response to this, the failure handling device 3 calculates the relevance between the new failure cause and the specified event through statistical processing using each relevance of the specified event, and adds the calculated relevance to the relevance definition information 313. An example of the statistical processing is, but is not limited to, selecting the median of relevance values ​​greater than 0.

[0081] For example, when a new failure cause A6 is added, it is assumed that events P1 and P4 are specified in the above-described relevance definition information 313 of Fig. 13. At this time, the failure handling device 3 calculates the relevance between the failure cause A6 and event P1 using the following formula (12). The failure handling device 3 also calculates the relevance between the failure cause A6 and event P4 using the following formula (13). The failure handling device 3 also sets the relevance between the failure cause A6 and the other events P2, P3, and P5 to "0."

[0082] 14 is a diagram showing an example in which a failure cause (assumption) A6 is added to the relevance definition information 313. That is, the following has been added to the row for the failure cause A6: relevance R(P1, A6) "1", relevance R(P2, A6) "0", relevance R(P3, A6) "0", relevance R(P4, A6) "2.5", and relevance R(P5, A6) "0".

[0083] By using a statistical process that selects the median value among existing relevance values ​​greater than 0, it is possible to use a highly effective positive value as the relevance and to eliminate the influence of extremely high relevance values, thereby ensuring the validity of the initial relevance value.

[0084] Seventh Embodiment When a definition for a new event is added to the fault handling rules 312, it is also necessary to add the degree of association between the corresponding event and each fault cause (hypothesis) to the relevance definition information 313. In this case, a user such as an operations person specifies, to the fault handling device 3, a fault cause that is considered to be related to (serves as evidence for) the event to be added from the relevance definition information 313. In response to this, the fault handling device 3 calculates the degree of association between the new event and the specified fault cause by statistical processing using each relevance for the specified fault cause, and adds the calculated degree of association to the relevance definition information 313. An example of the statistical processing is, but is not limited to, selecting the median value from among the relevance values ​​greater than 0.

[0085] For example, when a new event P6 is added, it is assumed that fault causes (assumptions) A1 and A4 are specified in the above-described relevance definition information 313 of Fig. 13. At this time, the fault handling device 3 calculates the relevance between the event P6 and the fault cause A1 using the following formula (14). The fault handling device 3 also calculates the relevance between the event P6 and the fault cause A4 using the following formula (15). The fault handling device 3 also sets the relevance between the event P6 and the other fault causes A2, A3, and A5 to "0."

[0086] 15 is a diagram showing an example in which an event P6 is added to the relevance definition information 313. That is, the following have been added to the column for the event P6: relevance R(P6, A1) "1", relevance R(P6, A2) "0", relevance R(P6, A3) "0", relevance R(P6, A4) "2", and relevance R(P6, A5) "0".

[0087] By using a statistical process that selects the median value among existing relevance values ​​greater than 0, it is possible to use a highly effective positive value as the relevance and to eliminate the influence of extremely high relevance values, thereby ensuring the validity of the initial relevance value.

[0088] The relevance level is preferably an integer (including positive and negative). By including negative values ​​in the relevance level, hypotheses (causes of failure) with low relevance to each occurrence event included in the set of occurrence events can be easily excluded from selection by setting the threshold value to "0."

[0089] Furthermore, when selecting candidate causes of a fault, by excluding from the selection candidates those with a certainty level below the threshold value of "0", it is possible to prevent the erroneous selection of a hypothesis (cause of fault) that has little correlation with the set of events that occur due to the fault and the implementation of countermeasures.

[0090] Furthermore, by updating the relevance (feedback) after the selected countermeasure is implemented, it is possible to select an appropriate countermeasure for the next event occurrence pattern. Conversely, by not performing recursive feedback within a single fault handling process, it is possible to prevent unpredictable chain failures caused by repeatedly implementing incorrect countermeasures.

[0091] Other Embodiments In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.

[0092] The present disclosure is not limited to the above-described embodiments, and may be modified as appropriate without departing from the spirit and scope of the present disclosure. In addition, the present disclosure may be implemented by appropriately combining the respective embodiments.

[0093] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes: (Supplementary Note A1) A failure handling device comprising: an identifying means for identifying a degree of association between each of a plurality of occurrence events that have occurred within a predetermined period and each failure cause, based on definition information that defines the degree of association between events that can occur in an information system and failure causes; a calculating means for calculating a certainty factor for each failure cause in a set of the plurality of occurrence events using the identified degree of association, based on the elapsed time since the occurrence of each of the plurality of occurrence events; a selecting means for selecting, from the plurality of failure causes, the one with the highest calculated certainty factor as a candidate failure cause; and an execution control means for causing the information system to execute a countermeasure content associated with the candidate failure cause. (Supplementary Note A2) The failure handling device according to Supplementary Note A1, further comprising: an updating means for updating the degree of association between the candidate failure cause and each of the plurality of occurrence events according to a result of executing the countermeasure content. (Supplementary Note A3) The fault handling device according to Supplementary Note A1 or A2, further comprising: a determination means for determining whether a sequence of occurrence of a plurality of events that may occur in the information system is included in fault handling rules that previously associate the sequence of occurrence of the plurality of events with fault causes and countermeasures; wherein the selection means selects the countermeasure from the fault handling rules based on a determination result by the determination means; and the execution control means causes the selected countermeasure to be executed on the information system. (Supplementary Note A4) The fault handling device according to Supplementary Note A3, wherein the calculation means calculates the certainty factor when the determination result indicates that the sequence of occurrence is not included in the fault handling rules; and the selection means selects the countermeasure associated with the candidate cause of the fault selected based on the certainty factor from the fault handling rules. (Supplementary Note A5) The fault handling device according to Supplementary Note A4, wherein the selection means selects the countermeasure associated with the sequence of occurrence from the fault handling rules when the determination result indicates that the sequence of occurrence is included in the fault handling rules.(Appendix A6) The fault handling device according to Appendix A1 or A2, wherein the calculation means calculates the importance for each occurrence event so that the shorter the elapsed time, the higher the value, and calculates the certainty using the calculated importance and the identified relevance. (Appendix A7) The fault handling device according to Appendix A1 or A2, wherein the selection means, if the highest of the calculated certainty levels is 0 or less, does not select the candidate fault cause and notifies the operations terminal of the information system that a fault cause in the set of the plurality of occurrence events cannot be identified. (Appendix A8) The fault handling device according to Appendix A1 or A2, wherein the selection means selects the candidate fault cause from the plurality of fault causes after excluding fault causes whose certainty level is calculated to be 0 or less. (Appendix B1) A fault handling system comprising: a storage device that stores definition information that defines the degree of association between an event that can occur in an information system and a fault cause, and history information of a plurality of events that have occurred in the information system over a predetermined period of time, and a fault handling device connected to the storage device, wherein the fault handling device comprises: identification means that refers to the storage device and identifies the degree of association between each of the plurality of events and each fault cause from the definition information, calculation means that calculates a certainty for each fault cause in a set of the plurality of events using the identified degree of association based on the elapsed time since the occurrence of each of the plurality of events included in the history information, selection means that selects from the plurality of fault causes the one with the highest calculated certainty as a candidate fault cause, and execution control means that causes the information system to execute a countermeasure corresponding to the candidate fault cause. (Appendix B2) The fault handling system according to Appendix B1, wherein the fault handling device further comprises update means that updates the degree of association between the candidate fault cause and each of the plurality of events in accordance with a result of executing the countermeasure.(Appendix C1) A failure handling method in which a computer identifies the degree of relevance between each of a plurality of occurrence events that have occurred during a predetermined period and each failure cause from definition information that defines the degree of relevance between events that can occur in an information system and failure causes, calculates a certainty degree for each failure cause in the set of the plurality of occurrence events using the identified relevance degree based on the elapsed time since the occurrence of each of the plurality of occurrence events, selects the failure cause with the highest calculated certainty degree from the plurality of failure causes as a candidate failure cause, and executes a response content associated with the candidate failure cause on the information system. (Appendix D1) A fault handling program that causes a computer to execute the following steps: a specification process that specifies the degree of association between each of a plurality of occurrence events that occurred during a specified period and each fault cause, based on definition information that defines the degree of association between events that can occur in an information system and the fault cause; a calculation process that calculates a certainty factor for each fault cause in a set of the plurality of occurrence events using the specified degree of association, based on the elapsed time since the occurrence of each of the plurality of occurrence events; a selection process that selects the fault cause with the highest calculated certainty factor from the plurality of fault causes as a candidate fault cause; and an execution control process that causes the information system to execute the countermeasures associated with the candidate fault cause.

[0094] Although the present invention has been described above with reference to the embodiments (and examples), the present invention is not limited to the above-described embodiments (and examples). Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0095] This application claims priority based on Japanese Patent Application No. 2022-205555, filed December 22, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0096] 100 Failure handling device 11 Identification unit 12 Calculation unit 13 Selection unit 14 Execution control unit 1000 Failure handling system 1 Information system 2 Monitoring tool 3 Failure handling device 4 Handling execution tool 5 Mail server 6 Operation terminal 310 Storage unit 311 Program 312 Failure handling rule 313 Association degree definition information 314 Event occurrence history 320 Memory 330 Communication unit 340 Control unit 341 Occurring event recording unit 342 Identification unit 343 Determination unit 344 Calculation unit 345 Selection unit 346 Execution control unit 347 Update unit 3121 Rule information 312m Rule information 351 Event occurrence pattern 361 Event ID 362 Event ID 36n Event ID 352 Cause of failure 353 Handling content 3141 History record 314j History record 371 History ID 372 Occurrence time 373 Detecting device 374 Event ID 375 Event content ts Judgment start time te Judgment end time t1 Occurrence time t2 Occurrence time t3 Occurrence time P1 Event P2 Event P3 Event P4 Event P5 Event P6 Event A1 Failure cause A2 Failure cause A3 Failure cause A4 Failure cause A5 Failure cause A6 Failure cause

Claims

1. an identification means for identifying the degree of association between each of a plurality of occurrence events that occurred during a predetermined period and each of the failure causes from definition information that defines the degree of association between an event that may occur in an information system and a failure cause; a calculation means for calculating a certainty factor for each fault cause in the set of the plurality of occurrence events using the determined relevance factor based on the elapsed time from the occurrence of each of the plurality of occurrence events; a selection means for selecting, from the plurality of fault causes, the one having the highest calculated certainty as a candidate fault cause; an execution control means for causing the information system to execute the countermeasures associated with the candidate causes of the failure; A fault handling device comprising:

2. The system further includes an update unit that updates the degree of association between the candidate cause of the fault and each of the plurality of occurring events in accordance with the result of executing the countermeasure. The fault handling device according to claim 1 .

3. The information processing system further includes a determination unit that determines whether or not the order of occurrence of the plurality of events that may occur in the information system is included in a failure response rule that previously associates the order of occurrence of the plurality of events with the cause of the failure and the content of the response, the selection means selects the response content from the fault response rules based on the determination result by the determination means, The execution control means causes the information system to execute the selected countermeasure.

3. The fault handling device according to claim 1 or 2.

4. the calculation means calculates the certainty factor when the determination result indicates that the failure handling rule does not include the occurrence order; The selection means selects the countermeasure content associated with the candidate cause of the fault selected based on the certainty factor from the fault countermeasure rules. The fault handling device according to claim 3 .

5. When the determination result indicates that the order of occurrence is included in the fault handling rule, the selection means selects a handling content associated with the order of occurrence from the fault handling rule. The fault handling device according to claim 4.

6. The calculation means For each occurrence event, calculate the importance so that the shorter the elapsed time, the higher the value; The certainty factor is calculated using the calculated importance and the specified relevance.

3. The fault handling device according to claim 1 or 2.

7. The selection means If the highest of the calculated certainties is equal to or less than 0, the candidate for the failure cause is not selected, and a notice is sent to the operation terminal of the information system that the failure cause in the set of the plurality of occurrence events cannot be identified.

3. The fault handling device according to claim 1 or 2.

8. a storage device that stores definition information that defines the degree of association between an event that may occur in an information system and a cause of a failure, and history information of a plurality of events that occurred in the information system over a predetermined period of time; a failure handling device connected to the storage device, The failure handling device a specifying means for specifying a degree of association between each of the plurality of occurrence events and each of the failure causes from the definition information by referring to the storage device; a calculation means for calculating a certainty factor for each fault cause in the set of the plurality of occurrence events using the determined relevance factor based on the elapsed time since the occurrence of each of the plurality of occurrence events included in the history information; a selection means for selecting, from the plurality of fault causes, the one with the highest calculated certainty as a candidate for the fault cause; an execution control means for causing the information system to execute the countermeasures associated with the candidate causes of the failure; A fault handling system comprising:

9. The computer Identifying the degree of association between each of a plurality of events that occurred during a predetermined period and each of the causes of the failure, based on definition information that defines the degree of association between events that may occur in the information system and the causes of the failure; calculating a certainty factor for each fault cause in the set of the plurality of occurrence events using the determined relevance factor based on the elapsed time since the occurrence of each of the plurality of occurrence events; selecting the cause of the fault with the highest calculated certainty from among the plurality of causes of the fault as a candidate cause of the fault; causing the information system to execute a countermeasure corresponding to the candidate cause of the failure; How to deal with the problem.

10. a process of identifying the degree of association between each of a plurality of events that occurred during a predetermined period and each of the failure causes based on definition information that defines the degree of association between events that may occur in an information system and the failure causes; a calculation process for calculating a certainty factor for each fault cause in the set of the plurality of occurrence events using the determined relevance factor based on the elapsed time from the occurrence of each of the plurality of occurrence events; a selection process for selecting, from the plurality of fault causes, the one with the highest calculated certainty as a candidate for the fault cause; an execution control process for causing the information system to execute the countermeasures associated with the candidate causes of the failure; A troubleshooting program that causes a computer to execute the following.