Countermeasure selection device, system, and countermeasure selection method

The countermeasure selection device addresses the challenge of selecting system failure recovery measures by automatically prioritizing them based on weighted criteria, enhancing system efficiency and reducing costs.

JP7672275B2Active Publication Date: 2025-05-07HITACHI IND EQUIP SYST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021071306
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-20
Publication Date
2025-05-07
Estimated Expiration
2041-04-20

AI Technical Summary

Technical Problem

Existing countermeasure selection devices for system failures do not adequately consider customer and operator-specific evaluation indicators such as cost, and require expertise to prioritize measures based on stability and safety outputs.

Method used

A countermeasure selection device that automatically determines recovery measures by using a computer system to manage countermeasure candidates, evaluate them based on weighted priority criteria, and output prioritized countermeasures for system failures.

Benefits of technology

Enables the automatic selection and prioritization of recovery measures that align with specific recovery requirements, improving system utilization and reducing operational costs, even for workers with limited knowledge or experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672275000001
    Figure 0007672275000001
  • Figure 0007672275000002
    Figure 0007672275000002
  • Figure 0007672275000003
    Figure 0007672275000003
Patent Text Reader

Abstract

To provide a measure selection device, a system, and a measure selection method that automatically select a measure that meets a recovery requirement and provide priority to output the measure.SOLUTION: A device for determining a recovery measure for a failure that occurs in an own device or the other device includes: a measure candidate selection unit that selects a recovery measure candidate for the failure that occurs; a measure evaluation value management unit that manages an evaluation value for each measure for one or more evaluation indexes, which are items that are emphasized in a recovery measure; a measure priority determination unit that determines, on the basis of one or more evaluation indexes, priority of the recovery measure candidate; and an output unit that outputs the recovery measure candidate with the determined priority provided.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a countermeasure selection device. [Background technology]

[0002] In recent years, systems have been provided that use communication functions to collect data on on-site devices in factories and perform remote monitoring and control of facilities and equipment. On the other hand, as the number of connected devices increases and systems become more complex, the amount of work required to take measures when a failure occurs increases, and there is a demand for automation of recovery measures. However, when there are multiple recovery measures for a failure that occurs, the optimal measure varies depending on the requirements of each system. For example, as a criterion for selecting a measure when a failure occurs in a part of a system, there are cases where time is prioritized to aim for early recovery, and cases where availability is prioritized to prioritize not affecting other system areas even if it takes time. In addition, there are cases where human costs are prioritized, where remote response is prioritized without dispatching workers to the site as much as possible. It is important to consider the recovery requirements for these various evaluation indexes and select and take appropriate measures in accordance with the requirements in terms of improving the system availability and reducing operation costs.

[0003] Therefore, there is a technology that supports the selection of recovery measures for an occurrence of a failure, taking into consideration the stability and security of each measure.Patent Document 1 (JP Patent Publication 2020-86474) describes a recovery support device that includes an index value calculation means for calculating a predetermined index value for a recovery work sequence that indicates a work procedure for recovering from an abnormality that occurs in a group of devices that constitute a communication network, based on the recovery work sequence, and an output means for outputting the index value calculated by the index value calculation means to a predetermined output destination (see claim 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2020-86474 A Summary of the Invention [Problem to be solved by the invention]

[0005] In the background art described in Patent Document 1, stability and safety are taken into consideration, but measures are not selected taking into consideration the inclinations of customers and operators that differ for each system, such as the above-mentioned evaluation index related to cost. In actual operation, it is desirable to select measures that flexibly consider these arbitrary evaluation indexes. In addition, the output of evaluation values ​​related to stability and safety encourages workers to select measures, but whether stability or safety should be emphasized depends on the system operation policy, etc. Therefore, a certain level of knowledge and work experience on the system is required to select the optimal measure from the combination of evaluation values. In order to enable even workers with little knowledge and experience to execute recovery measures, it is desirable to output not only the evaluation value of each measure but also the priority and ranking of the measures based on the viewpoint of the specified requirements, and to uniquely indicate the measures to be taken.

[0006] The present invention has been made in consideration of the above-mentioned problems, and aims to automatically select countermeasures in accordance with any recovery requirements for an occurring failure, and present the countermeasures to be taken in order of priority. [Means for solving the problem]

[0007] A representative example of the invention disclosed in the present application is as follows: A countermeasure selection device that determines a recovery countermeasure for a failure that occurs in its own device or another device, The present invention is configured by a computer having an arithmetic unit that executes a program and a storage unit that stores the program and data, and manages a list of candidate recovery measures for a failure by referring to the candidate recovery measures management information. a countermeasure candidate selection unit that selects a candidate recovery countermeasure for the failure that has occurred; and a countermeasure evaluation value management unit that manages evaluation values ​​for each countermeasure with respect to one or more evaluation indexes that are important matters in the recovery countermeasure; a recovery requirement management unit that manages recovery requirements including weighted values ​​indicating the priority of each of the evaluation indexes and devices to which the recovery requirements are applied; a countermeasure priority determination unit that determines priorities of the candidate recovery measures based on one or more of the evaluation indexes; an input unit that receives an input of an evaluation value for each of the measures and the evaluation index, a weighted value for each of the restoration requirements and the evaluation index, and an apparatus to which the restoration requirements are applied; an output unit that outputs the recovery measure candidates with the determined priorities assigned thereto; The input unit includes an area in which an evaluation value managed by the countermeasure evaluation value management unit can be selected for each countermeasure and evaluation index, an area in which a weight value managed by the recovery requirement management unit can be set for each recovery requirement and evaluation index, and an area in which a recovery requirement managed by the recovery requirement management unit can be set for each of the own device or the other device, and the countermeasure priority determination unit calculates a total value of the evaluation values ​​reflecting the weight values ​​in the recovery requirements applied to a device to be restored for each recovery measure candidate selected by the countermeasure candidate selection unit, and determines a priority such that the higher the calculated total value, the higher the priority of the recovery measure candidate. A countermeasure selection device comprising: Effect of the Invention

[0008] According to one aspect of the present invention, measures that are in accordance with recovery requirements can be automatically selected and output with priorities assigned. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]

[0009] [Figure 1] FIG. 2 is a diagram illustrating a system configuration and a hardware configuration of each device according to the first embodiment. [Diagram 2] 4 is a diagram illustrating an example of a configuration of a device information management table managed by a device information management unit of the management device according to the first embodiment. FIG. [Diagram 3] 4 is a diagram illustrating an example of the configuration of a countermeasure candidate management table managed by a countermeasure candidate selection unit of the management apparatus according to the first embodiment. FIG. [Figure 4] 4 is a diagram illustrating an example of the configuration of an evaluation value management table managed by a countermeasure evaluation value management unit of the management apparatus according to the first embodiment. FIG. [Diagram 5] 1 is a diagram illustrating an example of a configuration of a recovery requirement weight management table managed by a recovery requirement management unit of the management apparatus according to the first embodiment; [Figure 6] 1 is a diagram illustrating an example of a configuration of a recovery requirement application management table managed by a recovery requirement management unit of the management apparatus according to the first embodiment; [Figure 7] 11 is a flowchart of a countermeasure selection process when a failure occurs in the first embodiment. [Figure 8] FIG. 11 is a diagram showing an example of a screen display for outputting countermeasure candidates and priority information in the first embodiment. [Figure 9] 10A and 10B are diagrams illustrating examples of screen displays for setting information of various tables managed by the management device in the first embodiment. [Figure 10] 13 is a diagram illustrating a configuration example of a recovery requirement application management table managed by a recovery requirement management unit of the management apparatus according to the second embodiment. FIG. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of a countermeasure performance management table managed by a countermeasure evaluation value management unit of a management apparatus according to a third embodiment. [Figure 12A] FIG. 11 is a diagram showing an example of the configuration of an evaluation value management table managed by a countermeasure evaluation value management unit of a management apparatus according to a third embodiment. [Figure 12B] FIG. 13 is a diagram illustrating an example of a conversion function from a performance statistical value to an evaluation value in the third embodiment. [Figure 13] 13 is a flowchart of a countermeasure selection process when a failure occurs in the third embodiment. [Figure 14] FIG. 13 is a diagram illustrating a configuration example of a recovery requirement weight management table managed by a recovery requirement management unit of a management apparatus according to the fourth embodiment. [Figure 15] 13 is a flowchart of a countermeasure selection process when a failure occurs in the fourth embodiment. [Figure 16] FIG. 13 is a diagram illustrating a hardware configuration of an apparatus according to a fifth embodiment. [Figure 17] 13 is a flowchart of a countermeasure selection process when a failure occurs in the fifth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] First, an overview of the system according to the embodiment of the present invention will be described.

[0011] In a management device that determines recovery measures for a failure that has occurred, a countermeasure candidate selection unit manages countermeasure list information for each type of failure as management information, and a countermeasure evaluation value management unit manages evaluation values ​​for each countermeasure for any evaluation index such as the time required for the countermeasure and the cost, etc. Also, a recovery requirement management unit manages weighted values ​​obtained by weighting the priority for each evaluation index for each recovery requirement, and further manages recovery requirements applied to each managed device.

[0012] Then, when the occurrence of a failure is detected by receiving an alert indicating the occurrence of a failure, the countermeasure candidate selection unit first selects countermeasure candidates corresponding to the failure that has occurred. Next, the recovery requirement management unit refers to the recovery requirements applied to the managed device to be recovered, and further acquires weighted values ​​for each evaluation index defined in the recovery requirement. Next, the countermeasure priority determination unit of the management device determines the priority of each countermeasure candidate in accordance with the recovery requirement. Specifically, a total evaluation value reflecting the weighted value (each evaluation value multiplied by a weight corresponding to the weighted value) is calculated for each countermeasure candidate, and the total evaluation value is set as the countermeasure priority, and the priority order of each countermeasure candidate is determined in descending order of value. The higher the evaluation value of a countermeasure for the evaluation index to be prioritized, the higher the total evaluation value is calculated, so that a countermeasure in accordance with the recovery requirement can be selected with a high priority. Finally, the output unit outputs the countermeasure candidates together with the countermeasure priority (or priority order).

[0013] This countermeasure selection process automatically selects and outputs appropriate countermeasure candidates according to recovery requirements with any evaluation index to be prioritized. In addition, because prioritized countermeasure candidates are automatically output, even workers with little knowledge or experience can clearly determine which countermeasures should be prioritized, enabling appropriate failure recovery work.

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments described below do not limit the invention according to the claims, and all of the elements and combinations thereof described in the embodiments are not necessarily essential to the solution of the invention.

[0015] In the following description, information may be described using the expression "AAA table", but the information may be configured in any data structure. In other words, to indicate that the information does not depend on the data structure, the "AAA table" can be expressed as "AAA information".

[0016] A countermeasure selection method according to an embodiment of the present invention will be described below with reference to Figures 1 to 17. Example 1 will be described with reference to Figures 1 to 9, Example 2 with Figure 10, Example 3 with Figures 11 to 13, Example 4 with Figures 14 and 15, and Example 5 with reference to Figures 16 and 17.

[0017] <Example 1> In the first embodiment, a basic form of countermeasure selection processing for a occurring fault will be described. First, the system configuration and the configuration of each device will be described with reference to Fig. 1. Next, the tables, i.e., information, managed by the management device will be described with reference to Figs. 2 to 6. After that, the flow of the countermeasure selection processing will be described with reference to Fig. 7, and screen display examples related to output of countermeasure selection results and input of information will be described with reference to Figs. 8 and 9.

[0018] In the following description of the embodiments, when there is no need to separately describe the components, they will be described without a subscript (e.g., device 101), and when there is a need to separately describe the components, they will be described with a subscript (e.g., device 101-a).

[0019] First, the system configuration and the hardware configuration of each device in the first embodiment will be described with reference to FIG.

[0020] The system illustrated in FIG. 1 includes a plurality of devices 101 (101-a to 101-c) to be managed, and a management device 102. The devices 101 and the management device 102 are connected by wired or wireless communication, and the device 101 transmits data acquired from the installation site, operation information of the device itself, alerts, and the like to the management device 102. The management device 102 provides services using data received from the device 101, detects occurrence of a failure in the device 101 by receiving an alert, and selects and outputs measures for recovery. Note that FIG. 1 illustrates a configuration in which a system A consisting of devices 101-a and 101-b and a system B consisting of a device 101-c exist and are managed by one management device 102, but a configuration in which one management device 102 is provided for each system may also be used. The management device 102 may be installed at the same site as the device 101, or at a different location such as on the cloud.

[0021] Next, the hardware configuration of the device 101 will be described. As described above, the device 101 is a device that stores data acquired from the site and its own operation information in packets and transmits them, and has a communication function with the management device 102. The device 101 has various configurations depending on the type of data to be acquired. For example, the device 101 may be a temperature measuring device that measures the temperature at the site, a camera device that acquires images of the site, or an aggregation device that aggregates and transmits data from other devices present at the site. In FIG. 1, the device 101 has a communication I / F 111, a CPU 112, and a storage device 113.

[0022] For example, when transmitting / receiving packets to / from the management device 102 via wireless communication, the communication I / F 111 has a transmitting unit that converts digital data to / from a wireless signal and transmits the converted digital data to a wireless signal, and a receiving unit that extracts the digital data from the received wireless signal. When the device 101 transmits / receives packets not only to / from the management device 102 but also to / from other devices present at the site, the device 101 may be equipped with a plurality of communication I / Fs 111. These communication means may be any means, such as Ethernet (registered trademark), WiFi (registered trademark), a telephone network, or an optical fiber line.

[0023] The CPU 112 is a calculation device that executes various computer programs stored in the storage device 113, thereby realizing various functions of the device 101. Note that some of the processing performed by the CPU 112 executing the computer programs may be executed by another calculation device (for example, hardware such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC)).

[0024] The storage device 113 has a storage device constituted by, for example, a read-only semiconductor memory (ROM) and a storage element constituted by, for example, a rewritable semiconductor memory (RAM), and stores computer programs for implementing various processes, acquired data, etc. The storage device 113 may have a non-volatile storage medium (magnetic disk drive, non-volatile memory, etc.).

[0025] The application program 114 manages various settings such as data acquisition methods and transmission schedules, and executes acquisition processing by the CPU 112 connected via an internal bus, and transmission commands to the communication processing unit 115. In addition, the application program 114 may have a function of managing rules for generating an alert that notifies the occurrence of a fault in the device itself (e.g., generating an alert when the CPU utilization rate exceeds a threshold value), and executing a command to send an alert to the communication processing unit 115 when a fault occurs.

[0026] The communication processing unit 115 executes transmission and reception processing in communication. Specifically, it executes assembly processing of packets to be transmitted and packet analysis processing such as judging whether a received packet is addressed to the device itself. The device 101 may be an independent device or an embedded device. As described above, the device 101 can be configured in various ways according to the type of on-site data to be acquired, and may include, for example, a temperature sensor, a camera module, etc.

[0027] Next, a hardware configuration of the management device 102 will be described. The management device 102 has a communication I / F 121, a CPU 122, an input unit 123, an output unit 124, and a storage device 125. The communication I / F 121, the CPU 122, and the storage device 125 are the same in terms of hardware as the communication I / F 111, the CPU 112, and the storage device 113 of the device 101 described above. Also, the communication processing unit 127 is the same as the communication processing unit 115 of the device 101 described above.

[0028] The management device 102 is a computer system configured on one physical computer, or on multiple logically or physically configured computers, and may operate on a virtual computer constructed on multiple physical computer resources. For example, the application program 126, the communication processing unit 127, the device information management unit 128, the fault detection unit 129, the countermeasure candidate selection unit 130, the countermeasure evaluation value management unit 131, the recovery requirement management unit 132, and the countermeasure priority determination unit 133 may each operate on separate physical or logical computers, or may be combined to operate on a single physical or logical computer.

[0029] The input unit 123 is composed of, for example, a keyboard, a mouse, etc., and is used by an operator to input various operations and settings. The output unit 124 is composed of, for example, a liquid crystal display monitor, etc., and displays necessary screens and the results of various processes. However, in the case where input information is received from an external device via the communication I / F 121 and output information is provided to the external device, such as in a form in which another external device connected to the management device 102 remotely logs in to the management device 102, the input unit 123 and the output unit 124 do not need to be mounted on the management device 102.

[0030] The application program 126 is a program that provides a user with a service that utilizes collected data. For example, a program that provides an average value per unit time of on-site data (e.g., temperature) received from the device 101 executes data analysis processing such as calculating an average value from collected data values. In addition, the application program 126 may also include a program that remotely sets and remotely manages a data transmission schedule from the device 101.

[0031] The device information management unit 128 manages the system to which each device 101 belongs. In order to manage the system to which each device 101 belongs, the device information management unit 128 manages a device information management table 128a, which will be described later with reference to Fig. 2. However, when one management device 102 is provided for each system, the system to which the device 101 belongs is self-evident, so the device information management unit 128 can be omitted.

[0032] The fault detection unit 129 detects the occurrence of a fault in the device 101 based on information received from the device 101. The method of detecting a fault is not limited to a specific method, and any method may be used. For example, the fault may be detected by receiving an alert notifying the occurrence of a fault from the device 101, or by analyzing operation information received from the device 101.

[0033] The countermeasure candidate selection unit 130 selects countermeasure candidates for a failure detected by the failure detection unit 129. In order to select countermeasure candidates, the countermeasure candidate selection unit 130 manages a countermeasure candidate management table 130a, which will be described later with reference to FIG.

[0034] The countermeasure evaluation value management unit 131 manages the evaluation value by evaluation index for each countermeasure against a failure. The number and types of evaluation indexes are arbitrary, and for example, any evaluation index related to a predetermined recovery requirement, such as human cost and material cost incurred when executing the countermeasure, communication interruption time, and countermeasure required time, may be set. In order to manage the evaluation value of the countermeasure by the evaluation index, the countermeasure evaluation value management unit 131 manages an evaluation value management table 131a described later in FIG. 4.

[0035] The recovery requirement management unit 132 manages information related to recovery requirements, specifically, manages the priority of each evaluation index for each recovery requirement as a weighted value. In order to manage the weighted value information, the recovery requirement management unit 132 manages a recovery requirement weighted value management table 132a, which will be described later in Fig. 5. Furthermore, the recovery requirement management unit 132 manages a list of recovery requirements applied to each device 101 using a recovery requirement application management table 132b, which will be described later in Fig. 6.

[0036] The countermeasure priority determination unit 133 calculates an execution priority according to a predetermined recovery requirement for each countermeasure candidate selected by the countermeasure candidate selection unit 130. The priority is calculated using an evaluation value and a weighted value, and the details will be described later with reference to Fig. 7. After calculating the priority for each countermeasure candidate, the countermeasure priority determination unit 133 outputs the priority information together with the countermeasure candidate for the failure that has occurred via the output unit 124.

[0037] 2, a device information management table 128a managed by the device information management unit 128 of the management device 102 will be described. The device information management table 128a manages information indicating which system each device 101 belongs to. Fig. 2 illustrates the configuration of the device information management table 128a in the first embodiment.

[0038] Device ID 201 is an identifier of device 101. Specifically, device ID 201 is a field in which the address, host name, etc. of device 101 is entered, and the identifier of device 101 should conform to the method adopted by the system. When each device 101 is identified by an IP address, MAC address, or unique identifier, these identifiers may be entered. In the example of FIG. 2, the identifier of device 101 is indicated by the subscript in FIG. 1. System ID 202 is an identifier of the system to which device 101 entered in device ID 201 belongs. The identifier of the system should also conform to the method adopted by each system. In the example of FIG. 2, the identifier of the system is indicated by the system name shown in FIG. 1.

[0039] By referring to the device information management table 128a, the device information management unit 128 of the management device 102 can grasp the system to which each device 101 belongs, and in the example of FIG. 2, it can be determined that the device 101-a (device ID: a) and the device 101-b (device ID: b) belong to system A (system ID: System-A), and the device 101-c (device ID: c) belongs to system B (system ID: System-B). The device information management table 128a may be registered by an operator when constructing a system, or the system information stored in a packet transmitted by the device 101 may be acquired by a packet analysis process of the communication processing unit 127 and the information may be automatically registered. Also, as described above, when one management device 102 is provided for each system, the system to which the device 101 belongs is self-evident, so the device information management unit 128 can be omitted.

[0040] 3, a description will be given of a countermeasure candidate management table 130a managed by the countermeasure candidate selection unit 130 of the management device 102. The countermeasure candidate management table 130a manages a list of countermeasure candidates for each type of failure. FIG. 3 illustrates the configuration of the countermeasure candidate management table 130a in the first embodiment.

[0041] The fault ID 301 is an identifier for distinguishing the fault content. The naming rule for the fault identifier is arbitrary, and in the example of FIG. 3, a notation such as "Fault001" is adopted, but a specific fault content such as "radio interference occurrence" may be adopted as the identifier, for example, as shown in FIG. 3. The countermeasure ID 302 is an identifier for a countermeasure for recovering from the fault described in the fault ID 301. The naming rule for the countermeasure identifier is also arbitrary, and in the example of FIG. 3, a notation such as "Counter001" is adopted, but a specific countermeasure content such as "change of wireless channel" may be adopted as the identifier, for example, as shown in FIG. 3. In addition, as shown in FIG. 3, if there are multiple countermeasures for one fault, multiple countermeasure IDs 302 may be registered in the table for one fault ID 301.

[0042] By referring to the countermeasure candidate management table 130a, the countermeasure candidate selection unit 130 of the management device 120 can select a countermeasure candidate according to the fault that has occurred, and in the example of Fig. 3, when a fault caused by radio wave interference due to other devices is detected at the installation site of the device 101 that uses wireless communication, it can be determined that a countermeasure candidate such as "changing the wireless channel" or "changing the installation location" to avoid the interference, or "replacing the antenna" to transmit with a stronger output than the interfering radio wave should be selected. Note that the countermeasure candidate management table 130a may be used to register faults and countermeasures anticipated at the system construction stage, or may be used to register additional faults and countermeasures together with the corresponding countermeasures whenever a new fault is observed.

[0043] The evaluation value management table 131a managed by the countermeasure evaluation value management unit 131 of the management device 102 will be described with reference to Fig. 4. The evaluation value management table 131a manages evaluation values ​​for evaluation indexes related to recovery requirements for each countermeasure registered in the countermeasure candidate management table 130a. Fig. 4 shows an example of the configuration of the evaluation value management table 131a in the first embodiment.

[0044] The countermeasure ID 401 is an identifier of the countermeasure that is the same as the countermeasure ID 302 in the countermeasure candidate management table 130a in FIG. 3. When the countermeasures described in the countermeasure candidate management table 130a and the evaluation value management table 131a are the same, the identifiers described in the countermeasure ID are the same. The evaluation value 402 is an evaluation value for each evaluation index regarding the countermeasure described in the countermeasure ID 401. In FIG. 4, the evaluation indexes include "personnel cost" (e.g., dispatch cost of a worker), "physical cost" (e.g., cost of replacement parts), "communication interruption time" (interruption time of communication connecting the device 101 and the management device 102, or the device 101 and other devices), and "required time". However, as described above, the number, types, and names of the evaluation indexes to be managed are not limited, and may be arbitrarily defined according to the index to be considered as a recovery requirement. For example, an index such as "stability" may be set as in Patent Document 1.

[0045] In the example of FIG. 4, the evaluation value is defined as a range of "1 to 10" with a higher value indicating a higher evaluation, but the range of the evaluation value may be set arbitrarily. The method of setting the evaluation value is also arbitrary, and as described in Patent Document 1, the evaluation value may be calculated using machine learning based on past recovery information, or an original criterion may be set for each evaluation index to define the evaluation value. For example, in the case of "communication interruption time," no communication interruption may be assigned an evaluation value of 10 points, an interruption time of the order of several minutes may be assigned 7 points, an interruption time of the order of several hours may be assigned 4 points, and a day or more may be assigned 1 point.

[0046] By setting the evaluation value management table 131a, it is possible to manage which index each measure is superior in. For example, in FIG. 4, "wireless channel change" is superior in the index of "human cost" because it is not necessary to dispatch a worker to change the settings remotely, and it is accompanied by a temporary "communication interruption time" due to the restart of the device 101 accompanying the setting change. In addition, "installation location adjustment" is superior in the index of "communication interruption time" because it does not require the restart of the device 101 and does not cause communication interruption, but it is inferior to the measure "wireless channel change" in the index of "human cost" because it requires human labor to adjust the installation location and is accompanied by the dispatch of a worker or a request for work from a customer on-site. In this example, depending on whether "human cost" or "communication interruption time" is prioritized as a recovery requirement, which of the measures "wireless channel change" and "installation location adjustment" should be selected differs, and by referring to the evaluation value information, it is possible to compare the superiority or inferiority of measures according to the recovery requirements. The evaluation value management table 131a may be registered at the system construction stage, or may be registered or updated at any timing during the operation of the system.

[0047] The recovery requirement weight management table 132a managed by the recovery requirement management unit 132 of the management device 102 will be described with reference to Fig. 5. The recovery requirement weight management table 132a manages the priority of each evaluation index for each recovery requirement. Specifically, the priority of each evaluation index is defined as a weight for the evaluation value, and if the priority of the evaluation index is high, a high weight is set. Fig. 5 shows an example of the configuration of the recovery requirement weight management table 132a in the first embodiment.

[0048] The recovery requirement ID 501 is an identifier for distinguishing the recovery requirement. The naming rule for the identifier is arbitrary, and in the example of FIG. 5, a notation such as "Policy001" is adopted, but for example, a name indicating the policy of the recovery requirement, such as "cost priority" as shown in FIG. 5, may be registered. The weight value 502 expresses the priority of each evaluation index in the recovery requirement described in the recovery requirement ID 501 as a weight value for the evaluation value. In the example of FIG. 5, the weight value is defined in the range of "0 to 1", but the range of the weight value may be set arbitrarily.

[0049] For example, in the example of FIG. 5, in the recovery requirement "Policy002" that prioritizes the system availability, the weighting value for the evaluation index "communication outage time" is set higher than the other indexes, reflecting the recovery requirement that the top priority is given to suppressing system operation stoppage due to communication interruption. In the example of FIG. 5, the weighting value of the other evaluation indexes is set to "0.1", but if there is no need to consider other evaluation indexes at all, the weighting value may be set to "0". On the other hand, for example, when a penalty cost is paid to a customer according to the system operation downtime, it is desirable to consider the evaluation index "communication outage time" even in the recovery requirement "Policy001" that prioritizes cost. In that case, as illustrated in FIG. 5, by setting the weighting value of "communication outage time" a little higher, it is possible to set a recovery requirement that considers "communication outage time" while giving top priority to the indexes "personnel cost" and "physical cost" that are directly related to cost. By flexibly defining the detailed priority level that differs for each recovery requirement as a weighting value, any recovery requirement that matches the system operation policy and customer orientation can be expressed and managed in the recovery requirement weighting value management table 132a. The recovery requirement weight management table 132a may be registered at the system construction stage, or may be registered or changed at any timing during the operation of the system.

[0050] 6, a description will be given of the recovery requirement application management table 132b managed by the recovery requirement management unit 132 of the management device 102. The recovery requirement application management table 132b manages recovery requirements applied to each device 101. Fig. 6 shows an example of the configuration of the recovery requirement application management table 132b in the first embodiment.

[0051] The device ID 601 and the system ID 602 ​​are the identifier of the device 101 and the identifier of the system to which it belongs. As described in the explanation of the device information management table 128a in Fig. 2, the format of the identifier is arbitrary. Note that if the identifier of the device 101 is entered in the field of the device ID 601 and there is no duplication of the identifier of the device 101 between systems, the device 101 can be uniquely identified by the device ID 601, so the system ID 602 ​​may be left blank. Also, if one management device 102 is provided for each system, the system to which each device 101 belongs is self-evident, so the field of the system ID 602 ​​can be omitted.

[0052] The recovery requirement ID 603 is an identifier of a recovery requirement applied to the device 101 described in the device ID 601 and the system described in the system ID 602. The identifier described in the recovery requirement ID 603 is selected from the identifiers described in the recovery requirement ID 501 of the recovery requirement weight management table 132a in FIG. 5. When applying the same recovery requirement to all the devices 101 belonging to a certain system, as shown in FIG. 6, the system identifier may be specified in the system ID 602 ​​field and a wild card "*" may be specified in the device ID 601 column. This is particularly effective when many devices 101 belong to the same system and it is troublesome to enter the identifiers of all the devices 101 in the device ID 601. Of course, different recovery requirements may be applied to multiple devices 101 belonging to the same system as shown in FIG. 6.

[0053] The countermeasure selection process by the management device 102 will be described with reference to Fig. 7. As a general flow of the process, in the management device 102, when the fault detection unit 129 detects the occurrence of a fault in the device 101, the countermeasure candidate selection unit 130 selects countermeasure candidates for the occurred fault. Thereafter, the countermeasure priority determination unit 133 determines countermeasure priorities according to the recovery requirements for each countermeasure candidate, and the output unit 124 outputs the countermeasure candidates with the priority information attached. Fig. 7 is a flowchart of the countermeasure selection process in the first embodiment, and will be described in detail below.

[0054] In step S701 of Fig. 7, the fault detection unit 129 of the management device 102 detects the occurrence of a fault in the device 101. As described above, the fault detection method is arbitrary, such as detection based on the reception of an alert, and if the occurrence of a fault is detected (YES), the process proceeds to step S702. At this time, if the management device 102 is performing integrated management of multiple systems, the device information management table 128a of Fig. 2 managed by the device information management unit 128 is referenced, and system information to which the device 101 where the fault occurred belongs is obtained. If no fault is detected (NO), the process of step S701 is executed again after the lapse of an arbitrary time period.

[0055] In step S702, the countermeasure candidate selection unit 130 refers to the countermeasure candidate management table 130a in Fig. 3 and selects countermeasure candidates for the fault detected in step S701. For example, if the fault is "radio interference", three countermeasure candidates, "change wireless channel", "adjust installation location", and "replace antenna", are selected according to the countermeasure candidate management table 130a in Fig. 3. When the process of step S702 ends, the process proceeds to step S703.

[0056] In step S703, the recovery requirement management unit 132 refers to the recovery requirement application management table 132b in Fig. 6 and acquires a recovery requirement to be applied to the failure source device 101 detected in step S701. For example, if the failure source device 101 is device 101-a (device ID: a), the recovery requirement "Policy001 (cost priority)" is acquired in the example of Fig. 6. When the process of step S703 ends, the process proceeds to step S704.

[0057] In step S704, the recovery requirement management unit 132 refers to the recovery requirement weight management table 132a in Fig. 5 and acquires a weight value for each evaluation index in the applied recovery requirement acquired in step S703. In the above-mentioned example, since the applied recovery requirement is "Policy001", the weight values ​​of "Evaluation index 1 (human cost): 1, Evaluation index 2 (physical cost): 1, Evaluation index 3 (communication interruption time): 0.5, Evaluation index 4 (required time): 0.1" are acquired in the example of Fig. 5. When the process of step S704 ends, the process proceeds to step S705.

[0058] In step S705, the countermeasure priority determination unit 133 refers to the evaluation value management table 131a in Fig. 4 managed by the countermeasure evaluation value management unit 131, obtains the evaluation value for each countermeasure candidate selected in step S702, and calculates a total evaluation value reflecting the weight value obtained in step S704. Specifically, when there are N types of evaluation indexes, the evaluation value of the countermeasure related to evaluation index i (1≦i≦N) is xi, and the weight value in the applied restoration requirement is yi, the total evaluation value is defined by the following formula. Total evaluation value = (x1 × y1) + ... + (xi × yi) + ... + (xN × yN)

[0059] In the above example, the weighting value of the applied restoration requirement "Policy001" is "evaluation index 1:1, evaluation index 2:1, evaluation index 3:0.5, evaluation index 4:0.1", and the evaluation value of "Counter001 (wireless channel change)" in the example of FIG. 4 is "evaluation index 1:10, evaluation index 2:10, evaluation index 3:7, evaluation index 4:10", so the total evaluation value for "wireless channel change" is calculated as 24.5 (=10×1+10×1+7×0.5+10×0.1). Similarly, the total evaluation value of "Counter002 (installation location adjustment)" is calculated as 20.5 (=5×1+10×1+10×0.5+5×0.1), and the total evaluation value of "Counter003 (antenna replacement)" is calculated as 9.6 (=1×1+5×1+7×0.5+1×0.1). When the process of step S705 is completed, the process proceeds to step S706.

[0060] In step S706, the countermeasure priority determination unit 133 determines the countermeasure priority as the total evaluation value, and determines the priority order of each countermeasure candidate in descending order of the value. After determining the priority order, the countermeasure candidates selected in step S702 are output from the output unit 124 together with the determined priority order or countermeasure priority (total evaluation value). An example of the screen display related to the output will be described later with reference to FIG. 8.

[0061] In the above example, in the cost-prioritized restoration requirement "Policy001", among the three countermeasure candidates, "wireless channel change" with the highest total evaluation value is determined to have the first priority (countermeasure priority: 24.5), "installation location adjustment" with the next highest total evaluation value is determined to have the second priority (countermeasure priority: 20.5), and "antenna replacement" is determined to have the third priority (countermeasure priority: 9.6), and this result is output via the output unit 124. As a result, in this example, it is possible to determine that "wireless channel change" is the optimal countermeasure. Note that when "Policy002" in FIG. 5 is applied, the total evaluation value of "wireless channel change" is calculated to be 10, the total evaluation value of "installation location adjustment" is calculated to be 12, and the total evaluation value of "antenna replacement" is calculated to be 7.7, so that "installation location adjustment" can be determined to be the optimal countermeasure in the availability-prioritized restoration requirement "Policy002". In this way, by determining the countermeasure priority based on the weighted value and the evaluation value, countermeasures that conform to the applied restoration requirement can be flexibly selected. Furthermore, since the countermeasure candidates are output with a priority order, even an operator with little knowledge or experience can easily determine the countermeasure to be taken for the fault.When the process of step S706 ends, the countermeasure selection process of FIG.

[0062] An example of a screen display for outputting candidate countermeasures and priority information selected in the countermeasure selection process of Fig. 7 will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of a screen display output in step S706 of Fig. 7. The output unit 124 of the management device 102 outputs data for displaying a display screen 800 shown in Fig. 8. The display screen 800 includes a display area 801 for displaying fault information, and a display area 802 for displaying countermeasure information for the fault that has occurred.

[0063] Display area 801 displays fault information detected by fault detection unit 129 of management device 102. In the example of Fig. 8, information on device 101 from which the fault occurred, information on the system to which it belongs, and the fault details are displayed, and any other information, such as the time the fault was detected, may also be displayed as necessary.

[0064] Display area 802 displays countermeasure information for recovery from the failure information output to display area 801. In the example of Fig. 8, recovery requirement information applied to the failure source device 101 among the information managed by the recovery requirement management unit 132, and countermeasure candidates selected by the countermeasure candidate selection unit 130 with the priorities determined by the countermeasure priority determination unit 133 are displayed. Note that, although the priority of each countermeasure candidate is displayed in the example of Fig. 8, countermeasure priority (total evaluation value calculated in step S705 of Fig. 7) may also be displayed together in order to clearly indicate the degree of difference in priority between the priorities.

[0065] The operator can identify the details of the fault that has occurred and the device 101 and system that are the source of the fault by viewing the screen shown in Fig. 8. Furthermore, by referring to the prioritized countermeasure information, even an operator with no knowledge or experience can uniquely determine which countermeasure should be taken with priority. Note that the screen is not limited to the example shown in Fig. 8, and may be a screen that displays, for example, a predetermined number of candidate countermeasures with high priority (for example, one candidate countermeasure with the highest priority), and the display format of the fault information and countermeasure information is not limited to a specific method.

[0066] An example of a screen display for setting information in various tables managed by the management device 102 will be described with reference to Fig. 9. Fig. 9 shows an example of a setting screen. The output unit 124 of the management device 102 outputs data for displaying a display screen 900 shown in Fig. 9. The display screen 900 includes a display area 901 for setting an evaluation value for each measure, a display area 902 for setting a weighting value for each recovery requirement, and a display area 903 for setting a recovery requirement to be applied to each device.

[0067] The display area 901 is an area for setting the evaluation value management table 131a (FIG. 4) managed by the countermeasure evaluation value management unit 131 of the management device 102. When an evaluation value for each evaluation index is selected for each countermeasure, the input value is set in the evaluation value management table 131a. Note that measures for failure recovery and evaluation indexes to be considered in the recovery requirements can be added at any time by operating the add button 904. The add button 904 can be used to arbitrarily add new countermeasures and new evaluation indexes that arise in the course of system operation.

[0068] The display area 902 is an area for setting the recovery requirement weight management table 132a (FIG. 5) managed by the recovery requirement management unit 132 of the management device 102. When the weight value of each evaluation index is selected for each recovery requirement, the input value is set in the recovery requirement weight management table 132a. When an evaluation index is added using the add button 904 in the display area 901, the evaluation index column in the display area 902 is also automatically added. Also, the add recovery requirement button 904 is provided in the display area 902, making it possible to add recovery requirements during the operation of the system.

[0069] The display area 903 is an area for setting the recovery requirement application management table 132b (FIG. 6) managed by the recovery requirement management unit 132 of the management device 102. When the recovery requirement to be applied to each device 101 and each system is selected, the input value is set in the recovery requirement application management table 132b. Note that an add device and system button 904 is provided in the display area 903, making it possible to add a device 101 during the operation of the system.

[0070] Furthermore, it is also possible to manage the information of various tables not only on the screen of FIG. 9 but also by inputting and outputting external files. Therefore, the example screen display of FIG. 9 is provided with a file input button 905 and a file output button 906. Operating the file input button 905 loads various table values ​​saved in an external file onto the screen of FIG. 9. Operating the file output button 906 outputs the table values ​​input on the screen of FIG. 9 to an external file. This makes it easy to achieve linkage with external files in which table values ​​are stored. However, the inclusion of a file input / output function is optional.

[0071] As shown in Fig. 9, the setting screen for the information of various tables makes it easy to change the various table values ​​midway, and can flexibly handle the addition of the above-mentioned countermeasures, evaluation indexes, recovery requirements, and devices 101. Note that, in the screen display example of Fig. 9, display areas are provided for setting the evaluation value management table 131a of Fig. 4, the recovery requirement weight value management table 132a of Fig. 5, and the recovery requirement application table of Fig. 6, but similarly, display areas may be provided for setting the device information management table 128a of Fig. 2 and the countermeasure candidate management table 130a of Fig. 3. Also, in the screen display example of Fig. 9, a form in which the table values ​​are selected and set in a pull-down format is illustrated, but the set values ​​may be directly input, and various input formats can be adopted.

[0072] In this embodiment, the screens of Figures 8 and 9 are displayed and operated via the input unit 123 and output unit 124 of the management device 102. However, as mentioned above, for example, by remotely logging in to the management device 102, an external device may receive information via the communication I / F 121, output display data for displaying the screens of Figures 8 and 9 on the external device via the communication I / F 121, and display the screens on the display of the external device. In this case, the input unit 123 and the output unit 124 may be omitted.

[0073] As described above, according to this embodiment, when a failure occurs in the device 101, candidate countermeasures that conform to the recovery requirements can be automatically selected for the failure that has occurred, and candidate countermeasures with countermeasure priorities that reflect any recovery requirement can be output. This makes it possible to take appropriate countermeasures that conform to the recovery requirements, thereby improving the system operation rate and reducing the countermeasure costs required for recovery. Furthermore, since priority information is clearly indicated for each candidate countermeasure, even an operator with little knowledge or experience can easily determine the appropriate countermeasure.

[0074] <Example 2> In the first embodiment, a countermeasure selection process in which a single recovery requirement is applied to each device 101 has been described. On the other hand, there may be cases in which different recovery requirements should be applied as appropriate depending on the situation. For example, there may be cases in which a recovery requirement of "cost priority" is applied during the system construction period, and "availability priority" is applied after the system starts operating. In this example, it is desirable that the recovery requirement applied to the device 101 is automatically changed from "cost priority" to "availability priority" at the time the system starts operating. Therefore, in the second embodiment, assuming a change in recovery requirements according to the operation phase of the system, etc., a form in which the applied recovery requirement is changed with a change in time or an excess of a threshold value of a specific parameter as a trigger will be described.

[0075] In the second embodiment, the configuration of a recovery requirement application management table 132b managed by the recovery requirement management unit 132 of the management device 102 will be described with reference to Fig. 10. Note that various configurations and processes according to the second embodiment are the same as those of the first embodiment except for the configuration of the recovery requirement application management table 132b shown in Fig. 10, and therefore descriptions thereof will be omitted.

[0076] Fig. 10 shows an example of the configuration of the recovery requirement application management table 132b in the embodiment 2. The fields of the device ID 601, the system ID 602, and the recovery requirement ID 603 are the same as the fields of the recovery requirement application management table 132b in the embodiment 1 in Fig. 6, and the description of these will be omitted. The recovery requirement application management table 132b in the embodiment 2 has fields of a judgment parameter 1003, a judgment condition 1004, and a threshold value 1005, and when the conditions described in the field groups are satisfied, the recovery requirement described in the recovery requirement ID 1006 is applied.

[0077] The judgment parameter 1003 is information on a parameter to be judged for the recovery requirement to be applied. In the example of Fig. 10, the current time and the packet error rate are exemplified, but the parameter type and name to be written in the field are arbitrary. For example, a parameter such as a CPU usage rate may be used as the operation information of the device 101.

[0078] The judgment condition 1004 is a magnitude relationship with respect to a threshold value 1005. In addition to the magnitude relationships illustrated in Fig. 10, conditions such as "= (match)" and "≠ (mismatch)" can be set. The notation format in this field is arbitrary, and it may be written with an inequality sign or a character string such as "less than."

[0079] The threshold value 1005 is the threshold value of the judgment parameter 1003. In the example of FIG. 10, if the time is defined in the judgment parameter 1003, the threshold value may be defined in any format according to the parameter type, such as by entering time information in the field. In the example of FIG. 10, for example, if the current time is before "2020 / 10 / 01 00:00:00", the recovery requirement of "Policy001 (priority on cost)" is applied to the device 101-a (device ID: a), and after the time described in the threshold value 1005 is exceeded, the recovery requirement of "Policy002 (priority on availability)" is applied. Therefore, for example, if you want to change the recovery requirement to be applied before and after the system starts operating as described above, you can set the operation start time in the threshold value 1005 field. On the other hand, for example, the recovery requirement "Policy003 (balanced priority)" is applied to the device 101-b (device ID: b) while the packet error rate in communication with the management device 102 is within 50%, and "Policy004 (time priority)" is applied when the packet error rate exceeds 50%. In this way, when the packet error rate increases and a failure occurs at a level that interferes with data collection, the recovery requirement is changed to "time priority" to aim for early recovery. In this way, by setting a parameter related to the performance index in the judgment parameter 1003, it is possible to switch the recovery requirement according to the system quality. Of course, if it is desired to apply one recovery requirement at all times as in the first embodiment, the judgment parameter 1003, the judgment condition 1004, and the threshold value 1005 may be left undefined, as illustrated in FIG. 10.

[0080] The judgment parameters 1003 defined in the recovery requirement application management table 132b in the second embodiment may define the personnel required for recovery in addition to the time-related judgment conditions shown in FIG. 10 and the judgment conditions of the operating state of the device.

[0081] Incidentally, although other various configurations and processes are similar to those of the first embodiment, in step S703 of the countermeasure selection process of FIG. 7 in the second embodiment, when obtaining the recovery requirements to be applied to the failure source by referring to the recovery requirement application management table 132b of FIG. 10, the recovery requirements to be applied are not determined only from the device ID 1001 and the system ID 1002, but the application recovery requirements are determined in consideration of the judgment parameters 1003, the judgment conditions 1004, and the threshold value 1005.

[0082] As described above, according to this embodiment, the recovery requirements to be applied can be dynamically changed according to the operation phase and quality of the system, and appropriate measures can be selected in accordance with the changing recovery requirements. Of course, in the first embodiment, the operator can monitor the time and changes in specific parameters and change the applied recovery requirements using the screen of Fig. 9, but if the judgment conditions are registered in advance in the recovery requirement application management table 132b as in this embodiment, the applied recovery requirements can be automatically switched and the measures selection process can be executed without human intervention.

[0083] <Example 3> In the third embodiment, a form in which an evaluation value of a measure is set and updated based on the actual value of each evaluation index when the measure is actually executed will be described. When taking measures against an occurred failure, the actual values ​​of the cost and the time required may differ from the assumptions made when the evaluation value was established. In selecting measures that meet the recovery requirements, it is desirable to set an evaluation value that reflects the actual situation, and in the third embodiment, a form in which the actual values ​​of each evaluation index that occurs when the measure is executed are recorded and the evaluation value is determined based on the actual information will be described.

[0084] In the third embodiment, in addition to the evaluation value management table 131a, the countermeasure evaluation value management unit 131 of the management device 102 in Fig. 1 manages a countermeasure result management table 131b that manages the result value for each evaluation index generated when a countermeasure is executed. Of the tables managed by the countermeasure evaluation value management unit 131 of the management device 102 in the third embodiment, Fig. 11 shows the countermeasure result management table 131b, and Fig. 12 shows the evaluation value management table 131a. Also, the countermeasure selection process including the update of the evaluation value based on the result information will be described with reference to Fig. 13. Note that various configurations and processes according to the third embodiment are the same as those of the first or second embodiment except for the configurations and processes shown in Figs. 11 to 13, and therefore the description thereof will be omitted.

[0085] The countermeasure result management table 131b managed by the countermeasure evaluation value management unit 131 of the management device 102 will be described with reference to Fig. 11. When taking countermeasures against a occurred failure, an operator records the recovery result and the result value for each evaluation index in the countermeasure result management table 131b. Fig. 11 shows an example of the configuration of the countermeasure result management table 131b in the third embodiment.

[0086] The countermeasure ID 1101 is an identifier of a countermeasure implemented for the occurred failure. The format of the countermeasure identifier may be the same as the countermeasure ID 302 in the countermeasure candidate management table 130a in FIG. 3 or the countermeasure ID 401 in the evaluation value management table 131a in FIG.

[0087] The execution date and time 1102 is the date and time when the measure described in the measure ID 1101 was implemented. The date and time may be expressed in any format.

[0088] The recovery result 1103 is a field in which whether or not the countermeasure described in the countermeasure ID 1101 has actually been performed and the result has been recovered from the failure is recorded. In the example of Fig. 11, the recovery result 1103 is recorded as "o" or "x", but the format is arbitrary. The fields of execution date and time 1102 and recovery result 1103 are not essential. However, by providing the recovery result 1103 field and recording the recovery result, for example, if the recovery probability by each countermeasure is included in the evaluation index, the recovery probability can be calculated using the recovery result 1103 field.

[0089] The performance value 1104 is a field in which the performance value that actually occurs as a result of taking measures is recorded for each evaluation index related to the recovery requirement. In the example of Fig. 11, when recording the performance value in the performance value 1104 field, the unit information at the time of entry may be arbitrarily set for each evaluation index. The performance value 1104 may be a value that is aggregated for each device or system.

[0090] Each time a countermeasure is actually taken in the real environment, the recovery result and the performance value are recorded and accumulated in the countermeasure performance management table 131b in Fig. 11, so that an evaluation value that more reflects the actual situation can be set in the evaluation value management table 131a in Fig. 12, which will be described later. Note that the countermeasure performance management table 131b in Fig. 11 may record any information, such as the identifier of the device 101 that is the target of the countermeasure, or the identifier of the system to which the device 101 belongs.

[0091] 12A and 12B, an evaluation value management table 131a managed by the countermeasure evaluation value management unit 131 of the management device 102 will be described. The evaluation value management table 131a of the third embodiment manages performance statistics (performance average in the example of FIG. 12A) for each evaluation index for each countermeasure based on performance information recorded in the above-mentioned countermeasure performance management table 131b, and sets an evaluation value based on the performance statistics. Fig. 12A shows the configuration of the evaluation value management table 131a in the third embodiment, and Fig. 12B shows an example of a conversion function from performance statistics to an evaluation value.

[0092] First, the measure ID 401 in Fig. 12A is an identifier of the measure, and is the same as the measure ID 401 in the evaluation value management table 131a in the first embodiment in Fig. 4. The format of the measure identifier may be the same as the measure ID 302 in the measure candidate management table 130a in Fig. 3 and the measure ID 401 in the evaluation value management table 131a in Fig. 4.

[0093] The performance statistics 1202 are performance statistics for each evaluation index related to the countermeasure described in the countermeasure ID 1201. Specifically, the statistics are calculated based on the performance values ​​recorded in the countermeasure performance management table 131b in Fig. 11, and in the example of Fig. 12A, an average value for each evaluation index is calculated and recorded as the performance statistics. However, the information recorded in the field of the performance statistics 1202 is not limited to the average value, and statistical information suited to the purpose, such as a maximum value or a minimum value, may be used.

[0094] The evaluation value 1203 is an evaluation value for each evaluation index related to the countermeasure described in the countermeasure ID 1201. However, the evaluation value 1203 in the third embodiment is set based on the performance value described in the performance statistics 1202. In the example of FIG. 12A, the best performance statistics is the highest evaluation value (10 points), and the worst performance statistics is the lowest evaluation value (1 point), and the performance statistics are normalized in the range of "1 to 10 points" for each evaluation index and converted into an evaluation value. The evaluation values ​​described in FIG. 12A show an example of linear normalization (conversion) as shown by the solid line in FIG. 12B, but if it is desired to give a large inclination to the evaluation value for a performance value exceeding a certain threshold, a conversion function as shown by the dashed line in FIG. 12B may be adopted. Also, in the example of FIG. 12A, the best and worst performance statistics are calculated for each evaluation index and converted into an evaluation value, but if the evaluation axis (cost, etc.) is the same, such as "evaluation index 1 (personnel cost)" and "evaluation index 2 (physical cost)", the best and worst performance statistics may be set by combining multiple evaluation indexes. 12A, when the best and worst performance statistical values ​​are calculated using only evaluation index 1, the best is 967 yen and the worst is 37,000 yen, but when evaluation index 2 is combined, the best is 0 yen and the worst is 37,000 yen, and the performance values ​​of evaluation index 1 and evaluation index 2 may be converted to each evaluation value in the range of "0 yen to 37,000 yen." In this way, the conversion function from the performance statistical value to the evaluation value may be set arbitrarily.

[0095] A process of selecting a countermeasure when a fault occurs in the third embodiment will be described with reference to FIG. 13. In the third embodiment, a countermeasure candidate for the fault that occurred is selected, and priority information is ofThe method includes a procedure of recording the performance value that actually occurred after the countermeasure was implemented, and then recalculating and updating the evaluation value based on the performance value information. Fig. 13 is a flowchart of the countermeasure selection process, and the details will be described below. However, since steps S701 to S706 in Fig. 13 are the same as those in the first or second embodiment in Fig. 7, their explanation will be omitted, and only the processing in steps S1307 and S1308 after the output of countermeasure candidates has been completed will be described.

[0096] In step S1307, after an operator or the like has taken measures, the recovery results and the performance values ​​for each evaluation index are registered together with information about the measures taken in the measures performance management table 131b in Fig. 11 managed by the measures evaluation value management unit 131. The registration in step S1307 may be performed manually by the operator, or, for example, the management device 120 may actually measure the communication interruption time that occurs when the measures are taken and automatically register the actual measurement value in the measures performance management table 131b. When the process of step S1307 ends, the process proceeds to step S1308.

[0097] In step S1308, the performance values ​​and evaluation values ​​written in evaluation value management table 131a (FIG. 12) of countermeasure evaluation value management unit 131 are recalculated and updated. Specifically, by referring to countermeasure performance management table 131b updated in step S1307, the field of performance statistical value 1202 in evaluation value management table 131a in FIG. 12 is updated, and the evaluation value for each evaluation index of evaluation value 1203 is recalculated and updated based on the new performance statistical value. When the processing of step S1308 ends, the countermeasure selection processing of FIG. 13 ends.

[0098] The countermeasure selection process in Fig. 13 allows the evaluation value to be set based on the performance value at the time of countermeasure execution, making it possible to select countermeasures that better reflect the characteristics of the actual environment. In particular, when the performance value changes, for example, the time required for countermeasure execution gradually shortens due to environmental changes or an operator's improvement in countermeasure proficiency, the evaluation value can be set each time based on the performance value as in steps S1307 to S1308 in Fig. 13, so that an evaluation value that follows the fluctuation of the performance value can be set. Note that, in the example in Fig. 13, the performance value is registered and the evaluation value is updated in steps S1307 to S1308 every time after a countermeasure is executed, but the performance value may be registered and the evaluation value updated periodically, for example, once a week, at a timing that is not linked to the timing of countermeasure execution.

[0099] As described above, the countermeasure selection process according to the present embodiment sets evaluation values ​​that reflect actual results, and countermeasures that are more in line with actual characteristics can be selected. As described above, even if the actual results for each countermeasure fluctuate over time due to environmental changes, etc., the evaluation values ​​are updated to follow the fluctuations, so that countermeasures that always reflect the actual environmental characteristics can be selected.

[0100] <Example 4> In the first to third embodiments, the countermeasure priority determination unit 133 unconditionally determines countermeasure candidates with high total evaluation values ​​as high priority. However, as examples of recovery requirements, constraint conditions may be set such as "countermeasures whose total of human cost and physical cost exceeds 50,000 yen are excluded from candidates" or "countermeasures that involve communication interruption for one hour or more are excluded from candidates." Therefore, in the fourth embodiment, it is possible to set constraint conditions that countermeasure candidates should satisfy for each recovery requirement.

[0101] In the fourth embodiment, Fig. 14 is a diagram showing a recovery requirement weight management table 132a managed by the recovery requirement management unit 132 of the management device 102, and Fig. 15 is a flowchart of a countermeasure selection process when a failure occurs. Note that, since various configurations and processes related to the fourth embodiment are the same as those of the first to third embodiments except for the configurations and processes shown in Figs. 14 and 15, the description thereof will be omitted.

[0102] FIG. 14 shows an example of the configuration of the recovery requirement weight management table 132a in the fourth embodiment. In this embodiment, as in the first to third embodiments, the recovery requirement weight management table 132a manages the weight for each recovery requirement, and furthermore, in order to manage the constraint conditions that the candidate measures should satisfy, fields of a countermeasure feasibility determination index 1402 and a minimum evaluation value condition 1403 are newly provided. Countermeasures that satisfy the conditions described in the field group for the applied recovery requirement become the candidate measures, and even if the total evaluation value is excellent, if the condition is not met, the countermeasure is excluded from the candidate measures. The fields of the recovery requirement ID 501 and the weight value 502 are the same as the respective fields of the recovery requirement weight management table 132a in the first embodiment in FIG. 5, and the description of these will be omitted. The newly provided countermeasure feasibility determination index 1402 and the minimum evaluation value condition 1403 will be described below.

[0103] The countermeasure feasibility determination index 1402 is an evaluation index that is a criterion for determining the feasibility of a countermeasure candidate in the recovery requirement described in the recovery requirement ID 501. In the example of Fig. 14, it is described in a format such as "index 1", but the format is arbitrary as long as it indicates the type of evaluation index related to the recovery requirement. In addition, when determining the feasibility of a countermeasure candidate based on a combination of multiple evaluation indexes such as the above-mentioned "total value of human cost and physical cost", it is possible to define a combination of evaluation indexes such as "index 1 + index 2" as shown in Fig. 14.

[0104] The minimum evaluation value condition 1403 is an evaluation value that should be satisfied for one or more evaluation indexes described in the countermeasure feasibility judgment index 1402 in order to be a candidate countermeasure. For example, in the example of FIG. 14, the recovery requirement of "Policy001 (cost priority)" indicates that if the combined evaluation value of evaluation index 1 (personnel cost) and evaluation index 2 (physical cost) is not 13 or more, the recovery requirement is excluded from the candidate countermeasure. Note that, in the example of FIG. 14, the minimum evaluation value condition 1403 defines the threshold value regarding the feasibility of the countermeasure candidate as an evaluation value, but in cases where the actual value can be calculated backward from the evaluation value as in the third embodiment, a value regarding a specific evaluation index (such as money, time, etc.) may be defined. By managing the recovery requirement weight value management table 132a as shown in FIG. 14, any constraint condition that should be satisfied by each countermeasure can be set for each recovery requirement.

[0105] With reference to Fig. 15, the countermeasure selection process at the time of failure occurrence in the fourth embodiment will be described. In the fourth embodiment, after selecting countermeasure candidates for the occurred failure, when calculating the total evaluation value for each countermeasure candidate in the countermeasure priority determination unit 133, the recovery requirement weighted value management table 132a in Fig. 14 is referred to and countermeasures that do not satisfy the constraint conditions set in the applied recovery requirements are excluded from the candidates. Fig. 15 is a flowchart of the countermeasure selection process in the fourth embodiment, and the details will be described below. However, since steps S701 to S704 and step S706 in Fig. 15 are the same as steps S701 to S704 and step S706 in the first and second embodiments in Fig. 7, the explanation will be omitted and the process of excluding countermeasures that do not satisfy the constraint conditions in the newly provided step S1505 will be described.

[0106] In step S1505, a total evaluation value according to the recovery requirement is calculated for each countermeasure candidate, which is the same as step S705 in Fig. 7. However, at this time, the countermeasure priority determination unit 133 refers to the recovery requirement weight management table 132a in Fig. 14 managed by the recovery requirement management unit 132, and excludes countermeasures that do not satisfy the constraint conditions set in the recovery requirement to be applied from the candidates. For example, if the occurring fault is "Fault001 (occurrence of radio interference)" in Fig. 3, and three countermeasure candidates, "Counter001 to Counter003," are selected for this countermeasure, and the recovery requirement to be applied is "Policy001 (cost priority)" in Fig. 14, countermeasures with a combined evaluation value of evaluation index 1 (personnel cost) and evaluation index 2 (physical cost) of less than 13 are excluded from the candidates. 4, the countermeasure "Counter003 (antenna replacement)" is excluded from the countermeasure candidates because the combined evaluation value of evaluation index 1 and evaluation index 2 is 6, and the two countermeasure candidates "Counter001" and "Counter002" are selected. When the process of step S1505 ends, the process proceeds to step S1506, where countermeasure priorities are calculated for the remaining countermeasure candidates after excluding countermeasures that do not satisfy the constraint conditions, and output processing is executed.

[0107] Candidate measures that satisfy the constraint conditions set for each recovery requirement can be automatically output by the measure selection process in Fig. 15. For the sake of simplicity, the process of excluding measures that do not satisfy the constraint conditions has been described in the measure selection process in Fig. 7 in the first and second embodiments, but this process may be applied to the third embodiment by replacing step S1305 in Fig. 13 with step S1505 in Fig. 15.

[0108] As described above, by setting the constraint conditions according to this embodiment, measures that satisfy any constraint conditions set for each recovery requirement can be selected and output. In particular, in cases where communication interruption for a certain period of time or more would cause severe damage to the system, by setting the constraint conditions appropriately, only safe measures can be extracted and selected in terms of system operation.

[0109] <Example 5> In the first to fourth embodiments, the management device 102 determines a countermeasure against a failure occurring in the device 101, but the device 101 may determine the countermeasure by itself. The advantage of executing the countermeasure selection process in the management device 102 is that the processing load on the device 101 is reduced, while the device 101 determines the countermeasure by itself has the advantage that the countermeasure can be determined in a shorter time than if the management device 102 determines the countermeasure. Therefore, in the fifth embodiment, a description will be given of a configuration in which the device 101 determines a countermeasure against a failure occurring in the device itself.

[0110] In the fifth embodiment, Fig. 16 shows a hardware configuration of the device 101, and Fig. 17 is a flowchart of a countermeasure selection process by the device 101. Note that various configurations and processes of the fifth embodiment are the same as those of the first to fourth embodiments except for the configurations and processes shown in Figs. 16 and 17, and therefore the description thereof will be omitted.

[0111] The hardware configuration of the device 101 in the fifth embodiment will be described with reference to Fig. 16. The communication I / F 1601, CPU 1602, storage device 1605, application program 1606, and communication processing unit 1607 in Fig. 16 are the same as the communication I / F 111, CPU 112, storage device 113, application program 114, and communication processing unit 115 in Fig. 1 mounted on the device 101 in the first embodiment, respectively, except for the contents of the processing executed by the CPU 1602, and therefore description thereof will be omitted. Note that if the management device 102 does not exist and there is no need to transmit data acquired by the device 101 or operation information of the device itself, the communication I / F 1601 may not be mounted.

[0112] In the fifth embodiment, the device 101 detects a failure of its own device and executes processes from selecting countermeasure candidates to determining priority and outputting, so the device 101 of the fifth embodiment has an input unit 1603, an output unit 1604, a failure detection unit 1609, a countermeasure candidate selection unit 1610, a countermeasure evaluation value management unit 1611, a recovery requirement management unit 1612, and a countermeasure priority determination unit 1613, which are the same as the input unit 123, the output unit 124, the failure detection unit 129, the countermeasure candidate selection unit 130, the countermeasure evaluation value management unit 131, the recovery requirement management unit 132, and the countermeasure priority determination unit 133 of Fig. 1 mounted on the management device 102 of the first embodiment. The device information management unit 128 of Fig. 1 mounted on the management device 102 of the first embodiment is not mounted on the device 101 of the fifth embodiment because it manages a system to which other devices belong. Also, the failure detection unit 1609 of the device 101 differs from the failure detection unit 129 of the management device 102 in that it does not detect failures of other devices, but detects a failure that occurs in the device itself by an arbitrary method based on alert information generated by the device itself. The recovery requirement application management table 132b (FIG. 6 or FIG. 10) managed by the recovery requirement management unit 1612 of the device 101 may be in a form that manages only the recovery requirements that are applied to the device itself. Note that, in a case where input information is received from an external device or output information is provided to an external device via the communication I / F 1601, such as a form of remotely logging in to the device 101 from another external device, the input unit 1603 and the output unit 1604 do not need to be mounted on the device 101.

[0113] A countermeasure selection process by the device 101 when a failure occurs will be described with reference to Fig. 17. Fig. 17 is a flowchart of the countermeasure selection process of the fifth embodiment. The countermeasure selection process of the fifth embodiment is the same as the countermeasure selection process of Fig. 1, except that in step S1701, a failure in the device itself is detected, not a failure in another device.

[0114] In step S1701, the failure detection unit 1608 of the device 101 detects the occurrence of a failure in the device itself. If the occurrence of a failure is detected (YES), the process proceeds to step S1702. If no failure is detected (NO), the process of step S1701 is executed again after an arbitrary time has elapsed.

[0115] In step S1702, the countermeasure candidate selection unit 1610 refers to the countermeasure candidate management table 130a in Fig. 3 and selects a countermeasure candidate for the failure detected in step S1701. When the process of step S1702 ends, the process proceeds to step S1703.

[0116] In step S1703, the recovery requirement management unit 1612 acquires the recovery requirement applied to the own device by referring to the recovery requirement application management table 132b in Fig. 6 or Fig. 10. When the process of step S1703 ends, the process proceeds to step S1704.

[0117] In step S1704, the recovery requirement management unit 1612 refers to the recovery requirement weight management table 132a in Fig. 5 or 14, and acquires a weight for each evaluation index in the applied recovery requirement acquired in step S1703. When the process of step S1704 ends, the process proceeds to step S1705.

[0118] In step S1705, the countermeasure priority determination unit 1613 refers to the evaluation value management table 131a in Fig. 4 or 12A managed by the countermeasure evaluation value management unit 1611, acquires the evaluation value for each countermeasure candidate selected in step S1702, and calculates a total evaluation value reflecting the weight value acquired in step S1704. When the process of step S1705 ends, the process proceeds to step S1706.

[0119] In step S1706, the countermeasure priority determination unit 1613 determines the countermeasure priority as the total evaluation value, and determines the priority order of each countermeasure candidate in descending order of the value. After determining the priority order, the countermeasure candidates selected in step S1702 are output from the output unit 1604 together with the determined priority order or countermeasure priority (total evaluation value). When this process ends, the flowchart related to the countermeasure selection process in FIG. 17 ends.

[0120] The countermeasure selection process of Fig. 17 enables the device 101 to automatically select and output countermeasure candidates together with priority information for a failure that has occurred in the device itself. Note that, for the sake of simplicity, the countermeasure selection process by the device 101 has been described based on the countermeasure selection process of Fig. 7 in the first and second embodiments, but the process of the third or fourth embodiment may be applied to the countermeasure selection process of the device 101 by replacing step S701 in Fig. 13 or step S701 in Fig. 15 with step S1701 in Fig. 17.

[0121] As described above, this embodiment allows the device 101 to determine a countermeasure for a failure that occurs within the device itself, without the need for the management device 102. For example, since the device 101 can detect a failure within itself earlier than the management device 102 can detect a failure, as described above, the countermeasure can be determined in a shorter time from the occurrence of a failure than if the management device 102 were to determine the countermeasure.

[0122] The present invention is not limited to the above-described embodiments, and includes various modified examples and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the configurations described. Furthermore, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, a part of the configuration of each embodiment may be added, deleted, or replaced with another configuration.

[0123] In addition, each of the above-mentioned configurations, functions, processing units, processing means, etc. may be realized in hardware, for example by designing some or all of them as an integrated circuit, or may be realized in software by a processor interpreting and executing a program that realizes each function.

[0124] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC card, an SD card, or a DVD.

[0125] In addition, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In reality, it can be considered that almost all components are connected to each other. [Explanation of symbols]

[0126] 101: device, 102: management device, 111: communication I / F, 112: CPU, 113: storage device, 114: application program, 115: communication processing unit, 121: communication I / F, 122: CPU, 123: input unit, 124: output unit, 125: storage device, 126: application program, 127: communication processing unit, 128: device information management unit, 129: failure detection unit, 130: countermeasure candidate selection unit, 131: countermeasure evaluation value management unit, 132: recovery requirement management unit, 133: countermeasure priority determination unit

Claims

1. A countermeasure selection device that determines a recovery countermeasure for a failure that occurs in its own device or another device, The computer includes a processor that executes a program and a memory device that stores the program and data. Countermeasure candidate management information for managing a list of recovery measure candidates for the contents of the failure; a countermeasure candidate selection unit that refers to the countermeasure candidate management information and selects a recovery countermeasure candidate for the failure that has occurred; a countermeasure evaluation value management unit that manages evaluation values ​​for each countermeasure with respect to one or more evaluation indexes that are important matters in a recovery countermeasure; a recovery requirement management unit that manages recovery requirements including weighted values ​​indicating the priority of each of the evaluation indexes and devices to which the recovery requirements are applied; a countermeasure priority determination unit that determines priorities of the candidate recovery measures based on one or more of the evaluation indexes; an input unit that receives an input of an evaluation value for each of the measures and the evaluation index, a weighted value for each of the restoration requirements and the evaluation index, and an apparatus to which the restoration requirements are applied; an output unit that outputs the recovery measure candidates with the determined priorities; The input unit includes: an area in which an evaluation value to be managed by the measure evaluation value management unit can be selected for each measure and evaluation index; an area in which a weighting value managed by the restoration requirement management unit can be set for each restoration requirement and evaluation index; a region in which a recovery requirement managed by the recovery requirement management unit can be set for each of the own device or the other device; the countermeasure priority determination unit calculates, for each candidate recovery measure selected by the candidate recovery measure selection unit, a total value of the evaluation values ​​reflecting the weighted values ​​in the recovery requirements applied to the device to be restored, and determines the priority of the candidate recovery measure so that the higher the calculated total value, the higher the priority of the candidate recovery measure.

2. The countermeasure selection device according to claim 1, The countermeasure selection device according to claim 1, wherein the evaluation index includes an index prioritizing the cost required for failure recovery or an index prioritizing time.

3. The countermeasure selection device according to claim 1, The countermeasure selection device, wherein the restoration requirement management unit manages restoration requirements that are switched in response to predetermined conditions and applied to the device.

4. The countermeasure selection device according to claim 3, The countermeasure selection device, wherein the predetermined condition includes a time condition or an operating state condition of the device.

5. The countermeasure selection device according to claim 1, The measure evaluation value management unit Record the recovery results and the performance values ​​of the evaluation indicators when the measures are implemented, Calculating an average value of the performance values ​​for each of the measures and each of the evaluation indexes; a countermeasure selection device which calculates and registers an evaluation value for each countermeasure based on the average value.

6. The countermeasure selection device according to claim 1, the restoration requirement management unit manages a lower limit evaluation value that the evaluation index should satisfy for each of the restoration requirements; The countermeasure selection device, wherein the countermeasure priority determination unit excludes, from candidates, countermeasures whose evaluation index does not satisfy the lower limit evaluation value.

7. A device constituting the system; A system including a management device that determines a recovery measure for a failure that occurs in the device, The management device includes: The computer includes a processor that executes a program and a memory device that stores the program and data. Countermeasure candidate management information for managing a list of recovery measure candidates for the contents of the failure; a countermeasure candidate selection unit that refers to the countermeasure candidate management information and selects a recovery countermeasure candidate for the failure that has occurred; a countermeasure evaluation value management unit that manages evaluation values ​​for each countermeasure with respect to one or more evaluation indexes that are important matters in a recovery countermeasure; a recovery requirement management unit that manages recovery requirements including weighted values ​​indicating the priority of each of the evaluation indexes and devices to which the recovery requirements are applied; a countermeasure priority determination unit that determines priorities of the candidate recovery measures based on one or more of the evaluation indexes; an input unit that receives an input of an evaluation value for each of the measures and the evaluation index, a weighted value for each of the restoration requirements and the evaluation index, and an apparatus to which the restoration requirements are applied; an output unit that outputs the recovery measure candidates with the determined priorities assigned thereto; The input unit includes: an area in which an evaluation value to be managed by the measure evaluation value management unit can be selected for each measure and evaluation index; an area in which a weighting value managed by the restoration requirement management unit can be set for each restoration requirement and evaluation index; a region in which a recovery requirement managed by the recovery requirement management unit can be set for each of the devices; The system is characterized in that the countermeasure priority determination unit calculates a sum of the evaluation values ​​reflecting the weighted values ​​in the recovery requirements applied to the device to be restored, for each candidate recovery measure selected by the candidate countermeasure selection unit, and determines a priority such that the higher the calculated sum, the higher the priority of the candidate recovery measure.

8. A countermeasure selection method in which a countermeasure selection device determines a recovery countermeasure for a failure occurring in a device itself or another device, comprising: The countermeasure selection device includes a calculation device for executing a program and a storage device for storing the program and data, the storage device stores countermeasure candidate management information for managing a list of recovery countermeasure candidates for a failure content; The countermeasure selection method includes: an input step in which the computing device receives input of an evaluation value for each measure and each evaluation index, a weighted value for each restoration requirement and each evaluation index, and a device to which the restoration requirement is applied; a countermeasure candidate selection step in which the computing device refers to the countermeasure candidate management information and selects a recovery countermeasure candidate for the failure that has occurred; a countermeasure evaluation value management step in which the computing device manages the evaluation value for each countermeasure with respect to one or more evaluation indexes that are important matters in a recovery countermeasure; a recovery requirement management procedure in which the computing device manages the recovery requirements including weighted values ​​indicating priorities of the evaluation indexes and devices to which the recovery requirements are applied; a countermeasure priority determination step in which the computing device determines priorities of the recovery countermeasure candidates based on one or more evaluation indexes; an output step in which the computing device outputs the recovery measure candidates with the determined priorities assigned thereto; In the input step, the arithmetic device selects an evaluation value managed by the countermeasure evaluation value management step for each countermeasure and evaluation index, sets a weight value managed by the recovery requirement management step for each recovery requirement and evaluation index, and sets a recovery requirement managed by the recovery requirement management step for each of the own device or the other device; The countermeasure selection method is characterized in that, in the countermeasure priority determination procedure, the computing device calculates, for each recovery measure candidate selected in the countermeasure candidate selection procedure, a total value of the evaluation values ​​reflecting the weighted values ​​in the recovery requirements applied to the device to be restored, and determines the priority of the recovery measure candidate so that the higher the calculated total value, the higher the priority of the recovery measure candidate.

Citation Information

Patent Citations

  • Maintenance introduction method, system and program

    JP2004192153A

  • Trouble coping apparatus, troubleshooting method for information technology system, and program therefor

    JP2009238010A

  • Restoration planning management device, restoration planning program and restoration planning method

    JP2017199128A

  • Recovery support apparatus, recovery support method and program

    JP2020086474A