Determination program, determination method, and information processing apparatus

The determination program addresses the lack of failure prediction and component relationship consideration in existing monitoring systems by identifying high-risk components and alerting for redundancy loss, thereby reducing the likelihood of redundancy loss in storage apparatuses.

JP7694341B2Active Publication Date: 2025-06-18エフサステクノロジーズ株式会社
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021177930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-06-18
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

Existing monitoring apparatuses do not predict component-level failures in storage apparatuses and do not consider the relationships between components, leading to potential redundancy loss when failures occur.

Method used

A determination program that identifies components with higher failure rates and those with failure rates higher than similar components based on log data, and determines if redundancy is lost due to component failures, outputting alarms when necessary.

Benefits of technology

The solution reduces the possibility of redundancy loss in monitored devices by predicting potential failures and alerting support personnel, thereby enabling timely maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694341000001
    Figure 0007694341000001
  • Figure 0007694341000002
    Figure 0007694341000002
  • Figure 0007694341000003
    Figure 0007694341000003
Patent Text Reader

Abstract

To reduce possibility of loss of redundancy in a monitor target device.SOLUTION: A determination program according to the present invention causes a computer to carry out processing for specifying, from a database obtained by receiving log relating to each of a plurality of components in a monitor target device from a plurality of monitor target devices, a first component having a component failure rate higher than a predetermined threshold for each monitor target device, specifying a second component having a component failure rate responding to a failure factor higher than that of a component of the same kind from the database based on error information included in the log, determining whether or not redundancy disappears due to failure of the first component or the second component with respect to a component group having redundancy by two or more components in the monitor target device and having the first component or the second component, and if it is determined that redundancy disappears due to the failure of the first component or the second component, outputting an alarm regarding to the first component or the second component.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a determination program, a determination method, and an information processing apparatus.

Background Art

[0002] There is known a monitoring apparatus (information processing apparatus) that monitors a monitoring target apparatus such as a storage apparatus.

[0003] For example, when a failure occurs in a storage apparatus, the monitoring apparatus collects and analyzes a log related to the failure transmitted from the storage apparatus, and identifies components such as hardware in which the failure has occurred. Then, the monitoring apparatus notifies the support person of the information of the identified component, thereby enabling a prompt response to the occurred failure.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] The monitoring apparatus monitors the storage apparatus in units of components based on a log related to a failure that has occurred in the storage apparatus, and does not assume prediction of a component-level failure (failure prediction) that may occur in the future in the storage apparatus. Further, in the monitoring apparatus, the relationship between components in the storage apparatus is not considered.

[0006] Therefore, when a failure occurs in the storage apparatus that is the monitoring target, redundancy in the storage apparatus may be lost.

[0007] On one side, one of the objectives of the present invention is to reduce the possibility of redundancy loss in the device to be monitored.

Means for Solving the Problem

[0008] In one aspect, the determination program may cause a computer to execute the following processes. The processes may include identifying, for each of the monitored devices, a first component whose failure rate is higher than a predetermined threshold from a database obtained by receiving logs regarding each of a plurality of components included in the monitored device from a plurality of the monitored devices. Further, the processes may include identifying a second component whose failure rate according to the cause of failure is higher than that of the same type of components based on the error information included in the logs from the database. Furthermore, the processes may include determining, for a component group redundantly configured by two or more components in the monitored device and including the first component or the second component, whether redundancy is lost due to the failure of the first component or the second component. Also, the processes may include outputting an alarm regarding the first component or the second component when it is determined that the redundancy of the component group is lost due to the failure of the first component or the second component.

Effect of the Invention

[0009] On one side, it is possible to reduce the possibility of redundancy loss in the device to be monitored.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the embodiments described below are merely examples, and there is no intention to exclude various modifications and applications of technologies not explicitly described below. For example, the present embodiment can be implemented with various modifications without departing from its gist. In the drawings used in the following embodiments, parts with the same reference numerals represent the same or similar parts unless otherwise specified.

[0012] 〔1〕One embodiment 〔1-1〕Example of system configuration FIG. 1 is a block diagram showing a configuration example of a monitoring system 1 as an example of one embodiment. The monitoring system 1 may include a server 2, a plurality of storage devices 3, and a terminal 4.

[0013] Each of the plurality of storage devices 3 is an example of a device to be monitored and transmits logs to the server 2 at a predetermined timing.

[0014] The server 2 is an example of a monitoring device or a prediction device and performs a monitoring process on each of the plurality of storage devices 3 as a monitoring target. For example, when the server 2 acquires a log from the storage device 3, it may execute identification of replacement parts, prediction of occurrence of a failure, and analysis of the possibility of loss of redundancy based on the log. Then, the server 2 may notify the terminal 4 of the information on the identified replacement parts and the information based on the prediction result and the analysis result.

[0015] The terminal 4 is a terminal used by a support person who provides support for the storage device 3 to be monitored, for example, a maintenance service. The support person may perform maintenance of the storage device 3 according to the notification content through the terminal 4. Note that the maintenance service may be executed by the terminal 4 and the server 2.

[0016] Each of the server 2, the plurality of storage devices 3, and the terminal 4 may be communicably connected to each other via the network 1a. The network 1a may include, for example, one or both of a LAN (Local Area Network) and the Internet. Note that at least one of between the server 2 and the plurality of storage devices 3, between the server 2 and the terminal 4, and between the plurality of storage devices 3 and the terminal 4 may be communicably connected to each other via a network different from the network 1a.

[0017] In the following description, the storage device 3 is taken as an example of the device to be monitored, but the present invention is not limited thereto. The device to be monitored may be various computers (information processing devices) such as a server, a PC, and a communication device.

[0018] [1-2] Hardware configuration example The server 2 according to one embodiment may be a physical server or a virtual server (VM; Virtual Machine). Further, the functions of the server 2 may be realized by one computer or by two or more computers. Furthermore, at least a part of the functions of the server 2 may be realized by using HW (Hardware) resources and NW (Network) resources provided by a cloud environment.

[0019] FIG. 2 is a block diagram showing a hardware (HW) configuration example of a computer 10 that realizes the functions of the server 2 according to one embodiment. When a plurality of computers are used as the HW resources for realizing the functions of the server 2, each computer may have the HW configuration illustrated in FIG. 2.

[0020] As shown in FIG. 2, the computer 10 may illustratively include a processor 10a, a memory 10b, a storage unit 10c, an IF (Interface) unit 10d, an IO (Input / Output) unit 10e, and a reading unit 10f as its HW configuration.

[0021] Processor 10a is an example of an arithmetic processing unit that performs various controls and calculations. Processor 10a may be communicably connected to each block in computer 10 via bus 10i. Note that processor 10a may be a multi-processor including a plurality of processors, a multi-core processor having a plurality of processor cores, or a configuration having a plurality of multi-core processors.

[0022] Examples of processor 10a include integrated circuits (ICs) such as a CPU, MPU, GPU, APU, DSP, ASIC, and FPGA. Note that as processor 10a, a combination of two or more of these integrated circuits may be used. CPU is an abbreviation for Central Processing Unit, MPU is an abbreviation for Micro Processing Unit. GPU is an abbreviation for Graphics Processing Unit, APU is an abbreviation for Accelerated Processing Unit. DSP is an abbreviation for Digital Signal Processor, ASIC is an abbreviation for Application Specific IC, and FPGA is an abbreviation for Field-Programmable Gate Array.

[0023] Memory 10b is an example of HW that stores various data and information such as programs. Examples of memory 10b include one or both of a volatile memory such as DRAM (Dynamic Random Access Memory) and a non-volatile memory such as PM (Persistent Memory).

[0024] The storage unit 10c is an example of HW that stores information such as various data and programs. Examples of the storage unit 10c include magnetic disk devices such as HDDs (Hard Disk Drives), semiconductor drive devices such as SSDs (Solid State Drives), and various storage devices such as non-volatile memories. Examples of non-volatile memories include flash memories, SCMs (Storage Class Memories), ROMs (Read Only Memories), and the like.

[0025] The storage unit 10c may store a program 10g (determination program) that realizes all or part of the various functions of the computer 10.

[0026] For example, the processor 10a of the server 2 can realize the function as the server 2 (control unit 20 illustrated in FIG. 3) described later by expanding and executing the program 10g stored in the storage unit 10c in the memory 10b.

[0027] The IF unit 10d is an example of a communication IF that controls connections and communications between the server 2 and various networks including networks with each of the plurality of storage devices 3 and terminals 4. For example, the IF unit 10d may include an adapter compliant with a LAN such as Ethernet (registered trademark) or an optical communication such as FC (Fibre Channel). The adapter may support one or both of wireless and wired communication methods.

[0028] For example, the server 2 may be communicably connected to each of the plurality of storage devices 3 and terminals 4 via the IF unit 10d and the network. Note that the program 10g may be downloaded from the network to the computer 10 via the communication IF and stored in the storage unit 10c.

[0029] The IO unit 10e may include one or both of an input device and an output device. Examples of the input device include a keyboard, a mouse, a touch panel, and the like. Examples of the output device include a monitor, a projector, a printer, and the like. Further, the IO unit 10e may include a touch panel or the like in which an input device and a display device are integrated.

[0030] The reading unit 10f is an example of a reader that reads data and program information recorded on the recording medium 10h. The reading unit 10f may include a connection terminal or device to which the recording medium 10h can be connected or inserted. Examples of the reading unit 10f include an adapter compliant with USB (Universal Serial Bus) or the like, a drive device that accesses a recording disk, a card reader that accesses a flash memory such as an SD card, and the like. Note that a program 10g may be stored in the recording medium 10h, and the reading unit 10f may read the program 10g from the recording medium 10h and store it in the storage unit 10c.

[0031] Examples of the recording medium 10h include non-transitory computer-readable recording media such as magnetic / optical disks and flash memories. Examples of the magnetic / optical disk include a flexible disk, a CD (Compact Disc), a DVD (Digital Versatile Disc), a Blu-ray Disc, an HVD (Holographic Versatile Disc), and the like. Examples of the flash memory include semiconductor memories such as a USB memory and an SD card.

[0032] The above-described HW configuration of the computer 10 is an example. Therefore, an increase or decrease in HW (for example, addition or deletion of any block), division, integration in any combination, or addition or deletion of a bus, etc. within the computer 10 may be appropriately performed. Note that the computer that realizes the storage device 3 (for example, the controller of the storage device 3) may have the same hardware configuration as the above-described computer 10.

[0033] Functional configuration example [1-3] Next, an example of the functional configuration (software configuration) of the monitoring system 1 according to an embodiment will be described. FIG. 3 is a block diagram showing an example of the functional configuration of the monitoring system 1 according to an embodiment.

[0034] As shown in FIG. 3, the storage device 3 may illustratively include a log transmission unit 32 that transmits the log 31. For example, the controller of the storage device 3 may collect information on the components provided in the storage device 3 and accumulate it in the log 31. The collection of information on the components and the accumulation in the log 31 may be performed at various timings such as, for example, when the storage device 3 is started up, when the configuration is changed, at regular timings, and when a failure is detected.

[0035] FIG. 4 is a diagram showing an example of the log 31. As shown in FIG. 4, the log 31 may illustratively include items such as the "device ID" of the storage device 3, "configuration information" of the component, "status information", "operation information", and the "time" when the entry was created.

[0036] The "device ID" is an example of the identification information of the storage device 3. The "configuration information" may include information such as the part name (part number, model number, etc.) of the component, the part No. (identification information of the component), and the mounting position (component, slot, connector, etc.) of the component in the storage device 3. Further, the "configuration information" may include information such as the identification information of the other component when the component is configured redundantly with other components in the storage device 3.

[0037] Note that the "redundant configuration" may include one or both of a configuration in which a plurality of components operate in parallel or independently in preparation for the occurrence of a failure, and a configuration in which at least one component operates and the remaining components are in a standby state. Further, the "redundancy" may include one or both of hardware-level redundancy and software-level redundancy. As an example, the "redundancy" may include RAID (Redundant Arrays of Inexpensive Disks) in which hardware such as a plurality of storage devices is made redundant by software.

[0038] The "status information" may include information indicating whether the component is normal or abnormal, information such as vendor name, model name, lot, and FW (Firmware) version number. The "operation information" is information related to the operation of the components in the storage device 3, such as the operation time of the components and error information. The error information is various information related to the failures that have occurred in the component, for example, information such as the number of failures that have occurred, the occurrence frequency of collectable errors, and the occurrence frequency of drive timeouts. The number of failures that have occurred is the number of failures that have occurred in the same component.

[0039] Thus, the log 31 may include at least one entry containing an entry of information collected in a state where no failure has occurred for each of the plurality of components provided in the storage device 3.

[0040] The log transmission unit 32 transmits the log 31 accumulated in the storage device 3 to the server 2 at a timing including one or both of a predetermined timing, for example, a periodic timing such as once a week, and a timing at which the occurrence of a failure is detected.

[0041] As shown in FIG. 3, the server 2 may illustratively include a memory unit 21, a communication unit 22, a data formatting unit 23, a prediction unit 24, and a component to be maintained output unit 29. The communication unit 22, the data formatting unit 23, the prediction unit 24, and the component to be maintained output unit 29 are an example of a control unit 20, and may be realized by the processor 10a of the server 2 illustrated in FIG. 2 executing a program 10g developed in the memory 10b.

[0042] The memory unit 21 is an example of a storage area and stores various data used by the server 2. The memory unit 21 may be realized by, for example, a storage area possessed by one or both of the memory 10b and the storage unit 10c shown in FIG. 2.

[0043] As shown in FIG. 3, the memory unit 21 may illustratively store a monitoring target device DB (Database) 21a, a component failure rate table 21b, a failure cause table 21c, a redundancy number table 21d, and output information 21e. Hereinafter, for convenience, each of the information 21a to 21e stored in the memory unit 21 is represented in a table format, but it is not limited thereto, and at least one of these information 21a to 21e may be in various formats such as a DB (Database) or an array.

[0044] The communication unit 22 executes reception of the log 31 transmitted from the storage device 3 and transmission of a notification based on the analysis (determination) result by the prediction unit 24.

[0045] The data shaping unit 23 extracts and shapes information used for analysis by the prediction unit 24 from the log 31 received by the communication unit 22. For example, the data shaping unit 23 may generate or update the monitoring target device DB 21a based on the log 31.

[0046] FIG. 5 is a diagram showing an example of the monitoring target device DB 21a. As shown in FIG. 5, the monitoring target device DB 21a may illustratively include items such as "No.", "device ID", "component name", "redundancy configuration component No.", "mounting position", "alarm information", "alarm continuation count", "operation information", and "component information".

[0047] "No." is identification information for a component entry. "Device ID" is an example of identification information of the storage device 3. "Component name", "redundancy configuration component No.", and "mounting position" are information included in the configuration information of the log 31. "Operation information" is at least a part of the operation information including error information included in the log 31. "Component information" is at least a part of the configuration information and status information included in the log 31.

[0048] "Alarm information" indicates the level of an alarm (normal, caution, warning, emergency, etc.) corresponding to the state of the redundancy configuration by the component, and represents the level of the alarm that the server 2 notifies the terminal 4. The level of the alarm has the normal level as the lowest level, and the levels increase in the order of caution and warning, with the emergency level being the highest level. "Alarm continuation count" indicates the number of times that alarm information other than normal has been notified to the terminal 4 for the component.

[0049] For example, when the alarm information is "normal", the level of the alarm indicates that there is no possibility of redundancy loss even if the component fails, or that it is negligibly small, for example, indicating that alarm notification is unnecessary. When the alarm information is "caution", the level of the alarm indicates that although there is a low possibility of redundancy loss even if the component fails, it should be replaced if possible at the timing of the next regular maintenance or the like. When the alarm information is "warning", the level of the alarm indicates that there is a high possibility of redundancy loss when the component fails, so it should be replaced at the timing of the next regular maintenance or the like. When the alarm information is "emergency", the level of the alarm indicates that redundancy is lost when the component fails, so countermeasures should be considered at the time when the alarm is notified.

[0050] Note that the initial values of the alarm information and the alarm continuation count may be "normal" and "0", respectively.

[0051] The "component name" shown in FIG. 5 distinguishes components in lot units. A lot is the minimum unit for production management or sales management. Components may include factors that result in "defects" in lot units during manufacturing. Therefore, in the monitoring target device DB21a, components are distinguished in lot units in order to detect the bias of the failure rate in lot units according to the failure factors of the components by the method described later.

[0052] Hereinafter, part names distinguished by lot are denoted by adding capital letters of the alphabet as symbols, such as "Part A" to "Part F". On the other hand, as will be described later, part names not distinguished by lot are denoted by adding small letters of the alphabet as symbols, such as "Part a".

[0053] The data shaping unit 23 may identify the storage device 3 and parts from the log 31 received by the communication unit 22, and generate or update an entry in the monitoring target device DB21a for each of the storage device 3 and parts. For example, the data shaping unit 23 may set information based on the log 31 in items of the monitoring target device DB21a other than the alarm information and the number of consecutive alarms.

[0054] Note that the data shaping unit 23 may generate an entry including information other than the alarm information and the number of consecutive alarms in the monitoring target device DB21a based on, for example, the log 31 received by the communication unit 22 at regular timings, or the configuration information (not shown) of one or more storage devices 3 to be monitored.

[0055] Also, in the storage device 3, parts may be replaced in response to, for example, notification of an alarm to the terminal 4 described later, maintenance (repair), or occurrence of a failure of a part. In this case, when the communication unit 22 receives the log 31 regarding the part after replacement, the data shaping unit 23 may register the information of the part after replacement in place of the information of the part before replacement stored in the monitoring target device DB21a.

[0056] For example, the data shaping unit 23 identifies the entry of the part before replacement in the monitoring target device DB21a based on the device ID and the mounting position included in the log 31 of the part after replacement. Then, the data shaping unit 23 may update, for example, replace the part name, the operation information, and the part information of the identified entry with the information of the part after replacement. Further, the data shaping unit 23 may set the alarm information and the number of consecutive alarms of the entry to the initial values.

[0057] Based on the log 31 received by the communication unit 22, the prediction unit 24 may perform a specific process for replacement parts. For example, in the monitoring process, the prediction unit 24 may identify parts such as hardware where a failure has occurred, and notify the terminal 4 of the information on the identified parts.

[0058] Also, the prediction unit 24 according to an embodiment may perform a prediction process for the occurrence of component failures in the monitoring process. For example, in the prediction process, the prediction unit 24 updates the monitored device DB 21a based on the component failure rate table 21b, the failure cause table 21c, and the redundancy number table 21d. Then, based on the updated monitored device DB 21a, the prediction unit 24 predicts the possibility of redundancy loss when each component fails, and when there is a possibility of redundancy loss, notifies the terminal 4 of an alarm.

[0059] In one embodiment, "notification" may be performed by various methods. For example, "notification" may include at least one of sending a message to various addresses such as the email address of the terminal 4, displaying and outputting a message to an output device such as the monitor of the terminal 4, storing a message in a storage area accessible by the terminal 4 (for example, the memory unit 21 or an external storage), and the like.

[0060] As illustrated in FIG. 3, when focusing on the functions for performing the prediction process, the prediction unit 24 may include a component failure rate calculation unit 25, a component failure determination unit 26, a failure cause analysis unit 27, and a redundancy determination unit 28.

[0061] The component failure rate calculation unit 25 calculates the failure rate (failure occurrence rate) of a component based on the number of failures of each component included in the operation information of the monitored device DB 21a.

[0062] For example, the component failure rate calculation unit 25 may calculate the failure rate of a component by dividing the number of component failures by the total operating time of the component for each storage device 3 and for each component. The unit of the component for which the component failure rate calculation unit 25 calculates the failure rate may be a component in which at least one of, for example, the component name, vendor name, model name, and FW version number is the same. In the following description, it is assumed that the component failure rate calculation unit 25 calculates the failure rate for each component having the same component name (component number, model number).

[0063] Note that the failure rate is not limited to the number of failures / (total operating time), and may be calculated based on various information included in the operation information of the log 31, such as the number of failures, operating time, error or failure occurrence frequency, timeout occurrence frequency, etc.

[0064] The component failure rate calculation unit 25 may store, for example, the calculated failure rate in the component failure rate table 21b for each component name.

[0065] FIG. 6 is a diagram showing an example of the component failure rate table 21b. As shown in FIG. 6, the component failure rate table 21b may include, by way of example, items of "component name", "failure rate", and "failure rate (reference value)". The "failure rate (reference value)" is an example of a predetermined reference value, and is, for example, a theoretical value or a design value of the component failure rate. The reference value may be obtained from, for example, verification results at the time of design or a specification sheet (specification).

[0066] In the example of FIG. 6, the "component name" is a component name such as a component number or a model number, and is not distinguished in terms of lots. For example, "component a" may be a component name common to "component A" and "component B" with different lots shown in FIG. 5.

[0067] The component failure determination unit 26 determines whether the failure rate for each storage device 3 and each component calculated by the component failure rate calculation unit 25 exceeds the reference value in the component failure rate table 21b.

[0068] When the failure rate exceeds the reference value, the component failure determination unit 26 determines the component of the storage device 3 as a "component requiring attention", sets an alarm for the monitored device DB21a, notifies the terminal 4 of the alarm, and instructs the failure cause analysis unit 27 to analyze the failure cause. The component requiring attention is an example of a first component whose failure rate is higher than a predetermined threshold value.

[0069] For example, the component failure determination unit 26 sets "Attention" in the alarm information of the entry of the component requiring attention in the monitored device DB21a, and adds "1" to the alarm continuation count when something other than "Normal" is set in the alarm information.

[0070] FIG. 7 is a diagram showing an example of the monitored device DB21a when the failure rate X of component a of device A exceeds the reference value x. In this case, as shown in FIG. 7, the component failure determination unit 26 sets "Attention" in the alarm information of the entries of components A and B of device A (entries No. "1" and "2") in the monitored device DB21a and maintains the alarm continuation count at "0".

[0071] In addition, the component failure determination unit 26 notifies the terminal 4 of an alarm at the "Attention" level regarding the component requiring attention. For example, the alarm may be an alarm prompting replacement of component a (components A and B) if possible at the timing of the next regular maintenance of device A or the like. At this time, the component failure determination unit 26 may set a time stamp of the alarm notification date and time for the alarm information (or other additional items) of the component requiring attention.

[0072] The failure cause analysis unit 27 analyzes the failure cause for each lot of components and determines whether there is a bias in the failure rate for each lot according to the failure cause.

[0073] For example, the failure cause analysis unit 27 may calculate the failure rate for each failure cause and each lot based on the operation information of the monitored device DB21a, and store the calculated failure rate in the failure cause table 21c.

[0074] FIG. 8 is a diagram showing an example of the failure cause table 21c. As shown in FIG. 8, the failure cause table 21c may include items such as "component name", "failure cause", "failure rate", and "same type components" illustratively. The "component name" is the component name for each lot. The "failure cause" is the cause of the occurrence of a failure (error), and examples include various causes that can be obtained from the operation information (error information) of the log 31, such as software error, hardware error, media error, no response (timeout), etc. The "same type components" is the identification information of components that have the same component name as the component in the entry but different lots.

[0075] As shown in FIG. 8, in the failure cause table 21c, without distinguishing the storage device 3, focusing on the components of the same lot, the failure rate for each failure cause is set.

[0076] Note that the failure cause analysis unit 27 may determine the same type components for each component based on the monitored device DB21a or the log 31 and set them in the failure cause table 21c.

[0077] When there is a bias in the failure rate in the failure cause analysis unit 27, for example, for each failure cause, if the failure rate of the components of a certain lot is higher than the failure rate of the components of other lots with the same component name (as an example, n% or more higher; n is a real number of 0 or more, for example, "10"), the failure cause analysis unit 27 determines that the component is a "component to be warned". For example, the failure cause analysis unit 27 sets an alarm for the component to be warned in the monitored device DB21a and notifies the alarm to the terminal 4. The component to be warned is an example of a second component whose failure rate of the component corresponding to the failure cause is higher than that of the same type components.

[0078] For example, the failure cause analysis unit 27 sets "warning" in the alarm information of the entry of the component to be warned in the monitored device DB21a, and when something other than "normal" has already been set in the alarm information, adds "1" to the alarm continuation count.

[0079] FIG. 9 is a diagram showing an example of the monitoring target device DB21a when the failure rate of the failure cause y of component B is 10% or more higher than the failure rate of the failure cause y of component A, which is the same type of component. Note that since there is no entry for the failure rate of the failure cause y of component A, it is 0%. In this case, the failure cause analysis unit 27 sets "warning" in the alarm information of the entry (entry No. "2") of component B of device A of the monitoring target device DB21a. Note that the failure cause analysis unit 27 sets "1" in the alarm continuation count of the entry No. "2".

[0080] In addition, the failure cause analysis unit 27 notifies the terminal 4 of an alarm at the "warning" level regarding the warning target component. For example, the alarm may be an alarm prompting replacement of component B at the timing of the next regular maintenance of device A including component B. At this time, the failure cause analysis unit 27 may set a time stamp of the alarm notification date and time for the alarm information (or other additional items) of the warning target component.

[0081] In addition, the warning target component identified by the failure cause analysis preferably becomes a maintenance target component in a plurality of storage devices 3 including the component, regardless of a specific storage device 3.

[0082] Therefore, in addition to device A including the identified warning target component, the failure cause analysis unit 27 may also set an alarm for component B of device C including the warning target component. For example, the failure cause analysis unit 27 may set "warning" in the alarm information of the entry (entry No. "8") of component B of device C of the monitoring target device DB21a and maintain the alarm continuation count at "0".

[0083] In addition, the failure cause analysis unit 27 may transmit an alarm prompting replacement of component B at the timing of the next regular maintenance of device C to the terminal 4. In this case, a time stamp may be set in the monitoring target device DB21a.

[0084] In this way, according to the server 2, based on the operation information (error information) of the components, it is possible to identify warning target components that may have lot defects across a plurality of storage devices 3. Therefore, it is possible to efficiently or accurately perform failure prediction for each component with respect to a plurality of monitoring target devices.

[0085] Note that the failure cause analysis unit 27 was described as analyzing the failure causes of components in units of lots, but it is not limited to this. For example, the failure cause analysis unit 27 may analyze the failure causes of components in various units such as a lot group obtained by grouping a plurality of lots, model name, FW version number, etc. In these cases, the component names in the monitoring target device DB 21a and the failure cause table 21c may be distinguished in units such as the same lot group, model name, FW version number, etc.

[0086] Note that components with the same component name but different lots, lot groups, model names, or FW version numbers can be said to be "homogeneous components" or "equivalent components" to each other.

[0087] The redundancy determination unit 28 determines whether redundancy is lost due to the failure of a component determined as a component of concern by the component failure determination unit 26 or a component determined as a warning target component by the failure cause analysis unit 27.

[0088] For example, the redundancy determination unit 28 may obtain the redundancy number of the component of concern or the warning target component from the redundancy number table 21d, and determine whether redundancy is lost due to the failure of the component of concern or the warning target component based on the obtained redundancy number.

[0089] FIG. 10 is a diagram showing an example of the redundancy number table 21d. As shown in FIG. 10, the redundancy number table 21d may, for example, include items of "device ID", "part name", and "redundancy number". The "device ID" and "part name" are the device ID and part name shown in FIG. 5. The "redundancy number" is the number of redundancies of the part in the storage device 3, and is the total number of parts that play the role (function) of the part. For example, for parts A and B that are duplicated in device A in FIG. 5, the redundancy number of both is "2" in FIG. 10.

[0090] The redundancy number table 21d may be generated before the operation of the monitoring system 1, such as in the design or construction of the storage device 3 to be the monitoring target device, and stored in the memory unit 21.

[0091] For example, when there are N - 1 or more parts among the parts with a redundancy number of "N" (N is an integer of 2 or more) in the redundancy number table 21d, and "caution" or "warning" is set in the alarm information for the part in the monitoring target device DB21a, the redundancy determination unit 28 determines the part as an "urgent target part". For example, for the urgent target part, the redundancy determination unit 28 sets an alarm for the monitoring target device DB21a and notifies the alarm to the terminal 4.

[0092] For example, when the redundancy determination unit 28 sets "urgent" in the alarm information of the entry of the urgent target part of the monitoring target device DB21a, and when something other than "normal" is already set in the alarm information, "1" is added to the alarm continuation count.

[0093] FIG. 11 is a diagram showing an example of the monitoring target device DB21a when performing the redundancy determination process based on the state of FIG. 9. The redundancy determination unit 28 compares the monitoring target device DB21a (see FIG. 9) with the redundancy number table 21d (see FIG. 10).

[0094] For example, since the redundancy numbers of parts A and B of device A in the redundancy determination unit 28 are N = "2", and the number of settings of "Caution" or "Warning" in the alarm information is "2" ≥ (N - 1), the redundancy determination unit 28 determines parts A and B of device A as emergency target parts. In this case, the redundancy determination unit 28 sets "Emergency" in the alarm information of the entries (entries No. "1" and "2") of parts A and B of device A in the monitored device DB21a, and sets "1" and "2" respectively in the alarm continuation counts.

[0095] Also, for example, since the redundancy numbers of parts A and B of device C in the redundancy determination unit 28 are N = "2", and the number of settings of "Caution" or "Warning" in the alarm information is "1" ≥ (N - 1), the redundancy determination unit 28 determines part B with "Caution" or "Warning" set in the alarm information as an emergency target part. In this case, the redundancy determination unit 28 sets "Emergency" in the alarm information of the entry (entry No. "8") of part B of device C in the monitored device DB21a, and sets "1" in the alarm continuation count. Note that the redundancy determination unit 28 may also determine part A of device C with the alarm information being "Normal" as an emergency target part. In this case, the redundancy determination unit 28 sets "Emergency" in the alarm information of the entry (entry No. "7") of part A of device C, and maintains the alarm continuation count at "0".

[0096] Note that information on software-based redundancy configurations such as RAID may not be set in the redundancy number table 21d. This is because even if the RAID level is determined, the hardware (member disks) constituting the RAID is not fixed.

[0097] However, the RAID level and the number of member disks can be obtained from the configuration information included in the monitored device DB21a or the log 31. Therefore, for software-based redundancy configurations such as RAID, the redundancy determination unit 28 may calculate the redundancy number based on the monitored device DB21a or the log 31, and compare the redundancy number with the monitored device DB21a.

[0098] For example, when the RAID of a certain storage device 3 has 10 member disks and a RAID level of "RAID6", failures of up to 2 of the member disks are tolerated. In other words, it can be regarded as being triplicated (N = 3). In this case, when the alarm information of 2 (≥ N - 1) of the storage devices constituting the RAID is "Attention" or "Warning", the redundancy determination unit 28 may determine the storage device with the set alarm information as an emergency target component.

[0099] In addition, the redundancy determination unit 28 notifies the terminal 4 of an "Emergency" level alarm regarding the emergency target component. For example, the alarm may be an alarm indicating that when the emergency target component fails in device A having components A and B and device C having component B (and A), redundancy is lost, and thus countermeasures should be considered at the time when the alarm is notified. At this time, the redundancy determination unit 28 may set a time stamp of the alarm notification date and time for the alarm information (or other additional items) of the emergency target component.

[0100] In this way, the redundancy determination unit 28 determines whether redundancy is lost due to the failure of the attention target component or warning target component in a component group that is made redundant by two or more components in the storage device 3 and includes the attention target component or warning target component. And when the redundancy determination unit 28 determines that the redundancy of the component group is lost due to the failure of the attention target component or warning target component, it outputs an alarm "Emergency" regarding the attention target component or warning target component.

[0101] Thereby, it is possible to early notify the terminal 4 of an alarm prompting replacement of the attention target component or warning target component determined to have lost redundancy, in other words, the emergency target component, and reduce the possibility of the redundancy configuration disappearing.

[0102] When the maintenance target component output unit 29 receives a maintenance target component output request from the terminal 4 via the communication unit 22, it generates output information 21e indicating the maintenance target component and transmits the output information 21e to the terminal 4 via the communication unit 22.

[0103] For example, in the maintenance plan of the storage device 3, the terminal 4 sends an output request for parts to be maintained to the server 2. The output request may include the device ID of the device to be monitored and the maintenance execution time (for example, date and time).

[0104] The maintenance target part output unit 29 refers to the monitored device DB 21a, extracts the parts to be maintained from the parts of the device ID included in the output request, and stores them in the memory unit 21 as the output information 21e. Examples of the parts to be maintained include parts with a possibility of failure that may occur within a predetermined period (for example, within six months) close to the maintenance execution time among the parts of the device ID included in the output request. The parts with a possibility of failure may include parts that are the emergency target.

[0105] For example, the maintenance target part output unit 29 may extract an entry of a maintenance target part whose combination of alarm information and the number of consecutive alarms satisfies a predetermined condition from the parts of the specified device ID of the monitored device DB 21a (see FIG. 11) and set it in the output information 21e.

[0106] Examples of the predetermined condition include cases where the combination of alarm information and the number of consecutive alarms corresponds to any one of three warnings, two alerts, one emergency, or zero. The predetermined condition may be appropriately changed according to the setting of the predetermined period.

[0107] FIG. 12 is a diagram showing another example of the monitored device DB 21a. FIG. 13 is a diagram showing an example of the output information 21e generated by extracting the maintenance target parts of device A from the monitored device DB 21a shown in FIG. 12. FIG. 14 is a diagram showing an example of the output information 21e generated by extracting the maintenance target parts of device B from the monitored device DB 21a shown in FIG. 12.

[0108] For example, when the device ID of device A is specified in the output request, as shown in FIG. 12, component B mounted on the controller 2 of device A has "Attention" transmitted three times. Therefore, as illustrated in FIG. 13, the maintenance target component output unit 29 extracts the entry of component B of device A from the monitored device DB21a and sets it in the output information 21e.

[0109] Also, for example, when the device ID of device B is specified in the output request, as shown in FIG. 12, component D mounted on the controller 1-slot1 of device B has "Warning" transmitted twice. Therefore, as illustrated in FIG. 14, the maintenance target component output unit 29 extracts the entry of component D of device B from the monitored device DB21a and sets it in the output information 21e.

[0110] As described above, the server 2 according to one embodiment estimates the period during which each component may fail (the date from the time of maintenance), in other words, the lifespan of each component, based on the information regarding the failure of each component included in the monitored device. Thereby, in the monitoring system 1, appropriate preventive maintenance can be carried out, and the possibility of redundancy loss in the monitored device can be reduced.

[0111] [1-4] Operation Example Next, with reference to FIGS. 15 to 19, an operation example of the server 2 in the monitoring system 1 according to one embodiment will be described.

[0112] [1-4-1] Monitoring Process FIG. 15 is a flowchart for explaining an operation example of the monitoring process by the server 2. FIGS. 16 to 18 are flowcharts for explaining the operation examples of the failure rate determination process, the failure cause determination process, and the redundancy maintenance determination process shown in FIG. 15, respectively.

[0113] As illustrated in FIG. 15, the communication unit 22 of the server 2 collects the log 31 from the storage device 3 (step S1). For example, the data formatting unit 23 formats the log 31 received by the communication unit 22 and stores it in the monitored device DB21a.

[0114] The component failure rate calculation unit 25 and the component failure determination unit 26 execute a failure rate determination process for each storage device 3 based on the monitoring target device DB21a and the component failure rate table 21b for each component targeted by the log 31 received in step S1 (step S2).

[0115] The failure cause analysis unit 27 executes a failure cause determination process for each equivalent component based on the monitoring target device DB21a and the failure cause table 21c (step S3).

[0116] The redundancy determination unit 28 executes a redundancy maintenance determination process for each of the components requiring attention and the components targeted for warning determined in the failure cause determination process based on the monitoring target device DB21a and the redundancy number table 21d (step S4), and the process ends.

[0117] (Step S2: Failure Rate Determination Process) Next, an operation example of the failure rate determination process in step S2 of FIG. 15 will be described. As illustrated in FIG. 16, the component failure rate calculation unit 25 and the component failure determination unit 26 calculate the failure rate based on the monitoring target device DB21a (step S21), and set the calculated failure rate in the component failure rate table 21b.

[0118] The component failure determination unit 26 refers to the component failure rate table 21b and determines whether the failure rate exceeds the reference value (step S22). If the failure rate does not exceed the reference value (NO in step S22), the failure rate determination process ends.

[0119] If the failure rate exceeds the reference value (YES in step S22), the component failure determination unit 26 updates the setting of the alarm in the entry of the monitoring target device DB21a corresponding to the component (component requiring attention) (step S23). For example, the component failure determination unit 26 sets "Attention" in the alarm information, and if the alarm information was already other than "Normal" before the setting, adds "1" to the alarm continuation count.

[0120] The component failure determination unit 26 notifies the alarm "Attention" for the device ID of the entry in which "Attention" is set in the alarm information (step S24), and the failure rate determination process ends.

[0121] (Step S3: Failure cause determination process) Next, an operation example of the failure cause determination process in step S3 of FIG. 15 will be described. As illustrated in FIG. 17, the failure cause analysis unit 27 analyzes the failure cause for equivalent components, for example, for each lot, based on the monitored device DB 21a (step S31), calculates the failure cause and the failure rate for each lot, and sets the failure cause and the failure rate in the failure cause table 21c.

[0122] The failure cause analysis unit 27 refers to the failure cause table 21c and determines whether the failure rate corresponding to the failure cause is higher than that of other equivalent components, for example, components of other lots (step S32). If the failure rate corresponding to the failure cause is less than or equal to that of components of other lots (NO in step S32), the failure cause determination process ends.

[0123] If the failure rate corresponding to the failure cause is higher than that of components of other lots (YES in step S32), the failure cause analysis unit 27 updates the alarm setting of the entry in the monitored device DB 21a corresponding to the component (component to be warned) (step S33). For example, the failure cause analysis unit 27 sets "Warning" in the alarm information, and if the alarm information is already other than "Normal" before the setting, "1" is added to the alarm continuation count.

[0124] The failure cause analysis unit 27 notifies the alarm "Warning" for the device ID of the entry in which "Warning" is set in the alarm information (step S34).

[0125] In addition, the failure cause analysis unit 27 updates the alarm setting of the entry in the monitored device DB 21a corresponding to the component (same component) with the same part name as the component to be warned in other storage devices 3 (step S35), and the failure cause determination process ends.

[0126] (Step S4: Redundancy Maintenance Judgment Process) Next, an operation example of the redundancy maintenance process in step S4 of FIG. 15 will be described. As illustrated in FIG. 18, the redundancy determination unit 28 analyzes the redundancy maintainability of the component to be noted or the component to be warned based on the monitored device DB21a and the redundancy number table 21d (step S41).

[0127] The redundancy determination unit 28 determines whether the component to be noted or the component to be warned is in a redundant configuration (step S42). If it is not in a redundant configuration (NO in step S42), the redundancy maintenance judgment process ends.

[0128] If it is in a redundant configuration (YES in step S42), the redundancy determination unit 28 determines whether redundancy can be maintained when the component to be noted or the component to be warned fails (step S43). If redundancy can be maintained (YES in step S43), the redundancy maintenance judgment process ends.

[0129] If redundancy cannot be maintained (NO in step S43), the redundancy determination unit 28 updates the alarm setting of the entry of the monitored device DB21a corresponding to the component (step S44). For example, the redundancy determination unit 28 sets "urgent" in the alarm information, and if the alarm information was already other than "normal" before the setting, "1" is added to the alarm continuous count.

[0130] The redundancy determination unit 28 notifies the alarm "urgent" to the device ID of the entry in which "urgent" is set in the alarm information (step S45), and the redundancy maintenance judgment process ends.

[0131] [1-4-2] Output Process of Components to be Maintained FIG. 19 is a flowchart for explaining an operation example of the output process of components to be maintained by the server 2.

[0132] As illustrated in FIG. 19, the communication unit 22 of the server 2 receives an output request for a list of components to be maintained (output information 21e) from the terminal 4 (step S51).

[0133] The maintenance target component output unit 29 refers to the monitored device DB21a (step S52) and selects an unselected entry of the component with the device ID specified in the output request (step S53).

[0134] The maintenance target component output unit 29 determines whether the alarm state of the selected entry is "normal" (step S54). If it is "normal" (YES in step S54), the process proceeds to step S60.

[0135] If the alarm state of the selected entry is not "normal" (NO in step S54), the maintenance target component output unit 29 determines whether the alarm state of the selected entry is "caution" (step S55). If it is "caution" (YES in step S55), the maintenance target component output unit 29 determines whether the alarm continuation count of the selected entry is 3 or more (step S56). If the alarm continuation count is not 3 or more (NO in step S56), the process proceeds to step S60.

[0136] If the alarm continuation count of the selected entry is 3 or more (YES in step S56), the maintenance target component output unit 29 adds the information of the selected entry to the output information 21e (step S59), and the process proceeds to step S60.

[0137] In step S55, if the alarm state of the selected entry is not "caution" (NO in step S55), the maintenance target component output unit 29 determines whether the alarm state of the selected entry is "warning" (step S57). If it is "warning" (YES in step S57), the maintenance target component output unit 29 determines whether the alarm continuation count of the selected entry is 2 or more (step S58). If the alarm continuation count is not 2 or more (NO in step S58), the process proceeds to step S60.

[0138] When the number of times the alarm for the selected entry is “2” or more (YES in step S58), the maintenance target part output unit 29 adds the information of the selected entry to the output information 21e (step S59), and the process proceeds to step S60.

[0139] In step S57, when the alarm state of the selected entry is not “warning” (NO in step S57), the alarm state of the selected entry is “emergency”. In this case, the maintenance target part output unit 29 adds the information of the selected entry to the output information 21e (step S59), and the process proceeds to step S60.

[0140] In step S60, the maintenance target part output unit 29 determines whether there is an unselected entry for the part with the device ID specified in the output request in the monitored device DB21a. If there is an unselected entry (YES in step S60), the process proceeds to step S53.

[0141] If there is no unselected entry in the monitored device DB21a (NO in step S60), the maintenance target part output unit 29 transmits the output information 21e to the terminal 4 via the communication unit 22 (step S61), and the maintenance target part output process ends.

[0142] 〔2〕Others The technology according to the above-described embodiment can be implemented with the following modifications and changes.

[0143] For example, the communication unit 22, data formatting unit 23, prediction unit 24, component failure rate calculation unit 25, component failure determination unit 26, failure cause analysis unit 27, redundancy determination unit 28, and maintenance target part output unit 29 of the server 2 shown in FIG. 3 may be combined or divided in any combination.

[0144] Also, the information 21a to 21e stored in the memory unit 21 shown in FIG. 3 may be combined or divided in any combination.

[0145] Furthermore, the terminal 4 notified of an alarm from the storage device 3 may notify various messages to, for example, a support person or the like (support person, administrator or user of the storage device 3, etc.). As an example, the terminal 4 may notify a warning message prompting a data backup to an external storage to a support person or the like. Also, for example, when the terminal 4 receives an "urgent" alarm notification regarding the loss of RAID redundancy, if there is an available hot spare drive, the terminal 4 may notify a message instructing a support person or the like to transfer data from the suspected drive targeted by the alarm to the hot spare drive to maintain redundancy.

[0146] 〔3〕Supplementary Note Regarding the above embodiments, the following supplementary notes are further disclosed.

[0147] (Supplementary Note 1) From a database obtained by receiving logs for each of a plurality of components included in a monitoring target device from a plurality of the monitoring target devices, identify a first component whose failure rate of the component is higher than a predetermined threshold for each of the monitoring target devices, From the database, based on the error information included in the log, identify a second component whose failure rate of the component according to the failure cause is higher than that of the same type of component, For a component group made redundant by two or more components in the monitoring target device, and for the component group including the first component or the second component, determine whether redundancy is lost due to the failure of the first component or the second component, When it is determined that redundancy of the component group is lost due to the failure of the first component or the second component, output an alarm regarding the first component or the second component, A determination program that causes a computer to execute the process.

[0148] (Supplementary Note 2) The determination process includes a process of determining that redundancy is lost due to a failure of the first component or the second component when the number of the first components or the second components included in the component group with redundancy number N (N is an integer of 2 or more) is (N-1) or more based on information managing the redundancy number of each of a plurality of components included in the device to be monitored. The determination program according to Supplementary Note 1.

[0149] (Supplementary Note 3) The process of identifying the first component includes a process of outputting a first alarm regarding the identified first component. The process of identifying the second component includes a process of outputting a second alarm regarding the identified second component. Based on the output frequencies of the first alarm and the second alarm and the first component or the second component determined to have lost redundancy in the component group, a component to be subject to maintenance is determined. The determination program according to Supplementary Note 1 or Supplementary Note 2, which causes the computer to execute the process.

[0150] (Supplementary Note 4) The same-type components are components having the same component name and at least one of lot, lot group, model name, and firmware version number being different. The determination program according to any one of Supplementary Notes 1 to 3.

[0151] (Supplementary Note 5) From a database obtained by receiving logs regarding each of a plurality of components included in the device to be monitored from a plurality of the devices to be monitored, a first component having a component failure rate higher than a predetermined threshold is identified for each of the devices to be monitored. From the database, based on error information included in the log, a second component having a component failure rate according to a failure cause higher than that of the same-type components is identified. For a component group redundantly configured by two or more components in the device to be monitored and including the first component or the second component, it is determined whether redundancy is lost due to a failure of the first component or the second component. When it is determined that the redundancy of the component group is lost due to a failure of the first component or the second component, an alarm regarding the first component or the second component is output. A determination method executed by a computer.

[0152] (Appendix 6) The determination process is based on information managing the redundancy number of each of a plurality of components included in the monitoring target device, and when the number of the first component or the second component included in the component group with a redundancy number N (N is an integer of 2 or more) is (N - 1) or more, the process includes determining that the redundancy is lost due to a failure of the first component or the second component. The determination method according to Appendix 5.

[0153] (Appendix 7) The process of identifying the first component includes a process of outputting a first alarm regarding the identified first component. The process of identifying the second component includes a process of outputting a second alarm regarding the identified second component. Based on the output frequencies of the first alarm and the second alarm, and the first component or the second component for which it is determined that the redundancy of the component group is lost, a component to be subject to maintenance is determined. The determination method according to Appendix 5 or Appendix 6, executed by the computer.

[0154] (Appendix 8) The same type of components are components with the same component name and at least one of lot, lot group, model name, and firmware version number being different. The determination method according to any one of Appendices 5 to 7.

[0155] (Appendix 9) From a database obtained by receiving logs regarding each of a plurality of components included in the monitoring target device from a plurality of the monitoring target devices, a first component with a component failure rate higher than a predetermined threshold is identified for each of the monitoring target devices. Based on the error information included in the log, identify a second component from the database whose failure rate of components according to the cause of failure is higher than that of the same type of components. For a component group that is redundant with two or more components in the device to be monitored and includes the first component or the second component, determine whether the redundancy is lost due to the failure of the first component or the second component. When it is determined that the redundancy of the component group is lost due to the failure of the first component or the second component, output an alarm regarding the first component or the second component. An information processing apparatus including a control unit.

[0156] (Appendix 10) In the process of the determination, based on the information managing the redundancy number of each of the plurality of components included in the device to be monitored, when the number of the first component or the second component included in the component group with a redundancy number N (N is an integer of 2 or more) is (N - 1) or more, determine that the redundancy is lost due to the failure of the first component or the second component. The information processing apparatus according to Appendix 9.

[0157] (Appendix 11) The control unit In the process of identifying the first component, output a first alarm regarding the identified first component. In the process of identifying the second component, output a second alarm regarding the identified second component. Based on the output times of the first alarm and the second alarm and the first component or the second component for which it is determined that the redundancy of the component group is lost, determine the component to be subject to maintenance. The information processing apparatus according to Appendix 9 or Appendix 10.

[0158] (Appendix 12) The same type of components are components with the same part name and at least one of the lot, lot group, model name, and firmware version number being different. The information processing apparatus according to any one of Appendices 9 to 11.

Description of Symbols

[0159] 1 Monitoring system 1a Network 2 Server 20 Control unit 21 Memory unit 21a Monitored device DB 21b Component failure rate table 21c Failure cause table 21d Redundancy number table 21e Output information 22 Communication unit 23 Data formatting unit 24 Prediction unit 25 Component failure rate calculation unit 26 Component failure determination unit 27 Failure cause analysis unit 28 Redundancy determination unit 29 Maintenance target component output unit 3 Storage device 31 Log 32 Log transmission unit 4 Terminal

Claims

1. From a database obtained by receiving logs for each of a plurality of components included in a device to be monitored from a plurality of the devices to be monitored, identify a first component whose failure rate of the component is higher than a predetermined threshold value for each of the devices to be monitored, From the database, identify a second component whose failure rate of the component according to the cause of failure is higher than that of the same type of components based on the error information included in the log, For a component group that is redundant with two or more components in the device to be monitored and includes the first component or the second component, determine whether redundancy is lost due to a failure of the first component or the second component, When it is determined that redundancy of the component group is lost due to a failure of the first component or the second component, output an alarm regarding the first component or the second component, A determination program that causes a computer to execute processing.

2. The determining process includes, based on information managing the redundancy number of each of a plurality of components included in the device to be monitored, when the number of the first component or the second component included in the component group having a redundancy number N (N is an integer of 2 or more) is (N - 1) or more, a process of determining that redundancy is lost due to a failure of the first component or the second component, The determination program according to claim 1.

3. The process of identifying the first component includes a process of outputting a first alarm regarding the identified first component, The process of identifying the second component includes a process of outputting a second alarm regarding the identified second component, Based on the number of times the first alarm and the second alarm are output and the first component or the second component for which it is determined that redundancy of the component group is lost, determine a component to be maintained, The determination program according to claim 1 or claim 2 that causes the computer to execute processing.

4. The same type of parts are parts with the same part name, and at least one of the lot, lot group, model name, and firmware version number is different. The determination program according to any one of claims 1 to 3.

5. From a database obtained by receiving logs for each of a plurality of parts included in the monitoring target device from a plurality of the monitoring target devices, identify a first part whose failure rate is higher than a predetermined threshold for each monitoring target device. From the database, based on the error information included in the log, identify a second part whose failure rate according to the failure cause is higher than that of the same type of parts. For a component group that is redundant with two or more components in the monitoring target device and includes the first component or the second component, determine whether redundancy is lost due to the failure of the first component or the second component. When it is determined that the redundancy of the component group is lost due to the failure of the first component or the second component, output an alarm regarding the first component or the second component. A determination method in which a computer executes processing.

6. From a database obtained by receiving logs for each of a plurality of parts included in the monitoring target device from a plurality of the monitoring target devices, identify a first part whose failure rate is higher than a predetermined threshold for each monitoring target device. From the database, based on the error information included in the log, identify a second part whose failure rate according to the failure cause is higher than that of the same type of parts. For a component group that is redundant with two or more components in the monitoring target device and includes the first component or the second component, determine whether redundancy is lost due to the failure of the first component or the second component. When it is determined that the redundancy of the component group is lost due to the failure of the first component or the second component, output an alarm regarding the first component or the second component. An information processing apparatus including a control unit.

Citation Information

Patent Citations

  • Disk device and disk array control system

    JP1994051915A

  • Disk array device and control method therefor

    JP1999345095A

  • Array disk group maintenance management system, array disk group maintenance management device, array disk group maintenance management method, and array disk group maintenance management program

    JP2008171231A

  • Monitoring control network system

    JP2011138251A

  • Preventive maintenance instruction device, system, method and program

    JP2019036158A