Fault response system and fault response method

The failure response system automates the identification of necessary parts and procedures for storage device failures, addressing inefficiencies in existing systems by determining urgency and coordinating factory analysis and maintenance, thus enhancing response efficiency.

JP2025133582APending Publication Date: 2025-09-11HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024031618
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing systems struggle to efficiently address highly urgent failures in storage devices, as they often require manual identification of necessary parts and procedures, leading to increased effort and inefficiency, especially when knowledge of the required actions is not readily available.

Method used

A failure response system that includes a fault analysis system to determine urgency and countermeasures, requesting factory analysis for high-urgency failures and arranging maintenance personnel and parts for low-urgency failures, utilizing a storage device, component management system, maintenance personnel assignment system, and factory terminal for efficient response.

Benefits of technology

Enables efficient and appropriate measures for highly urgent storage device failures by automating the identification of necessary parts and procedures, reducing manual effort and improving response times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025133582000001_ABST
    Figure 2025133582000001_ABST
Patent Text Reader

Abstract

To efficiently execute appropriate countermeasures even for a highly urgent fault in a storage device.SOLUTION: A fault response system 1 includes a fault analysis system 10 comprising an arithmetic device configured to: hold information on countermeasures for faults in a device to be monitored: determine the presence or absence of a countermeasure that matches a new fault as a level of urgency based on new fault information and countermeasure information; when the new fault is determined to have a high urgency, request a factory system to analyze it; and when the new fault is determined to have a low urgency, request a predetermined system to arrange for maintenance personnel and components according to the countermeasure.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention generally relates to a technology for a failure response system and a failure response method, and more specifically to a technology that enables an appropriate response to be efficiently executed even for a highly urgent failure in a storage device. [Background technology]

[0002] In recent years, many companies have been faced with issues such as a rapid increase in data volume and the increasing complexity of storage environments. As a result, services that provide storage environment construction and operation in one integrated system have emerged, providing an ideal solution for these companies.

[0003] Here, we will outline how the company providing the service responds when a failure occurs in a storage device. When a failure of some kind occurs in a storage device, sensors installed in the storage device notify a predetermined management system of the failure. This management system is, for example, an operations management system at the company providing the service. An administrator in the maintenance department checks the contents of the notification in the operations management system and makes the necessary arrangements.

[0004] Based on the above report from the storage device (at the customer site), the administrator arranges for the dispatch of a maintenance technician. The dispatched maintenance technician then travels to the customer site and takes action such as replacing parts in the storage device that has failed. Generally, the above-described flow of actions is taken when a storage device failure occurs, but there is still room for improvement in terms of efficiency. Therefore, various conventional technologies have been proposed with the aim of improving the efficiency of responding to failures that accompany such system operations.

[0005] For example, Patent Document 1 discloses a technology relating to automated fault response that automatically analyzes computer faults, arranges for parts, and arranges for maintenance personnel. Specifically, this technology relates to an automated computer fault response system in a server system made up of a user computer, a communication network, and a server computer, the system having fault analysis means that, when a fault occurs in the user computer, receives fault information of the user computer sent via the communication network and performs a fault analysis of the user computer based on the fault information, parts management means that arranges for suspected parts required for repair service based on the results of the fault analysis by the fault analysis means and ships them to the user of the user computer, maintenance staff assignment means that receives the fault information and information on the results of the fault analysis and assigns a maintenance staff member required for repair service, and a maintenance staff portable terminal carried by the maintenance staff that displays fault response instruction information based on the fault information and information on the results of the fault analysis on a portable terminal carried by the assigned maintenance staff and repairs the computer based on the fault response instruction information. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-306360 Summary of the Invention [Problem to be solved by the invention]

[0007] As disclosed in the prior art, when a failure occurs in a storage device, maintenance personnel perform on-site work using the necessary parts and maintenance procedures identified through failure analysis. However, there may be cases where the necessary parts and maintenance procedures cannot be identified through failure analysis. In such cases, the maintenance department must request a failure analysis from the factory. However, determining the need for such a request requires considerable knowledge and effort. Furthermore, the effort required to make the request is also required. Furthermore, the series of arrangements for shipping faulty parts recovered from the customer site to the factory is manually performed by maintenance personnel, which contributes to an increase in the amount of effort required beyond the actual work of troubleshooting. In other words, even with the prior art, it is not possible to accurately address highly urgent failures for which knowledge of the necessary parts and maintenance procedures is not available, and to streamline the subsequent series of operations.

[0008] The present invention has been made in view of the above-mentioned problems, and aims to provide a technique that enables efficient and appropriate measures to be taken even in the case of a highly urgent failure in a storage device. [Means for solving the problem]

[0009] The present application includes multiple means for solving the above-mentioned problems, examples of which are as follows: To solve the above-mentioned problems, a failure response system according to one aspect of the present invention includes a failure analysis system including: a storage device that stores information on countermeasure cases for failures in monitored devices; a process for determining whether or not there is a countermeasure case that matches the new failure in terms of urgency, based on failure information reported about a new failure in the monitored device and information on the countermeasure case; a process for requesting a system in a factory that manufactures or repairs the monitored devices to analyze how to deal with the new failure, if it is determined as a result of the determination that the new failure has a high urgency; and a computing device that executes a process for requesting a predetermined system to arrange for maintenance personnel and parts, according to the countermeasure case that matches the new failure, if it is determined as a result of the determination that the new failure has a low urgency.

[0010] In addition, in order to solve the above-mentioned problems, a fault response method according to one aspect of the present invention is characterized in that an information processing device stores information on countermeasures for faults in a monitored device in a storage device, and determines whether or not there is a countermeasure for the new fault in terms of urgency based on fault information reported about a new fault in the monitored device and the countermeasure information; if the result of the determination shows that the new fault is of high urgency, requests a factory system responsible for manufacturing or repairing the monitored device to analyze how to deal with the new fault; and if the result of the determination shows that the new fault is of low urgency, requests a specified system to arrange for maintenance personnel and parts according to the countermeasure for the new fault. [Effects of the Invention]

[0011] According to the present invention, it is possible to efficiently take appropriate measures even in the case of a highly urgent failure in a storage device. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating an example of a network configuration of a failure response system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing device according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating an example of a functional configuration of a fault analysis system according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating an example of the functional configuration of a storage device at a customer site in this embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a functional configuration of a parts management system according to the present embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of a functional configuration of a maintenance worker assignment system according to the present embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a functional configuration of a maintenance personnel terminal according to the present embodiment. [Figure 8] FIG. 2 is a diagram illustrating an example of a functional configuration of a factory terminal according to the present embodiment. [Figure 9A] FIG. 3 is a diagram illustrating an example of the configuration of user information in the present embodiment. [Figure 9B] FIG. 2 is a diagram illustrating an example of the configuration of fault information according to the present embodiment. [Figure 9C] FIG. 10 is a diagram illustrating an example of the configuration of storage status information (customer) in this embodiment. [Figure 10A] FIG. 2 is a diagram showing an example of the configuration of past case information (analysis) in the present embodiment. [Figure 10B] FIG. 2 is a diagram showing an example of the configuration of fault report information (analysis) in this embodiment. [Figure 10C] FIG. 2 is a diagram illustrating an example of the configuration of storage status information (analysis) in this embodiment. [Figure 11A] FIG. 2 is a diagram showing an example of the configuration of maintenance procedure information (analysis) in the present embodiment. [Figure 11B] FIG. 4 is a diagram illustrating an example of the configuration of priority information in the present embodiment. [Figure 12A] FIG. 2 is a diagram illustrating an example of the configuration of maintenance staff information according to the present embodiment. [Figure 12B] FIG. 4 is a diagram illustrating an example of the configuration of maintenance time information according to the present embodiment. [Figure 12C] FIG. 2 is a diagram showing an example of the configuration of maintenance contract information according to the present embodiment. [Figure 13A] FIG. 2 is a diagram illustrating an example of the configuration of parts inventory information according to the present embodiment. [Figure 13B] FIG. 2 is a diagram illustrating an example of the configuration of base information according to the present embodiment. [Figure 13C] FIG. 2 is a diagram showing an example of the configuration of delivery information in this embodiment. [Figure 14] FIG. 2 is a diagram showing an example of the configuration of schedule information in the present embodiment. [Figure 15A] FIG. 2 is a diagram showing an example of the configuration of maintenance procedure information (factory) in the present embodiment. [Figure 15B] FIG. 2 is a diagram showing an example of the configuration of product information in the present embodiment. [Figure 15C] FIG. 2 is a diagram showing an example of the configuration of fault report information (factory) in the present embodiment. [Figure 16A]FIG. 2 is a diagram showing an example of the configuration of past case information (factory) in the present embodiment. [Figure 16B] FIG. 4 is a diagram showing an example of the configuration of survey information in the present embodiment. [Figure 17] FIG. 2 is a diagram showing a flow example (part 1) of a failure handling method according to the present embodiment. [Figure 18A] FIG. 10 is a diagram showing a flow example (part 2) of the failure handling method of the present embodiment. [Figure 18B] FIG. 10 is a diagram showing a flow example (part 3) of the failure handling method of the present embodiment. [Figure 18C] FIG. 10 is a diagram showing a flow example (part 4) of the failure handling method of the present embodiment. [Figure 19] FIG. 10 is a diagram showing a flow example (part 5) of the failure handling method of the present embodiment. [Figure 20] FIG. 10 is a diagram showing a flow example (part 6) of the failure handling method of the present embodiment. [Figure 21] FIG. 10 is a diagram showing a flow example (part 7) of the failure handling method of the present embodiment. [Figure 22] FIG. 10 is a diagram showing a flow example (part 8) of the failure handling method of the present embodiment. [Figure 23] FIG. 9 is a diagram showing a flow example (part 9) of the failure handling method of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] In the following description, a communication device may be one or more communication interface devices, which may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., an NIC and an HBA (Host Bus Adapter)).

[0014] In the following description, a "main storage device" refers to one or more memory devices, which are an example of one or more storage devices. At least one of the memory devices in the main storage device may be a volatile memory device or a non-volatile memory device.

[0015] In the following description, an "auxiliary storage device" may be one or more persistent storage devices, which are an example of one or more storage devices. The persistent storage device may typically be a non-volatile storage device, specifically, for example, a hard disk drive (HDD), a solid state drive (SSD), or a non-volatile memory express (NVMe) drive.

[0016] Furthermore, in the following description, a "processor" may refer to one or more processor devices. The at least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be another type of processor device such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may also be a processor core. The at least one processor device may also be a processor device in a broader sense, such as a hardware circuit that performs part or all of the processing (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).

[0017] In the following description, information that provides an output in response to an input may be described using expressions such as "xxx table" or "xxx database." However, this information may be data of any structure (for example, structured data or unstructured data), or may be a learning model such as a neural network, genetic algorithm, or random forest that generates an output in response to an input. Therefore, "xxx table" or "xxx database" may be referred to as "xxx information." In the following description, the structure of each database or table is an example, and one database or table may be divided into two or more databases or tables, or all or part of two or more databases or tables may be one database or table.

[0018] In the following description, processing may be described using a "program" as the subject. However, because a program is executed by a processor to perform a predetermined process using a storage device and / or an interface device, etc., as appropriate, the subject of the process may also be the processor (or a device such as a controller having the processor). A program may be installed in a device such as a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0019] In addition, in the following description, when describing elements of the same type without distinguishing between them, common parts of the reference symbols may be used, and when describing elements of the same type with distinction between them, reference symbols or element identifiers may be used.

[0020] <Network configuration of the failure response system> The present embodiment will be described below. Fig. 1 is a diagram showing an example of a network configuration that constitutes a failure response system 1 of the present embodiment. The failure response system 1 is composed of a failure analysis system 10 as its minimum configuration. However, as in the network configuration shown in the figure, in addition to the failure analysis system 10, at least one of a storage device 20, a component management system 30, a maintenance personnel assignment system 40, a maintenance personnel terminal 50, and a factory terminal 60 may also be included.

[0021] The failure response system 1 configured in this manner is, for example, a system that operates within a service that builds and manages a storage environment, and is based on the technical background of accurately and efficiently dealing with new failures for which past failure handling examples cannot be used as reference. The failure response system 1 in this embodiment is a system that can efficiently execute accurate responses even for high-urgency failures in storage devices.

[0022] In the network configuration shown in Fig. 1, the fault analysis system 10 may be a server device. The fault analysis system 10 is communicably connected to a storage device 20, a component management system 30, a maintenance staff assignment system 40, a maintenance staff terminal 50, and a factory terminal 60 via a network N. The network N may be the Internet, a local area network (LAN), a wide area network (WAN), a mobile phone network, or the like, but is not limited to these. The fault analysis system 10 is assumed to be a system operated by a service provider that manages and operates the storage environment. Of course, the fault analysis system 10 may also be operated by the customer who receives the service.

[0023] The fault analysis system 10 receives fault information from a storage device 20 currently in operation at a customer site and monitors the fault. It then determines the level of urgency of the fault and the corresponding countermeasures by sequentially making various judgments, such as determining whether the fault has been reported in the past and whether countermeasures, such as maintenance procedures, have been identified based on past case information (analysis) 15. The level of urgency and countermeasures obtained by the fault analysis system 10 are stored in maintenance procedure information (analysis) 18. The information on maintenance procedures, etc. stored here is notified to a parts management system 30 and a maintenance personnel assignment system 40, and is used to stock necessary replacement parts and secure maintenance personnel to perform maintenance. The specific configuration and functions of the fault analysis system 10 will be described in detail later.

[0024] The storage device 20 is a group of storage devices operating at a customer site that receives services for the storage environment and its operation and maintenance. Although not specifically shown, the storage device 20 has the configurations and functions typically found in storage devices, such as a power supply unit, controller, bus, and drives. Each of the configurations and functions is equipped with sensors that monitor its operating status. If a malfunction occurs in the configuration or function, these sensors notify the fault analysis system 10 via a network adapter or the like of the storage device 20.

[0025] The parts management system 30 is a system operated by the service provider company for inventory management of maintenance parts for the storage device 20. As will be described in detail later, the parts management system 30 has at least a function for managing the inventory status of each maintenance part, and a function for allocating and delivering predetermined maintenance parts in response to instructions from a maintenance technician or an administrator of the service provider company.

[0026] Furthermore, the maintenance personnel assignment system 40 manages information on each maintenance personnel of the storage device 20, contract information with the customer, and the like, and allocates an appropriate maintenance personnel in response to an assignment request from the fault analysis system 10, and arranges for the maintenance personnel to be dispatched and respond. The maintenance personnel assignment system 40 is assumed to be for a server device operated by the above-mentioned service provider company or a maintenance personnel dispatch company that has been outsourced by the service provider company. Of course, the maintenance personnel assignment system 40 may also be operated by the customer who receives the above-mentioned service.

[0027] The maintenance staff terminal 50 is an operation terminal for a maintenance staff member who performs maintenance work at a customer site. The maintenance staff member who operates this maintenance staff terminal 50 is the target of allocation by the maintenance staff assignment system 40. Therefore, the maintenance staff terminal 50 receives a dispatch request associated with the allocation from the maintenance staff assignment system 40, displays that information, and prompts the maintenance staff member to respond. The maintenance staff terminal 50 is assumed to be a general terminal device such as a PC (Personal Computer), smartphone, or tablet terminal.

[0028] The maintenance technician operates the maintenance technician terminal 50, checks the appropriate information such as the maintenance procedure and travel route indicated in the dispatch request, and begins moving to the customer site. The maintenance technician also performs appropriate maintenance procedures such as part replacement at the customer site and reports the results to the fault analysis system 10. Note that there may be cases where a fault in the storage device 20 is difficult to resolve even by applying the maintenance content specified in the existing manual. In such cases, the maintenance technician assignment system 40 assigns a technician and requests a technician to be dispatched after the factory terminal 60 analyzes the fault and formulates a new maintenance procedure. The information such as the maintenance procedure formulated in the factory terminal 60 is used as information for assigning a technician in the maintenance technician assignment system 40, and is then notified by the maintenance technician assignment system 40 to the maintenance technician terminal 50.

[0029] The factory terminal 60 is a terminal in a factory where development, manufacturing, repair, etc. of the storage device 20 is carried out. This factory terminal 60 performs a fault analysis in response to a request from the fault analysis system 10, and formulates a new maintenance procedure based on the results. The factory terminal 60 responds with information on the new maintenance procedure, etc. to the fault analysis system 10 and the maintenance personnel assignment system 40 via the network N. The factory terminal 60 may also receive input of the results of the fault analysis and maintenance procedure formulation by a knowledgeable manager or the like employed at the factory, and respond to the fault analysis system 10 and the maintenance personnel assignment system 40.

[0030] Note that data exchange between the failure analysis system 10 and each of the storage device 20, component management system 30, maintenance staff assignment system 40, maintenance staff terminal 50, and factory terminal 60 may be performed according to, for example, an API (Application Programming Interface) protocol. In this case, it is assumed that each device is pre-implemented with the functions and configurations for executing each process of requests and responses by the API.

[0031] <Hardware configuration of the failure analysis system> 2 is a diagram showing an example of the hardware configuration of an information processing device 100 corresponding to the fault analysis system 10 in this embodiment. However, the information processing device 100 can also be considered as an information processing device corresponding to each device other than the fault analysis system 10 that constitutes the fault response system 1. The information processing device 100 is configured such that an auxiliary storage device 101, a main storage device 102, a processor 104, an input device 105, an output device 106, and a communication device 107 are communicatively connected by an appropriate interface such as an internal BUS.

[0032] Of these, the auxiliary storage device 101 is a storage means configured with non-volatile storage elements as already described, and in this embodiment, it stores at least the following information: past case information (analysis) 15, fault flash information (analysis) 16, storage status information (analysis) 17, maintenance procedure information (analysis) 18, and priority information 19, which will be described later in Fig. 3. Specific data configuration examples of these pieces of information will be described later.

[0033] The main memory device 102 also serves as a storage means for storing programs 103 including an OS (Operating System) that controls the entire information processing device 100 and various applications. The processor 104 loads and executes the programs 103 in the main memory device 102 to implement required functions.

[0034] The input device 105 is a means for accepting input operations by an administrator or the like of the fault analysis system 10, and can be assumed to be UI (User Interface) devices such as a keyboard, mouse, microphone, etc. The output device 106 is a means for outputting the results of information processing to an administrator or the like of the fault analysis system 10, and can be assumed to be UI (User Interface) devices such as a display, speaker, touch panel, printer, etc.

[0035] The communication device 107 is a communication means configured by a communication chip compatible with the protocol of the network N. The communication device 107 accesses the network N and is connected to other devices constituting the failure response system 1 via this network N so as to be able to communicate with each other.

[0036] <Functional configuration of the fault analysis system> Next, the functional configuration of the fault analysis system 10 will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the functional configuration of the fault analysis system 10 in this embodiment. The fault analysis system 10 has functional units, namely, an urgency determination unit 11, a determination unit 12, a transmission / reception unit 13, and an information storage unit 14.

[0037] Each of these functional units is implemented by the processor 104 executing the program 103. In order for each of these functional units to function effectively and perform the necessary processing, the unit also has the following information: past case information (analysis) 15, fault alert information (analysis) 16, storage status information (analysis) 17, maintenance procedure information (analysis) 18, and priority information 19.

[0038] Of these, the urgency determination unit 11 is a functional unit that determines the urgency of maintenance regarding a failure in the storage device 20, for example, based on the failure information received from the storage device 20, past case information (analysis) 15, and failure flash information (analysis) 16. Note that the past case information, failure flash information, maintenance procedure information, etc. are distinguished by adding a term according to the equipment held, such as adding the term (analysis) to the end if they are held by the failure analysis system 10, or adding the term (factory) to the end if they are held by the factory terminal 60 (the same applies below).

[0039] The determining unit 12 is a functional unit that determines the maintenance procedure to be applied to the failure and the parts required for the maintenance, based on the failure information received from the storage device 20, the past case information (analysis) 15, and the maintenance procedure information (analysis). In this case, the condition is that the failure information is not included in the failure flash information (analysis) 16.

[0040] The transmitting / receiving unit 13 is a functional unit that transmits and receives data to and from other systems and devices, such as receiving fault information from the storage device 20 and transmitting an analysis request including the fault information to the factory terminal 60. The transmitting / receiving unit 13 controls the communication device 107 to execute the data transmission and reception process.

[0041] The information accumulation unit 14 is a functional unit that stores data obtained from other systems and devices, such as fault information received from the storage device 20, and analysis results and maintenance procedure information received from the factory terminal 60, in the auxiliary storage device 101. The information accumulation unit 14 controls the auxiliary storage device 101 and the communication device 107 to execute processes such as storing the data.

[0042] <Functional configuration of storage device at customer site> Next, the functional configuration of the storage device 20 at the customer site will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the functional configuration of the storage device 20 in this embodiment. The storage device 20 has the following functional units: a failure detection unit 21, a transmission / reception unit 22, and an information storage unit 23.

[0043] Each of these functional units is implemented by executing a predetermined program on a computing chip or the like provided in the storage device 20. In addition, in order for each of these functional units to function effectively and perform the necessary processing, the storage device 20 also has user information 24, fault information 25, and storage status information (customer) 26.

[0044] Of these, the fault detection unit 21 is a functional unit that cooperates with, for example, sensors installed in the configuration and functions of the storage device 20 to detect faults that occur in the storage device 20. These sensors are, for example, sensors that observe the rotation speed, surface temperature, processing log, etc. of units such as the power supply unit, controller, and drive of the storage device 20, as well as of their functions. The fault detection unit 21 compares the observed values ​​obtained from the sensors with predetermined reference values ​​to detect abnormal situations such as a functional shutdown, a decrease in processing speed, or damage, i.e., the occurrence of a fault.

[0045] The transmitter / receiver 22 is a functional unit that exchanges data with other systems and devices, such as sending fault information to the fault analysis system 10 and sending an analysis request including the fault information to the factory terminal 60. The transmitter / receiver 13 controls the communication device 107 to execute the data exchange process.

[0046] The information storage unit 23 is a functional unit that stores the information obtained by the transmission / reception unit 22 in fault information 25 and storage status information (customer) 26. The information storage unit 23 controls the auxiliary storage device 101 and the communication device 107 to execute the storage process.

[0047] <Functional configuration of parts management system> Next, the functional configuration of the component management system 30 will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the functional configuration of the component management system 30 in this embodiment. The component management system 30 has the following functional units: a transmitter / receiver unit 31, a component allocation unit 32, and a delivery arrangement unit 33.

[0048] Each of these functional units is implemented by executing a predetermined program on a computing chip or the like provided in the parts management system 30. In addition, in order for each of these functional units to function effectively and perform the necessary processing, the system also has parts inventory information 34, base information 35, and delivery information 36.

[0049] Of these, the transmitting / receiving unit 31 is a functional unit that transmits and receives data with other systems and devices, such as receiving a parts arrangement request from the failure analysis system 10 and transmitting information such as schedule information for delivery of replacement parts from a base to a customer site to the maintenance staff assignment system 40. The transmitting / receiving unit 31 controls the communication device 107 to execute the data transmission and reception process.

[0050] The parts allocation unit 32 is a functional unit that checks the parts inventory status at its own base based on parts inventory information 34 and allocates replacement parts to be used in maintenance work. If the replacement parts are not included in the inventory at its own base during this allocation process, it refers to base information 35 and requests the parts management system 30 at the other base to allocate the replacement parts.

[0051] Furthermore, the delivery arrangement unit 33 is a functional unit that executes the delivery procedures for the replacement parts by referring to delivery information 36 in order to deliver the replacement parts allocated by the parts allocation unit 32 to the customer site. The delivery information 36 will be described in detail later, but it stores information on delivery companies and the like used in delivering parts.

[0052] <Functional configuration of the maintenance technician assignment system> Next, the functional configuration of the maintenance staff assignment system 40 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the functional configuration of the maintenance staff assignment system 40 in this embodiment. The maintenance staff assignment system 40 has the following functional units: a transmitter / receiver unit 41, a maintenance staff allocation unit 42, and a collection arrangement unit 43.

[0053] Each of these functional units is implemented by executing a predetermined program on a computing chip or the like provided in the maintenance personnel assignment system 40. In addition, in order for each of these functional units to function effectively and perform the necessary processing, the system also has maintenance personnel information 44, maintenance time information 45, and maintenance contract information 46.

[0054] Of these, the transmitting / receiving unit 41 is a functional unit that transmits and receives data with other systems and devices, such as receiving a request for arranging a maintenance worker including failure information from the failure analysis system 10, receiving information such as schedule information for delivery of replacement parts from the parts management system 30, and transmitting maintenance details and customer site information to the maintenance worker terminal 50 of the assigned maintenance worker. The transmitting / receiving unit 41 controls the communication device 107 to execute the data transmission and reception process.

[0055] Furthermore, the maintenance personnel allocation unit 42 is a functional unit that identifies a maintenance personnel who has the appropriate skills to handle failures, such as replacing parts, erasing data, or disposing of parts, as specified in the maintenance contract information 46, in a situation where time can be secured for the specified maintenance procedures, based on the type of failure indicated in the failure information, location information of the customer site, and the like, as well as maintenance personnel information 44, maintenance time information 45, and maintenance contract information 46, and assigns a maintenance personnel to the failure case.

[0056] The collection arrangement unit 43 is a functional unit that identifies an appropriate time for collection of the defective part after the work time indicated by the maintenance time information 45 has elapsed, and executes delivery procedures to have a delivery company head to the location of the customer site indicated by the maintenance contract information 46 to collect the defective part at that time. It is preferable that the maintenance staff assignment system 40 acquires information about the delivery company from the parts management system 30.

[0057] <Functional configuration of maintenance personnel terminal> Next, the functional configuration of the maintenance staff terminal 50 will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the functional configuration of the maintenance staff terminal 50 in this embodiment. The maintenance staff terminal 50 has at least a functional unit of a transmitting / receiving unit 51.

[0058] This transmitting / receiving unit 51 is implemented by executing a predetermined program on a computing chip or the like provided in the maintenance terminal 50. In addition, in order for each of these functional units to function effectively and perform the necessary processing, it also has at least information on schedule information 52.

[0059] The transmitting / receiving unit 51 is a functional unit that transmits and receives data to and from other systems and devices, such as receiving various information about the storage device to be maintained, the details of the failure, the maintenance procedure, and the customer from the maintenance staff assignment system 40, reporting the failure response to the maintenance staff assignment system 40, and transmitting maintenance staff schedule information 52. The transmitting / receiving unit 51 controls the communication device 107 to execute the data transmission and reception process.

[0060] <Factory terminal functional configuration> Next, the functional configuration of the factory terminal 60 will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of the functional configuration of the factory terminal 60 in this embodiment. The factory terminal 60 has functional units, namely, an analysis unit 61 and a transmission / reception unit 62.

[0061] Each of these functional units is implemented by executing a predetermined program on a computing chip or the like provided in the factory terminal 60. In addition, in order for each of these functional units to function effectively and perform the necessary processing, the factory terminal 60 also has the following information: maintenance procedure information (factory) 63, product information 64, fault alert information (factory) 65, past case information (factory) 66, and investigation information 67.

[0062] Furthermore, in response to an analysis request transmitted from the failure analysis system 10, the analysis unit 61 analyzes the failure information contained in the analysis request based on the maintenance procedure information (factory) 63, product information 64, failure flash information (factory) 65, past case information (factory) 66, and investigation information 67, thereby functioning as a functional unit that identifies the cause of the failure and the maintenance procedure. Details of this analysis will be described later.

[0063] The transmitter / receiver 62 is a functional unit that exchanges data with other systems and devices, such as receiving a maintenance worker arrangement request including failure information from the failure analysis system 10, receiving replacement part delivery schedule information from the parts management system 30, and sending maintenance details and customer site information to the maintenance worker terminal 50 of the assigned maintenance worker. The transmitter / receiver 41 controls the communication device 107 to execute the data exchange process.

[0064] <Example of information configuration> Next, a specific example of the configuration of information managed by each device in the failure response system 1 of this embodiment and used in various processes will be described. Fig. 9A is a diagram showing an example of the configuration of user information 24 in this embodiment. The user information 24 is information managed in the storage device 20, and indicates information about users of the storage device 20, i.e., customers. This user information 24 is a collection of records in which values ​​such as device IDs, addresses, and contact information are linked using a user ID as a key.

[0065] Among these, the user ID is identification information that uniquely identifies a customer. The device ID is identification information that uniquely identifies the storage device 20 used by the customer. The address is information that indicates the location of the customer. The contact information is information that indicates the telephone number and email address of the customer.

[0066] FIG. 9B is a diagram showing an example of the configuration of the fault information 25 in this embodiment. The fault information 25 stores information about faults that have occurred in the storage device 20. This fault information 25 is a collection of records in which values ​​such as a fault ID, occurrence time, and fault content are linked using a device ID as a key. Among these, the device ID is identification information that uniquely identifies the storage device 20 in which a fault has occurred. The fault ID is assigned to each fault incident that has occurred and is identification information that uniquely identifies the fault incident. The occurrence time is information that indicates the time when the fault occurred. Furthermore, the fault content is information that indicates the specific content of the fault, such as a power supply unit down or drive damage. As already mentioned, the occurrence and content of a fault are detected by sensors provided in the storage device 20, and the corresponding information is stored in the fault information 25 as the above-mentioned values.

[0067] FIG. 9C is a diagram showing an example of the configuration of storage status information (customer) 26 in this embodiment. The storage status information (customer) 26 is information that stores information such as the operating status of the storage device 20. This storage status information (customer) 26 is a collection of records that link values ​​such as a device type ID and operating status with a device ID as a key. Of these, the device ID is identification information that uniquely identifies the storage device 20. The device type ID is information that indicates the type of parts, units, functions, etc. in the storage device 20. The operating status indicates the status of the storage device 20, and values ​​such as down or ready to operate are stored.

[0068] FIG. 10A is a diagram showing an example of the configuration of past case information (analysis) 15 in this embodiment. The past case information (analysis) 15 is a database that stores and maintains information on failure cases that have occurred in the past in the storage device 20 at each customer site. This past case information (analysis) 15 is a collection of records in which values ​​such as a failure ID, a device type ID, and a maintenance procedure number are linked using a past case ID as a key. Of these, the past case ID is identification information that uniquely identifies a past failure occurrence case. The failure ID is identification information for each failure that has occurred. Since one or more failures can occur in one case, the value set in the failure ID column can be single or multiple. The device type ID is identification information that indicates the type of unit or function in the storage device 20 in which the failure occurred. The maintenance procedure number is identification information for the maintenance procedure performed for the failure in the past case.

[0069] FIG. 10B is a diagram showing an example of the configuration of the fault flash information (analysis) 16 in this embodiment. The fault flash information (analysis) 16 is fault information extracted from the factory terminal 60, and is information about faults not included in the past case information (analysis) 15, and is a database that stores fault information with a high level of urgency. This fault flash information (analysis) 16 is a collection of records in which the values ​​of the device type ID and the fault ID are linked using the flash number as a key. Of these, the flash number is identification information that uniquely identifies the fault flash. The device type ID is identification information that uniquely identifies the unit or function of the storage device 20 in which the fault occurred. The fault ID is identification information that uniquely identifies the fault that has occurred.

[0070] FIG. 10C is a diagram showing an example of the configuration of storage status information (analysis) 17 in this embodiment. The storage status information (analysis) 17 is a database that collects, stores, and maintains status information of the storage device 20 at each customer site. This fault report information (analysis) 16 is a collection of records in which values ​​such as device ID, operating status, device type ID, address, and contact information are linked using a user ID as a key. Of these, the user ID is identification information that uniquely identifies a customer. The device ID is identification information that uniquely identifies a storage device 20. The operating status is identification information that uniquely identifies the operating status of the storage device 20. The device type ID is identification information that uniquely identifies the type of unit or function of the storage device 20. The address is a value that indicates the location of the customer, and the contact information is a value such as the customer's telephone number or email address.

[0071] FIG. 11A is a diagram showing an example of the configuration of maintenance procedure information (analysis) 18 in this embodiment. Maintenance procedure information (analysis) 18 is a database that stores and holds information on maintenance procedures obtained from the factory terminal 60. This maintenance procedure information (analysis) 18 is a collection of records in which values ​​such as maintenance procedures and part IDs are linked together using the maintenance procedure number as a key. Of these, the maintenance procedure number is identification information that uniquely identifies the maintenance procedure. Furthermore, the maintenance procedure is information that indicates the corresponding part in the manual in which the maintenance procedure is described. Furthermore, the part ID is identification information for a replacement part to be installed at a faulty location in the maintenance procedure.

[0072] FIG. 11B is a diagram showing an example of the configuration of priority information 19 in this embodiment. The priority information 19 is a table that defines values ​​based on the logic that, among units and functions of the storage device 20, those that have a greater impact when a failure occurs have a higher priority as a replacement target. This priority information 19 is a collection of records that link together device type IDs and priority values ​​using a failure ID as a key. Of these, the failure ID is identification information that uniquely identifies a failure. Furthermore, the device type ID is identification information that uniquely identifies the unit or function where the failure occurred. Furthermore, the priority is identification information that uniquely identifies the priority of part replacement when the failure occurs in that unit, etc.

[0073] FIG. 12A is a diagram showing an example of the configuration of maintenance personnel information 44 in this embodiment. The maintenance personnel information 44 is a database that stores and holds information about each maintenance personnel held by the maintenance personnel assignment system 40 for the purpose of assigning a maintenance personnel. This maintenance personnel information 44 is a collection of records that link together values ​​such as a maintenance personnel ID, name, work location, schedule, current location, available device type ID, and contact information. Of these, the maintenance personnel ID is identification information that uniquely identifies a maintenance personnel. The schedule is information that indicates the maintenance personnel's planned arrival and departure times and information about customer sites to be visited. The current location is current location information of the maintenance personnel that is acquired at regular intervals from the maintenance personnel terminal 50. The available device type ID is identification information that uniquely identifies the types of units and functions that the maintenance personnel can maintain.

[0074] 12B is a diagram showing an example of the configuration of a maintenance time report 45 in this embodiment. The maintenance time information 45 is a table that stores and holds information about the work time required for each maintenance procedure. This maintenance time information 45 is a collection of records that associate work time values ​​with maintenance procedure numbers. This work time can be assumed to be, for example, the average work time required when a certain number of maintenance personnel or more perform the target maintenance procedure.

[0075] FIG. 12C is a diagram showing an example of the configuration of maintenance contract information 46 in this embodiment. The maintenance contract information 46 is a database that stores and holds the details of the contract for customers who have concluded a service contract for the management and operation of a storage environment. This maintenance contract information 46 is a collection of records in which values ​​such as the maintenance contract, device ID, and contract date are linked using a user ID as a key. Of these, the user ID is identification information that uniquely identifies the customer. The maintenance contract is identification information that uniquely identifies the maintenance contract. This maintenance contract information specifies the conditions for performing maintenance procedures and the contents of the maintenance procedures (e.g., whether or not parts need to be replaced, whether or not the failed parts need to be destroyed or collected, etc.) for units and functions where a failure has occurred. The device ID is identification information that uniquely identifies the storage device 20 that is the subject of the maintenance contract.

[0076] FIG. 13A is a diagram showing an example of the configuration of parts inventory information 34 in this embodiment. Parts inventory information 34 is a database in which the parts management system 30 stores and maintains information indicating the inventory status of each part. This parts inventory information 34 is a collection of records in which values ​​such as part ID, inventory quantity, and minimum inventory quantity are linked together using a device type ID as a key. Of these, the device type ID is identification information that uniquely identifies the type of unit, etc. subject to inventory management. The part ID is identification information that uniquely identifies the part to be replaced in the unit, etc. The inventory quantity is information that indicates the inventory quantity of the part, and the minimum inventory quantity is information that indicates the inventory quantity required to maintain the unit, etc. without problems.

[0077] 13B is a diagram showing an example of the configuration of the base information 35 in this embodiment. The base information 35 is a table that stores and holds information about each base that stocks each part required for maintenance procedures. This base information 35 is a collection of records that link together values ​​such as a base ID and a location.

[0078] 13C is a diagram showing an example of the configuration of delivery information 36 in this embodiment. Delivery information 36 is a database that stores and holds information about delivery companies that ship parts from a base and deliver them to a customer site, or that collect faulty parts from a customer site and return them to a factory. This delivery information 36 is a collection of records that link together values ​​such as the name, business hours, and contact information of the delivery company.

[0079] FIG. 14 is a diagram showing an example of the configuration of schedule information 52 in this embodiment. The schedule information 52 is a database that stores and holds information about the work schedule of the maintenance technician in the maintenance technician terminal 50. This schedule information 52 is a collection of records that link together each value in the work schedule, such as the date, start time, end time, user ID, maintenance procedure number, and device ID. Of these, the user ID is identification information that uniquely identifies the customer who will perform the maintenance procedure. The maintenance procedure number is identification information that uniquely identifies the maintenance procedure to be performed in the storage device 20 of the customer. Furthermore, the device ID is identification information that uniquely identifies the storage device 20 that is the target of the maintenance procedure number.

[0080] Fig. 15A is a diagram showing an example of the configuration of maintenance procedure information (factory) 63 in this embodiment. Maintenance procedure information (factory) 63 is a database that stores and holds information about maintenance procedures formulated in the factory terminal 60. This maintenance procedure information (factory) 63 is a collection of records that link values ​​such as maintenance procedures and part IDs using a maintenance procedure number as a key. This maintenance procedure information (factory) 63 has the same data configuration as the maintenance procedure information (analysis) 18 shown in Fig. 11A, so a description thereof will be omitted.

[0081] FIG. 15B is a diagram showing an example of the configuration of product information 64 in this embodiment. Product information 64 is a database that stores and holds information about products developed, manufactured, and maintained in the factory. This product information 64 is a collection of records that link together values ​​such as a device type ID, product specifications, hardware configuration, software configuration, and manual. Of these, product specifications are specifications defined for a unit, etc. The hardware configuration is information about the hardware used in the unit, etc., and the software configuration is information about the software used in the unit, etc. Furthermore, the manual is identification information that uniquely identifies the manual, which describes usage instructions and maintenance procedures.

[0082] 15C is a diagram showing an example of the configuration of fault flash information (factory) 65 in this embodiment. The fault flash information (factory) 65 is information about a fault for which a maintenance procedure has not been determined, and has the same configuration as the fault flash information (analysis) 16 managed by the fault analysis system 10, so a description thereof will be omitted.

[0083] 16A is a diagram showing an example of the configuration of past case information (factory) 66 in this embodiment. Past case information (factory) 66 is a database that stores and holds information about failures that have occurred in the past and the maintenance procedures that have been applied to those failures. This past case information (factory) 66 has the same configuration as past case information (analysis) 15 managed by the failure analysis system 10, so a description thereof will be omitted.

[0084] FIG. 16B is a diagram showing an example of the configuration of investigation information 67 in this embodiment. The investigation information 67 is a database that stores and holds the results of analysis (investigation) of failed parts. This investigation information 67 is a collection of records in which values ​​such as the target part, investigation status, investigation result, user ID, device ID, and device type ID are linked using the investigation number as a key. Of these, the investigation status is the status of the failure analysis, and the investigation result is information such as the cause of the failure obtained by the failure analysis. The user ID is identification information that uniquely identifies the customer who is the user of the storage device 20. The device ID is identification information that uniquely identifies the storage device 20 that is the target of the failure analysis. The device type ID is identification information that uniquely identifies the unit, etc. that is the target of the failure analysis.

[0085] <Concept of troubleshooting> Here, we will outline how the company providing the above service responds when a failure occurs in a storage device. When a failure of some kind occurs in a storage device, sensors installed in the storage device notify the failure occurrence via the storage device 20 to a failure analysis system 10. This failure analysis system 10 makes the necessary arrangements for maintenance personnel, parts, etc. depending on the content of the notification.

[0086] Based on the above-mentioned report from the storage device 20 (at the customer site), the failure analysis system 10 arranges for the dispatch of a maintenance technician with the maintenance technician assignment system 40. A predetermined notification is sent from the maintenance technician assignment system 40 to the maintenance technician terminal 50 of the dispatched maintenance technician. The maintenance technician confirms this notification, goes to the customer site, and performs maintenance such as replacing parts in the storage device 20 that has the failure, following the maintenance procedures in the manual.

[0087] On the other hand, if the failure is difficult to handle using the maintenance procedures specified in the existing manual, the failure analysis system 10 that detects this will send an analysis request to the factory terminal 60, including a request to that effect. As already mentioned, the factory terminal 60 will then investigate the cause of the failure and formulate new maintenance procedures. The factory department is responsible for the development, manufacturing, and repair of the storage device 20, and has advanced knowledge and various analytical know-how. Therefore, it is possible to analyze the cause of new failure events that are not described in the manual and derive the necessary countermeasures. Therefore, it is also possible to manually identify the cause of the failure and formulate maintenance procedures, and then use these.

[0088] Replacement parts used by maintenance personnel at customer sites are managed in inventory at regional bases, for example. The parts management system 30 receives notification of the occurrence of the above-mentioned failure from the failure analysis system 10, and based on this notification, secures inventory of the target parts and arranges for delivery from the base nearest to the customer to the customer site.

[0089] After a maintenance technician replaces the faulty part with a working part on-site, the faulty part is collected and sent to the factory for inspection. Therefore, the maintenance technician assignment system 40 makes arrangements for the collection and delivery of such faulty parts. The faulty part collected at the factory is inspected, and the cause and details of the failure are notified to the maintenance technician terminal 50 and reported to the customer site. Furthermore, for data storage areas in the storage device 20, data is erased at the factory or destroyed on-site, for example, based on the maintenance contract. Parts destroyed on-site or parts that cannot be transported (e.g., lithium-ion batteries) are disposed of by a designated maintenance department.

[0090] <Troubleshooting method: Main flow> Next, the processing flow of the failure response method of this embodiment will be explained. Fig. 17 is a diagram showing an example flow (part 1) of the failure response method of this embodiment. Here, it is assumed that a failure has occurred at a certain customer site, and that this fact has been notified to the failure analysis system 10 from the storage device 20 or its monitoring device, etc. (S1). Therefore, the fault analysis system 10 determines the urgency of the fault (S2). This determination is made based on predetermined criteria, such as whether or not information about the fault has already been stored as a past case. Details of this will be described later based on the flow in FIG. 19.

[0091] If the result of the above determination is that the urgency of the failure is low (S2: low), the failure analysis system 10 requests the parts management system 30 to secure and replenish (additional order) replacement parts required for maintenance (S3).The failure analysis system 10 also notifies the maintenance personnel assignment system 40 of a request to assign a maintenance personnel, including the failure information, etc., and secures a maintenance personnel to deal with the failure (S4).

[0092] The maintenance personnel assignment system 40 determines whether to collect the faulty part to be replaced and arranges for the delivery company required for its delivery (S5). The maintenance personnel assignment system 40 also notifies the maintenance personnel terminal 50 of the assigned maintenance personnel of information regarding the maintenance procedure and handling of the faulty part (S6). The maintenance personnel performs maintenance work at the customer site based on this information. The maintenance personnel terminal 50 also notifies the failure analysis system 10 of the results (S7). The maintenance personnel collects the faulty part from the storage device 20 at the customer site and carries out the collection procedure by, for example, handing it over to the delivery company.

[0093] On the other hand, if the result of the determination in S2 above indicates that the urgency of the failure is high (S2: high), the failure analysis system 10 sends an emergency report to the factory terminal 60 (S8). This emergency report indicates that the failure that has occurred in the storage device 20 at the customer site requires an emergency, and the maintenance procedure has not yet been determined.

[0094] Meanwhile, the factory terminal 60 executes a process of analyzing the fault and formulating a maintenance procedure in response to the emergency report received from the fault analysis system 10 (S9). The fault analysis system 10 then notifies the maintenance personnel assignment system 40 of information such as the formulated maintenance procedure (S10). The factory terminal 60 then requests the parts management system 30 to secure and replenish the maintenance parts indicated in the information such as the formulated maintenance procedure (S11). In response to this request, the parts management system 30 executes predetermined processes for securing and ordering parts, taking into account the parts inventory status at each base.

[0095] <About the information update flow> 18A to 18C are assumed as cases of updating and registering maintenance procedures and fault alerts in the fault analysis system 10. Therefore, the flow of information registration for each case will be explained here. 18A, it is assumed that the factory terminal 60 receives a high-urgency notification from the storage device 20 (S15). In this case, the factory terminal 60 executes the processes of fault analysis and maintenance procedure formulation based on the fault information etc. indicated in the notification (S16). The factory terminal 60 also transmits the fault analysis results and maintenance procedure information obtained in S16 to the fault analysis system 10 (S17). Meanwhile, the fault analysis system 10 receives the information transmitted from the factory terminal 60, stores this in past case information (analysis) 15 and maintenance procedure information (analysis) 18 (S18), and terminates the process.

[0096] 18B shows the process in a situation where a new product has been released at the factory (S20). In this case, the factory terminal 60 formulates a new maintenance procedure for the new product (S21). The factory terminal 60 also transmits information about the formulated maintenance procedure to the fault analysis system 10 (S22). Meanwhile, the fault analysis system 10 receives the information about the maintenance procedure transmitted from the factory terminal 60, stores it in maintenance procedure information (analysis) 18 (S23), and ends the process.

[0097] 18C shows the processing in a situation where a product defect is discovered in a factory (S24). In this case, the factory terminal 60 creates a fault flash report regarding the product defect (S25). This case is based on the assumption that the processing is being performed at a time when maintenance procedures have not yet been established. The factory terminal 60 transmits the fault flash report to the fault analysis system 10 (S26). Meanwhile, the fault analysis system 10 receives the fault flash report transmitted from the factory terminal 60, stores it in the fault flash information (analysis) 16 (S28), and ends the processing.

[0098] <Urgency determination flow> 19 is a diagram showing a flow example (part 5) of the fault handling method of this embodiment, specifically showing the flow of urgency determination. Here, the details of the urgency determination process (S2) by the urgency determination unit 11 will be explained. In this case, the urgency determination unit 11 of the fault analysis system 10 searches for fault information in the fault flash report information (analysis) 16, and determines whether a corresponding incident has been registered (S31).

[0099] If the result of the above determination is that the fault information is found in the fault flash information (analysis) 16 (S31: Y), the urgency determination unit 11 determines that the urgency of the fault is high (S39) and ends the process. On the other hand, if the result of the above determination is that the fault information is not found in the fault flash information (analysis) 16 (S31: N), the urgency determination unit 11 determines whether the fault information includes multiple fault IDs, that is, whether multiple faults are occurring simultaneously (S32).

[0100] If the result of the above determination is that there are not multiple failure IDs (S32: N), the urgency determination unit 11 searches for the failure information in the past case information (analysis) 15 and determines whether a corresponding incident has been registered (S33).If the result of this determination is that the failure information is found in the past case information (analysis) 15 (S33: Y), the urgency determination unit 11 determines that the urgency of the failure is low (S36) and ends the process.

[0101] On the other hand, if the result of the determination in S33 above is that the fault information is not found in the past case information (analysis) 15 (S33: N), the urgency determination unit 11 determines that the urgency of the fault is high (S34) and ends the process. Also, if the result of the determination in S32 above is that there are multiple fault IDs (S32: Y), the urgency determination unit 11 searches the past case information (analysis) 15 for past cases that include all of the multiple fault IDs (S35).

[0102] If the search results in a hit for a past case that includes all of the multiple failure IDs (S35: Y), the urgency determination unit 11 determines that the urgency of the failure is low (S36) and terminates the process. On the other hand, if the search results in no hit for a past case that includes all of the multiple failure IDs (S35: N), the urgency determination unit 11 searches the past case information (analysis) 15 to determine whether a past case that includes each of the multiple failure IDs exists for each failure ID (S37).

[0103] If the search result shows that no past cases each including a plurality of failure IDs exist for each failure ID (S37: N), the urgency determination unit 11 determines that the urgency of the failure is high (S39) and ends the process. On the other hand, if the determination result shows that a past case each including a plurality of failure IDs exists for each failure ID (S37: Y), the urgency determination unit 11 determines whether the same priority exists among the plurality of failure IDs, that is, whether overlapping of priorities occurs (S38).

[0104] If the result of the above determination shows that there is overlap in priority between multiple failure IDs (S38: Y), the urgency determination unit 11 determines that the urgency of the failure is high (S39) and ends the process. On the other hand, if the result of the above determination shows that there is no overlap in priority between multiple failure IDs (S38: N), the urgency determination unit 11 determines that the urgency of the failure is low (S36) and ends the process.

[0105] <Example of a fault response flow (low urgency)> FIG. 20 is a diagram showing a flow example (part 6) of the failure response method of this embodiment, specifically showing a flow related to maintenance response when the urgency of the failure is low. In this case, it is assumed that a certain failure has occurred in the storage device 20 at the customer site in this embodiment. In this case, the storage device 20 at the customer site notifies the failure analysis system 10 of the failure observed by sensors and information related to the failure and its target, as well as the customer's identification information, as failure information (S40). Upon receiving this notification, the failure analysis system 10 analyzes the failure information (S41). In this analysis, it is assumed that the urgency determination process is performed by the urgency determination unit 11 described above. As a result of this analysis, it is assumed that the urgency of the failure is determined to be low.

[0106] The fault analysis system 10 notifies the maintenance personnel assignment system 40 of a request to assign a maintenance personnel, including the following information that has been determined based on the fault information: the customer's user ID, address, contact information, device ID (identification information of the storage device 20), fault content (identified from values ​​observed by sensors), and maintenance procedure (known maintenance procedure obtained by searching the fault information in past case information (analysis) 15) (S42).

[0107] Furthermore, the failure analysis system 10 notifies the component management system 30 of a request to secure component inventory, including information on the customer and the faulty component to be replaced as indicated by the maintenance procedure (S43). The component management system 30 receives this request (S44) and checks the inventory of the target component in the component inventory information 34 (S45). If the component management system 30 is unable to secure the necessary inventory at its own base, it refers to the base information 35 and requests component management systems 30 at other bases to secure inventory (S46).

[0108] Based on the base information 35 and delivery information 36, the parts management system 30 uses a delivery company to calculate the time required to ship the target part from a specific base to the customer site (where the failure has occurred), for example, using an external route search application. In other words, if the time required to deliver the part is calculated and delivery is to begin at a specific time, the estimated time of delivery completion can also be determined. Therefore, the parts management system 30 notifies the maintenance technician assignment system 40 of the information on the estimated time of delivery completion (S47).

[0109] Meanwhile, the maintenance staff assignment system 40 receives a request for assigning a maintenance staff, including the fault information, from the fault analysis system 10, and also receives information on the estimated delivery completion time of the target part from the part management system 30 (S48). The maintenance staff assignment system 40 selects a maintenance staff to assign to the fault based on the information obtained in S48, the maintenance staff information 44, and the maintenance time information 45 (S49).

[0110] This selection is performed, for example, by identifying a person who can travel to the target customer site within a certain time from the perspective of the current location and schedule and who can perform maintenance work on the unit indicated by the applicable device type ID, based on the values ​​of the work location, schedule, current location, and applicable device type ID indicated in the maintenance personnel information 44. Also, from the perspective of the schedule, a person is selected who can secure the work time required to perform the target maintenance procedure indicated in the maintenance time information 45, even taking into account the travel time from the current location to the customer site.

[0111] Furthermore, the maintenance personnel assignment system 40 notifies the maintenance personnel terminal 50 of the maintenance personnel of information such as the maintenance personnel identified up to this point, the maintenance procedures to be performed by the maintenance personnel, and the destination customer site (S50). Meanwhile, the maintenance personnel terminal 50 receives this notification (S51) and displays it on a display or the like for the maintenance personnel's confirmation. The maintenance personnel will move to the target customer site and perform the predetermined maintenance procedures there.

[0112] <Example of a fault response (urgency: high) flow> FIG. 21 is a diagram showing a flow example (part 7) of the failure response method of this embodiment, specifically showing a flow related to maintenance response when the urgency of the failure is high. In this case, it is assumed that a certain failure has occurred in the storage device 20 at the customer site in this embodiment. In this case, the storage device 20 at the customer site notifies the failure analysis system 10 of the failure observed by sensors and information related to the failure and its target, as well as the customer's identification information, as failure information (S40). Upon receiving this notification, the failure analysis system 10 analyzes the failure information (S41). In this analysis, it is assumed that the urgency determination process is performed by the urgency determination unit 11 described above. As a result of this analysis, it is assumed that the urgency of the failure is determined to be high.

[0113] The fault analysis system 10 notifies the factory terminal 60 of a request for fault analysis, including the information that has been determined based on the fault information, such as the customer's user ID, address, contact information, device ID (identification information of the storage device 20), and fault details (determined from values ​​observed by sensors) (S42).

[0114] Meanwhile, the factory terminal 60 receives the request from the fault analysis system 10 and executes fault analysis and maintenance procedure development (S43). This fault analysis and maintenance procedure development may be performed by an application such as a solver installed in the factory terminal 60, or the results may be obtained using an external service. The solver may be, for example, a judgment model that takes fault information as input and outputs information on the cause of the fault and the associated maintenance procedure. Such a judgment model may be generated by machine learning using sets of numerous fault information items examined and developed by knowledgeable personnel and information on the cause of the fault and the associated maintenance procedure as training data.

[0115] The factory terminal 60 notifies the parts management system 30 of a parts inventory request, including the parts to be replaced and the customer's user ID, indicated by the contents of the maintenance procedure obtained in S43 (S44). The parts management system 30 receives this inventory request notification (S48) and checks the inventory of the target parts in the parts inventory information 34 (S49). At this time, if the necessary inventory cannot be secured at its own base, the parts management system 30 refers to the base information 35 and requests parts management systems 30 at other bases to secure inventory (S50).

[0116] Based on the base information 35 and delivery information 36, the parts management system 30 uses a delivery company to calculate the time required to ship the target part from a specific base to the customer site (where the failure has occurred), for example, using an external route search application. In other words, if the time required to deliver the part is calculated and delivery is to begin at a specific time, the estimated time when delivery will be completed can also be determined. Therefore, the parts management system 30 notifies the maintenance technician assignment system 40 of the information on the estimated time when delivery will be completed (S51).

[0117] Meanwhile, the factory terminal 60 notifies the maintenance personnel assignment system 40 of a request for assigning a maintenance personnel, including the information on the maintenance procedure formulated in S43 and the fault information (sent in S42) (S45). The factory terminal 60 also sends the results of the fault analysis and the information on the maintenance procedure obtained in S43 to the fault analysis system 10 (S46). The fault analysis system 10 receives this information and stores it in past case information (analysis) 15 and maintenance procedure information (analysis) 18 (S48).

[0118] Meanwhile, the maintenance personnel assignment system 40 receives a request for assigning a maintenance personnel, including the failure information, from the failure analysis system 10, and also receives information on the estimated delivery completion time of the target part from the part management system 30 (S52). The maintenance personnel assignment system 40 selects a maintenance personnel to assign to the failure based on the information obtained in S52, the maintenance personnel information 44, and the maintenance time information 45 (S53).

[0119] This selection is performed, for example, by identifying a person who can travel to the target customer site within a certain time from the perspective of the current location and schedule and who can perform maintenance work on the unit indicated by the applicable device type ID, based on the values ​​of the work location, schedule, current location, and applicable device type ID indicated in the maintenance personnel information 44. Also, from the perspective of the schedule, a person is selected who can secure the work time required to perform the target maintenance procedure indicated in the maintenance time information 45, even taking into account the travel time from the current location to the customer site.

[0120] Furthermore, the maintenance personnel assignment system 40 notifies the maintenance personnel terminal 50 of the maintenance personnel of information such as the maintenance personnel identified up to this point, the maintenance procedures to be performed by the maintenance personnel, and the destination customer site (S54). Meanwhile, the maintenance personnel terminal 50 receives this notification (S55) and displays it on a display or the like for the maintenance personnel's confirmation. The maintenance personnel will move to the target customer site and perform the predetermined maintenance procedures there.

[0121] <Example of ordering flow for defective parts> 22 is a diagram showing an example flow (part 8) of the failure response method of this embodiment, specifically, a flow relating to ordering a faulty part and shipping a replacement part. In this case, for example, in the flow of FIG. 20 or 21, it is assumed that, following the step (S47 or S51) in which the part management system 30 notifies the maintenance staff assignment system 40 of information on the scheduled time of delivery completion, the maintenance staff assignment system 40 makes arrangements with a waste disposal company or the like to collect or dispose of the faulty part (S60). Such collection or disposal is applied in cases where the contractual contents defined in the maintenance contract information 46 require that the unit in question be recalled or disposed of in order to avoid risks such as data leakage in the storage device 20.

[0122] In this case, the maintenance technician who has performed a series of maintenance procedures at the customer site will be present until the disposal contractor arranged by the maintenance technician assignment system 40 completes the collection of the target parts, etc. Meanwhile, the disposal contractor will package the collected parts and deliver them to the factory. In response to this, the maintenance technician terminal 50 executes a process to report the completion of the fault response to the maintenance technician assignment system 40 in accordance with the maintenance technician's input (S61). Upon receiving this report, the maintenance technician assignment system 40 requests the fault analysis system 10 to send a fault log (S62).

[0123] The failure analysis system 10 transmits a detailed failure log to the factory terminal 60 (S63). The factory that receives the failed part performs a failure analysis of the failed part based on the detailed failure log and the failed part. At this time, the factory erases data from the failed part (e.g., a disk) in accordance with the provisions of the maintenance contract information 46. The factory terminal 60 transmits a customer explanation document (e.g., a document describing the cause and history of the failure) that is the result of this series of processes to the maintenance worker terminal 50 (S64). The factory terminal 60 also transmits a data erasure certificate to the customer's specified terminal, etc. (S65).

[0124] Meanwhile, the parts management system 30 executes an ordering process on the factory terminal 60 for a replacement part to replace the failed part (S66). The parts management system 30 also checks the inventory status of the replacement part at its own base using inventory information 34, and if the required inventory cannot be secured, it requests that another base secure inventory (S67). Upon receiving the order in S66, the factory delivers the target replacement part to the base of the parts management system 30.

[0125] <Example of defective parts collection flow> FIG. 23 is a diagram showing a flow example (part 9) of the failure response method of this embodiment, specifically, a flow relating to the collection and disposal of failed parts. In this case, only the differences from the flow example of FIG. 22 will be explained. Here, the failure analysis system 10 recognizes the need to dismantle and destroy the failed parts based on the maintenance contract information 46 with the customer, and notifies the factory terminal 60 of a request to do so (S72). Upon receiving this notification, the factory dismantles and destroys the failed parts that have already been received (S73). The remains of the dismantled and destroyed failed parts are then discarded at the factory, and the process ends.

[0126] Although one embodiment of the present invention has been described above, this is merely an example for explaining the present invention, and the scope of the present invention is not limited to this embodiment. The present invention can be implemented in various other forms.

[0127] The above description can be summarized as follows: The following summary may include supplementary explanations and explanations of variations of the above description.

[0128] In the fault response system of this embodiment, the storage device may further store fault alert information regarding faults that have been reported but for which a response has not yet been decided, and when determining the level of urgency, the computing device may compare the fault information reported regarding a new fault with the fault alert information, and if the fault information is included in the fault alert information, determine that the urgency of the new fault is high, and if, as a result of the comparison, the fault information is not included in the fault alert information and is not included in the response example, determine that the urgency of the new fault is high.

[0129] This allows for efficient determination of urgency by broadly classifying the urgency into high and low depending on whether or not a fault is included in the fault flash report information, which in turn allows for more accurate and efficient handling of high-urgency faults in storage devices.

[0130] Furthermore, in the fault response system of this embodiment, when determining the level of urgency, the computing device may compare the fault information reported regarding a new fault with the fault report information, and if the fault information is not included in the fault report information and the fault information indicates multiple faults, determine whether a handling example exists that includes all of the multiple faults at the same time, and if it is found that a handling example exists that includes all of the multiple faults at the same time, determine that the urgency of the new fault is low.

[0131] This allows efficient identification of low-urgency failures, and ultimately allows efficient and more accurate handling of high-urgency failures in storage devices.

[0132] Furthermore, in the failure response system of this embodiment, when determining the level of urgency, the computing device may determine whether a handling example exists that includes all of the multiple failures at the same time, and if it is found that no handling example includes all of the multiple failures at the same time, it may determine whether each of the multiple failures is included in any of the handling examples, and if it is found that there is one of the multiple failures that is not included in any of the handling examples, it may determine that the level of urgency of the new failure is high.

[0133] This makes it possible to accurately identify and deal with cases that require careful handling and require fault analysis in the factory, without any cases where multiple failures as described above have occurred simultaneously, and without any cases where any of these failures have occurred independently, i.e., where the priorities and mutual influences between the failures are unclear. Ultimately, it becomes possible to efficiently deal with more accurate responses even for failures in storage devices that require high urgency.

[0134] Furthermore, in the failure response system of this embodiment, the storage device further stores priority information that defines, for each part of the monitored device, the degree of impact on the device in the event of a failure as a priority, and the computing device, when determining the level of urgency, determines whether each of the multiple failures is included in any of the response cases, and if it is found that each of the multiple failures is included in any of the response cases, identifies the priority of each of the multiple failures using the priority information, and determines whether there are any failures with the same priority among the multiple failures, and if it is found as a result of the determination that there are any failures with the same priority among the multiple failures, determines that the urgency of the new failure is high.

[0135] This makes it possible to identify and carefully handle a case in which the priorities of the multiple (simultaneous) failures mentioned above have all occurred in the past but have not occurred simultaneously and have overlapping priorities, as a complex case in which it is not easy to determine the priorities between the failures, i.e., it is not possible to replace parts in an order according to the priorities. As a result, it becomes possible to efficiently take more appropriate measures even for failures in storage devices that have a high degree of urgency.

[0136] Furthermore, in the failure response system of this embodiment, the arithmetic device executes a process of notifying a maintenance personnel assignment system of information on a maintenance procedure corresponding to a response case that is suitable for the new failure, regarding the new failure that is found to have a low urgency, and a process of notifying a parts arrangement request to a parts management system that manages inventory information of parts for the monitored device, and the parts management system executes a process of checking the inventory of the parts and notifying the assignment system of delivery schedule information for delivering the parts to the location of the monitored device, and the assignment system holds information on each maintenance personnel, receives the maintenance procedure information and the delivery schedule information, identifies a maintenance personnel who is capable of performing the maintenance procedure and has a schedule that allows them to perform failure response work at the location, and executes a process of notifying the maintenance procedure and the parts delivery schedule information to the terminal of the maintenance personnel.

[0137] This allows for the simultaneous assignment of maintenance personnel and the securing of parts inventory when the failure is of low urgency, further improving operational efficiency, and ultimately enabling the efficient and accurate handling of storage device failures of high urgency.

[0138] Furthermore, in the failure response system of this embodiment, the arithmetic device notifies the factory system of the analysis request including failure information of the new failure regarding the new failure that is found to have a high level of urgency, and in response thereto, receives from the factory system information on a new maintenance procedure that is an analysis result regarding the new failure, associates the information with the new failure, and stores the information in the storage device as one of the response cases, and the factory system notifies the parts management system of the arrangement request including information on parts that need to be replaced, which are indicated by a maintenance procedure based on the failure analysis performed in the factory in response to the analysis request. and executes a process of notifying the assignment system of information about the maintenance procedure, and the parts management system executes a process of checking the inventory of the parts in response to the arrangement request and notifying the assignment system of delivery schedule information for when the parts are to be delivered to the location of the device to be monitored, and the assignment system receives the information about the maintenance procedure and the delivery schedule information, identifies a maintenance worker who is able to perform the maintenance procedure and has a schedule that allows them to perform work to respond to the failure at the location, and executes a process of notifying the maintenance worker's terminal of the maintenance procedure and the delivery schedule information for the parts.

[0139] This allows for automatic execution of an analysis request to the factory when the failure is deemed urgent, and streamlines the series of processes involved in arranging replacement parts and assigning maintenance personnel according to the maintenance procedures indicated by the analysis results, thereby enabling more accurate and efficient handling of urgent storage device failures. [Explanation of symbols]

[0140] N Network 1. Troubleshooting system 10. Fault Analysis System 11 Urgency Judgment Department 12 Judgment section 13 Transmitter / Receiver 14 Information storage unit 15 Past case information (analysis) 16 Breaking News (Analysis) 17 Storage status information (analysis) 18 Maintenance procedure information (analysis) 19 Priority information 20 Storage Devices 21 Fault detection unit 22 Transmitter / Receiver 23 Information storage unit 24 User Information 25. Fault Information 26 Storage Status Information (Customer) 30 Parts Management System 31 Transmitter / Receiver 32 Parts allocation department 33 Delivery Arrangements Department 34 Parts inventory information 35 Base Information 36 Shipping Information 40 Maintenance Personnel Assignment System 41 Transmitter / Receiver 42 Maintenance Personnel Allocation Department 43 Collection Arrangements Department 44 Maintenance personnel information 45 Maintenance time information 46 Maintenance contract information 50 Maintenance personnel terminal 51 Transmitter / Receiver 52 Schedule Information 60 Factory Terminal 61 Analysis Department 62 Transmitter / Receiver 63 Maintenance Procedure Information (Factory) 64 Product Information 65 Breaking News (Factory) 66 Past Case Information (Factory) 67 Survey Information 100 Information processing device 101 Auxiliary storage 102 Main storage 103 Programs 104 Arithmetic equipment 105 Input Device 106 Output Device 107 Communication equipment

Claims

1. a storage device that stores information on cases of dealing with failures in the monitored device; a computing device that executes a process of determining whether or not there is a countermeasure case that matches the new failure in terms of urgency, based on failure information reported regarding a new failure in the monitored device and information on the countermeasure case; a process of making a request to a factory system that manufactures or repairs the monitored device to analyze the details of the countermeasure case for the new failure, if it is determined as a result of the determination that the urgency of the new failure is high; and a process of making a request to a predetermined system to arrange for maintenance personnel and parts, according to the countermeasure case that matches the new failure, if it is determined as a result of the determination that the urgency of the new failure is low. A fault response system including a fault analysis system.

2. The storage device includes: Further retaining information on breaking news about failures that have been reported but for which no action has been taken, The computing device When determining the level of urgency, the fault information reported regarding a new fault is compared with the fault flash information, and if the fault information is included in the fault flash information, the level of urgency of the new fault is determined to be high; and if, as a result of the comparison, the fault information is not included in the fault flash information and is not included in the handling example, the level of urgency of the new fault is determined to be high.

2. The fault response system according to claim 1.

3. The computing device When determining the level of urgency, the fault information reported regarding a new fault is compared with the fault flash information, and if the fault information is not included in the fault flash information and indicates multiple faults, it is determined whether there is a handling example that simultaneously includes all of the multiple faults, and if it is found that there is a handling example that simultaneously includes all of the multiple faults, it is determined that the urgency of the new fault is low.

3. The failure response system according to claim 2.

4. The computing device When determining the level of urgency, if it is determined that there is no handling example that simultaneously includes all of the plurality of failures as a result of determining whether there is a handling example that simultaneously includes all of the plurality of failures, it is determined whether each of the plurality of failures is included in any of the handling examples, and if it is determined that there is a failure among the plurality of failures that is not included in any of the handling examples, it is determined that the level of urgency of the new failure is high.

4. The failure response system according to claim 3.

5. The storage device includes: further retaining priority information for each part of the device to be monitored, the priority information defining the degree of impact on the device in the event of a failure; The computing device When determining the level of urgency, it is determined whether each of the plurality of failures is included in any of the handling examples, and if it is found that each of the plurality of failures is included in any of the handling examples, it specifies the priority of each of the plurality of failures using the priority information; it is determined whether there are failures with the same priority among the plurality of failures, and if it is found as a result of the determination that there are failures with the same priority among the plurality of failures, it is determined that the urgency of the new failure is high.

5. The failure response system according to claim 4.

6. The computing device With regard to the new fault that is found to have a low urgency, a process is executed to notify a maintenance personnel assignment system of information on a maintenance procedure corresponding to a countermeasure example that is suitable for the new fault, and a process is executed to notify a parts management system that manages inventory information of parts for the device to be monitored of a request for arranging the parts, The parts management system executes a process of checking the inventory of the parts and a process of notifying the assignment system of delivery schedule information when delivering the parts to the installation location of the device to be monitored, and the assignment system holds information on each maintenance worker, receives the information on the maintenance procedure and the delivery schedule information, identifies a maintenance worker who is capable of performing the maintenance procedure and has a schedule that allows him or her to perform work to respond to a failure at the installation location, and executes a process of notifying the terminal of the maintenance worker of the maintenance procedure and the delivery schedule information for the parts.

2. The fault response system according to claim 1.

7. The computing device When the analysis request including the fault information of the new fault that is found to have a high urgency is notified to the factory system, new maintenance procedure information that is an analysis result of the new fault is received from the factory system, and the information is linked to the new fault and stored in the storage device as one of the handling examples; The factory system executes a process of notifying the parts management system of the arrangement request including information on parts that need to be replaced, which is indicated by a maintenance procedure based on a failure analysis performed at the factory in response to the analysis request, and a process of notifying the assignment system of information on the maintenance procedure, and the parts management system executes a process of checking the inventory of the parts in response to the arrangement request, and notifying the assignment system of delivery schedule information for delivering the parts to the installation location of the device to be monitored, and the assignment system receives the information on the maintenance procedure and the delivery schedule information, identifies a maintenance worker who is capable of performing the maintenance procedure and has a schedule that allows him or her to perform work to respond to the failure at the installation location, and executes a process of notifying the maintenance procedure and the delivery schedule information for the parts to a terminal of the maintenance worker.

7. The failure response system according to claim 6.

8. The information processing device storing information on cases of dealing with failures in the monitored device in a storage device; a process of determining whether or not there is a countermeasure case that matches the new failure based on fault information reported about the new failure in the monitored device and information about the countermeasure case; and If it is determined that the new fault has a high level of urgency as a result of the determination, a process of requesting a system of a factory that manufactures or repairs the monitored device to analyze how to deal with the new fault; If it is determined that the urgency of the new failure is low as a result of the determination, a process of making a request to a predetermined system to arrange for maintenance personnel and parts according to a countermeasure case that matches the new failure; A failure response method characterized by executing the above.

Citation Information

Patent Citations

  • Automated system for coping with computer fault and recording medium having fault coping automation program recorded thereon

    JP2001306360A