Preventive maintenance support system and preventive maintenance support method
The preventive maintenance support system optimizes cost performance by identifying equipment for replacement based on log information and importance, addressing inefficiencies in existing IT systems by minimizing unnecessary replacements and disruptions.
Patent Information
- Application Number
- JP2024088130
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-05-30
AI Technical Summary
Preventive maintenance in IT systems incurs additional costs and risks, as it requires investment in backup hardware resources and often disrupts operations, while the importance and redundancy of components are not adequately considered, leading to inefficient cost performance.
A preventive maintenance support system that identifies equipment needing replacement based on log information, considering redundancy and importance, and presents this information to an operation terminal, optimizing cost performance by minimizing unnecessary replacements.
The system optimizes cost performance by identifying devices requiring preventive maintenance, reducing unnecessary replacements, and minimizing operational disruptions, thereby enhancing the cost-effectiveness of maintenance operations.
Smart Images

Figure 2025182718000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a preventive maintenance support system and a preventive maintenance support method that support preventive maintenance by detecting signs of failure in monitored devices. [Background technology]
[0002] Traditionally, IT (Information Technology) systems that support social infrastructure have been required to operate normally. In recent years, the trend has shifted from the transition to cloud computing to a return to on-premise systems where important data can be managed within sight by owning in-house assets or signing up for private cloud services.
[0003] Until now, the thinking has been that servers and storage devices that make up IT systems should be maintained only after they break down. However, the idea of preventing business downtime by detecting signs of failure and replacing them in advance (preventive maintenance) even if it means paying extra costs is becoming more common.
[0004] However, preventive replacement based on failure prediction detection requires investment in backup hardware resources. For example, maintenance is often performed on holidays or late at night to avoid impacting IT systems. A dedicated maintenance system must be prepared to perform work during these specially set maintenance times. This type of preventive maintenance is usually not covered by maintenance service contracts. Today's IT systems combine and link various individual systems, widening the scope of the impact of a single failure. Furthermore, the working population is shrinking and aging. These factors make it increasingly difficult to prepare a dedicated maintenance system.
[0005] When considering a maintenance plan, the priority and cost to be spent will vary depending on the importance and redundancy of each individual system and device that makes up the IT system. (1) For example, the importance of suspected components (devices, etc., for which failures are predicted) and the extent of the impact of a failure of a suspected component based on the configuration of the IT system are not stored in a database. Therefore, when there are multiple suspected components, the priority of each component is not defined.
[0006] Furthermore, maintenance based on failure prediction detection does not necessarily require replacing the suspected component. It is also reasonable to consider mitigating the impact (by evacuating application software running on the suspected component or distributing the load to other devices) when considering optimizing system operation.
[0007] When replacing suspected components (hardware), the equipment may need to be shut down, and before the equipment is shut down, the IT system must be switched to a standby system. Depending on the IT system's design and the software used, there may be cases where the switchover requires a momentary interruption or the suspension of processing for a certain period of time, which is a prerequisite for the evacuation. Since the switchover requires labor hours and labor hours for informing related systems and users, increasing the number of times application software needs to be evacuated to perform preventive maintenance reduces cost-effectiveness. Therefore, it is necessary to present options for hedging against the risk of possible failure.
[0008] For example, Patent Document 1 states that "Information defining the configuration of businesses and O&M assets is registered in the business entity O&M asset definition DB. Information defining the systems belonging to O&M assets is registered in the O&M asset system definition DB," and "Predictive maintenance schedules can be planned to reflect the convenience of O&M assets and maintenance businesses." [Prior art documents] [Patent documents]
[0009] [Patent Document 1] Japanese Patent Application Publication No. 2019-133412 Summary of the Invention [Problem to be solved by the invention]
[0010] However, preventive maintenance is separate from regular maintenance (maintenance and replacement when a part breaks down), and involves inspecting parts and replacing them before they break down, so costs for parts and inspection work are required. Preventive maintenance incurs additional costs on top of regular maintenance, so the more preventive maintenance is performed, the higher the costs. It also increases the risk of replacing parts that are still in good working order with new ones.
[0011] Given the above situation, there was a demand for a method to optimize the cost performance when implementing preventive maintenance. [Means for solving the problem]
[0012] In order to solve the above problem, one aspect of the present invention is a preventive maintenance support system in which a computer having an arithmetic unit that executes a program and a storage device that stores the program executes processing to support preventive maintenance for monitored devices. The computing device acquires log information of the equipment installed in the monitored device, and identifies equipment to be replaced as a preventive replacement based on information contained in the log information indicating the status of the equipment and a reference value for signs of failure determined according to the redundancy and importance of the equipment, and presents the identified equipment on the operation terminal. [Effects of the Invention]
[0013] According to at least one aspect of the present invention, the computing device identifies devices that are subject to preventive replacement and presents the identified devices to the operation terminal, thereby optimizing cost performance when implementing preventive maintenance. Problems, configurations, and effects other than those described above will become apparent from the following description of the preferred embodiments of the invention. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a preventive maintenance support system according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of the hardware configuration of each processing unit that configures a preventive maintenance support system according to an embodiment of the present invention. [Figure 3] 1 is a flowchart illustrating an example of a procedure for preventive maintenance support processing by a preventive maintenance support system according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a monitoring target device according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of configuration information regarding path redundancy of a device mounted in a monitoring target device according to an embodiment of the present invention. [Figure 6] 10 is a diagram illustrating an example of configuration information regarding the business importance of devices mounted in a monitoring target apparatus according to an embodiment of the present invention. FIG. [Figure 7] FIG. 3 is a diagram showing an example of survey target data (survey target items) for each device included in the monitoring target device according to one embodiment of the present invention. [Figure 8] FIG. 10 is a diagram illustrating an example of thresholds used for detecting a failure sign in a monitored device according to an embodiment of the present invention. [Figure 9] 1 is a diagram illustrating the concept of a light intensity threshold of an SFP in a preventive maintenance support system according to an embodiment of the present invention. [Figure 10] FIG. 1 is a diagram illustrating an example of failure sign detection by a preventive maintenance support system according to an embodiment of the present invention. [Figure 11] 1 is a diagram showing a specific example of an IT system to which a preventive maintenance support system according to an embodiment of the present invention is applied. DETAILED DESCRIPTION OF THE INVENTION
[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, examples of modes for carrying out the present invention (hereinafter referred to as "embodiments") will be described with reference to the accompanying drawings. In this specification and the accompanying drawings, identical or similar components are given the same reference numerals, and redundant explanations may be omitted, or only explanations focusing on the differences may be given. Furthermore, when there are multiple identical or similar components, they may be described using the same reference numerals with different subscripts. Note that when there is no need to distinguish between these multiple components, the subscripts may be omitted in the description. The number of each component may be singular or plural unless otherwise specified. Furthermore, expressions such as "site," "component," and "equipment" are used to describe the elements that make up the monitored device, but these are interchangeable.
[0016] In the following embodiment, various information is described in table format, but the information may be in a data format other than a table format. Also, various names such as "XX information," "XX table," "XX list," and "XX list" are interchangeable.
[0017] <One embodiment> [Overall configuration of preventive maintenance support system] First, the overall configuration of a preventive maintenance support system according to an embodiment of the present invention will be described with reference to FIG. FIG. 1 is a diagram showing an example of the overall configuration of a preventive maintenance support system according to an embodiment of the present invention. The preventive maintenance support system 1 is roughly composed of an IT system 10, an analysis unit 20, and an operation terminal 30. The IT system 10 is a system that includes a monitored device 11 and a management unit 12 . The monitored device 11 is a device or system that is the target of failure sign detection. The management unit 12 is a processing unit (arithmetic device) that collects and manages data related to the monitored device 11. The management unit 12 periodically or at a predetermined timing collects data on the operation of each part of the monitored device 11 and stores the data as a log file (operation log 121). The management unit 12 also manages information related to the configuration of the monitored device 11 (configuration information 122).
[0018] The analysis unit 20 is a processing unit (arithmetic unit) that acquires data (operation log 121) on the operation of each part of the monitored device 11 provided by the management unit 12 and analyzes the data. The analysis unit 20 includes a graph creation unit 21 and a cost estimation unit 22. The graph creation unit 21 has a function of analyzing the operation log 121 acquired from the monitored device 11 and creating a graph 210 based on the analysis results. The cost estimation unit 22 calculates the cost of preventive maintenance (maintenance cost 220) for the monitored device 11 based on the analysis result (analysis data) of the operation log 121 by the graph creation unit 21. In other words, the cost estimation unit 22 has a function of calculating in advance the cost of implementing preventive replacement of equipment that is the target of preventive replacement, and presenting this cost to the operation terminal 30.
[0019] The operation terminal 30 is a device equipped with a display function that allows the operator to refer to the information processed by the analysis unit 20 (graph 210, maintenance cost 220).
[0020] The IT system 10, the analysis unit 20, and the operation terminal 30 can communicate with each other via a communication network. The management unit 12 communicates with the monitored device 11 via a communication network or a dedicated line. Although an example in which the IT system 10 is configured by the monitored device 11 and the management unit 12 has been shown, the IT system 10 may be configured so that the management unit 12 is not included therein.
[0021] [Hardware configuration of each device in the preventive maintenance support system] Next, the hardware configuration of each device constituting the preventive maintenance support system 1 will be described with reference to FIG. Fig. 2 is a diagram showing an example of the hardware configuration of each device that constitutes the preventive maintenance support system 1. The devices that constitute the preventive maintenance support system 1 correspond to the monitored device 11, the management unit 12, the analysis unit 20, and the operation terminal 30. The calculator 40 shown in Fig. 2 is an example of hardware used as a computer. Each device that constitutes the preventive maintenance support system 1 realizes preventive maintenance support performed by the devices shown in Fig. 1 working together when the calculator 40 (computer) executes a program.
[0022] The computer 40 includes a CPU (Central Processing Unit) 41, a ROM (Read Only Memory) 42, and a RAM (Random Access Memory) 43, all connected to a system bus. The computer 40 further includes a display device 44, an input device 45, a non-volatile storage 46, and a communication interface 47.
[0023] The CPU 41 reads out the program code of the software that realizes each function according to this embodiment from the ROM 42, loads it into the RAM 43, and executes it. Variables, parameters, etc. that arise during the calculation processing of the CPU 41 are temporarily written to the RAM 43, and these variables, parameters, etc. are read out as appropriate by the CPU 41. The functions of each device and terminal of the preventive maintenance support system 1 are realized by the CPU 41 executing the program code read out from the ROM 42. However, other processors such as an MPU (Micro Processing Unit) may be used instead of the CPU 41.
[0024] The nonvolatile storage 46 is an example of a recording medium, and is capable of storing data used by a program, data obtained by executing a program, and the like. For example, the nonvolatile storage 46 stores the introduction candidate equipment list 1000, an introduction area / introduction equipment list, and the evaluation results thereof. The nonvolatile storage 46 may also store some or all of the various data (information) shown in FIG. 1. The nonvolatile storage 46 may also store an OS (Operating System) and programs executed by the CPU 41. The nonvolatile storage 46 may be a hard disk drive (HDD), a solid state drive (SSD), an optical or magnetic disk medium, a semiconductor memory card, or the like.
[0025] A communication device such as a network interface card (NIC) is used as the communication interface 47. The communication interface 47 can transmit and receive various data to and from external devices via a communication network such as a connected LAN, a dedicated line, etc. The communication interface 47 is used to input various information (data) to each device and terminal of the preventive maintenance support system 1.
[0026] The display device 44 is a monitor such as a liquid crystal display, and displays a GUI (Graphical User Interface) screen, the results of arithmetic processing by the CPU 41, etc. The input device 45 generates an input signal in response to a user's operation and outputs it to the CPU 41. The input device 45 may be, for example, a mouse, a keyboard, a touch sensor, or the like, and the user can input information and instructions by operating the input device 45. The display device 44 and the input device 45 may be integrated into a touch panel.
[0027] [Preventive maintenance support processing procedure] Next, the preventive maintenance support process performed by the preventive maintenance support system 1 will be described with reference to FIG. FIG. 3 is a flowchart showing an example of the procedure of the preventive maintenance support process performed by the preventive maintenance support system 1.
[0028] First, the analysis unit 20 of the preventive maintenance support system 1 acquires information about the monitored device 11 of the IT system 10 (step S1). The information about the monitored device 11 includes configuration information about the redundancy of the devices that make up the monitored device 1 (see FIG. 5, which will be described later), and configuration information about the business importance (see FIG. 6, which will be described later). As an example, the processing of step S1 is performed when the analysis unit 20, which is a main function of the preventive maintenance support system 1, is delivered to the IT system 10.
[0029] Next, the analysis unit 20 acquires a combination of redundancy and business importance of each device constituting the monitored device 1 (see FIG. 8, which will be described later) (step S2).
[0030] Next, the analysis unit 20 acquires the operation log 121 of the monitoring target device 11 that is currently in operation from the management unit 12 (step S3). In the following description, the operation log may be abbreviated to "log".
[0031] Next, the analysis unit 20 extracts values (actual numerical values recorded in the operation log 121) corresponding to the survey target data for each part from the acquired operation log 121 (step S4). As shown in Fig. 7 described later, the survey target data corresponds to the survey target items, and the data used for analysis differs depending on the part (device) to be surveyed. For example, in the case of the SFP (Small Form-factor Pluggable) module, which is one of the standards for optical transceivers that connect optical fibers to communication devices, data on light intensity is used, so "light intensity" will be used as the example of data to be surveyed in the following explanation. The SFP module (hereafter referred to as "SFP") is an optical transceiver that converts between the electrical signals sent and received by communication devices and the optical signals that flow through optical fiber cables.
[0032] Next, the analysis unit 20 identifies the hardware location corresponding to the extracted investigation target data from FIG. 5, which will be described later (step S5).
[0033] Next, the analysis unit 20 extracts a threshold value for the data to be investigated from the combination of redundancy and business importance (FIG. 8) (step S6). The threshold value is a reference value for a failure sign, and is used, for example, to determine whether or not to replace a device.
[0034] Next, the analysis unit 20 determines whether the value of the investigation target data extracted from the operation log 121 exceeds a threshold value (step S7). The value of the investigation target data in the operation log is compared with a threshold value that is a reference value for a sign of SFP failure, and the part to be replaced is identified. If the value of the investigation target data does not exceed the threshold value (NO determination in step S7), the analysis unit 20 proceeds to the determination process of step S9.
[0035] If the value of the data to be investigated in the operation log 121 exceeds the threshold in step S7 (YES determination in step S7), the analysis unit 20 proceeds to the determination process in step S8.
[0036] If the determination in step S7 is YES, the analysis unit 20 identifies the device whose value of the survey target data exceeds the threshold as a device that needs to be replaced (step S8).
[0037] Next, if the determination is NO in step S7, or after the processing of step S8, the analysis unit 20 determines whether or not the investigation of all the data to be investigated has been completed (step S9).
[0038] If the determination in step S9 is NO, the analysis unit 20 proceeds to step S5, identifies the next hardware location corresponding to the extracted investigation target data, and repeats the subsequent processes.
[0039] On the other hand, if the judgment in step S9 is YES, the analysis unit 20 proceeds to step S2, and again obtains the combination of redundancy and business importance of each device during the next regular survey, and performs the survey from obtaining the operation log 121 in step S3. The reason for re-obtaining the combination of redundancy and business importance for each device at the time of the next investigation is that the state of each device may be different between the current investigation and the next investigation. If the state of the device, i.e., the value of the data being investigated in the operation log, is different, it may be necessary to change the threshold, and the results of comparing the value of the data being investigated with the threshold may also be different. In this embodiment, preventive maintenance support is performed by regularly inspecting devices that are subject to maintenance on a monthly or weekly basis.
[0040] While the monitored device 11 is in operation, the analysis unit 20 repeatedly executes the processes of steps S2 to S9 to determine which devices require preventive maintenance.
[0041] The analysis unit 20 notifies the operation terminal 30 of information indicating the equipment determined in step S9 to require replacement. At this time, the analysis unit 20 creates a graph 210 for the log of the equipment in question using the graph creation unit 21, calculates the cost required for maintaining (e.g., replacing) the equipment in question using the cost estimation unit 22, and notifies the operation terminal 30 of the graph 210 and the maintenance cost 220.
[0042] The system operator 480 checks the information displayed on the operation terminal 30 and dispatches a worker to the site where the monitored device 11 having the equipment to be replaced is located to replace the equipment, or makes a plan for replacement work for the equipment.
[0043] Preventive maintenance targets equipment that can still continue to operate. Therefore, there is no strong demand for immediate maintenance replacement. How preventive maintenance operations are defined is up to the customer. For example, operations could include immediately replacing the relevant equipment, or replacing a week's worth of relevant equipment all at once on a specific day of the week.
[0044] The effect of the configuration of the preventive maintenance support system 1 according to this embodiment is that data analysis from operation logs cannot be achieved on a public cloud, and can be said to be an advantage unique to private clouds and on-premise systems.
[0045] [Example of the internal structure of the monitored device] Next, a specific example of the internal structure of the monitored device 11 will be described with reference to FIG. FIG. 4 is a diagram illustrating an example of the hardware configuration of the monitored device 11. As shown in FIG.
[0046] The monitored device 11 has typical hardware that constitutes an IT system, such as a server, a Fiber Channel (FC) switch, storage, and a Local Area Network (LAN) switch. For example, Fig. 4 shows examples of a LAN switch 411, multiple servers 421-1, 421-2, 421-3, ..., 421-n (not shown), a spare server 421s, an FC switch 431, and storage 441 as devices that are targets of failure sign detection for the monitored device 11. The storage 441 has HDDs 451-1 to 451-n and SSDs 461-1 to 461-n as drives. The HDDs 451-1 to 451-n and the SSDs 461-1 to 461-n constitute a redundant array of inexpensive disks (RAID).
[0047] The multiple servers 421-1 to 421-n, the spare server 421s, the FC switch 431, and the storage 441 each have an SFP mounted on each port. The LAN switch 411 is connected to the plurality of servers 421-1 to 421-n and the spare server 421s via a LAN. A SAN (Storage Area Network) is connected between the multiple servers 421-1 to 421-n and the spare server 421s and the FC switch 431. Furthermore, the FC switch 431 and the storage 441 are also connected by a SAN.
[0048] 4, the LAN switch 411 is represented as "LAN switch 1," the server 421-n as "server n," the FC switch 431 as "FC switch 1," the storage 441 as "storage 1," the HDD 451-n as "HDDn," and the SSD 461-n as "SSDn," where n is a natural number.
[0049] The management unit 12 has one or more management servers 120, a database (DB) 470 that manages information, and a path connecting to the LAN switch 411 of the monitored device 11. The database 470 is a database shared by the management unit 12, the analysis unit 20, and the operation terminal 30. All tables and data described below are collectively managed by the database 470. However, in this embodiment, it is assumed that the database 470 is constructed on the management server 120, but the storage location of the database 470 is not limited to this example. Each management server 120 functions as one management unit 12.
[0050] It is known that the deterioration of the light-emitting element inside the SFP can cause problems with communication with the other device. Since the light-emitting element deteriorates over a few years, the SFP should be replaced in advance, taking into account the operating status (aging) of the LAN cable that uses optical fiber and the equipment equipped with the SFP.
[0051] The analysis unit 20 has one or more processing servers 200 that process data, a path connecting to the database 470, and a path connecting to the LAN switch 411 of the monitored device 11. Each processing server 200 functions as one analysis unit 20.
[0052] The operation terminal 30 has multiple terminal devices 300 that serve as user interfaces, a path connecting to the database 470, and a path connecting to the LAN switch 411 of the monitored device 11. Each terminal device 300 functions as one operation terminal 30. A system operator 480 operates the terminal devices 300 to input tables and data.
[0053] 5 and FIG. 6, which will be described later, the configuration information of the monitored device 11 is managed in a database, including the configuration of the hardware already installed in the monitored device 11 and the associated system names and business names. The system operator 480 adds information each time a device is installed or added.
[0054] [Configuration information about redundancy] Next, configuration information regarding the redundancy of the devices installed in the monitored device 11 will be described with reference to Fig. 5. Here, the redundancy of the path used for data transmission will be described as an example of the redundancy.
[0055] Fig. 5 is a diagram showing an example of configuration information regarding the path redundancy of the device mounted in the monitored device 11. The example in Fig. 5 shows configuration information when an SFP is the target of failure sign detection. A path redundancy table 500 in the figure shows configuration information relating to the path redundancy of each device installed in the monitored device 11. The path redundancy table 500 has the following items: SFP location, path redundancy, and connected server.
[0056] The SFP location item includes information indicating the location of the SFP installed in the monitored device 11. That is, the information indicates the location of the device (SFP, etc.) installed in the monitored device 11 that is the target of failure sign detection. The path redundancy item includes information indicating the path redundancy of a device (e.g., a server) to which an SFP is attached inside the monitored device 11. The path redundancy is an example of information indicating the redundancy of a device that is the target of failure sign detection. Note that if the device that is the target of failure sign detection is a storage drive or a server, the path redundancy can be replaced with simple redundancy. The connected server item includes information indicating the server in which the SFP is installed.
[0057] 5, for example, in the record in the first row of the path redundancy table 500, "Server 1 Port A" is stored in the SFP location, "1 path" is stored in the path redundancy, and "Server 1" is stored in the connected server. Also, in the record in the second row, "Server 2 Port A" is stored in the SFP location, "2 paths" is stored in the path redundancy, and "Server 2" is stored in the connected server.
[0058] [Configuration information about business importance] Next, configuration information regarding the business importance of the devices installed in the monitored device 11 will be described with reference to FIG.
[0059] Fig. 6 is a diagram showing an example of configuration information regarding the business importance of devices installed in the monitored device 11. The example of Fig. 6 shows configuration information when an SFP is the target of failure sign detection, similar to Fig. 5. The business importance table 600 in the figure shows configuration information related to the business importance of the server to which each device mounted on the monitored device 11 is connected. The business importance table 600 has the following items: connected server, used business, and business importance.
[0060] The connected server item includes information indicating the server in which the SFP is installed. The item of the service to be used includes information indicating the name of the service that uses the server described in the item of the connected server. The business importance item is the importance of the business described in the business usage item. In other words, the business importance is the importance of the business of the server to which the corresponding SFP is attached. For example, the business importance can be set as high, medium, or low. When the device targeted for failure sign detection is a storage drive or a server, the business importance is the business importance of the device itself targeted for failure sign detection (storage drive or server).
[0061] For example, if signs of failure are detected in multiple devices, when an operator decides which devices to replace as a preventive measure, the business importance level serves as reference information for determining the priority of the devices to replace as a preventive measure.
[0062] 6, for example, in the record in the first row of the business importance table 600, the connection server is set to "Server 1," the business used is set to "AAA business," and the business importance is set to "Low." Also, in the record in the second row, the connection server is set to "Server 2," the business used is set to "BBB business," and the business importance is set to "High."
[0063] In this way, for each SFP installed in the FC switch, a table is prepared in advance on the management server 120 that defines availability (redundancy) and importance based on server information using the path (FC path) that passes through that SFP.
[0064] [Survey data] Next, the survey target data (survey target items) for each device included in the monitored device 11 will be described with reference to FIG. 7 is a diagram showing an example of survey target data (survey target items) for each device included in the monitored device 11. This device corresponds to a server or the like to which a predictive detection target device (for example, an SFP) is attached.
[0065] The investigation target data table 700 in the figure has the following items: device type, detection target, investigation target data, and acquired log. The item of device type includes information indicating the device included in the monitoring target device 11. For example, this device may be an FC switch or storage. The detection target item includes information indicating devices that are targets of failure sign detection, such as SFPs, SSDs, and HDDs. The survey data items include data on the equipment being inspected. For example, for an SFP, the survey data would be the amount of light. For an SSD, the survey data would be the number of errors and the number of writes. The survey data uses indicators that can estimate the deterioration state of the equipment and devices over time. The acquired log items include information that identifies the log information (for example, operation log) from which information for each item of the corresponding record can be obtained. This information that identifies the log information can be a file path that indicates the location of the file that contains the log information, a URL that describes the location of the log information on the communication network, etc.
[0066] 7, for example, the record in the first row of the investigation target data table 700 stores "FC switch" as the device type, "SFP" as the detection target, "light intensity" as the investigation target data, and "xxx.Log" as the acquisition log. The record in the second row stores "storage" as the device type, "SSD" as the detection target, "number of errors" as the investigation target data, and "yyy.Log" as the acquisition log. The record in the third row stores "storage" as the device type, "SSD" as the detection target, "number of writes" as the investigation target data, and "zzz.Log" as the acquisition log.
[0067] [Threshold for detecting signs of failure] Next, the threshold value used for detecting a sign of failure in the monitored device 11 will be described with reference to Fig. 8. The threshold value used for detecting a sign of failure is a reference value for preventive replacement.
[0068] FIG. 8 is a diagram showing an example of setting thresholds used for detecting a sign of failure in the monitored device 11. In FIG. The threshold table 800 shown in FIG. 8 shows a list of thresholds defined by a combination of device redundancy (here, path redundancy) and business importance. When introducing a monitored device 11 into the IT system 10, the system operator 480 stores threshold information for the target device using the operation terminal 30. The system operator 480 is also permitted to periodically review the thresholds to account for aging degradation. For example, if it is necessary to take into account aging degradation of hardware due to the characteristics of the target device, the thresholds are periodically reviewed.
[0069] (path redundancy) The threshold table 800 stores thresholds for each combination of path redundancy and job importance. Redundancy can be thought of as "no redundancy," "duplication," or "multiplication (triple or more)." Without redundancy, there is a single point of failure. Without redundancy, failure itself is unacceptable, so it is necessary to prevent failures before they occur by detecting their symptoms with stricter thresholds. In the case of duplexing, a looser threshold is acceptable compared to when there is no redundancy. This is because the greater the multiplexing and redundancy, the less the impact of a single path failure.
[0070] The threshold table 800 shows "1 path," "2 paths," and "4 paths" as examples of path redundancy. In the case of 2 paths, if a failure occurs in one path, the performance of the corresponding device will decrease by 50%. In the case of 4 paths, if a failure occurs in three paths, the performance of the corresponding device will decrease by 25%.
[0071] (Business importance) Business importance is defined based on factors such as the availability required for the business. As an example, the definition could be based on the availability rate of the equipment (servers, etc.) to which the equipment targeted for failure sign detection is connected (large: 99.9999%, medium: 99.99%, small: 99%), or by business classification (large: accounting system, small: information system). Accounting system corresponds to transactional operations such as deposits and withdrawals in accounts and transfers between accounts in a financial system. Information system corresponds to operations that only affect a specific financial institution, such as managing account information in a financial system.
[0072] In Fig. 8, for example, when the path redundancy is "1 path", the threshold is set to "-a dBm or less" when the business importance is "high", "-d dBm or less" when the business importance is "medium", and "-g dBm or less" when the business importance is "low". Also, when the path redundancy is "2 paths", the threshold is set to "-b dBm or less" when the business importance is "high", "-e dBm or less" when the business importance is "medium", and "-h dBm or less" when the business importance is "low". Here, a <d<g、b<e<hである。
[0073] In this way, the business importance is information that indicates the relative position of the equipment that is the target of failure sign detection within the monitored device 11. Rather than uniformly evaluating all the equipment within the monitored device 11, weighting is performed according to business importance to differentiate the maintenance response for the equipment, and effects such as improved cost performance are expected. For example, even if thorough maintenance is performed on equipment that is only slightly affected in the event of a failure, the cost-effectiveness is low, but weighting can prevent such an inconvenience.
[0074] [Concept of threshold] Furthermore, the concept of thresholds in the preventive maintenance support system 1 will be explained using FIG. 9, taking the light intensity threshold of an SFP as an example. 9 is a diagram showing the concept of the light intensity threshold of the SFP in the preventive maintenance support system 1. The horizontal axis of the illustrated graph 900 represents the light intensity [dbm], and the vertical axis represents the number of operating SFPs. σ represents the standard deviation.
[0075] The distribution of SFP light intensity and operational count follows a roughly Gaussian distribution. For normal SFPs, the light intensity is distributed around -3 dBm. Approximately 68% of SFPs fall within a ±1 σ range around -3 dBm. Approximately 95% of SFPs fall within a ±2 σ range around -3 dBm, and approximately 99% fall within a ±3 σ range. The probability of SFP failure increases when the light intensity exceeds approximately -5 dBm. Therefore, as an example, it is recommended to set -adBm to -idBm in the threshold table 800 within the range of -3 dBm to -5 dBm. Among these, the threshold can be set as follows: "+1 σ ± α" for high business importance, "+2 σ ± α" for medium business importance, and "+1 σ ± α" for low business importance. The value of α will be explained later.
[0076] If the light intensity threshold is too high, the number of SFP replacements will increase (SFPs that are still in good working order will also be subject to replacement). Therefore, it is desirable to determine the replacement threshold by considering the balance between the costs incurred and the impact of failure. As an example, by periodically sliding the setting range to a stricter threshold in accordance with the introduction time (aging deterioration) of the monitored device 11, it is possible to prevent failures of important IT systems (monitored device 11) while reducing costs.
[0077] When setting the threshold for light intensity, the standard deviation σ can be used as is to determine the threshold. Alternatively, rather than using the standard deviation σ itself, a range can be used to use values around it. The closer the light intensity is to the mode value, the more SFP replacements will be required (increasing costs). Depending on the customer's budget, there may be cases where they cannot invest that much in component replacements, so it is desirable to be able to adjust the threshold within a certain range. For example, if the budget is exceeded, it is possible to reduce costs by setting a looser threshold value.
[0078] Incidentally, SFPs with a short remaining lifespan can be included as replacement targets without changing the threshold, since the light intensity gradually deteriorates. However, the number of failures is expected to increase with age. In response to this, by widening the range of failure signs in failure sign detection (setting a stricter value), it is possible to take dynamic measures such as preventive replacement before an actual failure occurs. For example, even within the same low importance range (±3σ), the left side of Figure 9 shows a stricter sliding direction. Therefore, it is possible to set the α value further to the left within the "low" importance setting range.
[0079] [Table Registration] The registration of each table shown in FIGS. 5 to 7 in database 470 can be performed by system operator 480. When an IT system 10 including a monitored device 11 is introduced, the hardware configuration of the monitored device 11, the business operations that use (are assigned to) the device, and the parts that are the targets of symptom detection are selected. The system operator 480 operates the operation terminal 30 (terminal device 300) to input this system configuration information into the database 470.
[0080] For example, configuration information procured through the equipment purchasing process is registered in the path redundancy table 500 in FIG. 5 and the business importance table 600 in FIG. In the survey target data table 700 shown in FIG. 7, areas where failure sign detection technology has been established and where failure sign detection is desired to be performed are registered. 8, an initial threshold is set when failure sign detection begins, and operation begins. When adjusting the threshold due to deterioration over time, the system operator 480 can update the threshold as needed.
[0081] [Calculate Cost] The cost of preventive maintenance when a certain threshold is adopted can be calculated using the following formula (1). By presenting the cost of preventive maintenance to the customer, the customer can consider for themselves whether the cost is reasonable and profitable in line with their circumstances (budget).
[0082] A=B*C*(D+E) (1) A: Estimated total cost B: Total number of monitored ports C: Distribution probability according to the threshold For example, if the threshold is -3.8 dBm, approximately 0.5% of all ports are affected. D: Replacement parts procurement cost E: Replacement work cost
[0083] The budget that can be allocated to preventive maintenance varies depending on the sales and costs (cumulative of normal maintenance costs and operational man-hours, etc.) incurred by the device equipped with the equipment that is the target of predictive detection (for example, server 2 in Figure 6) and the monitored device 11, etc. The budget that can be allocated to preventive maintenance is up to the customer.
[0084] [Example of failure prediction detection] 10 is a diagram showing an example of failure sign detection by the preventive maintenance support system 1. Failure sign detection and identification of equipment to be replaced will be described using the server 421-2 (server 2, FIG. 4) as an example.
[0085] For example, in the monitored device 11 shown in Figure 4, which is composed of servers 1 to 3, a spare server, an FC switch 1, storage 1, SFPs installed in their external ports, and FC cables connecting the SFPs, log information from FC switch 1 is periodically acquired to detect signs of hardware failure.
[0086] Communication between devices (SFPs) is achieved by light passing through the FC cable. Therefore, if the light intensity weakens, problems will occur with mutual communication. The light-emitting element of the SFP deteriorates over time, and the amount of light emitted will gradually weaken with continued operation. Therefore, in this embodiment, preventive maintenance using failure sign detection uses a threshold value for the amount of light to perform preventive replacement before communication is completely disrupted.
[0087] First, the analysis unit 20 acquires the log information "xxx.Log" of the detection target device of the monitoring target device 11 that is in operation from the management unit 12 (process (1): corresponding to step S3).
[0088] Next, the analysis unit 20 extracts values (actual numerical values recorded in the log information) corresponding to the data to be investigated for each part from the acquired log information "xxx.Log" (process (2), corresponding to step S4). In this example, the light intensity (RX: received light, TX: emitted light) of port A of server 2 connected to port 2 can be determined from the information on port 2 of FC switch 1 written in the log information "xxx.Log." In the example of FIG. 10, the light intensity is written as "-2.5 dBm (564.5 μW)."
[0089] Next, the analysis unit 20 identifies the hardware location corresponding to the extracted investigation target data (light intensity) from FIG. 5 (corresponding to process (3), step S5). For example, specify server 2 port A. Note that table 1020 shown in Fig. 10 is a table that combines the path redundancy table 500 and the task importance table 600 in Fig. 5.
[0090] Next, the analysis unit 20 refers to the threshold value table 800 in FIG. 8 and extracts a threshold value (-b dBm in the figure) for the data to be investigated from the combination of path redundancy and task importance (process (4), corresponding to step S6). In this example, the analysis unit 20 can calculate the standard deviation σ by aggregating the light intensity for each SFP mounted on the FC switch 1 and each opposing SFP on the connected server 2 from the acquired log information "xxx.Log." Then, from the calculated standard deviation σ, the analysis unit 20 determines a reference value (threshold value) for a sign of SFP failure for each business importance (e.g., high) and path redundancy (e.g., 2 paths) of the device (here, server 2) on which the SPF, the device to be detected, is mounted.
[0091] Next, the analysis unit 20 compares the periodically acquired light intensity information of the SFP with the reference values determined for the business importance and path redundancy of the device (here, server 2) in which the device to be detected is installed (process (5), corresponding to step S7).The analysis unit 20 then targets the SFP whose light intensity is below the reference value for preventive replacement. For example, if the RX Power value is lower than the threshold value, the SFP at port A of server 2 is identified as the SFP to be replaced.
[0092] As described above, the preventive maintenance support system (preventive maintenance support system 1) of this embodiment is a preventive maintenance support system in which a computer (computer 40) having an arithmetic unit (CPU 41, analysis unit 20) that executes a program and a storage device (ROM 42 or non-volatile storage 46) that stores the program executes processing to support preventive maintenance for a monitored device (monitored device 11). The computing device is configured to acquire log information (e.g., operation log 121) of equipment (e.g., SPF) installed in the monitored device, and identify equipment to be replaced as a preventive replacement target based on information (e.g., light intensity) that indicates the status of the equipment contained in the log information and a reference value (threshold value) for a failure sign that is determined according to the redundancy of the equipment (e.g., path redundancy) and the importance of the equipment (e.g., business importance), and present the identified equipment on the operation terminal (operation terminal 30).
[0093] In the preventive maintenance support system having the above configuration, the calculation device (analysis unit 20) identifies equipment to be replaced as a preventive replacement target using a reference value of a failure sign determined according to the redundancy and importance of the equipment, and presents the identified equipment on the operation terminal. Then, the system operator 480 checks the information presented on the operation terminal before deciding on preventive maintenance, thereby optimizing the cost performance when implementing preventive maintenance.
[0094] [Specific example] Next, a specific example of an IT system to which the preventive maintenance support system 1 is applied will be described with reference to FIG.
[0095] 11 is a diagram showing a specific example of an IT system to which the preventive maintenance support system 1 is applied. In this example, the device targeted for failure sign detection is a server, not an SFP. The data to be investigated for a server can include the number of errors and the number of writes. It can also include other error information output by the server or data indicating minor abnormalities that still allow the server to operate. Alternatively, from the perspective of detecting signs of failure, the total operating time of the server since its installation can also be used as the data to be investigated.
[0096] 11, the IT system is composed of a production data center 1110 and a disaster recovery data center 1120. If a failure occurs in the production data center 1110, the IT system is configured to switch to the disaster recovery data center 1120. Note that the management unit 12 and analysis unit 20 of FIG. 4 are not shown in FIG. 11.
[0097] The production data center 1110 includes two servers 1111 and 1112, a spare server 1113, and a storage 1114. The storage 1114 includes HDDs 1115 and 1116, and a spare HDD 1117.
[0098] The disaster recovery data center 1120 includes two servers 1121 and 1122, a spare server 1123, and a storage 1124. The storage 1124 includes HDDs 1125 and 1126, and a spare HDD 1127.
[0099] The analysis unit 20 periodically acquires log information from the production data center 1110 about the HDDs of each server and storage. Furthermore, the analysis unit 20 refers to system information 1130 consisting of a system demand level 1131 and redundancy 1132. The system demand level 1131 corresponds to the information on the business importance in the business importance table 600 shown in Fig. 6. Furthermore, the redundancy 1132 corresponds to the information on the path redundancy in the path redundancy table 500 shown in Fig. 7. Here, when the analysis unit 20 detects a failure sign in the server 1111 or 1112 that is the failure sign detection target, the analysis unit 20 notifies the system operator 480 (operation terminal 30) (process (1)).
[0100] For example, the analysis unit 20 notifies the system operator 480 of appropriate measures that are predefined according to the suspected part (the equipment that is the target of failure sign detection) as the system information 1130. Furthermore, the system operator 480 determines how to handle the situation, such as switching within the production data center 1110, switching to the disaster recovery data center 1120, or waiting and seeing, depending on the redundancy and importance of devices such as servers.
[0101] For example, in the production data center 1110, the server 1111 or 1112 can be switched to the standby server 1113 (process (2)). Alternatively, the production data center 1110 can be stopped and the disaster recovery data center 1120 can be operated to switch the data center (process (2)'). After confirming the detection of a sign of failure, the system operator 480 can also arrange for maintenance parts to be replaced as a preventative measure (process (3)). Then, after the preventive replacement of the maintenance parts, the system operator 480 can restart the server 1111 or 1112 in the production data center 1110 and switch the operating entity back to the production data center 1110 (process (4)).
[0102] In this embodiment, the level of response is determined according to the redundancy and importance of the equipment, such that preventive maintenance is not performed on systems with low importance, but rather a wait-and-see approach is taken, and truly important systems are switched to disaster recovery systems, etc. This makes it possible to optimize the total cost of preventive maintenance as a whole.
[0103] Basically, the system operator 480 checks the information (redundancy, importance) provided by the analysis unit 20 and displayed on the operation terminal 30, and determines what action to take. Note that the analysis unit 20 may be configured to associate the threshold value table with the switching action content and propose the action to the customer through the operation terminal 30.
[0104] When the system operator 480 confirms the detection of multiple signs of failure, he or she can check the system information 1130 (system importance, redundancy) displayed on the operation terminal 30 and determine the priority.
[0105] As described above, the present invention is not limited to the above-described embodiments, and various other modifications and applications are possible without departing from the spirit of the invention as set forth in the claims. For example, the above-described embodiments have been described in detail and specifically to clearly explain the present invention, and are not necessarily limited to those including all of the components described. Furthermore, it is also possible to add, replace, or delete other components to or from part of the configuration of each embodiment.
[0106] Furthermore, the above-described configurations, functions, processing units, etc. may be partially or entirely realized in hardware, for example, by designing them as integrated circuits, etc. As the hardware, a broad processor device such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit) may be used.
[0107] Furthermore, each component of the preventive maintenance support system according to the above-described embodiment may be implemented in any hardware as long as the respective hardware can transmit and receive information to each other via a network. Furthermore, the processing performed by a certain processing unit may be realized by a single piece of hardware, or may be realized by distributed processing using multiple pieces of hardware.
[0108] In the above-described embodiment, the control lines and information lines are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are interconnected.
[0109] In this specification, processing steps describing chronological processing include not only processing performed chronologically in the order described, but also processing that is not necessarily performed chronologically but is performed in parallel or individually (for example, processing by objects). Furthermore, the processing order of processing steps describing chronological processing may be changed as long as it does not affect the processing results. [Explanation of symbols]
[0110] 1...Preventive maintenance support system, 10...IT system, 11...Monitored device, 12...Management unit, 20...Analysis unit, 21...Graph creation unit, 22...Cost estimation unit, 30...Operation terminal, 210...Graph, 220...Maintenance cost, 40...Computer, 41...CPU, 42...ROM, 46...Non-volatile storage, 470...Data table, 480...System operator, 500...Path redundancy table, 600...Business importance table, 700...Investigation target data table, 800...Threshold value table
Claims
1. A preventive maintenance support system in which a computer having an arithmetic unit that executes a program and a storage device that stores the program executes a process that supports preventive maintenance for a monitored device, The computing device acquires log information of the equipment installed in the monitored device, and identifies equipment to be replaced as a preventive replacement target based on information indicating the status of the equipment included in the log information and a reference value of a failure sign determined according to the redundancy and importance of the equipment, and presents the identified equipment to an operation terminal. Preventive maintenance support system.
2. The importance of the device is the importance of the device in the monitored device to which the device is attached, or the importance of the device itself. The preventive maintenance support system according to claim 1 .
3. The computing device calculates in advance the cost of implementing preventive replacement of the device that is the preventive replacement target, and presents the calculated cost to the operation terminal. The preventive maintenance support system according to claim 2 .
4. the device installed in the monitored device is an optical transceiver that converts between electrical signals transmitted and received by communication devices and optical signals transmitted through a cable, The information representing the device status is the amount of light output by the optical transceiver. The preventive maintenance support system according to claim 1 .
5. the device installed in the monitoring target device is a storage drive, The information indicating the device status is the number of errors or the number of writes of the drive. The preventive maintenance support system according to claim 1 .
6. 1. A preventive maintenance support method in which a computer having an arithmetic unit that executes a program and a storage device that stores the program executes a process that supports preventive maintenance for a monitored device, A process in which the arithmetic device acquires log information of a device installed in the monitoring target device; a process in which the computing device identifies a device to be replaced as a preventive replacement target based on information indicating the state of the device included in the log information and a reference value of a failure sign determined according to the redundancy and importance of the device; and a process in which the arithmetic device presents the device to be replaced to an operation terminal. Preventive maintenance support methods.
Citation Information
Patent Citations
Monitoring apparatus, monitoring method, and program
JP2015164005A
Storage control device and control program
JP2017068754A
Storage system and memory device fault recovery method
WO2014132373A1
Maintenance planning device and maintenance planning method
JP2019133412A