Storage system
By using power-adjustable CPUs and control programs, the storage system maintains performance by reallocating power to redundant controllers, addressing the issue of blocked components.
Patent Information
- Application Number
- JP2023217108
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-12-22
AI Technical Summary
Existing storage systems experience performance degradation when components become blocked due to failures or maintenance, leading to overloaded controllers and reduced system performance.
Incorporating CPUs with adjustable power consumption and power control programs to dynamically allocate power to redundant controllers, optimizing power consumption and performance to maintain system efficiency.
The solution effectively suppresses performance degradation by reallocating power to normal controllers, ensuring high performance even when components are blocked.
Smart Images

Figure 2025100029000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a storage system, and is suitable for application to a storage system related to a technique for changing performance by changing supplied power, for example.
Background Art
[0002] Generally, in a storage system, a plurality of controllers are mounted for redundancy. When one of the controllers executing a certain process becomes blocked due to a failure or the like, the other controller can take over the process and continue the I / O process. However, as a result of taking over the process, the other controller may become overloaded, and the performance of the entire storage system may deteriorate.
[0003] In recent years, there has been a demand for a technique for optimizing the power allocation of a storage system and maintaining high performance while suppressing power consumption. For example, Patent Document 1 discloses a technique for controlling the power of a storage drive and dynamically optimizing the power allocation according to the load. Patent Document 1 determines the power allocation based on the configuration of a flash drive as the storage drive and the power supply capacity of the storage system.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the technology described in Patent Document 1, for example, it does not assume a state where some components constituting the storage system are blocked due to maintenance work or the occurrence of a failure. If some of these components are blocked, the performance of the entire storage system may decrease.
[0006] The present invention has been made in consideration of the above points, and intends to propose a storage system capable of suppressing a decrease in the performance of the entire storage system even after some components are blocked.
Means for Solving the Problems
[0007] In order to solve such problems, in the present invention, in a storage system having a plurality of storage drives that provide a data storage capacity and a plurality of storage controllers that execute data writing or reading processes with the storage drives, each of the plurality of controllers includes at least a component including a CPU (Central Processing Unit) whose performance can be changed by changing the amount of power supplied, and a memory storing a power control program for controlling a target value of power consumption of the component. When the power control program is executed by the CPU, in response to detecting a blocked controller among the plurality of controllers, a function of raising the target value of power consumption for the components included in the controller is executed in a normal controller constituting a redundant system for the controller that caused the blockage.
Effects of the Invention
[0008] According to the present invention, it is possible to suppress a decrease in the performance of the entire storage system even after some components are blocked.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Mode for Carrying Out the Invention
[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. FIG. 1 is a block diagram showing a configuration example of a storage system 100 according to a first embodiment. The storage system 100 includes one node 103 having a plurality of controllers 104 and a drive box 113 that stores a drive 115 which is a storage device. The node 103 includes, for example, two controllers 104. The drive box 113 includes a plurality of drives 115. In the present embodiment, the plurality of drives 115 are grouped by, for example, RAID (Redundant Array of Independent Disks), and a parity group is formed for each plurality of drives 115 belonging to each group.
[0011] The node 103 includes a FAN 112 in addition to the controller 104. The drive box 113 also includes a FAN 112. The controller 104 has a function of providing a volume to be read from and written to by the host terminal 101.
[0012] The controller 104 includes a memory data backup drive 106, a memory 107, a CPU 108, a power supply unit 109, an environmental microcomputer (hereinafter also abbreviated as "environmental MIC") 110, a front-end interface (hereinafter also abbreviated as "FE I / F") 105, and a back-end interface (hereinafter also abbreviated as "BE I / F") 111.
[0013] The drive 115 is, for example, an SSD (Solid State Drive) using a flash memory as a storage medium, an HDD (Hard Disk Drive) using a magnetic disk as a storage medium, or the like. The memory is, for example, a semiconductor memory such as a DRAM (Dynamic Random Access Memory). The memory data backup drive is, for example, a drive such as an SSD and is used to back up the contents of the memory, for example, when there is an external power loss.
[0014] The power supply unit 109 supplies power to the storage system 100. FAN 112a cools one controller 104. FAN 112b cools the other controller 104. FANs 112p to 112z independently cool each drive 115. In this embodiment, when there is no need to particularly distinguish these FANs 112a, 112b, 112p to 112z, they are simply collectively referred to as "FAN 112". The environmental MIC 110 acquires environmental information. Here, the environmental information is, for example, the temperature or power consumption of the controller 104. Also, the temperature of the drive 115 is acquired by each temperature sensor 116 mounted on each drive 115.
[0015] The FE I / F 105 is, for example, a Fibre Channel HBA (Host Bus Adapter) or a NIC (Network Interface Controller). The BE I / F 111 is, for example, a SAS (Serial Attached SCSI) HBA, a PCI Express (hereinafter abbreviated as "PCIe") adapter, or a NIC.
[0016] Each controller 104 and the drive 115 are connected by, for example, a switch 114. Furthermore, the CPUs 108 of the multiple controllers 104 are connected to each other by an interconnect such as PCIe. The CPUs 108 may also be connected to each other via, for example, a PCIe switch. The storage system 100 is connected to a storage area network (SAN) 102 such as Fibre Channel or Ethernet. Furthermore, the host terminal 101 is also connected to the SAN 102. The SAN 102 may include a switch or the like. Furthermore, multiple hosts may be connected to the SAN 102.
[0017] Fig. 2 is a diagram showing an example of programs and information stored in the memory 107 of the controller 104 shown in Fig. 1. The memory 107 has, as control programs, for example, a power control program 200, a cooling performance control program 201, a temperature control program 202, a fault monitoring program 203, and a load monitoring program 204.
[0018] The memory 107 has, as information used by the above-mentioned control programs, a system information management table 205, a power mode table 206, a parity group configuration information table 207, a temperature management table 208, and a blocked state management table 209. Note that the memory 107 may also have other information such as control programs and control data.
[0019] The power control program 200, for example, works in conjunction with a cooling performance control program 201 to optimize the power and performance of the storage system 100 in response to a blocked state or an uneven load.
[0020] As described above, the storage system 100 according to this embodiment includes a plurality of controllers 104. The plurality of controllers 104 include a CPU 108 as an example of a component whose performance can be changed by changing the supplied power, and a power control program 200 that controls the power consumption of the CPU 108. This component may be other than the CPU 108 described above. When one of the plurality of controllers 104 is blocked, the power control program 200 supplies as much power as possible to the CPU 108 of the normal other controller 104 that is the handover destination of the process being executed by the blocked one controller 104.
[0021] In other words, the storage system 100 of this embodiment includes a plurality of storage drives 115 that provide a data storage capacity, and a plurality of storage controllers 104 that execute data writing or reading processes with the storage drives 115. Each of these controllers 104 has a component (including at least the CPU) whose performance can be changed by changing the supplied power amount, and a memory 107 that stores a power control program 200 that controls the target value of the power consumption of the component 108. When the power control program 200 is executed by the CPU 108, in response to the detection of a blocked controller among the plurality of controllers 104, a normal controller 104 that constitutes a redundant system for the blocked controller 104 (that is, the normal other controller 104 that is the handover destination of the process being executed by the blocked one controller 104), a function of raising the target value of the power consumption for the component (for example, the CPU 108) included in the controller 104 is executed. Note that the detection of the blocked controller may be implemented by a program different from the power control program 200.
[0022] The power control program 200 monitors the load on the CPU 108 of one of the controllers 104 when the one controller 104 is not blocked. When a part of the CPUs 108 with a high load is detected within the one controller 104, instead of restricting the power supply and performance to other CPUs 108 that can tolerate a performance degradation, it supplies as much power as possible to the part of the CPUs 108 with a high load. In other words, when the power control program 200 is executed by the CPU 108, the plurality of controllers 104 monitor the loads of the components (e.g., the CPU 108) each has. When a controller 104 having a component exceeding a predetermined load criterion is detected, the target value of the power consumption for the component exceeding the load criterion within that controller 104 is increased, and for other components, if there is a component below the predetermined load criterion for that component, the function of reducing the target value of the power consumption and the performance of that component is executed.
[0023] After the blockage of one of the controllers 104 or the high load on the part of the CPUs 108 is resolved, the power control program 200 restores the power supply and performance to the CPU 108 of the one controller 104 to their original states. In other words, when the power control program 200 is executed by the CPU 108, in response to the recovery of the blocked controller 104, the function of reducing the target value of the power consumption for the recovered controller 104 and the controller 104 constituting the redundant system is executed. Also, a new target value of the power consumption is set for the recovered controller 104. Further, when the power control program 200 is executed by the CPU 108, in response to the load of the component (e.g., the CPU 108) for which the target value of the power consumption has been increased falling below a predetermined criterion, the function of reducing the target value of the power consumption and the performance of that component is executed.
[0024] As described above, the storage system 100 according to the present embodiment includes a FAN 112 that cools the controller 104 and a cooling performance control program 201 that controls the cooling performance of the FAN 112. When the cooling performance control program 201 receives an instruction to change the cooling performance from the power control program 200 or the temperature control program 202, it controls the power supply unit 109 and the FAN 109 to change the cooling performance. More specifically, when the power control program 200 causes one of the controllers 104 to become blocked, and the power supply to the CPU 108 of the other normal controller 104 or some of the CPUs 108 with high load is increased, the cooling performance control program 201 changes the power supplied to the CPU 108 of the other controller 104 or some of the CPUs 108, and in conjunction with this, strengthens the cooling performance of the FAN 112. Or, when there is no margin to strengthen the cooling performance of the FAN 112, it limits the power supplied to one of the controllers 104. In other words, the storage system 100 according to the present embodiment includes a cooling device (FAN 112) that cools each controller 104 and a cooling performance control program 201 stored in the memory 107 that controls the cooling performance of the FAN 112. When the cooling performance control program 201 is executed by the CPU 108, a function is executed to increase the output of the FAN 112 provided for the controller 104 whose target power consumption value has been increased or the controller 104 including components whose target power consumption value has been increased. Further, when the cooling performance control program 201 is executed by the CPU 108, a function to monitor the output of the FAB 112, and when it is detected by the monitoring function that the output of the FAN 112 has reached a predetermined standard, a function to reduce the target power consumption value set for the controller 104 provided with the detected FAN 112 is executed.
[0025] The storage system 100 according to this embodiment includes a FAN 112 that cools the controller 104 and a cooling performance control program 201 that controls the FAN 112. The cooling performance control program 201 restores the cooling performance of the FAN 112 to its original state in conjunction with the restoration of the power supply and performance to the CPU 108 of one of the controllers 104. In other words, when the cooling performance control program 201 is executed by the CPU 108, in response to detecting the controller 104 with a reduced target power consumption or the controller 104 including a component with a reduced target power consumption, the function of reducing the output of the FAN 112 provided for the controller 104 is executed.
[0026] The temperature control program 202 periodically monitors the temperature of the system and records the temperature status in the temperature management table 208. When the temperature of each component is likely to reach the threshold value, the temperature control program 202 sends an instruction to change the cooling performance to the cooling performance control program 201. The threshold value mentioned here is, for example, a temperature set with a margin with respect to the limit temperature that leads to the occurrence of a device failure.
[0027] The storage system 100 according to this embodiment includes nodes 103 each including a plurality of controllers 104. When one controller 104 malfunctions and a process handover is to be performed, the power control program 200 preferentially selects, as a process handover destination, a controller 104 among the plurality of controllers 104 that has a margin in power supply and load. When there is no margin in power supply and load for all of the plurality of controllers 104, the power control program 200 selects a normal other controller 104 belonging to the same node 103 as the malfunctioning one controller 104. In other words, the storage system 100 is a system including a plurality of nodes 103 each including a plurality of controllers 104. When the power control program 200 is executed by the CPU 108, a function is executed to allocate, as the malfunctioning controller 104 and the normal controller 104 constituting the redundant system, a controller 104 among the controllers 104 included in any of the nodes 103, whose target power consumption can be increased and whose load is below a predetermined standard. Alternatively, when the power control program 200 is executed by the CPU 108 and it is detected that the target power consumption of any controller included in any node has reached the upper limit where it can be increased or the load has reached a predetermined standard, a function is executed to allocate a normal controller 104 included in the same node 103 as the malfunctioning controller 104.
[0028] The storage system 100 according to this embodiment includes a plurality of drives 115 whose performance can be changed by changing the supplied power. When some of the plurality of drives 115 are blocked, this power control program 200 supplies as much power as possible to other drives 115 with high load, or when some of the plurality of drives 115 with high load are detected, instead of limiting the power supply and performance to other drives 115 where there is no problem even if the performance is degraded, it supplies as much power as possible to some of the drives 115 with high load. In other words, the power control program 200 is a power control program that further controls the target value of the power consumption of the storage drive 115. When the power control program is executed by the CPU 108 and a blocked drive 115 is detected among the plurality of storage drives 115, a function of raising the target value of the power consumption for at least one or more other storage drives is executed. Further, when the power control program 200 is executed by the CPU 108 and a storage drive 115 exceeding a predetermined load criterion is detected among the plurality of storage drives 115, the target value of the power consumption for the storage drive 115 exceeding the criterion is raised, and the target value of the power consumption and its performance for other storage drives 115 below the predetermined load criterion are lowered. This enables an appropriate response to an increase in load. Note that the blockage monitoring of the storage drive 115 may be implemented by a program different from the power control program.
[0029] The storage system 100 according to this embodiment includes FANs 112p to 112z (hereinafter also collectively referred to as "FAN 112") as an example of a cooling mechanism capable of cooling the drive 115, and a cooling performance control program 201 for controlling the cooling performance of the FAN 112. This cooling performance control program 201 strengthens the cooling performance of the FAN 112 in conjunction with the change in the power supplied to the drive 115 by the power control program 200, or limits the power supplied to some of the drives 115 when there is no margin to strengthen the cooling performance of the FAN 112. In other words, the storage system 100 of this embodiment has a cooling mechanism (FAN 112) capable of cooling the storage drive 115 and a cooling performance control program 201 stored in the memory 107 for controlling the output of the FAN 112. When the cooling performance control program 201 is executed by the CPU 108, a function of increasing the output of the FAN 112 provided for the storage drive 115 whose target power consumption value has been increased is executed. Further, when the cooling performance control program 201 is executed by the CPU 108, a function of monitoring the output of the FAN 112 and, when it is detected by the monitoring function that the output of the FAN 112 has reached a predetermined standard, a function of reducing the target power consumption value set for the storage drive 115 provided with the detected FAN 112 is executed.
[0030] The above power control program 200 restores the power and performance of some of the drives 115 or other drives 115 to their original states after some of the blocked drives 115 among the plurality of drives 115 are restored or after the high load on other drives 115 is eliminated. In other words, when the power control program 200 is executed by the CPU 108 and it is detected that the blocked storage drive 115 has been restored, a function is executed to lower the target value of the power consumption for the storage drive 115 among the plurality of storage drives 115 for which the target value of the power consumption has been increased in response to the blockage of the storage drive 115. Also, a new target value of the power consumption is set for the restored storage drive 115. Further, when it is detected that the load on a storage drive 115 for which the target value of the power consumption has been increased in response to exceeding a predetermined load criterion has fallen below the criterion, a function is executed to lower the target value of the power consumption and its performance for the storage drive 115.
[0031] As described above, the storage system 100 according to the present embodiment has, as an example of a cooling mechanism capable of cooling the drive 115, a FAN 112 and a cooling performance control program 201 for controlling the cooling performance of the FAN 112. This cooling performance control program 201 restores the cooling performance of the FAN 112 to its original state in conjunction with the restoration of the power and performance of the drive 115 by the power control program 200. In other words, when the cooling performance control program 201 is executed by the CPU 108, a function is executed to lower the output of the FAN 112 provided for the storage drive 115 in response to detecting the storage drive 115 for which the target value of the power consumption has been lowered.
[0032] On the other hand, the above-described failure monitoring program 203 periodically monitors and adds up the failure counts of each component and monitors the occurrence status of failure blockages. Also, the failure monitoring program 203 monitors the signs of failure blockages and immediately sends out a response instruction when a failure blockage occurs or when the occurrence of a failure blockage is predicted.
[0033] The load monitoring program 204 monitors the load on each component of the storage system 100 and records the load status in the system information management table 205. The load at this time is, for example, the load on the CPU 108 for the controller 104 and the load during I / O processing for the drive 115.
[0034] FIG. 3 is a diagram showing an example of the system information management table 205 shown in FIG. 2. The system information management table 205 stores information regarding the components mounted in the storage system 100.
[0035] The system information management table 205 manages the mounting positions, device identification information, status, power consumption, load, and power mode of each component (device). In this embodiment, the “~ column” indicates the entry in the column corresponding to “~” in the table-form data. For example, in the case of the “power mode column”, it represents the entry in the column “power mode” in the system information management table 205.
[0036] The mounting position of each component is, for example, among the FE I / F 105, BE I / F 111, CPU 108, and memory 107 mounted on the controller 104, all those mounted on controller #0 (in this embodiment, “#” indicates an identification number) are shown as the mounting position “0”. The mounting positions of other components are also managed as consecutive numbers based on the actual component configuration.
[0037] The device identification information is information used to discriminate which component on the control program side. The status is displayed as “Active” if it is in operation and “Inactive” if it is not operating normally for some reason such as when a failure occurs or it is a spare. The power consumption column indicates how much power each component is currently consuming.
[0038] For example, for each component within the controller 104, the power consumption is that which can be obtained in the environment MIC 110, and is a value different from the rated power consumption. However, for the power supply unit, it indicates the power supplied from, for example, an outlet or the like, rather than the actually consumed power. That is, the power consumption column of the power supply unit records the maximum power available for use by the controller.
[0039] In the load column, the load monitoring program 204 records how much processing load each component currently has. The power mode column is information managed by the power mode table 206 and is an indicator showing the magnitude of the supplied power and the operating performance.
[0040] FIG. 4 is a diagram showing an example of the power mode table 206 shown in FIG. 2. The power mode table 206 manages the power modes of the components (devices) mounted on the storage system 100.
[0041] The power mode table 206 manages the mounted devices, power modes, power consumption, and processing capabilities. The mounted device column corresponds to the device information column shown in the system information management table 205. The power mode indicates that '0' represents the state with the most supplied power, and '3' represents the state with the least power. The power consumption indicates the power consumed in each power mode and corresponds to the target value or upper limit value of the power that can be consumed set for each device. Therefore, changing the power mode corresponds to raising or lowering the target value or upper limit value of the power that can be consumed set for each device. The processing capability indicates how much performance can be exhibited when operating with the power consumption corresponding to each power mode.
[0042] Note that in this embodiment, among the components (devices), there may be those that have a difference in the number of corresponding power modes or do not have multiple power modes. Also, in this embodiment, although there are two items, power consumption and processing capability, it may include other information necessary for selecting the power mode, such as constraints (functions are restricted) when using the power mode.
[0043] FIG. 5 is a diagram showing an example of the parity group configuration information table 207 shown in FIG. 2. The parity group configuration information table 207 manages the parity group information of a plurality of drives 115. That is, the parity group configuration information table 207 manages which drive 115 each drive 115 mounted in the drive box 113 forms a parity group with. Note that "SPARE" of the parity group indicates that it is unused.
[0044] FIG. 6A is a diagram showing an example of the temperature management table 208 shown in FIG. 2. The temperature management table 208 manages the temperature information regarding the controller 104 and the drive 115. The temperature management table 208 manages the mounted device, the mounting position, the device identification information, the current temperature, the temperature threshold (with margin), and the temperature threshold (limit value).
[0045] Regarding the mounting position and the device identification information in the temperature management table 208, they are the same as those in the system information management table 205 described above. The current temperature indicates the temperature acquired from the environmental MIC 110 in the case of the controller 104, and indicates the temperature acquired from the temperature sensor 116 in the case of the drive 115.
[0046] The temperature threshold (limit value) indicates the limit temperature threshold (limit value) at which a failure may occur. The temperature threshold (with margin) column indicates a value with a margin with respect to the limit temperature threshold (limit value) at which a failure may occur. Note that in this embodiment, the margin is assumed to be about 70%, but actually it is assumed to be close to the temperature at which the storage system 100 issues an alert.
[0047] FIG. 6B is a diagram showing an example of the occlusion state management table 209 shown in FIG. 2. The occlusion state management table 209 manages the occlusion states of the controller 104 and the drive 115. The occlusion state management table 209 manages the mounted device, the mounting position, the device identification information, the failure count, and the occlusion state.
[0048] Among the occlusion state management tables 209, the mounting position and device identification information are the same as those in the system information management table 205. The failure count is a variable used by the control program to determine the presence or absence of a failure, and when it exceeds a certain threshold, it is determined that a failure has occurred. When the failure count exceeds the threshold, it is determined that a failure has occurred, and "failure occlusion" is recorded in the occlusion state column. Note that this process is performed by the failure monitoring program 203 in the control program. Also, regardless of the failure count, if a failure has already occurred and the occlusion state has been reached, or if a sign of failure occlusion is detected, the occlusion state column is similarly updated to failure occlusion. In addition, maintenance occlusion is also prepared as a state other than failure occlusion. If it is not in either occlusion state, the occlusion state column indicates normal.
[0049] Figure 7 is a flowchart showing an example of the procedure for occlusion processing during maintenance. The occlusion processing during maintenance is a process of occluding a part of the controller 104 in the node 103 accompanying the maintenance occlusion operation, and is executed by the power control program 200.
[0050] Here, the maintenance occlusion operation refers to an operation such as logically disconnecting the connection to one of the controllers 104, or an operation such as cutting off the power supply to the controller and changing the connection in order to perform the replacement process of the controller. Therefore, during the maintenance occlusion operation, in the storage system 100 as shown in FIG. 1, it operates with a single controller 104 configuration.
[0051] In this embodiment, the controller to be occluded is also referred to as the "occluded controller", and a normal controller belonging to the same node as the occluded controller is also referred to as the "redundant controller".
[0052] First, the power control program 200 reads the occlusion state management table 209 from the memory 107 and updates the occlusion state column of the controller 104 scheduled to perform the maintenance occlusion operation to the "maintenance occlusion state" (step S701). After that, the person in charge of the maintenance occlusion operation, that is, the maintenance staff, cuts off the power supply to the occlusion controller (step S702). However, if it is not necessary to cut off the power supply due to the maintenance occlusion operation, the power supply does not have to be cut off.
[0053] The power control program 200 calls the cooling performance control program 201 and enhances the cooling performance corresponding to the redundant controller 104 belonging to the same node 103 as the occlusion controller 104 (step S703). This is a countermeasure against heat generation due to the improvement of the processing performance of the redundant controller 104, and is implemented by the power control program 200 instructing the cooling performance control program 201. The details of the cooling performance enhancement process will be described later.
[0054] The power control program 200 sets the highest possible state, for example, power mode "0" or "1", for the CPU 108 of the redundant controller 104 that performs the failover process (step S704).
[0055] Power mode "0" is a mode that allows the operating performance to be increased so that the processing capacity is maximized, and accordingly, the power consumption is allowed to increase to the maximum. When the power consumption of the CPU 108 is increased, the power control program 200 increases the operating frequency to the maximum to improve the performance.
[0056] As a means of increasing the operating performance of the CPU 108, the number of cores constituting the CPU 108 may be increased. Also, when the read / write processing of the memory 107 increases with the improvement of the operating performance of the CPU 108, the power supply to the memory 107 may also be increased. Similarly, when the communication volume with the host terminal 101 or the drive 115 increases with the improvement of the operating performance of the CPU 108, the power supply of the FE I / F 105 and the BE I / F 111 may also be increased.
[0057] On the one hand, the power mode "1" is a mode that supplies power so that the processing capacity can be exerted up to about 80%. The power modes of the above-described components are determined with reference to the load recorded in the system information management table 205 and the processing capacity of the power mode table 206. If there is a component whose power mode has been changed, the power control program 200 updates the power mode column of the system information management table 205 to the power mode at the time of recording.
[0058] After the change of the power mode, a failover process is performed when the controller 104 becomes blocked (step S705). The failover process is a process for having the normal other controller 104 take over the processing of the blocked one controller 104 when one controller 104 becomes blocked. In a state where the one controller 104 is blocked, since there is only one controller, there is a risk that the processing performance will simply be halved as it is.
[0059] In contrast, in the present embodiment, the power control program 200 improves the operation performance of the redundant controller 104 that performs the failover process, thereby enabling performance improvement in a state where there is only one controller. Therefore, during the execution of the failover process, in addition to maintaining the power mode at the highest level, as long as the redundant controller 104 continues to operate alone even after the execution of the failover process, by maintaining the operation power mode of the redundant controller 104 at "0" or "1", the operation performance in a state where there is only one controller is maintained at a high level.
[0060] FIG. 8 is a flowchart showing an example of the procedure for sudden occlusion processing. The sudden occlusion processing is processing that occurs when a sudden failure occurs in one of the controllers 104 and the system enters an occlusion state after going through the operations during occlusion. Here, an event in which a sudden failure occurs and the system becomes occluded can be, for example, a situation where communication with a device is suddenly interrupted during the operation of the storage system 100, such as a device failure. Note that the event in which a sudden failure occurs and the system becomes occluded may be caused by something other than the above-described example. Note that the operations during occlusion are executed by the power control program 200.
[0061] The power control program 200 first refers to the occlusion state management table 209 to check the occlusion state of the controller 104 (step S801). Subsequently, the power control program 200 determines whether only one of the plurality of controllers 104 in the node 103 is in an occluded state (both controllers 104) or whether there is no occluded controller 104 (step S802). If it is determined that there is no occluded controller 104, or if both controllers are occluded, the power control program 200 ends the sudden occlusion processing.
[0062] On the other hand, if only one of the controllers 104 is occluded, the power control program 200 proceeds to the power mode switching process (step S803). In the power mode switching process, the power control program 200 first calculates the power that can be additionally supplied to the redundant controller 104 (step S803).
[0063] In the process of calculating the available power, the power control program 200 first refers to the system information management table 205 to obtain the power consumption of the redundant controller 104. At the same time, the power control program 200 checks the current supplied power recorded in the power consumption column of the power supply unit 109 and also obtains the maximum power that can be supplied to the redundant controller 104.
[0064] Subsequently, the power control program 200 refers to the power mode table 206 and calculates the power consumption taking into account the increase in the power mode of the controller 104 (step S804). The power control program 200 compares the calculated power consumption with the power that can be additionally supplied to the redundant controller 104, and determines whether the power mode can be increased (step S805). When the power control program 200 determines that the power mode can be increased, it executes step S806. On the other hand, when it determines that the power mode cannot be increased, it does not change the power mode and executes the failover process (step S810).
[0065] The power control program 200 calls the cooling performance control program 201 and enhances the cooling performance corresponding to the redundant controller located in the same node 103 as the closed controller 104 (step S806). This is a countermeasure against heat generation due to the improvement of processing performance. Details of this process will be described later.
[0066] The power control program 200 changes the setting of the maximum power mode obtained by calculation (for example, power mode "0") for the CPU 108 of the redundant controller 104 that performs the failover process (step S807).
[0067] In this embodiment, when the read / write processing of the memory 107 increases with the improvement of the operation performance of the CPU 108, the power supplied to the memory 107 may also be increased. Similarly, when the communication volume with the host terminal 101 or the drive 115 increases with the improvement of the operation performance of the CPU 108, the power supplied to the FE I / F 105 and the BE I / F 111 may also be increased. If there is a component whose power mode has been changed, the power mode column in the system information management table 205 is updated to the power mode at the time of recording.
[0068] The power control program 200 performs failover processing when the controller 104 malfunctions (step S810). Failover processing refers to the process of having the normal redundant controller 104 take over the processing being executed by the malfunctioning controller 104 when one of the controllers 104 malfunctions. In the state where the one controller 104 is malfunctioning, since there is only one controller, there is a risk that the simple processing performance may be halved as it is. Also, due to the system configuration, there are no redundant paths remaining, so there is a risk that the fault tolerance may decrease as it is.
[0069] To address these two problems, the power control program 200 can improve the overall performance of the storage system in a state where there is only one controller and reduce the risk related to fault tolerance by quickly completing the malfunction work by improving the operating performance of the normal redundant controller 104 that performs failover processing. Therefore, in this embodiment, during the execution of the failover processing, in addition to maintaining the power mode at the highest level, after the failover processing is completed, as long as only the redundant controller 104 continues to operate, the operating power mode of the redundant controller 104 is maintained in the state of "0" or "1" to maintain the operating performance in a state where there is only one controller at a high level.
[0070] FIG. 9 is a flowchart showing an example of the procedure of the recovery processing. The recovery processing is a process of recovering the malfunction state of the malfunctioning controller 104. Here, recovery refers to the operation of ending the processing at the time of malfunction and making the malfunctioning controller 104 available normally.
[0071] This recovery processing shows the flow until the processing that was failovered to the redundant controller 104 before the above-described malfunction processing is failbacked to the recovered one controller 104. Note that this recovery processing is executed by the power control program 200. This recovery processing may be automatically executed after the maintenance staff finishes physically replacing the controller 104, or may be activated by an external input based on the instruction of the maintenance staff.
[0072] First, the power control program 200 updates the occlusion state column of the one controller 104 in the occlusion state management table 209 to the "normal state" (step S901). Then, the power control program 200 refers to the system information management table 205, checks the load status of the CPU 108 in the redundant controller 104 (step S902), and determines whether the load is higher than a predetermined standard (step S903). In step S903, the state where the load is high, that is, the standard for determining the state where the load is high, is defined as the state where the load of the CPU 108 is, for example, 80% or more.
[0073] When the CPU load in the redundant controller 104 is less than 80%, the power control program 200 executes step S904, while when the load of the CPU is 80% or more, the power control program 200 executes step S907. Even when the load of the CPU is less than 80%, if it is expected that the processing will become high in the future, step S907 may be executed.
[0074] In steps S904 to S906, a process of returning the power mode raised at the time of occlusion to the normal (state) is performed. First, the power control program 200 calls the cooling performance control program 201 to return the cooling performance of the FAN 112 of the redundant controller 104 to normal (step S904).
[0075] Subsequently, the power control program 200 returns the power modes of the CPU 108, memory 107, FE I / F 105, and BE I / F 111 of the redundant controller to normal (step S905). The power control program 200 operates the power modes of the respective components of the controller 104 to be restored in the "normal" state (step S906).
[0076] Here, the "normal" state shown in steps S904, S905, and S906 is a state where an appropriate power mode is set for the current processing load. For example, when the current processing load is 40% and there is a CPU 108 with a power mode of "1", since the processable performance is 80% for example, it is considered that the operating performance is high with respect to the actual processing load.
[0077] In this case, even if the power mode of the CPU 108 is lowered to "2", it is considered that the current process can be sufficiently completed, so the power mode is lowered to "2". However, for example, when reducing the supply power of the CPU 108, the processing load may relatively increase due to a decrease in operating performance. Therefore, the power mode is lowered one step at a time. By setting to the minimum necessary power mode in this process, it becomes surplus power for improving performance during congestion.
[0078] In steps S907 to S908, the power control program 200 starts the operation of the controller 104 to be recovered while maintaining the high power mode of the redundant controller 104. First, the power control program 200 refers to the system information management table 205 and calculates the power that can be supplied to the recovery controller 104 (step S907). Then, based on the calculable supply power obtained by the calculation, the power control program 200 sets the minimum necessary power mode that can be set for each component of the recovery controller 104 and operates it (step S908).
[0079] In the recovery process of the controller 104, if there is a component whose power mode has been changed, the power control program 200 updates the power mode column of the system information management table 205 to the power mode at the time of recording (step S909).
[0080] The power control program 200 fails back the processing of the redundant controller 104 to the recovered controller 104 (step S910).
[0081] FIG. 10 is a flowchart showing an example of the procedure of the power mode change process. The power mode change process is a process of changing the power mode when the load on the CPU 108 of the controller 104 is high with respect to a predetermined standard. The power mode change process is executed by the power control program 200.
[0082] First, the power control program 200 refers to the system information management table 205 and checks the load status of the CPU 108 in the controller 104 (step S1001). Then, the power control program 200 refers to the power mode table 206 and acquires the information of the power mode (step S1002). In step S1003, the power control program 200 determines whether there is a CPU 108 with a high load in the controller 104 (step S1003).
[0083] When there is a CPU 108 with a high load, the power control program 200 executes step S1004, while when there is no CPU 108 with a high load, the power control program 200 executes step S1013. The criterion for determining the high load state in step S1002, that is, the state where the load is high, is assumed to be a state where the load of the CPU 108 is, for example, 80% or more.
[0084] Steps S1004 to S1006 show the power mode switching process when only the CPU 108 of one of the controllers 104 among the CPUs 108 stored in the controller 104 in the node 103 has a high load.
[0085] Based on the load information of the CPU 108 acquired in step S1001 and the information of the power mode acquired in step S1002, the power control program 200 determines whether there is room in the power mode of the low-load CPU 108, that is, the CPU 108 that does not satisfy the criterion for determining the high-load state (step S1004).
[0086] When it is determined that there is margin, the power control program 200 executes step S1005, while when it is not determined that there is margin, step S1006 is executed. In the present embodiment, the criterion for determining that there is margin is, for example, the case where the power mode capable of achieving the current processing load at the minimum is lower than the current power mode. That is, for example, when the current processing load is 40% and there is a CPU 108 with a power mode of 1, since the processable performance is 80% for example, it is considered that the operation performance is high with respect to the actual processing load. In this case, even when the power mode of the CPU 108 is lowered to "2", it is considered that the current processing can be sufficiently completed, so the power mode is lowered to "2".
[0087] In step S1005, the power control program 200 lowers the power mode of the low-load CPU 108. However, for example, when reducing the supply power of the CPU 108, since the operation performance may decrease and the processing load may relatively increase, the lowering of the power mode is carried out step by step.
[0088] In step S1006, the power control program 200 calculates the power that can be supplied to each component of the controller 104. The load status of the CPU 108 in the system information management table 205 acquired in step S1001 is used for this calculation. The power control program 200 calculates the surplus power that can be supplied to the controller 104 by calculating the difference between the maximum power that can be supplied to the power supply unit and the total power consumption of the components in operation.
[0089] Based on the obtained surplus power, the power control program 200 determines whether it is possible to raise the power mode of the high-load CPU 108 (step S1007). In this determination, the power control program 200 refers to the power mode table 206 and compares whether the increment in power consumption associated with raising the power mode falls within the range that can be supplied by the surplus power. If the power control program 200 can raise the power mode, it executes step S1008. On the other hand, if it is difficult to raise the power mode, it ends this power mode change process.
[0090] In steps S1008 to S1011, the power control program 200 raises the power mode of the high-load CPU 108. First, the power control program 200 calls the cooling performance control program 201 to enhance the cooling performance of the FAN 112 of the controller 104 (step S1008). Next, the power control program 200 raises the power mode of the high-load CPU 108 to the highest possible state (step S1009). The power mode at this time is determined based on the calculated surplus power.
[0091] After raising the power mode of the CPU 108, if there are other components within the controller 104 that require high performance with the improvement of the CPU performance, the power mode raising process for the component is also performed (step S1010). The power control program 200 updates the power mode column of the system information management table 205 (step S1011) and ends this power mode change process. The components for which the power mode column is updated are those that have had their power mode raised among steps S1008, S1009, and S1010. In each process from step S1008 to step S1011, if there are multiple high-load CPUs 108, the same process is performed for the multiple CPUs.
[0092] Steps S1012 to S14 are processes executed when the determination in step S1002 is NO, indicating that there are no blocked components in the controller 104 and there is no high-load CPU 108. First, in step S1012, the power control program 200 determines whether there is a margin in the processing power of each CPU 108. This determination is the same as the content performed in step S1004. If the power control program 200 determines that there is a margin in step S1012, it executes step S1013. On the other hand, if it does not determine that there is a margin, it ends the power mode change process.
[0093] In step S1013, the power control program 200 returns the power mode of each component in the controller 104 to "normal". "Normal" is a state where an appropriate power mode is set for the current processing load. The power mode reduction process of this process is the same as that in step S1006. Note that if it is assumed that the processing load will increase in the future or the processing load will become high due to the reduction of the power mode, the implementation of this process may not be necessary. After returning the power mode of each component, the power control program 200 also returns the power mode of FAN 112 to normal by the cooling performance control program 201 (step S1014).
[0094] FIG. 11 is a flowchart showing an example of the procedure of the cooling performance control process. The cooling performance control process is executed by the cooling performance control program 201. The cooling performance control program 201 is called and executed from the power control program 200 or the temperature control program 202. The cooling performance control program 201 controls the supply power of FAN 112 and each component to control the cooling performance.
[0095] The cooling performance control program 201 determines whether to enhance the cooling performance of the component (or conversely, return the cooling performance to "normal") in step S1101. Note that enhancing the cooling performance means increasing the output of the cooling device (FAN112 in this embodiment). This determination result depends on the program (power control program 200 or temperature control program 202) that calls the cooling performance control program 201. When the cooling performance control program 201 enhances the cooling performance, it executes step S1102, while when it does not enhance the cooling performance, it executes step S1108.
[0096] In step S1102, the cooling performance control program 201 refers to the system information management table 205 and checks the power mode of FAN112 corresponding to the component to be cooled. The components to be cooled are, for example, the controller 104, the drive 115, etc., and the corresponding FAN112s are mounted on the node 103 and the drive box 113 respectively. Subsequently, the cooling performance control program 201 determines whether it is possible to enhance the cooling performance of FAN112, that is, whether it is possible to increase the power mode (step S1103). When the cooling performance control program 201 can enhance the cooling performance, it executes step S1104, while when it cannot enhance the cooling performance, it executes step S1107.
[0097] In step S1104, the cooling performance control program 201 increases the power supplied to FAN112 and sets the power mode to "0". By this process, the cooling performance control program 201 increases the rotation speed of FAN112 and enhances the cooling performance (step S1105). When the power mode of FAN112 is changed, the cooling performance control program 201 updates the power mode column of the corresponding FAN in the system information management table 205 to the power mode at the time of recording.
[0098] When the supply power of FAN112 has already reached its maximum and it is not possible to enhance the cooling performance by controlling the rotation speed, the cooling performance control program 201 suppresses the heat generated by changing the power mode of the component (step S1107). Here, the reduction range of the power mode may be one step or more. When the power mode of the component is changed, the cooling performance control program 201 updates the power mode column of the component in the system information management table 205 to the power mode at the time of recording.
[0099] In step S1108, the cooling performance control program 201 returns the power mode of FAN112 corresponding to the component to "normal". The "normal" at this time is the state of the power mode "1" or "2". When the cooling performance control program 201 changes the power mode of FAN112, it updates the power mode column of the FAN in the system information management table 205 to the power mode at the time of recording.
[0100] When the cooling performance control program 201 changes the power mode of FAN112 or the component, it updates the power mode column in the system information management table 205 to the power mode at the time of recording.
[0101] When there is no need to enhance the cooling performance, the cooling performance control program 201 returns the power mode of FAN112 to "normal". In this case, the cooling performance control program 201 does not need to strictly determine the power mode of FAN112, and the power mode may increase or decrease according to the temperature state of the system. When the power mode of FAN112 is changed, the cooling performance control program 201 updates the power mode column of the FAN112 in the system information management table 205 to the power mode at the time of recording.
[0102] FIG. 12 is a flowchart showing an example of the procedure of the temperature management process. The temperature management process is executed by the temperature control program 202. The temperature management process monitors the temperature regularly and asynchronously, for example, and manages the temperature by calling the cooling performance control program 201.
[0103] The temperature control program 202 first refers to the temperature management table 208 to check the temperature threshold of the controller 104 or the drive 115 (step S1201).
[0104] The temperature control program 202 checks the temperature of the controller 104 from the environmental MIC 110 installed in the controller 104. Alternatively, the temperature of the drive 115 is obtained from the temperature sensor built into the drive 115 (step S1202).
[0105] The temperature control program 202 compares the acquired current component temperature with the temperature threshold (step S1203). As a result of the comparison, if the component temperature does not exceed the temperature threshold, the process ends. If the acquired component temperature exceeds the temperature threshold, the cooling performance control program 201 is instructed to enhance the cooling performance (step S1204).
[0106] FIG. 13 is a flowchart showing an example of the procedure of the failure monitoring process. The failure monitoring process is executed by the failure monitoring program 203. In the failure monitoring process, for example, the failure status is monitored periodically and asynchronously, and the update of the failure occlusion state is performed.
[0107] The failure monitoring program 203 first refers to the occlusion state management table 209 to check the failure count of each component in the storage system 100 (step S1301). The failure count here can be, for example, the number of read errors, which is the number of data that could not be read when attempting to read. Also, when showing details for each drive type, examples of failure counts are as follows. That is, for example, if drive 115 is an HDD (Hard Disk Drive), the failure count can be the number of sectors with alternative sector processing completed. Note that a sector is the unit for reading and writing data on an HDD. When a bad sector is found, a process of replacing it with a spare is performed. Also, for example, if drive 115 is an SSD (Solid State Drive), the failure count can be the number of failed blocks. Note that a block is the unit for deleting data in NAND flash and indicates the number of blocks that can no longer be read or written.
[0108] Subsequently, the failure monitoring program 203 determines whether there is a component that has been blocked due to a failure (hereinafter also referred to as "failure occlusion") (step S1302). If the failure monitoring program 203 determines that there is one, it executes the failure occlusion time process (step S800), while if it determines that there is none, it executes step S1303.
[0109] After step S1303, the occurrence of events for which failures and failure counts should be added is monitored. First, in step S1303, the failure monitoring program 203 checks the failure count. In step S1304, the failure monitoring program 203 checks whether the failure count exceeds the failure occurrence threshold. If the failure monitoring program 203 determines that the failure count exceeds the failure occurrence threshold, it executes step S1305, while if it determines that the failure count does not exceed the failure occurrence threshold, it executes step S1306.
[0110] In step S1305, the failure monitoring program 203 updates the failure status column in the blockage status management table 209 to "failure blockage". The failure monitoring program 203 executes the processing at the time of failure blockage.
[0111] In step S1306, the failure monitoring program 203 monitors the failure situation. This failure monitoring program 203 determines whether an event in which the failure count is incremented has occurred during the failure monitoring (step S1307). When the event occurs, the failure monitoring program 203 executes step S1308, while when the event does not occur, it executes step S1309.
[0112] In step S1308, the failure count in the blockage status management table is incremented. On the other hand, in step S1309, the failure monitoring program 203 determines whether a sign of failure or a sudden failure occurs. When a sign of failure or a sudden failure occurs, the failure monitoring program 203 executes step S1305, while when not, it ends this failure monitoring process.
[0113] FIG. 14 is a flowchart showing an example of the procedure of the load monitoring process. The load monitoring process is executed by the load monitoring program 204. In this load monitoring process, for example, the load status of each component is monitored periodically. The "load" mentioned here is, for example, the load status of the CPU 108 for the controller 104, and for example, the I / O processing load for the drive 115. The load information to be monitored is not limited to these, and for example, load information about other components or other items may be used.
[0114] First, the load monitoring program 204 acquires the load of each component (step S1401). The information regarding the load obtained here (also referred to as "load information") is recorded in the load column of the system information management table 205 (step S1402).
[0115] (2) Second Embodiment The storage system according to the second embodiment has substantially the same configuration and operation as the storage system according to the first embodiment. Therefore, descriptions of the same configuration and operation will be omitted, and hereinafter, mainly the different points will be described.
[0116] FIG. 15 is a flowchart showing an example of the procedure of drive blockage processing for blocking a part of the drive accompanying the replenishment work. The drive blockage processing is executed by the power control program 200. In the drive blockage processing, the case of blocking a part of the drive 115 accompanying the maintenance work is shown.
[0117] Here, the maintenance blockage work refers to, for example, work for logically disconnecting the connection to a certain drive 115, or work such as cutting off the power supply to the drive 115 or changing the connection for performing the replacement process of the drive 115.
[0118] First, the power control program 200 reads out the blockage state management table 209 from the memory 107, and updates the blockage state column of the drive 115 to be blocked (hereinafter also referred to as "blocked drive") planned to perform the maintenance blockage work in the table to "maintenance blockage" (step S1501). Then, the person in charge of performing the maintenance blockage work, that is, the maintenance staff, cuts off the power supply to the blocked drive 115 (step S1502). However, if it is not necessary to stop the power supply to the drive 115 accompanying the maintenance blockage, the power supply does not have to be cut off.
[0119] After cutting off the power supply, the power control program 200 refers to the parity group configuration information table 207, and selects the drive belonging to the same parity group as the blocked drive 115 as the "performance-required drive" (step S1503). Note that the "performance-required drive" is a drive expected to have a high load as the read or write amount increases due to the processing at the time of blockage, and one or more drives may be selected.
[0120] Subsequently, the power control program 200 calls the cooling performance control program 201 to enhance the cooling performance of FAN 112 corresponding to the performance-required drive 115 (step S1504). Next, the power control program 200 sets the power mode of the performance-required drive 115 to the highest possible state, for example, power mode "0" or "1" (step S1505).
[0121] Power mode "0" is a mode that raises the operating performance so that the processing capacity is maximized and allows the power consumption to increase to the maximum accordingly. For example, power mode "1" is a power mode that supplies power so that the processing capacity can be exerted up to about 80%, and in power mode "0", power is supplied so that the processing capacity can be exerted 100%.
[0122] The power modes of the above-described components are determined with reference to the load information recorded in the system information management table 205 and the processing capacity of the power mode table 206. When there is a component whose power mode has been changed, the power control program 200 updates the power mode column in the system information management table 205 to the power mode at the time of recording (step S1506).
[0123] However, when there are multiple performance-required drives 115 and it is difficult to uniformly raise the power mode, it is not necessary to raise the power mode, or the power that can be supplied to each drive can be calculated, and the performance can be improved as much as possible. When the power that can be supplied is calculated, the power mode column in the system information management table 205 records the smallest power mode defined in the power mode table 206 that does not exceed the current supply power. That is, for example, when the drive 115 is an SSD and the supply power is 17W, the power mode is recorded as "1".
[0124] FIG. 16 is a flowchart showing an example of the procedure for sudden occlusion processing. FIG. 16 corresponds to FIG. 8 in the first embodiment. The sudden occlusion processing is executed by the power control program 200. The sudden occlusion processing is processing when a failure occurs in the drive 115 and it becomes an occlusion state. Here, "failure occlusion" refers to a situation where communication with a device is suddenly interrupted during the operation of the storage system 100, such as a component failure. Note that when a failure occurs, the situation may be other than the examples shown above.
[0125] The power control program 200 first refers to the occlusion state management table 209 to confirm the occlusion state of the drive 115 (step S1601). The power control program 200 determines whether there is an occluded drive among the drives 115 (step S1602). If there is an occluded drive 115, the power control program 200 proceeds to the power mode switching process (step S1603). On the other hand, if there is no occluded drive 115, the power control program 200 ends the sudden occlusion processing.
[0126] After step S1603, the process of switching the power mode for performance improvement during the failure occlusion process is shown. First, in step S1603, the power control program 200 refers to the parity group configuration information table 207 and selects the drive 115 belonging to the same parity group as the occluded drive 115 as the "performance required drive". Subsequently, in step S1604, the power control program 200 refers to the system information management table 205 to acquire the power consumption and power mode of the performance required drive 115.
[0127] Further, the power control program 200 refers to the power mode table 206 and obtains the power consumed in each power mode. In step S1605, the power that can be supplied to the performance-required drive is calculated. In the calculation process of the available power, the power supplied to the power supply unit 109 of the drive box 113 is confirmed, and the maximum power that can be supplied to the performance-required drive 115 is calculated. Subsequently, the power control program 200 calculates the power consumption taking into account the increment of the power mode of the performance-required drive 115 being raised. The power control program 200 compares this value with the additional power that can be supplied to the performance-required drive 115 (step S1606). If the power control program 200 determines that the power mode can be raised, it executes step S1607. On the other hand, if it determines that the power mode cannot be raised, it does not change the power mode and performs the operation during failure occlusion (step S1610).
[0128] In step S1607, the power control program 200 calls the cooling performance control program 201 to enhance the cooling performance corresponding to the performance-required drive 115.
[0129] Next, the power control program 200 changes the power mode of the performance-required drive 115 to the maximum setting obtained by calculation (step S1608). If there is a component whose power mode has been changed, the power control program 200 updates the power mode column in the system information management table 205 to the power mode at the time of recording.
[0130] After changing the power mode, the power control program 200 performs the operation during failure occlusion of the drive 115 (step S1610).
[0131] Figure 17 is a flowchart showing an example of the procedure of the recovery process. The recovery process is a process for recovering the occlusion state of the occluded drive 115. Here, recovery means, for example, that the process during occlusion is completed and the drive 115 is in a state where it can be set to a communicable state again. The power control program 200 controls the supply power of the drive 115 in the drive box 113 via the switch 114.
[0132] After the occlusion drive 115 recovers, the power control program 200 starts the recovery operation. This recovery operation may be automatically performed after the maintenance personnel finish the physical replacement of the controller 104, or may be activated by an external input according to the instructions of the maintenance personnel. First, the occlusion status column of the drive 115 in the occlusion status management table 209 is updated to "normal" (step S1701). Then, the power control program 200 refers to the system information management table 205, checks the load status of the performance-required drive 115 (step S1702), and determines whether the load is in a high state (step S1703).
[0133] In the determination of step S1703, for example, the state where the processing is high is defined as the state where the CPU 108 load is 80% or more. The power control program 200, for example, executes step S1704 when the load of the performance-required drive 115 is less than 80%, and executes step S1707 when it is 80% or more. Also, even when it is less than 80%, step S1707 may be executed if it is expected that the processing will become high in the future.
[0134] In steps S1704 to S1706, a process of returning the power mode raised during occlusion to the normal state is performed. First, the power control program 200 calls the cooling performance control program 201 and returns the cooling performance of the FAN 112 of the redundant controller 104 to "normal" (step S1704). Subsequently, the power control program 200 returns the power mode of the performance-required drive 115 to "normal" (step S1705). After the end of step S1705, the power control program 200 operates the power mode of the recovering drive 115 in "normal" (step S1706).
[0135] Here, the "normal" state shown in step S1704, step S1705, and step S1706 corresponds to the state of "1" or "2" as a power mode in principle. The determination of this power mode may be made according to the processing load of the CPU 108, or the power mode may be uniformly set to "2". In the recovery process of this drive 115, if there is a component whose power mode has been changed, the power mode column in the system information management table 205 is updated to the power mode at the time of recording (step S1707).
[0136] In steps S1708 to S1709, the operation of the recovery drive 115 is started while maintaining the power mode of the performance requirement drive 115 at a high level. First, the power control program 200 refers to the system information management table 205 and calculates the power that can be supplied to the recovery drive 115 (step S1708). Then, based on the calculated available power, the power control program 200 sets and operates the minimum necessary power mode that can be set for the recovery drive (step S1708).
[0137] FIG. 18 is a flowchart showing an example of the procedure of the power mode change process. The power mode change process is a process executed by the power control program 200 when the load of the drive 115 is biased among components. The power control program 200 controls the supply power of the drive 115 in the drive box 113 via the switch 114.
[0138] The power control program 200 first checks the parity group configuration information table 207 stored in the CPU 108 and obtains the parity group information for selecting one parity group (step S1801). Subsequently, the power control program 200 refers to the system information management table 205 and obtains the load of each drive 115 within the parity group (step S1802). Also, the power control program 200 refers to the power mode table 206 and obtains the information on the power mode as well (step S1803). The power control program 200 determines whether there is a drive 115 with a high load among the drives 115 of the parity group by using the information on the power mode (step S1804). If the power control program 200 determines that there is one, it executes step S1805, while if it determines that there is none, it executes step S1813.
[0139] In step S1805, the power control program 200 determines whether there is one or more drives 115 within the parity group that do not have a high load. Here, in this embodiment, "not having a high load" means, for example, a state where the processing load of the drive 115 is less than 80%. If the power control program 200 determines that there is one or more drives with a low load, it executes step S1806, while if it does not determine that there is one or more drives with a low load, it determines step S1808.
[0140] In step S1806, the power control program 200 determines whether there is a margin in the processing capacity of the low-load drive 115 determined in step S1805. If the power control program 200 determines that there is a margin in the processing capacity, it executes step S1807, while if it determines that there is no margin in the processing capacity, it executes step S1808. In this embodiment, the criterion for determining that there is a margin is that the power mode capable of achieving the current processing load at the minimum is lower than the current power mode.
[0141] That is, for example, when the current processing load is 40% and there is a drive 115 with a power mode of "1", it means that approximately half of the processing load is applied to 75% of the processable performance. In this case, for example, even if the power mode of the CPU 108 is lowered to "2" and set to a state where 50% of the performance can be exhibited, it is considered that the current processing can be sufficiently completed, so it is determined that the power mode can be lowered to "2".
[0142] In step S1807, the power control program 200 lowers the power mode of the low-load drive 115. However, for example, when reducing the supply power of the CPU 108, since the operation performance may decrease and the processing load may relatively increase, the lowering of the power mode is to be carried out one step at a time.
[0143] In step S1808, the power control program 200 calculates the power that can be supplied to the high-load drive 115 and the corresponding FAN 112. For this calculation, the power consumption information in the system information management table 205 obtained in step S1802 is used. The power control program 200 calculates the sum of the maximum power that can be supplied to the power supply unit and the total power consumption of the operating drives, and calculates the difference between the two to obtain the surplus power that can be supplied to the drive 115. Based on the obtained surplus power, the power control program 200 determines whether the power mode of the high-load drive 115 can be raised (step S1809).
[0144] This determination is made by the power control program 200 referring to the power mode table 206 and comparing whether the increase in power consumption associated with raising the power mode falls within the range that can be supplied by the surplus power. When the power mode can be raised, step S1810 is executed, while when it is difficult to raise the power mode, this process ends.
[0145] In steps S1810 to S1812, power mode boosting processing for the high-load CPU 108 is performed. First, the power control program 200 calls the cooling performance control program 201 to enhance the cooling performance of the FAN 112 of the controller 104 (step S1810). Then, the power control program 200 boosts the power mode of the high-load CPU 108 to the highest possible state (step S1811). The power mode at this time is determined based on the calculated surplus power. After the end of step S1811, the power control program 200 updates the power mode column of the system information management table 205 (step S1812) and ends this process. In each process from step S1808 to step S1812, if there are multiple high-load CPUs 108, the process is performed for the multiple CPUs 108.
[0146] Steps S1813 to S1815 indicate the case where there is no blockage in the drive 115 and there is no high-load drive 115. First, in step S1813, the power control program 200 determines whether there is a margin in the processing capacity of each drive 115. Step S1813 is the same as the content performed in step S1806. If it is determined in step S1813 that there is a margin in the processing capacity of each drive 115, step S1814 is executed. On the other hand, if it is not determined that there is a margin in the processing capacity of each drive 115, this process ends.
[0147] In step S1814, the power control program 200 returns the power mode of the drive 115 to "normal". "Normal" is a state where an appropriate power mode is set for the current processing load. Step S1814 is the same as the content performed in step S1807. Note that if it can be assumed that the processing load will increase in the future or the processing load will become high due to the lowering of the power mode, the power control program 200 may not execute step S1814. After the power control program 200 returns the power mode of each component, the power mode of the FAN 112 is also returned to "normal" by the cooling performance control program 201 (step S1815).
[0148] The storage system 100 according to the first and second embodiments (hereinafter also collectively referred to as "the present embodiments") as described above includes a plurality of storage drives 115 that provide data storage capacity, and a plurality of storage controllers 104 that execute data writing or reading processes with the storage drives 115. Each of these controllers 104 has a component (including at least a CPU) whose performance can be changed by changing the amount of power supplied, and a memory 107 that stores a power control program 200 for controlling the target value of the power consumption of the component 108. When the power control program 200 is executed by the CPU 108, in response to detecting a blocked controller among the plurality of controllers 104, a normal controller 104 that constitutes a redundant system for the controller 104 that caused the blockage (that is, the normal other controller 104 as the successor of the process executed by one of the blocked controllers 104) executes a function of raising the target value of the power consumption for the component (for example, the CPU 108) included in the controller 104.
[0149] In this way, even when one of the controllers 104 is blocked due to maintenance work, occurrence of a failure, etc. in a part of the CPUs 108 that make up the storage system 100, the power allocated to the CPUs 108 of the remaining normal redundant controllers 104 can be optimized, and the performance degradation of the entire storage system in a state where one of the controllers 104 is blocked can be suppressed.
[0150] (3) Third Embodiment The storage system according to the third embodiment has substantially the same configuration and operation as the storage system according to the first or second embodiment. Therefore, the description of the same configuration and operation will be omitted, and hereinafter, the description will mainly focus on the differences. In the storage system according to the third embodiment, although the reference numerals attached to its respective components are different from the reference numerals attached to the components of the storage system according to the first or second embodiment, each component of the storage system according to the third embodiment has substantially the same configuration as each component of the storage system according to the first embodiment and performs the same operation.
[0151] FIG. 19 is a diagram showing a configuration example of a storage system 1900 according to the third embodiment. The storage system 1900 includes two nodes 1903 with node numbers #0 and #1, and a drive box 1913 that stores a drive 1915 which is a storage device. That is, the storage system 1900 adopts a multi-node configuration that stores two nodes 1903.
[0152] Each node 1903 includes, for example, two controllers 1904. Specifically, the node 1903 with node number #0 includes, for example, two controllers 1904 with controller numbers #0 and #1. The node 1903 with node number #1 includes, for example, two controllers 1904 with controller numbers #2 and #3. In addition to the controller 1904, the node 1903 is also provided with FANs 1912a to 1912d.
[0153] FANs 1912a - 1912d are provided corresponding to each controller 1904 and independently cool each controller 1904. Also, the drive box 1913 is equipped with FAN 1912b. FANs 1912p - 1912z are provided corresponding to each drive 1915 within the drive box 1913 and independently cool each drive 1915. In this embodiment, when there is no need to particularly distinguish between FANs 1912a - 1912b and FANs 1912p - 1912z, they are simply collectively referred to as "FAN 1912".
[0154] The controller 1904 has a function of providing a volume for data reading and writing to the host terminal 1901. The controller 1904 includes a CPU 1908, a memory 1907, a memory data backup drive 1906, a power supply unit 1909, an environmental microcomputer (environmental MIC) 1910, a front - end interface (FE I / F) 1905, and a back - end interface (BE I / F) 1911.
[0155] The drive 1915 is, for example, an SSD (Solid State Drive) using a flash memory as a storage medium, an HDD (Hard Disk Drive) using a magnetic disk as a storage medium, etc. The memory 1907 is, for example, a semiconductor memory such as DRAM (Dynamic Random Access Memory). The memory data backup drive 1906 is, for example, a drive such as an SSD and is used to back up the content of the memory 1907, for example, when there is an external power loss.
[0156] The power supply unit 1909 supplies power to the storage system 1900. The environmental MIC 1910 monitors the surrounding environment. Here, the environment refers to, for example, the temperature status of the controller 1904. The FE I / F 1905 is, for example, a Fibre Channel HBA (Host Bus Adapter) or a NIC (Network Interface Controller). The BE I / F 1911 is, for example, a SAS (Serial Attached SCSI) HBA, a PCI Express (hereinafter referred to as PCIe) adapter, or a NIC.
[0157] Each controller 1904 and drive 1915 are connected, for example, by a switch 1914. Also, the CPUs 1908 of the plurality of controllers 1904 are connected by an interconnect such as PCIe. Note that the CPU 1908 and the CPU 1908 may be connected via a PCIe switch, for example.
[0158] The storage system 1900 is connected to a storage area network (SAN) 1902 such as Fibre Channel or Ethernet, for example. The host terminal 1901 is also connected to the SAN 1902.
[0159] The SAN 1902 may include a switch or the like. Also, a plurality of hosts may be connected to the SAN 1902. Also, although not shown in this storage configuration, the connection between the controllers 1904 is performed via an interconnect switch.
[0160] Figure 20 is a flowchart showing an example of the procedure of the sudden blockage process in the third embodiment. The sudden blockage process is executed by the power control program 200. The sudden blockage process is a process when a failure occurs in the controller 1904 and it enters a blocked state.
[0161] The power control program 200 first refers to the occlusion state management table 209 from the CPU 1908 to check the occlusion state of the controller 1904 (step S2101). Subsequently, the power control program 200 refers to the system information management table 205 to obtain the processing load and power mode of each component within the controller 1904. Also, the power control program 200 obtains the power consumption and achievable processing capacity when each power mode is set from the power mode table 206 (step S2102).
[0162] In step S2003, the power control program 200 changes the processing according to the occlusion state of the controller 1904 within the node 1903. When some of the controllers 1904 within the node 1903 are occluded, the power control program 200 executes step S2100. On the other hand, when all of the controllers 1094 within the node are occluded, the sudden occlusion process ends and is not targeted for performance improvement in this embodiment.
[0163] In step S2100 (failover destination controller selection process), steps S2004 to S2010, in order to improve the performance of the entire storage system in the state where some of the above-mentioned controllers 1904 are occluded, a process of changing the supplied power is performed. In step S2100, the power control program 200 selects the failover destination controller 1904. Details of the failover destination controller selection process will be described later. After the power control program 200 selects the failover destination controller 1904 in step S2100, in step S2004, it checks the mounting position of the selected controller 1904. When the selected controller 1904 is in the same node as the occluded controller 1904, the power control program 200 executes step S2005. On the other hand, when it is in another node, it executes step S2012.
[0164] In step S2005, the power control program 200 calculates the power that can be supplied to the selected controller 1904. In step S2006, the power control program 200 uses the calculated available power to check whether the power mode of the selected controller 1904 can be raised. If the power control program 200 can raise the power mode, it executes step S2007; if it cannot raise the power mode, it executes step S2011 (failover processing). Since the failover processing shown in step S2011 is the same as the failover processing in the first embodiment and the failover processing according to the second embodiment described above, the description thereof is omitted.
[0165] In step S2007, the power control program 200 instructs the cooling performance control program 201 to enhance the cooling performance of FAN1912 corresponding to the selected controller 1904. Subsequently, in step S2008, the power control program 200 raises the power mode of CPU1908 in the selected controller 1904 to the highest possible state.
[0166] Also, in step S2009, the power control program 200 may raise the power modes of the components in the selected controller 1904 according to the performance improvement of CPU1908. If there are components whose power modes have been changed, the power control program 200 changes the power mode column in the system information management table 205 to the recorded state.
[0167] In steps S2012 to S2013, it is the processing when the selected controller 1904 is in another node 1903. Based on the load information of the CPU 108 acquired in step S2001 and the information on the power mode acquired in step S2002, the power control program 200 determines whether there is margin in the power mode of the CPU 108 in another controller belonging to the same node 1903 as the selected controller 1904 (step S2012). When the power control program 200 determines that there is margin, it executes step S2013, while when it determines that there is no margin, it executes step S2005 described above.
[0168] In step S2013, the power control program 200 reduces the power mode of the CPU 1908 of another controller belonging to the same node 1903 as the selected controller 1904.
[0169] Figure 21 is a flowchart showing an example of the procedure of the failover destination controller selection process. The failover destination controller selection process is executed by the power control program 200. In the failover destination controller selection process, in order to increase the supply power and improve the performance when performing the failover process, the power control program 200 selects a controller 1904 that satisfies the condition that the CPU 1908 has a low load and there is room to increase the power mode (hereinafter referred to as the "condition") as the failover destination.
[0170] On the other hand, when there is no controller 1904 that meets the condition, the power control program 200 selects another controller 1904 in the node 1903 having the blocked controller 1904 as the failover destination.
[0171] The reason for preferentially selecting the controller 1904 is that the cache in the memory 1907 is duplicated within the node 1903, and recovery from a failure is faster within the same node 1903 as the blocked controller 1904. However, if the load on all normal controllers 1904 is high and the operating state already has a high supply power, it is assumed that no improvement in operating performance can be expected.
[0172] In the failover destination controller selection process, the power control program 200 first refers to the system information management table 205 to obtain the load and power mode of the CPU 1908 of the controller 1904 that can be a candidate for the failover destination (step S2101). Subsequently, the power control program 200 makes a determination according to the obtained load and power mode of the CPU 1908 as follows, and selects the controller 1904 as the failover destination.
[0173] Specifically, in step S2102, the power control program 200 checks the power mode of the CPU 108 of another controller within the same node 1903 as the blocked controller 1904. When the power control program 200 determines that there is room for increase, it executes step S2103, while when there is no room for increase, it executes step S2104. Here, having room for increasing the power mode means, for example, that the power mode is not "0".
[0174] In step S2103, the power control program 200 selects the controller 1904 within the same node 1903 as the blocked controller 1904 as the failover destination.
[0175] In step S2104, the power control program 200 checks the processing load on the CPU 1908 of another controller within the same node 1903 as the congestion controller 1904. If the load is high, the power control program 200 executes step S2105; if the load is not high (i.e., there is a margin), the power control program 200 executes step S2104. Here, a high load means, for example, that the load on the CPU 1908 exceeds 80%. Also, even if the load is less than 80%, if it is possible to predict that the load will exceed 80% in the future, step S2105 may be executed.
[0176] In step S2105, the power control program 200 changes the target controller 1904 to a controller 1904 belonging to a node 1903 different from the congestion controller 1904, and checks the power mode of the changed controller 1904. The power control program 200 checks the power mode of the changed controller 1904. If there is room to raise the power mode in any of the controllers 1904, the power control program 200 executes step S2107; if there is no room to raise the power mode, the power control program 200 executes step S2106.
[0177] In step S2106, the power control program 200 checks the processing load on the controller 1904 belonging to a node 1903 different from the congestion controller 1904. If the load on any of the controllers 1904 is high, the power control program 200 executes step S2103; if the load on any of the controllers 1904 is not high (i.e., there is a margin), the power control program 200 executes step S2107.
[0178] In step S2107, among the controllers belonging to a node 1903 different from the congestion controller 1904, the one with a greater margin for performance improvement is selected as the failover destination. However, a greater margin for performance improvement means that there is a greater margin to raise the power mode, or since the power modes are the same, the processing load is in a lower state.
[0179] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, each element described in parallel in the present embodiment may be in a mode in which at least one of the elements is connected in series to another element.
Industrial Applicability
[0180] The present invention can be applied to, for example, a storage system related to a technology for changing performance by changing supplied power.
Explanation of Signs
[0181] 100... Storage system, 108... CPU, 115... Drive, 116... Temperature sensor, 200... Power control program, 201... Cooling performance control program, 202... Temperature control program, 203... Failure monitoring program, 204... Load monitoring program, 205... System information management table, 206... Power mode table, 207... Parity group configuration information table, 208... Temperature management table, 209... Blocking state management table, 1900... Storage system, 1907... CPU, 1915... Drive, 1916... Temperature sensor
Claims
1. A plurality of storage drives that provide a data storage capacity, In a storage system having a plurality of storage controllers that perform data writing or reading processing with the storage drives, Each of the plurality of controllers, A component including at least a CPU (Central Processing Unit) whose performance can be changed by changing the amount of power supplied, A memory storing a power control program for controlling a target value of power consumption of the component, When the power control program is executed by the CPU, In response to detecting a blocked controller among the plurality of controllers, a function of raising the target value of power consumption for the components included in the controller is executed in a normal controller that constitutes a redundant system for the controller that caused the blockage. A storage system characterized by the above.
2. When the power control program is executed by the CPU, The plurality of controllers monitor the loads of the components each has, and when a controller having a component that exceeds a predetermined load standard is detected, the target value of power consumption for the component that exceeds the standard in the controller is raised, and for other components, the target value of power consumption and its performance for components that are below the predetermined load standard for the component are lowered. A function is executed. The storage system according to claim 1, characterized by the above.
3. A cooling device for cooling the controller, A cooling performance control program stored in the memory for controlling the cooling performance of the cooling device, When the cooling performance control program is executed by the CPU, A function of raising the output of the cooling device provided for the controller for which the target value of power consumption has been raised or the controller including the component for which the target value of power consumption has been raised is executed. The storage system according to claim 2, characterized by the above.
4. When the cooling performance control program is executed by the CPU, A function for monitoring the output of the cooling device, and when it is detected by the monitoring function that the output of the cooling device has reached a predetermined standard, a function for reducing the target value of power consumption set for the controller provided with the detected cooling device is executed. The storage system according to claim 3, characterized in that.
5. When the power control program is executed by the CPU, In response to the recovery of the blocked controller, a function for reducing the target value of power consumption for the recovered controller and the controller constituting the redundant system is executed. The storage system according to claim 4, characterized in that.
6. When the power control program is executed by the CPU, In response to the load of the component whose target value of power consumption has been increased falling below a predetermined standard, a function for reducing the target value of power consumption and performance for the component is executed. The storage system according to claim 5, characterized in that.
7. When the cooling performance control program is executed by the CPU, In response to detecting a controller with a reduced target power consumption value or a controller including a component with a reduced target power consumption value, a function for reducing the output of the cooling device provided for the controller is executed. The storage system according to claim 6, characterized in that.
8. A plurality of nodes including the plurality of controllers are provided, When the power control program is executed by the CPU, As the normal controller constituting the redundant system with the blocked controller, a function for allocating a controller whose target value of power consumption can be increased and whose load is below a predetermined standard from among the plurality of controllers is executed. The storage system according to claim 1, characterized in that.
9. When the power control program is executed by the CPU, As the blocked controller and the normal controller constituting the redundant system, when it is detected that the target value of the power consumption has reached the upper limit that can be increased or the load has reached the predetermined standard in any of the controllers included in the plurality of nodes, a function of allocating a normal controller included in the same node as the blocked controller is executed. The storage system according to claim 8, characterized in that.
10. The power control program is further a power control program that controls the target value of the power consumption of the storage drive, When the power control program is executed by the CPU, When a blocked drive is detected among the plurality of storage drives, a function of increasing the target value of the power consumption for at least one or more other storage drives is executed. The storage system according to claim 1, characterized in that.
11. When the power control program is executed by the CPU, When a storage drive exceeding a predetermined load standard is detected among the plurality of storage drives, the target value of the power consumption for the storage drive exceeding the standard is increased, and the target value of the power consumption and its performance for other storage drives below the predetermined load standard are decreased. A function is executed. The storage system according to claim 10, characterized in that.
12. A cooling mechanism capable of cooling the drive, A cooling performance control program stored in the memory for controlling the output of the cooling mechanism, Having, When the cooling performance control program is executed by the CPU, A function of increasing the output of the cooling device provided for the storage drive whose target value of power consumption has been increased is executed. The storage system according to claim 11, characterized in that.
13. When the cooling performance control program is executed by the CPU, A function of monitoring the output of the cooling device, and when it is detected by the monitoring function that the output of the cooling device has reached a predetermined standard, a function of decreasing the target value of the power consumption set for the storage drive provided with the detected cooling device is executed. The storage system according to claim 12, characterized in that.
14. When the power control program is executed by the CPU, when it is detected that the blocked storage drive has recovered, a function of reducing the target value of power consumption for the storage drive among the plurality of storage drives whose target value of power consumption has been increased in response to the blockage of the storage drive, and when it is detected that the load has fallen below the reference for the storage drive whose target value of power consumption has been increased in response to exceeding the predetermined load reference, a function of reducing the target value of power consumption and its performance for the storage drive is executed The storage system according to claim 13, characterized in that.
15. When the cooling performance control program is executed by the CPU, a function of reducing the output of the cooling mechanism provided for the storage drive is executed in response to detecting the storage drive whose target value of power consumption has been reduced The storage system according to claim 14, characterized in that.
Citation Information
Patent Citations
Power saving device by controller or disk control
JP2008250945A
System management methods and computer systems
JP2015531091A
Control apparatus, control method, and control program
JP2017084148A
Control device and control program
JP2020030670A
Information processing system, information processing program, information processing method, and information processing device
JP2020106927A