Computer system and storage management method

The management device in the remote copy system dynamically allocates resources for immediate failover and subsequent enhancement, addressing performance and cost challenges in disaster recovery scenarios.

JP2025094568APending Publication Date: 2025-06-25HITACHI VANTARA LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023210206
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing remote copy systems face challenges in maintaining host I/O processing performance during failover while minimizing hardware costs and avoiding recovery time violations, particularly when using cloud-based secondary sites.

Method used

A management device dynamically controls resource allocation at the secondary site, enabling immediate failover for critical volumes and subsequent resource enhancement, using scaling up or out to meet recovery requirements without excessive pre-allocated hardware.

Benefits of technology

This approach ensures compliance with recovery time objectives while reducing hardware costs and power consumption by optimizing resource usage during normal and failover operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094568000001_ABST
    Figure 2025094568000001_ABST
Patent Text Reader

Abstract

To provide a computer system that prevents violation of a recovery requirement when failing over to a secondary site in an event of a failure, while suppressing a cost of hardware at the secondary site during a normal operation.SOLUTION: A computer system includes a primary site storage system, a secondary site storage system, and a management device. The management device can change a resource of the secondary site storage system, and performs, when a failure occurs in the primary site, a failover in which a corresponding secondary volume takes over a business operation of a primary volume and controls to enhance a resource of the secondary site storage system, and controls the failover so that there are a secondary volume that is failed over and starts operating before the resource is increased, and a secondary volume that is failed over and starts operating after the resource is increased.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer system and a storage management method.

Background Art

[0002] There is a remote copy function as a technology for replicating a storage system between a plurality of geographically separated data centers in order to continue business even in the event of a disaster. In a storage system composed of a plurality of nodes and equipped with a remote copy function, the site that processes business applications during normal times is called the primary site, and when a site-wide failure occurs on the primary site side and the storage system stops, the site that is switched from the primary site and operated is called the secondary site.

[0003] For example, in Patent Document 1, in a configuration in which the secondary site is composed of a plurality of storage devices, when forming a remote copy pair between the primary site and the secondary site, the storage device on the secondary site side is selected so as to satisfy the performance and capacity requirements of the primary site, and a technique for constructing a pair is disclosed. This patent document states that "in an environment where the performance and capacity of the storage devices are different between the primary site that configures the remote copy and the recovery site, a configuration for reducing the load on the storage devices of the recovery site is provided. Based on the performance information and free capacity information of the storage at the recovery site, a second volume that forms a remote copy with the first volume is arranged in a storage device at the recovery site that can satisfy the performance requirements at the time of failure of the first volume provided to the host at the primary site. Further, a volume group is configured for each second volume arranged in one recovery site storage device and the first volume that forms a remote copy with the second volume, and further, an extended volume group including a plurality of second volumes in which data is replicated and recorded according to the write order to the first volume and the first volume is set in a management method of a computer system."

Prior Art Documents

Patent Documents

[0004] [Patent Document 1] WO2016 / 194096A1 Summary of the Invention [Problem to be solved by the invention]

[0005] In the above-mentioned Patent Document 1, in order to suppress a decrease in host I / O processing performance after switching (failover) from the primary site to the secondary site, the secondary site is provided with sufficient hardware to withstand remote copy processing and host I / O processing after failover. However, during normal operation, the secondary site only performs remote copy processing, and during normal operation, the hardware becomes excessive, which increases the introduction cost of the secondary site. Furthermore, if pay-as-you-go hardware such as that provided by a cloud vendor is used for the secondary site, the operating cost during normal operation increases by the amount of the excess hardware.

[0006] It is possible to configure the secondary site's hardware configuration to be capable of withstanding only remote copy processing during normal operation, and dynamically change the secondary site's hardware when the system state is switched due to failover and failback. However, in this case, dynamic changes to the hardware require time to be made, which may violate recovery requirements such as RTO (Recovery Time Objective). The present invention aims to prevent violation of recovery requirements such as RTO during failover to the secondary site when a failure occurs, while suppressing the cost of hardware at the secondary site during normal operation. [Means for solving the problem]

[0007] To achieve the above object, one of the typical computer systems of the present invention includes a primary site storage system that constructs a primary site providing a plurality of primary volumes to a host, a secondary site storage system that is connected to the primary site storage system via a network and constructs a secondary site providing a plurality of secondary volumes with remote copy set for the plurality of primary volumes, and a management device that manages the primary site storage system and the secondary site storage system. The management device can change the resources of the secondary site storage system. When a failure occurs in the primary site, the management device performs a failover in which the operations of the primary volumes are taken over by the corresponding secondary volumes, and controls to enhance the resources of the secondary site storage system, and controls the failover so that there are secondary volumes that perform failover and start operating before the enhancement of the resources and secondary volumes that perform failover and start operating after the enhancement of the resources. Also, one of the typical storage management methods of the present invention is a storage management method by a computer system including a primary site storage system that constructs a primary site providing a plurality of primary volumes to a host, a secondary site storage system that is connected to the primary site storage system via a network and constructs a secondary site providing a plurality of secondary volumes with remote copy set for the plurality of primary volumes, and a management device that manages the primary site storage system. The management device can change the resources of the secondary site storage system. When a failure occurs in the primary site, the management device performs a failover in which the operations of the primary volumes are taken over by the corresponding secondary volumes, and controls to enhance the resources of the secondary site storage system, and controls the failover so that there are secondary volumes that perform failover and start operating before the enhancement of the resources and secondary volumes that perform failover and start operating after the enhancement of the resources.

Advantages of the Invention

[0008] According to the present invention, it is possible to prevent a violation of recovery requirements during failover to a secondary site in the event of a failure while suppressing the cost of the hardware of the secondary site during normal operation.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Modes for Carrying Out the Invention

[0010] Embodiments will be described with reference to the drawings. Note that the embodiments described below do not limit the invention according to the claims, and not all of the elements and combinations thereof described in the embodiments are essential for the solution means of the invention.

[0011] In the following description, information may be described in terms of an "AAA table", but the information may be represented in any data structure. That is, in order to indicate that the information is independent of the data structure, the "AAA table" can be referred to as "AAA information".

[0012] Also, in the following description, the program may be the main body of the operation for the description of the process. However, the program is executed by a processor (e.g., a CPU), and in order to perform the defined process while appropriately using a storage resource (e.g., a memory) and / or a communication interface device (e.g., a NIC (Network Interface Card)), the main body of the process may be the processor. The process described with the program as the main body of the operation may be a process performed by a processor or a computer (system) having the processor.

[0013] Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0014] Also, in the following description, "VOL" is an abbreviation for a logical volume, and may be a logical storage device. VOL may be a physical volume (a volume based on a physical storage device) or a virtual volume.

[0015] Also, in the following description, an "instance" refers to a virtual computer (virtual machine) configured by software using resources on one or more physical computers.

[0016] In the following description, an "instance specification" is determined by a combination of specification values of resources such as CPU frequency, number of cores, memory speed, memory capacity, and network interface (I / F) bandwidth, and indicates the type of instance configuration. Note that the specification values may be CPU frequency, number of cores, memory speed, memory capacity, and network I / F bandwidth, or other values.

[0017] In the following description, "remote copy" is a function that replicates I / O data for a volume from a host to a volume at one site in order to ensure redundancy with a remote copy pair composed of a primary volume (PVOL) that is a volume at the primary site 10 and a secondary volume (SVOL) that is a volume at the secondary site 20. The process of volume replication to one site may be synchronous or asynchronous with respect to the I / O processing for the volume from the host.

Embodiment

[0018] FIG. 1 is an overall configuration diagram of a storage management system according to an embodiment. A storage management system 1 corresponding to the computer system of the claims includes a primary site 10 disposed on-premises, a secondary site 20 disposed in the cloud, and a disaster recovery management site 30, and the three sites are interconnected via an external network 400.

[0019] The primary site 10 includes one or more storage systems 100, one or more host computers 101, and a management device 102. The storage system 100, the host computer 101, and the management device 102 are connected via an internal network 103. The internal network 103 may be, for example, a LAN (Local Area Network), a WAN (Wide Area Network), or the like.

[0020] The storage system 100 is a device that provides a storage area for reading and writing data to and from the host computer 101. This storage area may be represented as volumes or LUNs. The storage system 100 may be a physical computer or a virtual computer.

[0021] The management device 102 is a computer used by a system administrator to manage the entire primary site 10. The management device 102 may be a physical computer or a virtual computer. The management device 102 acquires information from the storage system 100 and the host computer 101 by a program and displays the information via a user interface (GUI (Graphical User Interface), CLI (Command Line Interface)). The management device 102 has a function of transmitting an instruction input by the system administrator to the storage system 100 and the host computer 101 via the user interface. The management device 102 may have a function of automatically transmitting an optimal instruction to the storage system 100 and the host computer 101 without an instruction from the system administrator based on the information acquired from the storage system 100 and the host computer 101. Note that the functions of the management device 102 may be realized in any of the storage systems 100.

[0022] The host computer 101 is a computer that transmits read / write requests (hereinafter, appropriately referred to as I / O (Input / Output) requests) to the storage system 100 in response to requests from user operations or application programs (for example, file server programs, database server programs). The host computer 101 may be a physical computer or a virtual computer.

[0023] The secondary site 20 includes a storage cluster 2000, one or more host computers 201, and a management device 202. The storage cluster 2000 includes one or more storage nodes 200. The storage cluster 2000 may be called a storage system or a distributed storage system.

[0024] The storage node 200, the host computer 201, and the management device 202 are connected via an internal network 203. The internal network 203 may be, for example, a LAN (Local Area Network), a WAN (Wide Area Network), or the like.

[0025] The storage node 200 is a device that provides a storage area for reading and writing data to and from the host computer 201. This storage area may be expressed as a volume or a LUN. The storage node 200 may be a physical computer or a virtual computer.

[0026] The management device 202 is a computer used by a system administrator to manage the secondary site 20. The management device 202 may be a physical computer or a virtual computer. The management device 202 acquires information from the entire storage cluster 2000, the storage node 200, and the host computer 201 by a program, and displays the information via a user interface (GUI (Graphical User Interface), CLI (Command Line Interface)). The management device 202 has a function of transmitting an instruction input by the system administrator via the user interface to the entire storage cluster 2000, the storage node 200, and the host computer 201. The management device 202 may have a function of automatically transmitting an optimal instruction to the storage cluster 2000, the storage node 200, and the host computer 201 based on the information acquired from the storage cluster 2000, the storage node 200, and the host computer 201 without going through an instruction by the system administrator. Note that the function of the management device 202 may be realized in any of the storage nodes 200.

[0027] The host computer 201 is a computer that transmits read / write requests (hereinafter, these are appropriately referred to as I / O (Input / Output) requests) to the storage cluster 2000 in response to requests from user operations and application programs (for example, file server programs, database server programs). The host computer 201 may be a physical computer or a virtual computer. The host computer 201, for example, when a plurality of storage nodes 200 form a cluster, a multi-path is set up between the storage nodes that form the cluster. Note that any service can be used for the setting of the multi-path.

[0028] The disaster recovery management site 30 includes a disaster recovery management device 300. The disaster recovery management device 300 is connected to the primary site 10 and the secondary site 20 via an external network 400, and can perform operation management operations on each device in the primary site 10 and each device in the secondary site 20. The disaster recovery management device 300 may be a physical computer or a virtual computer.

[0029] Next, the storage node 200 will be described in detail. FIG. 2 is a configuration diagram of a storage node according to an embodiment. The storage node 200 includes one or more instances 210 and one or more storage devices 220. The instance 210 is a virtual computer configured by software using the resources of a physical computer in the cloud. The instance 210 may be a virtual machine.

[0030] Instance 210 includes a CPU 211, a memory 212, and a network I / F 213. The amounts of resources of the CPU 211, the memory 212, and the network I / F 213 in the instance 210 are amounts of resources corresponding to a predetermined instance specification. The CPU 211 is a virtual CPU to which a physical CPU of a cloud physical computer is virtually allocated. The CPU 211 performs processes such as access control to the storage device 220 based on programs and management information stored in the memory 212. The memory 212 is a virtual memory to which a physical memory of a cloud physical computer is virtually allocated. The memory 212 stores programs executed by the CPU 211 and management information referred to or updated by the CPU 211. The network I / F 213 is an I / F for communicating with the storage device 220, other storage nodes 200, the management device 202, the host computer 201, and the disaster recovery management device 300 via the network 400.

[0031] The storage device 220 is a physical or virtual storage device, and typically may be a non-volatile storage device. The storage device 220 may be, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage device 220 stores user data used by the host computer 201.

[0032] Next, the disaster recovery management device 300 will be described in detail. FIG. 3 is a configuration diagram of a disaster recovery management device according to an embodiment. The disaster recovery management device 300 includes a CPU 310, a memory 320, and a network I / F 330. The CPU 310 performs processes for controlling the storage system 100 and host computer 101 at the primary site 10, and the storage cluster 2000 and host computer 201 at the secondary site 20, based on programs and management information stored in the memory 320. The memory 320 stores programs executed by the CPU 310 and management information referred to or updated by the CPU 310. The network I / F 330 is an I / F for communicating with the storage system 100 and host computer 101 at the primary site 10, and the storage cluster 2000 and host computer 201 at the secondary site 20 via the network 400.

[0033] Next, the configuration of the memory 320 of the disaster recovery management device 300 will be described. FIG. 4 is a configuration diagram of the memory of the disaster recovery management device according to an embodiment. The memory 320 of the disaster recovery management device 300 stores a program 3200 and a management table 3300. The program 3200 includes a recovery requirement management program 3210, a performance information collection program 3220, a status monitoring program 3230, and a failover processing program 3240. The management table 3300 includes a recovery requirement management table 3310, an instance spec management table 3320, a storage node operation information management table 3330, and a remote copy pair management table 3340.

[0034] The recovery requirement management program 3210 collects recovery requirement information for each volume on the storage system 100 and the storage cluster 2000 based on the input by the administrator of the disaster recovery management device 300 and the information of the user's application program provided by the host computer and the storage system, and records the data in the recovery requirement management table 3310. The recovery requirement information is information indicating when to recover when a failure occurs in the storage system 100 of the primary site 10 and a failover is performed to the storage cluster 2000 of the secondary site 20. The recovery requirement information may be information with a clear time specification such as RTO, or may be information without specifying a clear time such as whether to perform immediate recovery.

[0035] The performance information collection program 3220 collects performance information from the storage node 200 and records the data in the storage node operation information management table 3330. The status monitoring program 3230 acquires the status of the primary site from the primary site 10 and detects whether the primary site 10 is in a down state. If the primary site 10 is in a down state, a failover instruction to the secondary site 20 is sent.

[0036] The failover processing program 3240 starts processing triggered by the switchover to the secondary site due to a failure at the primary site, and based on the information in the recovery requirement management table 3310, the storage node operation information management table 3330, and the remote copy pair management table 3340, determines the timing of resource enhancement such as scale-out and scale-up of the storage cluster 2000 and the timing of failover of each volume to meet the recovery requirements for each volume, and sends instructions for resource enhancement and failover to the storage cluster 2000.

[0037] Next, the recovery requirement management table 3310 will be described. FIG. 5 is a configuration diagram of a recovery requirement management table according to an embodiment. The recovery requirement management table 3310 is a table that manages the requirements for when to recover when failing over to the storage cluster 2000 of the secondary site 20 when a failure occurs in the storage system 100 of the primary site 10 of the storage management system 1. The recovery requirement management table 3310 stores entries for each remote copy pair composed of the volume of the storage system 100 and the volume of the storage cluster 2000. An entry in the recovery requirement management table 3310 includes fields of ID 3311, remote copy pair ID 3312, primary site VOL ID 3313, secondary site VOL ID 3314, and recovery requirement 3315.

[0038] In ID 3311, an identification number for each entry in the recovery requirement management table 3310 is stored. In the remote copy pair ID 3312, an identification number of the remote copy pair is stored, which is associated with the remote copy pair ID 3341 in the remote copy pair management table 3340. In the primary site VOL 3313, an identification number of the volume of the storage system 100 of the primary site 10 that constitutes the remote copy pair is stored. In the secondary site VOL 3314, an identification number of the volume of the storage cluster 2000 of the secondary site 20 that constitutes the remote copy pair is stored. In the recovery requirement 3315, recovery requirement information is stored. The recovery requirement information may be information with a clear time specification such as RTO, or may be information without specifying a clear time such as whether to perform immediate recovery.

[0039] Next, the instance spec management table 3320 will be described. FIG. 6 is a configuration diagram of an instance spec management table according to an embodiment. The instance specification management table 3320 is a table that manages the instance specifications available for the instance 210, and stores entries for each instance specification. The entries in the instance specification management table 3320 include fields of an instance specification ID 3321, a cost 3322, a CPU specification 3323, a memory capacity 3324, and a network bandwidth 3325.

[0040] The instance specification ID 3321 stores an identification number (instance specification ID) that uniquely identifies the instance specification corresponding to the entry. The cost 3322 stores the usage fee for the instance specification, for example, the price per hour. The CPU specification 3323 stores the frequency and the number of cores of the CPU allocated by the instance specification corresponding to the entry. The memory capacity 3324 stores the capacity of the memory (memory capacity) allocated by the instance specification corresponding to the entry. The network bandwidth 3325 stores the bandwidth of the network I / F (network bandwidth) allocated by the instance specification corresponding to the entry. Note that the instance specifications in the instance specification management table 3320 may include a plurality of instance specifications that differ only in the value of one of the resources (for example, the number of CPU cores).

[0041] Next, the storage node operation information management table 3330 will be described. FIG. 7 is a configuration diagram of a storage node operation information management table according to an embodiment. The storage node operation information management table 3330 is a table for managing information on the current state of the storage nodes 200 included in the storage cluster 2000. The storage node operation information management table 3330 stores an entry for each storage node 200. An entry in the storage node operation information management table 3330 includes fields for a node ID 3331, a node state 3332, a free capacity 3333, a CPU usage rate 3334, a memory usage rate 3335, a communication bandwidth usage rate 3336, and an instance spec ID 3337.

[0042] The node ID 3331 stores the identification number (node ID) of the storage node 200 corresponding to the entry. The node state 3332 stores the state of the storage node 200 corresponding to the entry. The node state 3332 may be information such as "normal" or "abnormal", for example. The free capacity 3333 stores the total free capacity of the storage devices 230 of the storage node 200 corresponding to the entry. The CPU usage rate 3334 stores the usage rate of the CPU 211 of the storage node 200 corresponding to the entry. The memory usage rate 3335 stores the usage rate of the memory 212 of the storage node 200 corresponding to the entry. The communication bandwidth usage rate 3336 stores the usage rate of the communication bandwidth at the communication I / F 240 of the storage node 200 corresponding to the entry. The instance spec ID 3337 stores the instance spec ID corresponding to the instance spec of the instance 210 of the storage node 200 corresponding to the entry, and is associated with the instance spec ID 3321 in the instance spec management table 3320.

[0043] An entry in the storage node operation information may include performance information of each storage node as time-series information. In that case, for example, a field such as the information acquisition time may be prepared.

[0044] Next, the remote copy pair management table 3340 will be described. FIG. 8 is a configuration diagram of a remote copy pair management table according to an embodiment. The remote copy pair management table 3340 is a table that manages information regarding the state of remote copy pairs, and stores an entry for each remote copy pair. An entry in the remote copy pair management table 3340 includes fields of a remote copy pair ID 3341, a primary site VOL ID 3342, a secondary site VOL ID 3343, a secondary site VOL destination node ID 3344, and a recovery state 3345.

[0045] The remote copy pair ID 3341 stores an identification number of the remote copy pair corresponding to the entry. The primary site VOL ID 3342 stores an identification number of the primary site VOL of the remote copy pair corresponding to the entry. The secondary site VOL ID 3343 stores an identification number of the secondary site VOL of the remote copy pair corresponding to the entry. The secondary site VOL destination node ID 3344 stores an identification number of the storage node 200 where the secondary site VOL of the remote copy pair corresponding to the entry is arranged. The recovery state 3345 stores information on the recovery state of the remote copy pair corresponding to the entry. Examples of the information on the recovery state may include contents such as "during failover", "failover completed", "moving secondary site VOL to another node", "preparing secondary site VOL destination node", and "processing not yet performed", and the state of the recovery process may be expressed by other contents.

[0046] Next, an overview from the steady state of the storage management system 1 to after the occurrence of a primary site failure will be described. FIG. 9 is a diagram for explaining an overview from the steady state of the storage management system 1 according to an embodiment to after the occurrence of a primary site failure. The upper part of FIG. 9 is an overview of the steady state of the storage management system 1, and the lower part is an overview when a failure occurs in the storage system 100 of the primary site 10 and a failover occurs to the storage cluster 2000 of the secondary site 20. In the storage system 100 of the primary site 10 placed on-premises, a PVOL which is a primary site VOL is defined, and in a steady state, I / O is possible from the host computer 101. In the storage node 200 of the storage cluster 2000 of the secondary site 20 placed in the cloud, an SVOL which is a secondary site VOL is defined, a remote copy pair with the PVOL is formed, and as the data on the PVOL is updated due to I / O to the PVOL from the host computer 101 in the steady state, data is transferred from the storage system 100 to the storage node 200 to update the data of the SVOL of the remote copy pair paired with the PVOL.

[0047] Each PVOL and each SVOL in FIG. 9 describe the PVOL and the SVOL so as to correspond to an example of an entry in the recovery requirement management table 3310 in FIG. 5 and an example of an entry in the remote copy pair management table in FIG. 8. For example, the VOL with the primary site VOL ID3313 being 1 in the entry with the ID3311 of the recovery requirement management table 3310 corresponds to the PVOL1 500a of the storage system 100, and the VOL with the secondary site VOL ID3314 being 1 corresponds to the SVOL1 510a of the storage node 200a. Regarding the recovery requirement 3315 as well, the flow of the process is described by associating the overview description in FIGS. 9 and 10 with the recovery requirement 3315 of each entry in the recovery requirement management table 3310.

[0048] In addition, the specifications of the instance 210 of the storage cluster 2000 of the secondary site 20 in the steady state only need to be, for example, performance that can withstand data transfer by remote copy and performance that can withstand I / O from the host computer 201 for the SVOL that requires immediate failover to the storage cluster 2000 of the secondary site 20 after a failure occurs in the storage system 100 of the primary site 10. The determination as to whether the instance specifications are performance that can withstand I / O from the host computer 201 for the SVOL that requires immediate failover may be made, for example, by collecting time-series data on the I / O volume and I / O performance from the host computer 101 to the storage system 100 in the steady state from the storage system 100 and estimating the achievable I / O performance while referring to the product specifications and spec sheets of the storage node 200. This estimation may be realized as a program on the disaster recovery management device 300.

[0049] After a failure occurs in the primary site 10, the SVOL that requires immediate failover immediately performs failover to enable I / O from the host computer 201. The determination that the SVOL requires immediate failover may be made, for example, when the information on the recovery requirement 3315 in the recovery requirement management table 3310 is "immediate failover required", or when the information on the recovery requirement 3315 has an RTO value set and the time of that RTO is shorter than the resource enhancement time of the storage cluster 2000.

[0050] Next, an overview from after a failure occurs in the primary site of the storage management system 1 to after scale-out of the secondary site will be described. FIG. 10 is a diagram for explaining an overview from after a failure occurs in the primary site of the storage management system 1 according to an embodiment to after scale-out of the secondary site.

[0051] The upper part of FIG. 10 shows an overview after a failure occurs in the storage system 100 of the primary site 10 and a failover to the storage cluster 2000 of the secondary site 20. The lower part shows an overview after enhancing the resources of the storage cluster 2000 (in the example of the figure, increasing the storage nodes 200 for scale - out), then moving the SVOL to the enhanced - resource storage nodes 200 and performing a failover. Regarding the overview after a failure occurs in the primary site shown in the upper part of FIG. 10, it is the state after the processing in the lower part of FIG. 9. Since the overview of the processing up to this state overlaps with the description in FIG. 9, it is omitted. Regarding the PVOL on the storage system 100 in the lower part of FIG. 10, it has the same defined state as the storage system 100 in the upper part of FIG. 10, but is omitted in the figure.

[0052] After all the failovers of the SVOLs that require immediate failover are completed, the processing in the lower part of FIG. 10 is entered. If there are SVOLs for which failover has not been performed and the existing built - in storage nodes 200 do not have sufficient performance to operate those SVOLs, the resources of the storage cluster 2000 are enhanced. The method of enhancing resources can be achieved by scale - up, which changes the instance spec of the storage node 200 to a higher - spec one, or by scale - out, which increases the storage nodes 200 in the storage cluster 2000. In the example of FIG. 10, an example of achieving resource enhancement by scale - out by adding storage nodes 200d and storage nodes 200e is described.

[0053] After resource enhancement, the SVOL is moved to the enhanced - resource storage nodes 200. Any method can be used for moving the SVOL. For example, the internal functions of the storage cluster 2000 can be used. After moving the SVOL, the moved SVOL is failovered to make it possible to perform I / O from the host computer 201. Until all the failovers of the SVOLs are completed, resource enhancement, SVOL movement, and SVOL failover processing are performed.

[0054] Resource enhancement, SVOL movement, and SVOL failover can be processed in parallel if there is no dependency between the SVOL and its destination storage node 200. For example, referring to the lower figure in FIG. 10, the failover process of SVOL3 510c, the movement of SVOL4 510d to storage node 200d, and the addition process of storage node 200e can be processed in parallel.

[0054] The failover order of the SVOL may be determined based on the recovery requirements. For example, it may be failover in the order of shorter RTO.

[0055] Which SVOL 510 is to be moved to which storage node 200 may be determined using, for example, an algorithm that performs bin packing based on the surplus performance of the storage node 200 and the I / O performance information required for the SVOL 510. The I / O performance required for the SVOL 510 may be based on, for example, the I / O performance information from the host computer 101 for the PVOL 500 that forms a remote copy pair with the SVOL 510 during normal operation, or may use the QoS setting value specified by the user.

[0056] Next, the recovery requirement management process by the recovery requirement management program 3210 will be described. FIG. 11 is a flowchart of the recovery requirement management process according to an embodiment.

[0055] The process of the recovery requirement management program 3210 is started, for example, by an instruction from the user.

[0057] The recovery requirement management program 3210 acquires recovery requirement information for each remote copy pair (step S4001). The recovery requirement information is information indicating when to recover when a failure occurs in the storage system 100 at the primary site 10 and a failover is performed to the storage cluster 2000 at the secondary site 20. The recovery requirement information may be information with a clear time specification such as RTO, or may be information without a clear time specification such as whether to perform immediate recovery.

[0058] The method for obtaining restoration requirement information may be obtained, for example, from the input of a user or an administrator, or by collecting the requirement information of the application program of the host computer that uses the volume of the remote copy pair managed by another program.

[0059] Next, the restoration requirement management program 3210 stores the restoration requirement information obtained in step S4001 in the restoration requirement management table 3310 (step S4002). When storing in the restoration requirement management table 3310, the restoration requirement 3315 may store the information obtained in step S4001 as it is, or may process the information, for example, round the RTO value, and then store it.

[0060] Next, the performance information collection process by the performance information collection program 3220 will be described. FIG. 12 is a flowchart of the performance information collection process according to an embodiment. The performance information collection process is executed periodically, for example, by the performance information collection program 3220. The performance information collection program 3220 obtains the performance information of the storage node 200 (step S4101). The method for obtaining performance information here may be, for example, to send a performance information acquisition request to each storage node 200 of the storage cluster 2000, and correspondingly, cause each storage node 200 to send performance information; or send a performance information acquisition request to the representative storage node 200 of the storage cluster 2000, and cause the representative storage node 200 to obtain the performance information of each storage node 200 and send the summarized information; or send a performance information acquisition request to the management device 202 of the secondary site 20, and cause the management device 202 to obtain the performance information of each storage node 200 and send the summarized information. Note that the performance information may be obtained in one communication or multiple communications.

[0061] Next, the performance information collection program 3220 stores the performance information obtained in step S4101 in the storage node operation information management table 3330 (step S4102). When storing the performance information in the storage node operation information management table 3330, the information obtained in step S4101 may be stored as it is. For example, if the obtained information is not a usage rate such as memory usage or communication bandwidth usage and cannot be stored as it is, the information may be processed and then stored in this step. The instance spec ID 3337 may store the instance spec of the storage node 200 in this storage node operation information management table 3330 when constructing the storage cluster 2000, or may be stored in a table separately prepared on the memory of the disaster recovery management device 300 or the management device 202, and the information may be obtained therefrom and stored in the storage node operation information management table 3330.

[0062] Next, the state monitoring process by the state monitoring program 3230 will be described. FIG. 13 is a flowchart of the state monitoring process according to an embodiment. The state monitoring process is executed periodically, for example, by the state monitoring program 3230. The execution period of the state monitoring process may be the minimum time required for disaster recovery requirements, and may be executed at a period in seconds, for example.

[0063] The state monitoring program 3230 acquires the state of the primary site (step S4201). The method for acquiring the state of the primary site here may be, for example, to send a state acquisition request to the storage system 100 of the primary site 10 and, in response, cause the storage system 100 to send state information, or the storage system 100 may periodically send state information to the disaster recovery management device 300, or the state information may be sent via the management device 102.

[0064] Next, based on the status information acquired in step S4201, the status monitoring program 3230 determines whether the storage system 100 at the primary site 10 is in a down state (S4202). If it is determined that the storage system 100 is in a down state (step S4202: YES), the process proceeds to step S4203. If it is determined that the storage system 100 is not in a down state (step S4202: NO), this status monitoring process ends. Next, the status monitoring program 3230 instructs the failover processing program 3240 to start processing (step S4203).

[0065] Next, the failover processing by the failover processing program 3240 will be described. FIG. 14 is a flowchart of the failover processing according to an embodiment. The failover processing is executed by the failover processing program 3240 that has received an instruction from the status monitoring program 3230.

[0066] The failover processing program 3240 issues a failover instruction to the storage cluster 2000 at the secondary site 20 so as to perform a failover of the secondary site VOL that requires an immediate failover (step S4301). To determine whether it is a secondary site VOL that requires an immediate failover, for example, the recovery requirement 3315 in the recovery requirement management table 3310 may be referred to. After issuing the failover instruction, it may or may not wait for the completion of the failover processing. If it does not wait for the completion of the failover processing after the failover instruction, for example, a process for monitoring the failover processing status may be incorporated. After the failover instruction, the recovery status 3345 in the remote copy pair management table 3340 is updated.

[0067] Next, the failover processing program 3240 determines whether there is a secondary site VOL for which failover processing has not been performed (step S4302). To determine whether there is a secondary site VOL for which failover processing has not been performed, for example, the recovery status 3345 in the remote copy pair management table 3340 may be referred to. If it is determined that there is no secondary site VOL for which failover processing has not been performed (step S4302: NO), the processing of the failover processing program 3240 ends. If it is determined that there is a secondary site VOL for which failover processing has not been performed (step S4302: YES), the process proceeds to step S4303.

[0068] Next, the failover processing program 3240 selects the secondary site VOL that has not had failover processing performed and has the most stringent recovery requirements in the recovery requirement management table (step S4303). As the secondary site VOL with the most stringent recovery requirements, for example, the secondary site VOL with the smallest RTO value for the recovery requirements may be selected. The secondary site VOL selected in step S4303 shall correspond to the "selected secondary site VOL" in steps S4304 to S4308.

[0069] Next, the failover processing program 3240 determines whether there is surplus performance in the destination storage node of the selected secondary site VOL (step S4304). If it is determined that there is surplus performance in the destination storage node of the selected secondary site VOL (step S4304: YES), the process proceeds to step S4308. If it is determined that there is no surplus (step S4304: NO), the process proceeds to step S4305. To determine whether there is surplus performance in the destination storage node of the selected secondary site VOL, for example, it may be determined based on the surplus performance of the storage node 200 and the I / O performance information required for the selected secondary site VOL. If there is surplus performance in the storage node 200 required to realize the I / O performance required for the selected secondary site VOL, it may be determined that there is surplus. For the surplus performance of the storage node 200, the information in the storage node operation information management table 3330 may be used. As for the I / O performance information required for the secondary site VOL, the I / O performance information of the primary site VOL that forms a remote copy pair with the secondary site VOL in the steady state may be used, or the QoS value set for the secondary site VOL may be used.

[0070] Next, the failover processing program 3240 determines whether there is a storage node 200 that has surplus performance required for the operation of the selected secondary site VOL (step S4305). If it is determined that there is a storage node 200 that has surplus performance required for the operation of the selected secondary site VOL (step S4305: YES), the process proceeds to step S4307. If it is determined that there is no such storage node (step S4305: NO), the process proceeds to step S4306. To determine whether there is a storage node 200 that has surplus performance required for the operation of the selected secondary site VOL, for example, the determination may be made based on the surplus performance of the storage node 200 and the I / O performance information required for the selected secondary site VOL. If there is a storage node 200 that has surplus performance necessary to realize the I / O performance required for the selected secondary site VOL, it may be determined that there is such a storage node. For the surplus performance of the storage node 200, the information in the storage node operation information management table 3330 may be used. As for the I / O performance information required for the secondary site VOL, the I / O performance information of the primary site VOL that forms a remote copy pair with the secondary site VOL in the steady state may be used, or the QoS value set for the secondary site VOL may be used.

[0071] Next, the failover processing program 3240 issues an instruction to enhance resources to the storage cluster 2000 of the secondary site 20 (step S4306). The method of enhancing resources may be realized by scaling up to change the instance specifications of the storage nodes 200 to higher specifications, or may be realized by scaling out to increase the storage nodes 200 of the storage cluster 2000. In the case of the storage cluster 2000 that enables both scaling out and scaling up, for the determination of which resource enhancement method to select, the information on the required time for each of scaling out and scaling up may be used, or the information on the performance impact on the volume of the secondary site VOL during the operation after failover due to each of the processes of scaling out and scaling up may be used, or these pieces of information may be combined and comprehensively determined. For the selection of the instance specifications of the storage nodes 200 during resource enhancement, the information in the instance specification management table 3320 may be used. When there are multiple selectable instance specifications, for the determination of which instance specification to select, the information on the I / O performance required for the volume of the secondary site VOL where failover is not completed and the I / O performance that can be provided by the corresponding storage nodes 200 for each instance specification may be prepared and used for the determination.

[0072] When scaling out, initial settings of nodes and movement of volumes are required. Therefore, the required time for scaling out generally becomes longer than the required time for scaling up. On the other hand, scaling up cannot be executed on the nodes including the volume during the failover process. Therefore, if the nodes including the volume to be failed over after resource enhancement do not include the volume for immediate failover, resource enhancement is performed by scaling up, and if the nodes including the volume to be failed over after resource enhancement include other volumes for immediate failover, resource enhancement is performed by scaling out, so that the time required until all failovers are completed can be shortened.

[0073] After issuing an instruction to enhance resources for the storage cluster 2000 of the secondary site 20, it is possible to wait for the completion of the resource enhancement process or not. If not waiting for the completion of the resource enhancement process, for example, when determining the secondary site VOL in steps S4302 and S4303, the secondary site VOL with the recovery status of "resource enhancement in progress" can be added as an option. As the priority, after the secondary site VOL for which failover processing has not been performed, the secondary site VOL with the status of "resource enhancement in progress" can be selected.

[0074] Next, the failover processing program 3240 issues an instruction to move the secondary site VOL to the storage node 200 that has the surplus performance required for the operation of the selected secondary site VOL (step S4307). The movement of the secondary site VOL may wait for the completion of the process or not. If not waiting for the completion of the movement process of the secondary site VOL, for example, when determining the secondary site VOL in steps S4302 and S4303, the secondary site VOL with the recovery status of "moving to another storage node" can be added as an option. As the priority, after the secondary site VOL for which failover processing has not been performed, the secondary site VOL with the status of "moving to another storage node" can be selected.

[0075] Next, the failover processing program 3240 issues an instruction to the storage cluster 2000 of the secondary site 20 to perform a failover on the selected secondary site VOL (step S4308). After issuing the failover instruction, it is possible to wait for the completion of the failover process or not. If not waiting for the completion of the failover process after the failover instruction, for example, the failover processing status can be monitored and the process of updating the recovery status 3345 in the remote copy pair management table 3340 can be incorporated. Finally, when there is no secondary site VOL for which failover processing has not been performed in step S4302, the processing of the failover processing program 3240 is terminated.

[0076] In this embodiment, when the primary site 10 recovers from a failure and fails back from the secondary site 20, any method may be adopted. For example, during the failure of the primary site 10, the differential information between the PVOL immediately after the occurrence of the failure due to data update by I / O from the host computer 201 to the SVOL of the storage node 200 of the secondary site 20 that has failed over is managed, and the differential information at the time of failback to the storage system 100 is transmitted to the storage system 100, and the data contents of the PVOL and the SVOL are synchronized, and a failback may be performed.

[0077] As described above, the disclosed computer system includes a primary site storage system (100) that constructs a primary site that provides a plurality of primary volumes to a host (host computer 101), a secondary site storage system (2000) that is connected to the primary site storage system via a network 400 and constructs a secondary site that provides a plurality of secondary volumes with remote copy set for the plurality of primary volumes, and a management device (disaster recovery management device 300) that manages the primary site storage system, and is a computer system (storage management system 1), wherein the management device can change the resources of the secondary site storage system, and when a failure occurs in the primary site, the management device performs a failover in which the operations of the primary volumes are taken over by the corresponding secondary volumes, and controls to enhance the resources of the secondary site storage system, and controls the failover so that there are a secondary volume that fails over and starts operating before the enhancement of the resources and a secondary volume that fails over and starts operating after the enhancement of the resources. With this configuration and operation, compared with the case where excessive resources are prepared for the hardware of the secondary site during normal operation, it is possible to prevent a violation of the recovery requirements at the time of failover to the secondary site while suppressing costs and power consumption.

[0078] In addition, the management device holds recovery requirement information indicating recovery requirements for the plurality of secondary volumes, and based on the surplus resources of the secondary site storage system and the recovery requirement information, determines for each of the secondary volumes whether to perform a failover and start operation before or after the enhancement of the resources. Therefore, it is possible to directly specify the primary volume that always prioritizes failover over resource enhancement.

[0079] In addition, the recovery requirement information specifies a target recovery time for the secondary volume, and the management device determines whether to perform a failover before or after the enhancement of the resources based on the surplus resources, the target recovery time, and the time required for the enhancement of the resources. Therefore, it is possible to promptly execute resource enhancement while reliably avoiding RTO violations.

[0080] In addition, the secondary site storage system has a plurality of nodes, the enhancement of the resources includes scaling up to improve the performance of the nodes and scaling out to increase the number of the nodes, and the management device determines whether to perform scaling up or scaling out during the enhancement of the resources. Therefore, it is possible to enhance resources in a manner according to the situation.

[0081] In addition, when one node of the computer system has a secondary volume that performs a failover and starts operation before the enhancement of the resources and a secondary volume that performs a failover and starts operation after the enhancement of the resources, the computer system performs scaling out and moves the secondary volume that performs a failover and starts operation after the enhancement of the resources to the node added by the scaling out and operates it. Therefore, even when there is a mixture of a volume that fails over before resource enhancement and a volume that fails over after resource enhancement, it is possible to shorten the time required until all failovers are completed.

[0082] In addition, when one node of the computer system does not have a secondary volume that fails over and starts operating before the enhancement of the resource, but has a secondary volume that fails over and starts operating after the enhancement of the resource, the computer system operates the secondary volume that fails over and starts operating after the enhancement of the resource after performing the scale-up. Therefore, when scale-up, which takes less time than scale-out, is possible, scale-up is prioritized, and the time required until all failovers are completed can be shortened.

[0083] Also, before the occurrence of the failure, the management device controls the node that arranges the secondary volume that fails over and starts operating before the enhancement of the resource so that the node has sufficient surplus resources for the secondary volume to fail over and operate. Therefore, it is possible to reduce costs and power consumption while avoiding a situation where the input / output performance deteriorates when a failover occurs.

[0084] Note that the present invention is not limited to the above-described embodiments, and includes various modifications. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, not only deletion of such a configuration is possible, but replacement and addition of configurations are also possible.

Explanation of Signs

[0085] 1: Storage management system, 10: Primary site, 20: Secondary site, Storage systems 100, 200: Storage nodes, 300: Disaster recovery management device, 400: Network, 2000: Storage cluster, 3210: Recovery requirement management program, 3240: Failover processing program, 3310: Recovery requirement management table

Claims

1. A positive-site storage system that constructs a positive site that provides a plurality of positive volumes to a host, A secondary-site storage system that is connected to the positive-site storage system via a network and constructs a secondary site that provides a plurality of secondary volumes with remote copies set for the plurality of positive volumes, A management device that manages the positive-site storage system and the secondary-site storage system, A computer system having: The management device can change the resources of the secondary-site storage system, When a failure occurs in the positive site, the management device Performs a failover in which the operations of the positive volumes are taken over by the corresponding secondary volumes, and controls to enhance the resources of the secondary-site storage system, A computer system characterized in that the failover is controlled such that there are a secondary volume that fails over and starts operating before the enhancement of the resources and a secondary volume that fails over and starts operating after the enhancement of the resources.

2. The computer system according to claim 1, The management device Holds recovery requirement information indicating recovery requirements for the plurality of secondary volumes, Based on the surplus resources of the secondary-site storage system and the recovery requirement information, determines for each of the secondary volumes whether to fail over and start operating before or after the enhancement of the resources A computer system characterized by this.

3. The computer system according to claim 2, The recovery requirement information specifies a target recovery time for the secondary volume, The management device determines whether to cause a failover before or after the enhancement of the resources based on the surplus resources, the target recovery time, and the time required for the enhancement of the resources. A computer system characterized by this.

4. The computer system according to claim 1, The secondary-site storage system has a plurality of nodes, The enhancement of the resources includes a scale-up that improves the performance of the nodes and a scale-out that increases the number of the nodes, The management device determines whether to perform the scale-up or the scale-out during the resource enhancement A computer system characterized by this.

5. The computer system according to claim 4, When one node has a secondary volume that fails over and starts operating before the enhancement of the resource and a secondary volume that fails over and starts operating after the enhancement of the resource, perform the scale-out, move the secondary volume that fails over and starts operating after the enhancement of the resource to the node added by the scale-out, and operate it. A computer system characterized by the above.

6. The computer system according to claim 5, When one node does not have a secondary volume that fails over and starts operating before the enhancement of the resource and has a secondary volume that fails over and starts operating after the enhancement of the resource, perform the scale-up and then operate the secondary volume that fails over and starts operating after the enhancement of the resource. A computer system characterized by the above.

7. The computer system according to claim 1, Before the occurrence of the failure, the management device controls the node that arranges the secondary volume that fails over and starts operating before the enhancement of the resource so that the node has sufficient surplus resources for the secondary volume to fail over and operate. A computer system characterized by the above.

8. A storage management method by a computer system including a primary site storage system that constructs a primary site that provides a plurality of primary volumes to a host, a secondary site storage system that is connected to the primary site storage system via a network and constructs a secondary site that provides a plurality of secondary volumes with remote copy set for the plurality of primary volumes, and a management device that manages the primary site storage system, The management device can change the resources of the secondary site storage system, When a failure occurs in the primary site, the management device Performs a failover in which the operation of the primary volume is taken over by the corresponding secondary volume, and controls to enhance the resources of the secondary site storage system, A storage management method characterized by controlling the failover so that there are a secondary volume that fails over and starts operating before the enhancement of the resource and a secondary volume that fails over and starts operating after the enhancement of the resource.

Citation Information

Patent Citations

  • Computer system and management method for computer system

    WO2016194096A1