Computer system and method for managing the same
The computer system management method addresses the challenge of optimizing storage node arrangements for remote copy migration to the cloud by using a management computer to group and allocate volumes across nodes, achieving a balanced configuration that meets performance and cost requirements.
Patent Information
- Application Number
- JP2023201014
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
AI Technical Summary
When migrating a remote copy configuration from an on-premises data center to a cloud-based SDS system, determining the optimal arrangement of storage nodes for secondary and journal volumes is challenging, requiring a balance of performance and cost while ensuring failover capabilities and specific remote copy constraints.
A computer system management method that uses a management computer to group multiple volumes, including secondary and journal volumes, and control their arrangement across multiple nodes in a storage cluster, optimizing node allocation based on performance, capacity, and cost considerations.
This approach allows for a low-cost, high-performance configuration of the secondary site on the cloud, ensuring that operation requirements are met and migration design complexities are minimized.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer system and a computer system management method.
Background Art
[0002] In an on-premises data center-to-data center or hybrid cloud operation mode combining an on-premises data center and the cloud, remote copy is used. As the remote copy, a storage system described in Patent Document 1 (Japanese Patent Application Laid-Open No. 2005-18506) is known. Patent Document 1 describes that "the first storage system stores information regarding updates to data stored in the first storage system as a journal. Specifically, the journal is composed of a copy of the data used for the update and update information such as a write command at the time of update. Further, the second storage system acquires the journal via a communication line between the first storage system and the second storage system. The second storage system holds a copy of the data held by the first storage system, and uses the journal to update the data corresponding to the data of the first storage system in the order of data update in the first storage system."
[0003] For the computer system for remote copy, the updated data written from the host server to the positive volume at the positive site (business site) is copied and stored in the journal volume (master journal volume). This data is copied to the journal volume (restore journal volume) at the secondary site (backup site) asynchronously with the I / O of the positive volume. In this way, the data written from the host to the positive volume at the positive site is transferred to the secondary site asynchronously with the write request. By writing the data from the restore journal volume to the secondary volume, the remote copy is completed.
[0004] In recent years, Software Defined Storage (SDS), which is constructed by implementing storage control software on general-purpose server devices (hereinafter referred to as storage nodes), has attracted attention. Since SDS does not require dedicated hardware and has high scalability, its demand is increasing. As an information processing system using SDS, there is known an information processing system in which a plurality of storage nodes each having one or more SDSs implemented therein are combined to form one cluster, and the cluster is provided as one storage device to a host device (hereinafter referred to as a host).
[0005] For example, Patent Document 2 (Japanese Patent Application Laid-Open No. 2019-185328) discloses an information processing system that provides a virtual logical volume (virtual volume) from a plurality of storage nodes on which SDS is implemented. Patent Document 2 states that "a control unit arranged in a storage node and set to an active mode for processing requests from a compute node, and a control unit arranged in another storage node and set to a passive mode for taking over processing when a failure occurs in the control unit or the like, inquire about the configuration of the redundancy group composed of the storage nodes, and based on the inquiry result, set a plurality of paths from the own compute node to the volume associated with the redundancy group. At this time, the priority of the path connected to the storage node where the active mode control unit is arranged is set to be the highest, and the priority of the path connected to the storage node where the passive mode control unit is arranged is set to be the next highest."
[0006] In the information processing system of Patent Document 2, the storage controllers (control software) are redundantly configured in an active-standby combination across multiple storage nodes, and the combination of these storage controllers is configured in a daisy chain between the storage nodes. Further, in the information processing system of Patent Document 2, it is possible to set a plurality of paths (multi-paths) from the upper device to the storage node, set a first-priority path for the storage node where the active control software exists, and set a second-priority path for the storage node where the standby control software exists. By setting such multi-paths, when a failure occurs in the active control software and the standby control software is promoted to active by the failover function, the information processing system of Patent Document 2 can continue the IO execution to the virtual volume by switching the path to be used according to the priority, and thus can realize a redundant configuration of the path.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0008] In an environment where a remote copy configuration is set up between storage devices in an on-premises data center, for the purpose of reducing costs at the secondary site, etc., the volume at the secondary site may be migrated to a storage device on the cloud. On-premises, the storage device at the secondary site is a single storage device. However, when migrating to the cloud, for example, migrating to the SDS described in Patent Document 2 on the cloud, when it is desired to continue the same operation as the storage device before migration, it is difficult to determine how many storage nodes with what specifications should be prepared and on which storage nodes the secondary volume and journal volume (restore volume) at the secondary site should be placed. Generally, it is conceivable to place each volume so that the free capacities are averaged as much as possible considering the free capacity of each storage node. Or, it is conceivable to place each volume so that the performance loads of each volume are balanced.
[0009] However, it is necessary to determine an arrangement that satisfies not only the performance requirements and capacity usage tendencies between on-premises data centers before migration, but also that there was a single storage before migration and multiple nodes after migration, and that satisfies the failover of the active standby control software and the constraints specific to the remote copy function. In addition, since cost is emphasized in the cloud, it is necessary to construct a configuration that is as inexpensive as possible.
[0010] Therefore, in the present invention, when constructing a secondary site for remote copy on the cloud, an object is to propose a low-cost configuration while satisfying the operation requirements before migration and the constraints specific to the remote copy function, and eliminating the need for migration design.
Means for Solving the Problems
[0011] To achieve the above object, one of the typical computer systems of the present invention includes a storage system that constructs a positive site providing one or more positive volumes for a host, a storage cluster connected to the storage system via a network and having a plurality of nodes, and a management computer. When constructing a secondary site having a secondary volume with remote copy set for the positive volume of the positive site in the storage cluster, the management computer manages a plurality of volumes including the secondary volume as a group, and based on the group, controls a plurality of volumes including the secondary volume to be arranged on the plurality of nodes of the storage cluster. Also, one of the typical computer system management methods of the present invention is a management method of a computer system including a storage system that constructs a positive site providing one or more positive volumes for a host, a storage cluster connected to the storage system via a network and having a plurality of nodes, and a management computer. When the management computer constructs a secondary site having a secondary volume with remote copy set for the positive volume of the positive site in the storage cluster, the method includes a step of managing a plurality of volumes including the secondary volume as a group, and a step of controlling the management computer to arrange a plurality of volumes having the secondary volume on the plurality of nodes of the storage cluster based on the group.
Advantages of the Invention
[0012] According to the present invention, when constructing a secondary site for remote copy on the cloud, a configuration that balances performance and cost can be proposed. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Mode for Carrying Out the Invention
[0014] In the following description, an "interface device" may be one or more interface devices. The one or more interface devices may be at least one of the following. · One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface device is an interface device for at least one of an I / O device and a remote display computer. The I / O interface device for the display computer may be a communication interface device. At least one I / O device may be either an input device such as a user interface device, for example, a keyboard and a pointing device, or an output device such as a display device. · One or more communication interface devices. The one or more communication interface devices may be one or more of the same type of communication interface devices (for example, one or more NICs (Network Interface Cards)) or two or more different types of communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).
[0015] Also, in the following description, "memory" is one or more memory devices which are an example of one or more storage devices, and may typically be a main memory device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.
[0016] Also, in the following description, a "persistent storage device" may be one or more persistent storage devices that are an example of one or more storage devices. A persistent storage device may typically be a non-volatile storage device (e.g., an auxiliary storage device), specifically, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), an NVME (Non-Volatile Memory Express) drive, or an SCM (Storage Class Memory).
[0017] Also, in the following description, a "storage device" may be at least a memory among a memory and a persistent storage device.
[0018] Also, in the following description, a "processor" may be one or more processor devices. At least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may also be a circuit that is an aggregate of gate arrays (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)) described by a hardware description language that performs part or all of the processing, which is a processor device in a broad sense.
[0019] In the following description, the information obtained as output for the input may be described using expressions such as "xxx table". However, such information may be data of any structure (for example, it may be structured data or unstructured data), or it may be a learning model typified by a neural network, genetic algorithm, or random forest that generates output for the input. Therefore, "xxx table" may be referred to as "xxx information". Also, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be combined into one table.
[0020] In the following description, the processing may be described with "program" as the subject. However, since the program is executed by a processor to perform the defined processing while appropriately using a storage device and / or an interface device, the subject of the processing may be the processor (or a device or system having the processor). The program may be installed from a program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable recording medium (for example, a non-transitory recording medium). Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0021] In the following description, ID is adopted as an example of the identification information of an element, but the identification information may be any information that can identify the element, such as a name.
[0022] In the following description, when describing elements of the same type without distinction, common reference signs among the reference signs are used, and when describing elements of the same type by distinction, reference signs may be used. For example, when not distinguishing multiple nodes, they are described as "node 135", and when distinguishing each node, they are described as "node 135-1", "node 135-2", and "node 135-3".
[0023] Hereinafter, some embodiments of the present invention will be described with reference to the drawings. In the following description, the logical volume is denoted as "Vol".
Example
[0024] FIG. 1 is a block diagram showing the configuration of a computer system 10 according to Embodiment 1 of the present invention. The computer system 10 includes an on-premises data center 100-1 at the primary site, an on-premises data center 100-2 at the secondary site, and a public cloud 150 at the secondary site. These are connected by a network 50. It also has a storage system 120-1 and a storage system 120-2. Since asynchronous remote copy is operated between these storage systems, the storage system 120-1 is called the primary storage system, and the storage system 120-2 is called the secondary storage system. Furthermore, it has a storage cluster 130 as a secondary storage cluster. The computer system 10 further includes a management computer 140, a host 110-1 (Host 1), a host 110-2 (Host 2), and a host 110-3 (Host 3). The host 110-1, the primary storage system 120-1, and the management computer 140 are located in the on-premises data center 100-1 at the primary site. The host 110-2 and the secondary storage system 120-2 are located in the on-premises data center 100-2 at the secondary site. The host 110-3 and the secondary storage cluster 130 are located in the public cloud 150 at the secondary site. Note that the management computer 140 is not limited to the on-premises data center 1 and may be located in either the on-premises data center 2 or the public cloud 150.
[0025] The network 50 may be redundant. The network may be Ethernet / FibreChannel / wireless. The configuration and connection relationship of the switch shown in the figure are only examples, and there may be an increase or decrease in the number of notations. It may also be divided by protocol. Internal networks 14-1 (14-2, 14-3) exist within on-premises data centers 100-1, 100-2 and public cloud 150. Although they are described together for convenience, they may be divided into a data input / output network and a management computer data network. The network protocol may be TCP / IP, FC, or other protocols.
[0026] Storage systems 120-1 (120-2) include storage devices 22-1 (22-2) which are data storage devices, data input / output ports 23-1 (23-2), management ports 24-1 (24-2), memories 21-1 (21-2), and processors 20-1 (20-2) connected to these elements. The storage device may be a physical HDD, SSD, etc. Ports 23-1 (23-2) perform interface processing for data input / output between storage systems 120-1 (120-2) and hosts 110-1 (110-2), and may be, for example, HBAs. In the figure, the data input / output ports between the host and the storage system and the data input / output ports between the primary and secondary storage systems are shown as the same port, but they may be separate ports. Ports 24-1 (24-2) perform interface processing for management data input / output between storage management computers 140, and may be, for example, NICs.
[0027] Storage cluster 130 is composed of a plurality of nodes, namely node 135-1 (node 1), node 135-2 (node 2), and node 135-3 (node 3). Each node includes a storage device 32-1 which is a data storage device, a data input / output port 33-1, a management port 34-1, a memory 31-1, and a processor 30-1 connected to these elements. Although the details of nodes 135-2 and 135-3 are omitted, they have the same configuration as node 135-1. The storage device may be a physical HDD, SSD, etc., or a virtual device. Port 33-1 performs interface processing of data input and output between the storage system 120-1 and the host 110-3, and may be, for example, an HBA. In the figure, the data input / output port between the host and the storage system and the data input / output port between the primary and secondary storage systems are shown as the same port, but they may be separate ports. Port 34-1 performs interface processing of management data input and output between the management computer 140, and may be, for example, a NIC. In the figure, three nodes are described, but any number of nodes may be used as long as there is one or more.
[0028] The storage system 120-1 (120-2) and the storage cluster 130 provide one or more logical volumes capable of reading and writing data to the host 110-1 (110-2, 110-3). The storage system 120-1 (120-2) and the storage cluster 130 receive an I / O command (e.g., a write command or a read command) specifying a logical volume from the host 110-1 (110-2, 110-3) and process the I / O command. The storage system 120-1 (120-2) and the storage cluster 130 read and write data at the address position specified in the write command or read command in the specified logical volume in response to the write command or read command input from the host 110-1 (110-2, 110-3).
[0029] The management computer 140 has a port 43-1 connected to the network 14-1, a storage device 42-1, a memory 41-1, and a processor 40-1 connected to these elements. The processor 40-1 controls the operation of the port 43-1 and performs overall control of the management computer 140 by executing predetermined processing using various programs and data stored in the memory 41-1 and the storage device 42-1. The management computer 140 manages the configuration information and operation information of the entire storage system 120-1 (120-2) and the storage cluster 130.
[0030] Host 110-1 (110-2, 110-3) has a port 13-1 (13-2, 13-3. Figures are omitted) connected to the internal network 14-1 (14-2, 14-3), a storage device 12-1 (12-2, 12-3. Figures are omitted), a memory 11-1 (11-2, 11-3. Figures are omitted), and a processor 10-1 (10-2, 10-3. Figures are omitted) connected to these elements. The port 13-1 (13-2, 13-3) performs interface processing of data input and output between the storage system 120-1 (120-2) and the storage cluster 130. The processor 10-1 (10-2, 10-3) controls the operation of the port 13-1 (13-2, 13-3) and performs overall control by executing predetermined processing using various programs and data stored in the memory 11-1 (11-2, 11-3) and the storage device 12-1 (12-2, 12-3).
[0031] The number of each storage system, storage cluster, and host is described as one for convenience, but there may be multiple units. The number of processors, memories, storage devices, and ports of each storage system, storage cluster, and management computer is described as one for convenience, but there may be multiple units. The management computer may be included in the storage system, host, or storage system.
[0032] FIG. 2 shows a functional block diagram of asynchronous remote copy performed between the primary storage system 120-1 at the primary site and the secondary storage system 120-2 at the secondary site. Asynchronous remote copy is a function of transferring data written from a host to the storage system at the business site (primary site) to the storage system at the backup site (secondary site) asynchronously with the write request.
[0033] The remote copy system consisting of the on-premises data center 100-1 at the primary site and the on-premises data center 100-2 at the secondary site has remote copy pairs consisting of groups (also called consistency groups, CTGs) 180-1, 180-2, and 180-3.
[0034] In group 180-1, a primary volume PVOL1 (160-1) and a (master) journal PJNL1 (190-1) are set. A journal is also referred to as a journal group. The journal PJNL1 (190-1) includes a (master) journal volume (JVOL1) (170-1).
[0035] And in the on-premises data center 100-2 at the secondary site of group 180-1, a secondary volume SVOL1 (160-2) and a (restore) journal SJNL1 (190-2) are set. The journal SJNL1 (190-2) includes a (restore) journal volume (JVOL4) (170-2). The updated data written from host 110-1 to the primary volume (PVOL1) is copied and stored in the primary journal volume (JVOL1). The updated data (journal data) of the journal volume (JVOL1), which is asynchronous to the I / O to the primary volume (PVOL1), is copied to the secondary journal volume (JVOL4) of the on-premises data center 100-2 at the secondary site.
[0036] The updated data of the secondary journal volume (JVOL4) is written to the secondary volume (SVOL1) of the on-premises data center 100-2 at the secondary site. In group 180-1, a remote copy pair of PVOL1-JVOL1-JVOL4-SVOL1 is set.
[0037] The number of journal volumes included in one journal (journal group) 190-1 (190-2) is described as one for convenience, but multiple may be included. The number of PVOL and SVOL included in one group 180-1 is described as one for convenience, but multiple may be included. The number of groups 180 is described as three for convenience between the main storage system and the sub-storage system, or between the main storage system and the sub-storage cluster, but any number of one or more may be used. Although the illustration of group 180-X and journal (journal group) (190-X) in the storage cluster is omitted, it is similarly defined in the storage cluster.
[0038] Figure 3 is a functional block diagram of the storage structure of the memory 21-1 (21-2) of the storage system 120-1 (120-2). The storage control program 201 is software that receives requests from the host, performs I / O processing, and stores data in the drive. It is also a program that implements overall storage control such as snapshots, compression, deduplication, and remote copy. Detailed description is omitted in this patent. The storage configuration management program 202 is a program for managing its configuration such as creating and deleting volumes and journals. The operation information management program 203 is a program for measuring, storing, and managing the operation information of the CPU, volume, etc. The device-side configuration management table 204 is a table for managing the configuration information of the volumes, journals, etc. of the storage system. It is updated by the storage configuration management program 202. The device-side operation information management table 205 is a table of the operation information of the CPU and memory of the storage system, and is updated by the operation information management program 203.
[0039] Figure 4 is a diagram showing an example of the layout configuration of the storage control software in the storage cluster 130. The storage control program 301-1 is configured to be redundant in an active-standby combination across multiple storage nodes 135-X. The active storage control program is set to the active state (active mode) that can receive IO requests from the host, and the standby storage control program is set to the standby state (standby mode) that does not receive IO requests from the host.
[0040] Also, in the storage cluster 130 in the computer system 10 according to the present embodiment, a redundancy group (storage controller group) 136 combining the above active-standby storage control programs 301-1 and 301-2 is configured in a daisy chain at each node. In FIG. 4, the active storage control program 301-1 and the standby storage control program 301-2 included in one redundancy group 136 are each one, but in the present embodiment, one redundancy group 136 may be composed of three or more storage control programs 301-1 and 301-2 (more specifically, one active storage control program 301-1 and two or more standby storage control programs 301-2). And although detailed description is omitted, when a failure occurs in the active storage control program 301-1 and the standby storage control program 301-2 is promoted to active by failover control, the volume operating in the storage control program 301-1 of node 1 is configured to appear to operate in the promoted storage control program 301-2. Thereby, normally each storage system operates at each node, and when a node fails, multiple storage systems operate at a specific node.
[0041] FIG. 5 shows an example of programs and data stored in the memory 31-1 of the storage cluster 130. As described with reference to FIG. 4, each storage system operates at each node. On the other hand, since it has a daisy-chain configuration, active and standby storage control programs of different redundancy groups (storage controller groups) operate at each node. Therefore, each node has Active management information and Standby management information for the storage control program that mainly operates at each node. The content of each management information is the same as each program and management table included in the memory 11-1 of the storage system 120-1 described with reference to FIG. 3. In addition, it has a cluster control program 307 for controlling so that each storage system operates in cooperation, and a control program management table 308 showing the relationship between the storage controller (storage control program) in each storage system and each node. The cluster control program 307 and the control program management table 308 are information common to all nodes, and may be held in one representative node, or may be redundantly held in synchronization in a plurality of nodes. The figure shows a configuration example of storing in synchronization at each node.
[0042] FIG. 6 shows an example of programs and data stored in the memory 41-1 in the management computer 140. The configuration management program 401 is a program for managing the configuration of the storage to be managed. The information collection / updating program 402 is a program for collecting / updating configuration information and operation information from the storage to be managed. The placement calculation / proposal program 403 is a program for searching for and proposing the number of nodes, node specifications, and resource placement configurations that result in the lowest cost while satisfying the constraints of the UR specific information.
[0043] The management computer side configuration management table 405 is a table showing the configuration of the storage to be managed and is updated by the configuration management program 401. The content holds the same information as the device side configuration management table 204 stored in the memory of each storage system and the device side configuration management table 304 stored in the memory of the storage cluster, for the number of each storage system and storage cluster. The management computer-side operation information management table 406 is a table of the operation information of the CPU and memory of the storage system to be managed, and is updated by the information collection / updating program 402. The content holds the same information as the device-side operation information management table 205 stored in the memory of each storage system and the device-side operation information management table 305 stored in the memory of the storage cluster, for the number of each storage system and storage cluster.
[0044] The cost table 407 shows the server and drive specifications and their costs. It is held by collecting user input or information provided by the cloud provider. The group performance / capacity / node correspondence table 408 is information indicating how much performance and capacity are required for each group of storage resources such as a consistency group, and which node should be placed on. The required node specification table 409 is information indicating the specifications of the nodes required to allocate the resources of the group described in the group performance / capacity / node correspondence table 408.
[0045] Figure 7 shows a configuration example of the device-side configuration management table 204. The device-side configuration management table 204 is composed of a volume management table 204-1, a journal management table 204-2, a pair management table 204-3, a remote path management table 204-4, a device physical configuration management table 204-5, and device upper limit specification information 204-6. The volume management table 204-1 has a volume resource ID, capacity, its attributes, and a node. The attributes include those of a normal volume accessed from a host and those used as a journal volume. The node indicates information on which node (in the case of a storage cluster, the node where the active storage controller exists) the volume exists. The journal management table 204-2 holds the ID of the journal resource, the ID (VOLID) of the journal volume included in the journal, the capacity of the journal volume (JVOL capacity), its status (Normal, Error), the usage rate of the journal volume, and information about the node. The node indicates information about which node the journal exists on (in the case of a storage cluster, the node where the active storage controller exists).
[0046] The pair management table 204-3 holds the ID of the remote copy pair (resource ID), the ID (PVOLID) of the primary volume included in the pair, the ID (PJNLID) of the primary journal, the ID (SVOLID) of the secondary volume, the ID (SJNLID) of the secondary journal, the group ID when managing multiple pairs in a group, and the status indicating the state of the pair.
[0047] The remote path management table 204-4 holds the ID of the remote path (resource ID), the ID of the storage device at the primary site where the remote path is set (Initiator device ID), the port ID of the storage device at the primary site (Initiator Port ID), the ID of the storage device at the secondary site (Target device ID), the port ID of the storage device at the primary site (Target Port ID), and information about the path group ID for grouping multiple remote paths. The device physical configuration management table 204-5 holds information about the device ID, node ID, and Port ID. The node ID holds information about one node in the case of a storage system or one or more nodes in the case of a storage cluster. The Port ID holds one or more connected PortIDs for each node. The device upper limit specification information 204-6 holds information about the item and the upper limit. The item describes the type of resource (e.g., journal, volume, remote path, pair) for providing the functions of the storage device. The upper limit is the upper limit specification information that the resource has. The example in Figure 7 shows that a maximum of 4 journals can be arranged for one node.
[0048] FIG. 8 shows a configuration example of the device - side operation information management table 205. The device - side operation information management table 205 is composed of device operation information 205 - 1, volume operation information 205 - 2, and inter - journal operation information 205 - 3. The device operation information 205 - 1 indicates the ID of the CPU and the memory respectively, the type of metric, the time, and the value of the metric at that time, as one of the resources of the storage system. The Write Pending Rate of the memory indicates the ratio of data that accumulates because there is a lot of Write data and cannot be written to the memory. This is one of the indicators showing that the load of the storage system is high. The metric is not limited to the values described here. For example, it may be the response time of the volume, the drive operation rate for each drive, the cache hit rate of the memory, the CPU, memory, and drive operation rates for each node in the storage cluster, etc. The device operation information 205 - 1 records values every minute, but the time interval is not limited to this.
[0049] The volume operation information 205 - 2 includes the ID of the volume, information indicating the type of metric (operation information), the time, and the value of the operation information at each time. The volume operation information 205 - 2 records values every minute, but the time interval is not limited to this. The inter - journal operation information 205 - 3 records the ID of the group (consistency group) that holds the primary and secondary journals, the ID of the journal (PJNLID) of the primary - side site (primary storage system 120 - 1), the ID of the journal (SJNLLID) of the secondary - side site (secondary storage system 120 - 2), information indicating the type of metric, and the value at each time. The inter - journal operation information 205 - 3 records values every minute, but the time interval is not limited to this.
[0050] FIG. 9 shows a configuration example of the control program management table 308. The control program management table 308 stores information on the storage controller ID, status, SCG ID, node ID, managed capacity, and free capacity. The storage controller ID indicates the identifier assigned to each storage control program 301 (in other words, the storage controller). The status indicates the status of the storage control program 301 to be managed. For example, not only the active and standby statuses, but also a dead status indicating a failure occurrence may be prepared. In addition to standby indicating a standby state, a passive status indicating a waiting state may be prepared, or other values may be used.
[0051] The SCG ID indicates the storage controller group ID to which the storage control program 301 to be managed belongs. The storage controller group ID is an identifier assigned to each redundancy group (storage controller group) 136 formed by including a combination of active-standby storage control programs 301. The node ID indicates the identifier of the storage node 135-X on which the storage control program 301 to be managed is arranged. The managed capacity indicates the total value of the capacity (pool) managed in units of the redundancy group (storage controller group) 136 to which the storage control program 301 to be managed belongs. The pool is used by allocating a storage area from the drive. The free capacity indicates the capacity that can be created by the storage controller group to which the storage control program 301 to be managed belongs.
[0052] Figure 10 shows a configuration example of the cost table 407. The cost table consists of server cost information 407-1 and drive cost information 407-2. The server cost information 407-1 has information on Type ID, number of CPUs, CPU type, memory, storage, network, and cost. The Type ID indicates the server type ID. The number of CPUs indicates the number of CPUs in one server. The CPU type is information indicating the clock speed and specifications of the CPU. Memory indicates the memory capacity. Storage has information meaning the capacity when included in the server, or external when not included. Other information may be noted. Network records the maximum available network bandwidth. Cost records the fee when using this server. In the figure, it records the fee per month, but it is not limited to this, and the cost for other periods or a fixed fee instead of a per-volume charge may be recorded. The information held in the server cost information 407-1 is not limited to these, and it may have other information indicating the server cost.
[0053] The drive cost information 407-2 has information on Type ID, type, maximum IOPS, throughput, and cost. The Type ID indicates the drive type ID. The type records the drive type. The maximum IOPS indicates the maximum IOPS per drive. Throughput indicates the maximum throughput per drive. Cost indicates the fee when using this drive. In the figure, it records the fee per 1 GiB per month, but it is not limited to this, and the cost for other periods or a fixed fee instead of a per-volume charge may be recorded. The information held in the drive cost information 407-2 is not limited to these, and it may have other information indicating the drive cost. The information in the cost table 407 may be set by the user in advance or obtained and set from the information provided by the public cloud vendor.
[0054] Figure 11 shows a configuration example of the group performance·capacity·node correspondence table 408. The group performance, capacity, and node correspondence table 408 holds information on GroupID, JNLID, the temporary ID of the node after migration, the required maximum capacity, and the maximum total value of the volume performance of the group at each time. The GroupID is an identifier for a set that needs to be placed together when migrating to the storage cluster at the secondary site. In Embodiment 1, the journal group corresponds to the consistency group (CTG) itself. What actually migrates to the storage cluster at the secondary site are the resources on the secondary site side within the CTG. The JNLID is the journal ID on the secondary site side included in that set. The temporary ID of the node after migration is an ID indicating to which node this group will be stored after migration. It is updated as needed during the placement calculation. The required maximum capacity indicates the capacity required for all resources (normal volume, journal volume) within that group. The maximum total value of the volume performance of the group at each time is the maximum value of the performance required for all volumes at a certain time. Although only IOPS and throughput are described in the figure, other information may be described. Also, not only the maximum value but also the minimum latency to be satisfied may be described. In the figure, values are recorded every minute, but the time interval is not limited to this.
[0055] Figure 12 shows a configuration example of the required node specification table 409. The required node specification table 409 holds information on the temporary ID of the node after migration, server specifications, drive specifications, and drive capacity. Information on the required server and drive specifications for each temporary ID of the node after migration is described using the TypeID information described in the cost table.
[0056] Hereinafter, an example of the processing performed in this embodiment will be described. The outline of the processing is shown in FIG. 13, and the detailed flow is shown in FIGS. 14 to 16. First, FIG. 13 shows an example of the overall flow in Embodiment 1. Before the process of FIG. 13 starts, it is assumed that the information collection and update program 402 collects the information of the device - side configuration management table 204 from the storage systems 120 - 1 and 120 - 2 to be managed and updates the management computer - side configuration management table 405. Also, it is assumed that the information collection and update program 402 periodically collects the information of the device - side operation information management table 305 from the storage systems 120 - 1 and 120 - 2 to be managed and updates the management computer - side operation information management table 406. Also, it is assumed that the information in the cost table 407 is set by the storage administrator. Assuming that those processes are completed, first, the flowchart of FIG. 13 starts by the user's access to the management computer 140.
[0057] When the user designates the positive - side storage system and the negative - side storage system to be migrated, the placement calculation and proposal program 403 calculates the total performance value and the required capacity for each volume in the CTG related to the journal for each journal of the negative - side storage system at each time (S1100). Details are shown in FIG. 14. Next, the placement calculation and proposal program 403 sizes the hardware that can satisfy the original requirements (the originally operating performance) with one node even after migrating the maximum total performance load and the used capacity to the negative - side storage cluster, and calculates the required cost (S1200). Details are shown in FIG. 15.
[0058] Finally, the placement calculation and proposal program 403 calculates for each group whether aggregation is possible by considering the limitations specific to the copy function such as performance upper limit, capacity upper limit, and the number of journal groups on another node, and then explores and proposes the specifications of the server and drive, and the number of nodes that result in the lowest combined cost (S1300). Details are shown in FIG. 16. Through the overall flow above, the specifications of the nodes, the number of nodes, and the placement of journals and volumes required in the negative - side storage cluster after migration are determined.
[0059] FIG. 14 shows an example of a flow for calculating the maximum total performance value and required capacity of all groups in Embodiment 1. First, the placement calculation and proposal program 403 executes the processes described in S1102 to S1104 for all journals in the loop process of S1101. The placement calculation and proposal program 403 creates a row in the group performance, capacity, node correspondence table 408 corresponding to each journal. The ID of the group (CTG ID) to which the journal belongs and the ID of the journal group are set in the created row. Also, the transition node temporary ID is assigned and set with an unused value (S1102). Next, the placement calculation and proposal program 403 calculates the total value of the operation information of all volumes (normal volume, journal volume) belonging to one group for each time. The calculated result is set as the maximum total volume performance value of the group for each time in the group performance, capacity, node correspondence table 408. At this time, the period (cycle) and time interval for calculating the operation information can be set by the user in advance, the same as the time interval in the management computer side operation information management table 406, or a fixed value determined by the system (e.g., period = 1 day, time interval = 1 minute). Any period (cycle) and time interval may be used (S1103). Next, the placement calculation and proposal program 403 calculates the total value of the used capacity of all volumes (normal volume, journal volume) belonging to one group and updates the value of the required maximum capacity in the group performance, capacity, node correspondence table 408. Note that the used capacity of each volume calculated at this time may be not only the current value but also the capacity usage history stored as the operation information of the volume, and a predicted capacity may be entered as the used capacity in the future (e.g., one year later) based on the capacity usage history trend. The future period at this time may be set by the user in advance or any fixed value determined by the system. The above processes are executed for all journal groups, and the flowchart of FIG. 14 ends.
[0060] FIG. 15 shows an example of a flow for sizing and cost calculation of hardware that satisfies the conditions of the maximum total performance load and capacity in Embodiment 1. First, the placement calculation and proposal program 403 searches for the maximum total performance load and capacity across all groups using the information in the group performance, capacity, and node correspondence table 408 (S1201).
[0061] Next, the placement calculation and proposal program 403 estimates (sizes) the server specifications and drive specifications that satisfy the maximum total performance load retrieved in S1201, and describes the required server specifications and drive specifications in the first row of the required node specification table 409 (S1202). There are several implementation methods for the sizing method implemented at this time. As an example, for throughput, for each hardware part (memory, processor, drive, etc.), model formulas for throughput and IOPS are created in advance. The model formula for each part determines the model in view of the internal I / O processing content (one I / O command processing, execution processing of storage functions such as copy, user data transfer). Then, the minimum value of the throughput of each part is calculated as the throughput of one node as a model. In this example, the model formula is created from the result of cumulative calculation from internal processing, but actually, several patterns of I / O may be issued with the combined specifications of several parts and the model may be determined from the experimental results. Using such a model, the hardware specifications that satisfy the values of the required requirements (IOPS, throughput, etc.) are estimated (sized). In the case of a public cloud, the network bandwidth between the primary site and the secondary site may be predetermined, in which case the network bandwidth may be calculated as fixed. The above sizing method is an example, and other sizing methods may be used. Note that in the storage cluster of this embodiment, the active and standby storage control programs are in a consecutive configuration, and when a node failure occurs, the standby storage control program of another node corresponding to the active storage control program of that node is activated and elevated, and two active storage control programs operate on one node at the same time. Therefore, it is considered that the performance required for one node is the total performance considering the case where the active storage control program of another node migrates by failover. That is, it can be said that the above sizing method is the sizing required for each active storage control program when the active storage control program operates with no node failures occurring. Regarding the description of the necessary node specifications in S1202 in Table 409, using the server cost information 407-1 and drive cost information 407-2 in the cost table 407, identify the Type of the server and the Type of the drive necessary to satisfy the sizing result, and record the resulting TypeID in the necessary node specification table 409.
[0062] Next, the placement calculation and proposal program 403 records the value of the drive with the maximum capacity retrieved in S1201 in the drive capacity of the first row of the necessary node specification table 409 (S1203). Next, the placement calculation and proposal program 403 records the information of the post-transition node temporary ID described in the group performance, capacity, and node correspondence table 408 in the necessary node specification table 409, and sets all nodes to the same server specifications, drive specifications, and drive capacity as the first row (S1204). Next, the placement calculation and proposal program 403 calculates the cost required for all nodes using the information in the cost table 407 and the necessary node specification table 409 (S1205). Thus, the flowchart in FIG. 15 ends.
[0063] FIG. 16 shows an example of a flow for searching and proposing the server-drive specifications and the number of nodes with the lowest cost in Embodiment 1. The flow is different depending on whether optimization is considered based on performance criteria or capacity criteria. In FIG. 16, the placement is calculated so as to meet the performance upper limit determined by the flow up to FIG. 15, and in some cases, a method of placement calculation based on performance criteria that allows each node to exceed the node capacity determined by the flow up to FIG. 15 is used. Basically, reducing the number of nodes is often important in terms of cost, and when both the performance and capacity upper limits must be met, it is highly likely that the number of nodes cannot be reduced, so such a flow is considered. Also, since generally the cost increases as the performance of the server increases more than that of the drive, in this embodiment, optimization based on performance criteria is described as an example of the embodiment.
[0064] First, the placement calculation and proposal program 403 executes the processes from S1302 to S1308 for all groups (the resources (journals, volumes) of the sub-storage systems within the group). First, in S1302, the placement calculation and proposal program 403 selects groups in ascending order of the maximum total performance value in a certain period in the group performance, capacity, and node correspondence table 408 (S1302).
[0065] Next, the placement calculation and proposal program 403 checks whether there is a node for which the total value of the maximum performance at each time does not exceed the upper limit of the total performance maximum value of each node, assuming that the currently selected group is temporarily moved to another node (S1303). Specifically, the check is performed using the following example. For example, in the group performance, capacity, node correspondence table 408 of FIG. 11, for the sake of simplicity of explanation, only the example of IOPS is first described. The group with GroupID G3 has a maximum IOPS value of 100 during the period from 10:00:00 to 10:01:00, which is the smallest among all groups. This group G3 is currently assumed to be placed on node NN3. If this is moved to another node, for example, node NN2 (that is, co-resided with the group of G2), the total IOPS value at 10:00:00 will be 400. Also, the total IOPS value at 10:01:00 will be 200 + 100 = 300. Looking at the maximum value at each time of a single node, in the current FIG. 11, G1 scheduled to be placed on NN1 has the maximum value, and the total value of 300 at the time of the temporary move earlier does not exceed this 550. Therefore, it can be determined that the performance can be covered by the server spec S1 and the drive spec D1 of the node described in the required node spec table 409. On the other hand, if the group G3 is moved to NN1 where the group G1 is scheduled to be placed, the upper limit of the total performance maximum value will be exceeded. This means that the group G3 cannot be moved to G1. In this way, it is checked whether there is a node for which the total value of the maximum performance at each time does not exceed the upper limit of the total performance maximum value of each node. The above example is described only for IOPS. However, it is not limited to this, and the same is done for the throughput described in FIG. 11 to check whether there is an overall value that does not exceed the upper limit. Furthermore, not only the performance of the storage system but also whether the upper limit of the network bandwidth between the primary site and the secondary site is exceeded may be checked.
[0066] Next, in S1304, the placement calculation and proposal program 403 determines whether there is a node that does not exceed the upper limit of the maximum performance value through the check in S1303. If the result in S1304 is negative (=NO), it returns to S1301, then selects the group with the smaller total maximum performance value next, and repeats the process from S1302. If the result in S1304 is positive (=YES), it proceeds to S1305.
[0067] In S1305, the placement calculation and proposal program 403 checks whether, assuming the group is temporarily moved to another node, it exceeds the upper limit of the device upper limit spec information. This means checking whether it exceeds the functional upper limit of the storage system or storage cluster when placed on one node, as described in the device upper limit spec information (the same as the device upper limit spec information 204-6 in the device-side configuration management table) in the management computer-side configuration management table 405. More specifically, when considering moving G3 to NN2 in S1303, there will be two journals in one node (more precisely, within the management range of one storage control program). Since this is 4 or less in the device upper limit spec information (the same as the device upper limit spec information 204-6 in the device-side configuration management table), in this case, S1305 results in a negative result (=NO). If S1305 is positive (=YES), it returns to S1301. If S1305 is negative (=NO), it proceeds to S1306.
[0068] Next, in S1306, the placement calculation and proposal program 403 checks whether, assuming the group is temporarily moved to another node, the required capacity exceeds the upper limit. As a more specific example, when moving the G3 group in Fig. 11 to NN2, the required capacity is 100 + 50 = 150, which does not exceed the upper limit of 200. Therefore, the result is negative (=NO). If S1306 is positive (=YES), it proceeds to S1307. If S1306 is negative (=NO), it proceeds to S1308.
[0069] Regarding S1307, in this case, up to S1306, although the performance upper limit is not exceeded, the pattern is such that the capacity upper limit is exceeded in terms of the arrangement. In this case, the arrangement calculation and proposal program 403 increases the drive capacity by the excess capacity for the drives that are hypothetically to be moved in the necessary node specification table 409. At S1308, the arrangement calculation and proposal program 403 deletes the row of the post-migration node temporary ID where the group that is to be migrated was originally hypothetically arranged in the necessary node specification table 409. Further, it updates the post-migration node temporary ID of the group that is to be migrated in the group performance, capacity, node correspondence table 408. As a more specific example, considering that the group G3 with the post-migration node temporary ID of NN3 is migrated to the node with the post-migration node temporary ID of NN2, it deletes the row in the necessary node specification table 409 where the post-migration node temporary ID is NN3, and updates the post-migration node temporary ID of G3 in the group performance, capacity, node correspondence table 408 to NN2. After completing the above processing, it returns to S1301 and executes the same processing for all groups.
[0070] When the processing is executed for all groups, the number of necessary nodes, node specifications (server specifications, drive specifications), and the arrangement of the groups are determined based on the performance criteria. At S1309, the arrangement calculation and proposal program 403 recalculates the cost from the information in the necessary node specification table 409, the group performance, capacity, node correspondence table 408, and the cost table, and presents the calculated cost, the number of nodes, and the group arrangement to the user.
[0071] As described above, based on the copy function requirements such as performance requirements, capacity requirements, and restrictions on the number of journals for each group, it is possible to search for and propose server / drive specifications and the number of nodes that result in lower costs. As described previously, this embodiment is a flow determined by performance criteria. As an alternative form, when considering capacity criteria, the first flow in FIG. 16 first checks the capacity, and within the range that satisfies the capacity upper limit, checks whether it falls within the performance upper limit. Note that it is also possible to implement both the performance criteria flow and the capacity criteria flow and search for the configuration that results in the fewest number of nodes. Also, the performance criteria in FIG. 16 and the example of the placement search based on the aforementioned capacity criteria show examples of algorithms based on simple dynamic programming, but more advanced combinatorial optimization calculations may be performed to more precisely calculate the optimal placement and combination of the number of nodes.
Embodiment
[0072] Embodiment 2 will be described. At that time, the differences from Embodiment 1 will be mainly described, and the description of the common points with Embodiment 1 will be omitted or simplified.
[0073] In Embodiment 1, the unit for calculating node placement is a consistency group (CTG) unit including a journal, and basically, the movement and placement of one journal group and the set of associated secondary volumes are calculated. However, in storage operation, not only the volumes of the consistency group including the journal of remote copy, but also volumes related to those volumes, or volumes that are not related within the storage device but should be operated together or placed on the same storage in view of the operation purpose of the upper application, may exist. With only Embodiment 1, there is a possibility that those volume groups are dispersed and placed on different nodes. Also, there may be volume groups that were originally stored in a single storage system but are physically dispersed internally, such as by dividing parity groups, and it is desired to physically disperse them and operate them even after migration.
[0074] Therefore, in Embodiment 2, not only journals and their related volumes, but also volumes related to sub-volumes and snapshots, or volumes that are not related directly within the storage device but are used by the same upper-level application and for which continuous operation within the same storage is desired, are considered in a pattern showing node placement in an extended group concept. Furthermore, it is considered that not only performance, capacity, and storage function limits during node placement, but also the condition that they should not be placed on the same node (the same storage control program) operationally, can be taken into account when calculating the group placement.
[0075] Since the configuration of the basic computer system is the same in Embodiment 1 and Embodiment 2, the description is omitted. As a difference in configuration, since the information managed by the management computer increases, the difference in that information and the flow using it will be described.
[0076] FIG. 17 shows an example of programs and data stored in the memory 41-1 in the management computer 140 in Embodiment 2. The difference from Embodiment 1 is that a group management table 410 is added.
[0077] The group management table 410 is a management table that stores group information, which is an extension of the journal group (CTG) in Embodiment 1. It is updated by the configuration management program 401. This information may be set with values by the user as a prior update instruction input. Also, the components of the group may be automatically set based on the relationship of copy pairs within the storage system and other related information of internal functions.
[0078] FIG. 18 shows a configuration example of the group management table 410 in Embodiment 2. The group management table 410 holds information on GroupID, group components, and group-specific requirements. The GroupID is an identifier for the group. The group components indicate the identifiers of the resources that are the components included in this group. As examples, journals and volumes (and snapshots as a type of volume) are described, but it is not limited to these and other information may be set. The group-specific requirements indicate the special requirements specific to that group. As examples, Affinity requirements and specific performance are described. The Affinity requirement is a requirement that indicates the information when there is a condition that a certain group may / may not coexist with another group. Specifically explained with the example in FIG. 18, it means that the groups GR1 and GR2 should not be placed on the same node even if they meet the requirements for performance and capacity when performing placement calculations. Regarding specific performance, in previous placement calculations, the criterion was to determine whether the total performance of all volumes for each group, which was required for the secondary-site resources between the original on-premises storage systems, was met. However, when migrating from on-premises to the cloud, the performance requirements do not necessarily match. Therefore, when changing the requirements, the specific performance is described with the performance requirements. These describe requirements such as IOPS and throughput. Although omitted in the figure, performance requirements for each time may also be described.
[0079] Next, the processing flow of Embodiment 2 will be described. Basically, it only changes from what was done in the group (equivalent to CTG) in Embodiment 1 to processing in the extended group unit of FIG. 18, and the major flow does not change. However, regarding the placement calculation flow, since there are changes, the changes will be shown.
[0080] FIG. 19 shows an example of a flow for searching and proposing the number of nodes for the server-drive specifications with the lowest cost in Embodiment 2. In FIG. 19, similar to FIG. 16, an example of the optimization flow based on performance criteria is described. First, the placement calculation and proposal program 403 executes the processes from S2302 to S2309 for all groups (not only journals and CTGs, but also the groups described in FIG. 18). First, in S2302, the placement calculation and proposal program 403 selects groups in ascending order of the maximum total performance value in a certain period in the group performance, capacity, and node correspondence table 408 (S2302).
[0081] Next, the placement calculation and proposal program 403 checks whether there is a node whose total maximum performance value at each time does not exceed the upper limit of the total performance maximum value of each node or the specific performance described in the group management table 410, assuming that the currently selected group is temporarily moved to another node (S2303). Specifically, only the differences in Embodiment 1 are shown. If no specific performance is described in the group management table 410, the same upper limit of the total performance maximum value as in the embodiment is checked. If specific performance is described in the group management table 410, the value of the specific performance is checked instead of the upper limit of the total performance maximum value. Otherwise, it is the same as in the embodiment.
[0082] Next, in S2304, the placement calculation and proposal program 403 determines whether there is a node that does not exceed the upper limit of the maximum performance value or the specific performance as a result of the check in S2303. If the result in S2304 is negative (=NO), it returns to S2301, selects the next group with a smaller total performance maximum value, and repeats the process from S2302. If the result in S2304 is positive (=YES), it proceeds to S2305.
[0083] In S2305, the placement calculation and proposal program 403 checks whether there are nodes that meet the group-specific requirements other than the intrinsic performance. Specifically, taking the example of FIG. 18, consider the Affinity requirement. When considering moving GR1 to another node, even if the total performance maximum value upper limit or the upper limit of the intrinsic performance is met when moving to the node where GR2 exists, since the Affinity requirement is not met, it is determined that there are no nodes that meet the group-specific requirements other than the intrinsic performance. If the result in S2305 is negative (=NO), the process returns to S2301. If the result in S2305 is positive (=YES), the process proceeds to S2306.
[0084] In S2306, the placement calculation and proposal program 403 checks whether the upper limit of the device upper limit spec information is exceeded when the group is temporarily moved to another node. This means checking whether the function upper limit of the storage system or the storage cluster is exceeded when placed on one node, as described in the device upper limit spec information (the same as the device upper limit spec information 204-6 in the device-side configuration management table) in the management computer-side configuration management table 405. If the result in S2306 is positive (=YES), the process returns to S2301. If the result in S2306 is negative (=NO), the process proceeds to S2307.
[0085] Next, in S2307, the placement calculation and proposal program 403 checks whether the required capacity exceeds the upper limit when the group is temporarily moved to another node. If the result in S2307 is positive (=YES), the process proceeds to S2308. If the result in S2307 is negative (=NO), the process proceeds to S2309.
[0086] Regarding S2308, in this case, up to S2307, it is a pattern where the performance upper limit is met but the capacity upper limit is exceeded. In this case, the placement calculation and proposal program 403 increases the drive capacity by the amount of the excess capacity for the drive capacity to which the group is to be moved in the temporary required node spec table 409.
[0087] In S2309, the placement calculation and proposal program 403 deletes the row of the post-migration node temporary ID where the group decided to migrate was originally temporarily placed in the required node specification table 409. Further, the post-migration node temporary ID of the group decided to migrate is updated in the group performance, capacity, and node correspondence table 408. When the above processing is completed, the process returns to S2301 and the same processing is executed for all groups.
[0088] When processing is executed for all groups, the required number of nodes, node specifications (server specifications, drive specifications), and group placement based on performance criteria are determined. In S2310, the placement calculation and proposal program 403 recalculates the cost from the information in the required node specification table 409, the group performance, capacity, and node correspondence table 408, and the cost table, and presents the calculated cost, number of nodes, and group placement to the user.
[0089] As described above, for each group of resources including CTG or more that includes a journal and its paired volume, it is possible to search for and propose server / drive specifications and the number of nodes that result in lower costs, taking into account copy function requirements such as performance requirements, capacity requirements, and limitations on the number of journals, as well as group-specific requirements.
[0090] Although several embodiments have been described above, these are examples for explaining the present invention and are not intended to limit the scope of the present invention only to these embodiments. The present invention can be implemented in various other forms. For example, in at least one of Embodiments 1 and 2, although the migration destination of the secondary site is described as a public cloud, it may also be a private cloud.
[0091] In addition, for example, in Embodiments 1 and 2, the relationship between the primary site and the secondary site between the on-premises data center 1 and the on-premises data center 2 is described as a 1:1 relationship. However, a configuration with a 1:N relationship in which there are multiple on-premises data centers 2 as secondary sites with respect to the on-premises data center 1 as the primary site may also be possible. Further, the destination public cloud (or private cloud) may also use multiple sites instead of one site, and may be handled by storage clusters of multiple sites. In addition, for example, in Embodiments 1 and 2, the storage system of the secondary site has a single-node configuration, but a configuration of a storage cluster (multiple nodes) may also be possible. In addition, in Embodiment 1, a process of reducing the number of nodes was shown on the premise that the number of nodes has a greater impact on cost than the specifications of the nodes. However, it is also possible to compare the cost of a configuration in which the specifications are reduced without changing the number of nodes with the cost of a configuration in which the number of nodes is reduced, and select the one with the lower cost.
[0092] As described above, the disclosed computer system 10 includes a storage system 120 that constructs a primary site that provides one or more primary volumes to the host 110, a storage cluster 130 that is connected to the storage system via a network and has a plurality of nodes, and a management computer 140. When constructing a secondary site having a secondary volume in which remote copy is set with the primary volume of the primary site in the storage cluster 130, the management computer 140 manages a plurality of volumes including the secondary volume as a group, and based on the group, controls a plurality of volumes including the secondary volume to be arranged on a plurality of nodes of the storage cluster. Specifically, it controls a plurality of volumes that are within the same group and should be arranged at the secondary site to be arranged on the same node within the storage cluster. According to this configuration and operation, when constructing a secondary site for remote copy on the cloud, a configuration that achieves both performance and cost can be proposed.
[0093] In addition, a main journal volume for copying and storing updated data of the main volume is configured at the main site, and at the secondary site, a secondary journal volume for storing updated data that is copied from the main journal volume via the network and is for writing to the secondary volume is configured. The remote copy is performed by copying the updated data of the main volume to the secondary volume via the main journal volume and the secondary journal volume. The management computer 140 manages a combination of the main volume, the main journal, the secondary volume, and the secondary journal volume related to the same remote copy as the same group. Therefore, it is possible to propose an inexpensive configuration while arranging the secondary volume and the secondary journal volume in the same node.
[0094] In addition, the management computer 140 calculates the load amounts of the secondary journal volume and the secondary volume in the same group based on the operation information of the main volume, and determines the node for arranging the secondary journal volume and the secondary volume based on the calculated load amounts of the secondary journal volume and the secondary volume and the information of the nodes of the storage cluster. The management computer 140 calculates the usable capacities of the secondary journal volume and the secondary volume in the same group based on the operation information of the main volume, and determines the node for arranging the secondary journal volume and the secondary volume based on the calculated load amounts and usable capacities of the secondary journal volume and the secondary volume and the information of the nodes of the storage cluster. A computer system characterized by this. Therefore, the performance actually required for the node can be surely ensured.
[0095] In addition, the information of the nodes of the storage cluster includes cost information, and the management computer 140 determines the nodes for arranging the secondary journal volume and the secondary volume based on the calculated load amounts of the secondary journal volume and the secondary volume and the information of the nodes of the storage cluster including the cost. Therefore, a configuration with low cost can be explored.
[0096] In addition, the management computer 140 determines the nodes for arranging the secondary journal volume and the secondary volume so that the number of the nodes for arranging the secondary journal volume and the secondary volume in a plurality of groups is reduced. Therefore, while ensuring performance, the number of nodes can be reduced, and the cost of the storage cluster can be lowered. In addition, an item upper limit, which is the upper limit of the items of resources, is defined for the nodes, and the items of resources include at least any one of the number of journal volumes, the number of secondary volumes, the number of remote copies, and the number of pairs. The management computer 140 determines the nodes for arranging the secondary journal volume and the secondary volume based on the calculated load amounts of the secondary journal volume and the secondary volume, the information of the nodes of the storage cluster including the cost, and the upper limit. Therefore, within the range where the resources meet the predetermined conditions, the cost of the storage cluster can be lowered.
[0097] In addition, in the disclosed system, the secondary volume and the secondary journal volume are each made redundant by an active storage control program and a standby storage control program arranged on different nodes. When a failure occurs in the active storage, the standby storage control programs on different nodes take over and operate the secondary volume and the secondary journal volume, and a plurality of active or standby storage control softwares can operate on the same node. The management computer 140 determines a node for arranging the secondary journal volume and the secondary volume based on the total load amount of the active storage control program and the storage control program that has changed from standby to active. With this configuration and operation, it is possible to construct a secondary site that can guarantee performance even when a failover occurs.
Explanation of Signs
[0098] 10: Computer system, 120-1: Primary storage system, 120-2: Secondary storage system, 130: Secondary storage cluster, 140: Management computer
Claims
1. A storage system that constructs a primary site that provides one or more primary volumes to a host, a storage cluster having a plurality of nodes connected to the storage system via a network, and an administrative computer, wherein the computer system has: when constructing a secondary site having a secondary volume with a remote copy set for the primary volume of the primary site in the storage cluster, the administrative computer manages a plurality of volumes including the secondary volume as a group, and controls, based on the group, a plurality of volumes including the secondary volume to be arranged on a plurality of nodes of the storage cluster. A computer system characterized by that.
2. The computer system according to claim 1, characterized in that it controls a plurality of volumes that are in the same group and should be arranged in the secondary site to be arranged on the same node in the storage cluster. A computer system characterized by that.
3. The computer system according to claim 1, wherein a primary journal volume for copying and storing update data of the primary volume is configured in the primary site, a secondary journal volume for storing update data that is copied from the primary journal volume via the network and is for writing to the secondary volume is configured in the secondary site, the remote copy is performed by copying the update data of the primary volume to the secondary volume via the primary journal volume and the secondary journal volume, and the administrative computer manages a combination of the primary volume, the primary journal, the secondary volume, and the secondary journal volume related to the same remote copy as the same group. A computer system characterized by that.
4. The computer system according to claim 3, wherein the administrative computer calculates the load amounts of the secondary journal volume and the secondary volume in the same group based on the operation information of the primary volume, and based on the calculated load amounts of the secondary journal volume and the secondary volume and the information of the nodes of the storage cluster, determines the nodes on which the secondary journal volume and the secondary volume are to be arranged. A computer system characterized by that.
5. The computer system according to claim 4, The management computer calculates the usable capacity that the secondary journal volume and the secondary volume in the same group can use based on the operation information of the positive volume, and determines the node for arranging the secondary journal volume and the secondary volume based on the calculated load and usable capacity of the secondary journal volume and the secondary volume, and the information of the nodes of the storage cluster. A computer system characterized by this.
6. The computer system according to claim 5, The information of the nodes of the storage cluster includes cost information, The management computer determines the node for arranging the secondary journal volume and the secondary volume based on the calculated load of the secondary journal volume and the secondary volume and the information of the nodes of the storage cluster including the cost. A computer system characterized by this.
7. The computer system according to claim 6, A computer system characterized by determining the node for arranging the secondary journal volume and the secondary volume so that the number of nodes for arranging the secondary journal volume and the secondary volume in a plurality of groups is reduced.
8. The computer system according to claim 7, An item upper limit, which is the upper limit of the resource item, is defined for the node, The resource item includes at least any one of the number of journal volumes, the number of secondary volumes, the number of remote copies, and the number of pairs, The management computer determines the node for arranging the secondary journal volume and the secondary volume based on the calculated load of the secondary journal volume and the secondary volume, the information of the nodes of the storage cluster including the cost, and the upper limit. A computer system characterized by this.
9. The computer system according to claim 4, For the secondary volume and the secondary journal volume, they are made redundant by an active storage control program and a standby storage control program arranged on different nodes. When a failure occurs in the active storage, the standby storage control program on a different node takes over and operates the secondary volume and the secondary journal volume. Multiple active or standby storage control softwares can operate on the same node, The management computer determines a node for arranging the secondary journal volume and the secondary volume based on the total load of the active storage control program and the storage control program changed from standby to active. A computer system characterized by the above.
10. A method for managing a computer system having a storage system that constructs a primary site that provides one or more primary volumes to a host, a storage cluster that is connected to the storage system via a network and has a plurality of nodes, and a management computer, When the management computer constructs a secondary site having a secondary volume with remote copy set for the primary volume of the primary site in the storage cluster, a step of managing a plurality of volumes including the secondary volume as a group; A step in which the management computer controls to arrange a plurality of volumes having the secondary volume on a plurality of nodes of the storage cluster based on the group; A method for managing a computer system, characterized by including the above.
Citation Information
Patent Citations
Storage system
JP2005018506A
Information processing system and path management method
JP2019185328A