Information processing system and control method for information processing system

JP7927033B2Active Publication Date: 2026-09-30HITACHI VANTARA LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024105359
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-09-30
Estimated Expiration
2044-06-28

AI Technical Summary

Benefits of technology

【0008】 本発明によれば、コピー先ボリュームの分散を抑制して性能を向上できる。上記した以外の課題、構成及び効果は以下の実施の形態の説明により明らかにされる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927033000001
    Figure 0007927033000001
  • Figure 0007927033000002
    Figure 0007927033000002
  • Figure 0007927033000003
    Figure 0007927033000003
Patent Text Reader

Abstract

To improve performance by suppressing dispersion of a copy destination volume.SOLUTION: Setting a first journal group including a primary volume and a first journal volume and a second journal group including a secondary volume and a second journal volume in a consistency maintaining group, and performing journal processing while ensuring consistency of replication between the first journal group and the second journal group; The second journal group is arranged in each node in which the secondary volume of the consistency holding group is arranged, and the secondary volume is arranged in each node of the second storage system on the basis of the use state of resources of each node of the second storage system and the number of nodes in which the secondary volume in the consistency holding group is arranged.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention relates to an information processing system and a control method for an information processing system. [[Background Art]]

[0002] Disaster recovery (DR) technology is known, which multiplexes and stores data at a remote site (secondary site) in preparation for data loss at a primary site when a large-scale disaster such as an earthquake or fire occurs. As storage operation in hybrid cloud environments progresses, an increasing number of cases are constructing a DR environment for data of on-premises storage systems on the cloud.

[0003] Distributed storage systems such as SDS (Software Defined Storage) are also used on the cloud. A DR environment is constructed in this distributed storage system. A distributed storage system consists of a large number of nodes, and when creating a volume, it automatically selects a node with available capacity to create the volume. In order to level the load on nodes by considering the IO load on each volume and the specifications of the nodes, there is the technology disclosed in Patent Document 1. Patent Document 1 describes that "In a distributed storage system 1, a volume classifier 300 classifies a plurality of volumes into a plurality of groups based on the load fluctuation cycle of each volume. A processor (resource classifier 400) calculates a total load obtained by summing the loads of a plurality of volumes on the same node in a group for each time period, and calculates a group load based on the peak of the total load. A processor (rebalancer 500) of any node calculates the group load of a destination node when a candidate volume to be moved in rebalancing for moving volumes between nodes is moved from a source node to the destination node, determines the volume to be moved in rebalancing and the destination volume based on the calculated group load of the destination node, and executes rebalancing." [[Prior Art Documents]] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2021-197010 [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] However, the prior art disclosed in Patent Document 1 does not consider the case where a disaster recovery (DR) environment is built between on-premises storage and a distributed storage system using DR technology. In a disaster recovery (DR) environment between an on-premises storage system and a distributed storage system, a DR relationship is established by creating a volume (destination volume) on one of the nodes of the distributed storage system for a volume (source volume) located on the on-premises storage system. However, if the destination volume is created distributed across multiple nodes of the distributed storage system, the number of journal groups for managing the update order increases, leading to decreased processing efficiency and reduced performance.

[0006] Therefore, the present invention aims to improve performance by suppressing the distribution of destination volumes, reducing the number of nodes, and lowering the overhead of journal processing. [Means for solving the problem]

[0007] To achieve the above objective, a typical information processing system of the present invention comprises: a first storage system having a node that provides a primary volume to a host; and a second storage system having a plurality of nodes that hold secondary volumes which are copies of the primary volume. In this information processing system, there are a plurality of pairs of primary volumes and secondary volumes, and a consistency-maintaining group is defined to replicate the data in the plurality of primary volumes, which is written to the plurality of primary volumes with consistency, to the plurality of secondary volumes while ensuring consistency. The replication includes a first journal volume provided on the same node as the primary volume and a second journal volume provided on the same node as the secondary volume. The journaling is performed using a method, and within the consistency-maintaining group, a first journal group including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set up, the journaling is performed while ensuring the consistency of the replicas between the first journal group and the second journal group, the second journal group is set up for each node where the secondary volume of the consistency-maintaining group is located, and the secondary volume is set up on each node of the second storage system based on the resource usage of each node of the second storage system and the number of nodes where the secondary volume is located in the consistency-maintaining group. Furthermore, one representative control method for an information processing system of the present invention is a control method for an information processing system comprising: a first storage system having a node that provides a primary volume to a host; and a second storage system having a plurality of nodes that hold secondary volumes which are copies of the primary volume, wherein there are a plurality of pairs of primary volumes and secondary volumes, and a consistency-maintaining group is configured to replicate the data in the plurality of primary volumes, which is written to the plurality of primary volumes with consistency, to the plurality of secondary volumes while ensuring consistency, and the replication is configured to include a first journal volume provided on the same node as the primary volume and a second journal volume provided on the same node as the secondary volume The journaling is performed using a method, and within the consistency-maintaining group, a first journal group including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set up, the journaling is performed while ensuring the consistency of the replication between the first journal group and the second journal group, the second journal group is set up for each node where the secondary volume of the consistency-maintaining group is located, and the secondary volume is set up on each node of the second storage system based on the resource usage of each node of the second storage system and the number of nodes where the secondary volume is located in the consistency-maintaining group. [Effects of the Invention]

[0008] According to the present invention, performance can be improved by suppressing the dispersion of the copy destination volume. Other problems, configurations, and effects will be clarified by the following description of the embodiments. [Brief explanation of the drawing]

[0009] [Figure 1]This is a block diagram showing an example of the hardware configuration of the information processing system 1 according to Embodiment 1 of the present invention. [Figure 2] This diagram shows the software configuration. [Figure 3] This figure shows the control information stored in memory 222. [Figure 4] This figure shows an example of volume information 410. [Figure 5] This figure shows an example of CTG information 420. [Figure 6] This figure shows an example of DR management information 430. [Figure 7] This figure shows an example of node information 440. [Figure 8] This figure shows an example where the destination volume is distributed across multiple nodes. [Figure 9] This flowchart shows an example of the process steps for selecting a destination volume. [Figure 10] This figure (Part 1) shows an example where the destination volume is distributed across multiple nodes. [Figure 11] This figure (part 2) shows an example where the destination volume is distributed across multiple nodes. [Figure 12] This figure (part 3) shows an example where the destination volume is distributed across multiple nodes. [Figure 13] This flowchart shows an example of the processing steps for moving a destination volume. [Figure 14] This flowchart shows an example of the processing steps for determining when to perform data migration. [Figure 15] This flowchart shows an example of the procedure for another process: selecting the destination volume. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the following embodiments are examples for illustrating the present invention, the present invention is not limited to the embodiments, and all application examples that conform to the idea of the present invention are included in the technical scope of the present invention. The present invention can also be implemented in various other forms. Unless particularly limited, each component may be plural or singular.

[0011] In the following description, the information of the present invention is described with the expression "~ information", but such information may be represented by, for example, a data structure such as "table", "list", "DB (database)", or any other form. In addition, when describing the content of each piece of information, the expressions "identification information", "identifier", "name", "ID" may be used, and these may be replaced with each other.

[0012] In addition, in the following description, processing performed by executing a program may be described. The program is executed by at least one or more processors (e.g., CPU) to perform predetermined processing while appropriately using storage resources (e.g., memory) and / or interface devices (e.g., communication ports). Therefore, the subject of processing may be the processor. Similarly, the subject of processing performed by executing a program may be a controller, apparatus, system, computer, node, storage system, storage apparatus, server, management computer, client, or host that includes the processor. The subject of processing performed by executing a program (e.g., a processor) may include a hardware circuit that performs part or all of the processing, and may also be modularized. For example, the subject of processing performed by executing a program may include a hardware circuit that performs encryption and decryption, or compression and decompression. Various programs may be installed in each computer via a program distribution server or a storage medium. The processor operates according to the program to function as a functional unit that implements a predetermined function. Apparatuses and systems including a processor are apparatuses and systems including these functional units. In addition, "read / write processing" may be described as "read / write processing" or "update processing" etc.

[0013] Further, in each drawing, common components are denoted by the same reference numerals. When elements of the same type are described without being distinguished in each drawing, reference signs or common numbers in the reference signs are used, and when elements of the same type are described while being distinguished, reference signs of the elements may be used, or IDs assigned to the elements may be used in place of the reference signs.

Example

[0014] FIG. 1 is a block diagram showing an example hardware configuration of an information processing system 1 according to Example 1 of the present invention. As shown in FIG. 1, the information processing system 1 is a disaster recovery system that provides a disaster recovery configuration. The information processing system 1 includes a primary site 100 and a secondary site 200, each of which is connected to each other via a network 10 (typically an IP (Internet Protocol) network). In the present embodiment, a disaster recovery configuration using an on-premises-based storage system (on-premises storage) for the primary site 100 and a distributed storage system for the secondary site 200 will be described. However, it is sufficient that at least the secondary site 200 is a distributed storage system, and the essence of the present invention does not change even if the primary site 100 is a distributed storage system. Further, for the sake of simplicity of description, "disaster recovery" may be abbreviated as DR (Disaster Recovery). The distributed storage system may be a cloud-based storage system (cloud storage) located in a cloud.

[0015] The primary site 100 is a storage system that provides application-based services to users (customers) in a normal state, and is formed of a so-called on-premises storage system. Specifically, the primary site 100 includes a server system 110, a storage controller 120, and a storage device 130.

[0016] The server system 110 has a processor 111, memory 112, and a network interface (I / F) 113, and is connected to the network 10 via I / F 113. The storage controller 120 has memory 122, a front-end network interface (I / F) 124, a back-end storage interface (I / F) 123, and a processor 121 connected to them. The storage controller 120 is connected to the server system 110 via I / F 124 and to the storage device 130 via I / F 123. The storage device 130 is a storage device that physically stores data. The memory and processor in the server system 110 and the storage controller 120 are redundant.

[0017] Memory 122 stores information and one or more programs. The processor 121 executes these one or more programs to provide storage space to the server system 110 (described here as a logical volume (for example, a volume created from a storage pool that virtualizes the capacity of storage device 130), but the essence of the invention does not depend on this), and processes I / O (Input / Output) requests such as write requests and read requests from the server system 110. For example, the server system 110 receives a write request or read request specifying a volume from a host device used by a user (customer) and sends it to the storage controller 120. The storage controller 120 then responds to the write request or read request by reading or writing data to the corresponding volume in the storage device 130.

[0018] The storage controller 120 and the storage device 130 may be configured as a single storage system. Examples include high-end storage systems using RAID (Redundant Array of Independent (or Inexpensive) Disks) technology and storage systems using flash memory.

[0019] The configuration of the storage device 130 may be, for example, a distributed storage system. Furthermore, the storage device 130 may be a Hyper-Converged Infrastructure (HCI) storage system, for example, a system that functions as a host system that issues I / O requests (e.g., an execution entity of an application that issues I / O requests (e.g., a virtual machine or container)) and as a storage system that processes said I / O requests (e.g., an execution entity of storage software (e.g., a virtual machine or container)). This is just one example of the configuration of the storage device 130 and is not limited thereto.

[0020] The secondary site 200 is a disaster recovery site (DR site) that holds data from the primary site 100 in preparation for data loss at the primary site 100 in the event of a large-scale disaster such as an earthquake or fire, and uses the held data to restore data and services at the primary site 100 in the event of a failure at the primary site 100. In the information processing system 1 according to this embodiment, the secondary site 200 is typically a distributed storage system. Alternatively, the secondary site 200 may be a distributed storage system belonging to a public cloud. Alternatively, the secondary site 200 may be a cloud storage system belonging to a public cloud and serving as the basis for cloud storage services provided by a cloud vendor. Examples of cloud storage services include AWS (Amazon Web Services) (registered trademark), Azure (registered trademark), and Google Cloud Platform (registered trademark). The cloud storage system used for the secondary site 200 may be a storage system belonging to another type of cloud (e.g., a private cloud) instead of a public cloud.

[0021] Secondary site 200 consists of a distributed storage system and is connected to network 10. In the information processing system 1 shown in Figure 1, data transfer between the storage device 130 at the primary site 100 and the distributed storage system at the secondary site 200 is performed directly via network 10. However, this is not the only method; data transfer may be performed using other network paths or lines. Furthermore, the type of network or line is not essential to the present invention.

[0022] A distributed storage system consists of multiple storage computers, each containing storage devices and processors, connected to one another via a network. Each computer is also called a node within the network. Each computer that makes up a distributed storage system is specifically called a storage node, and each computer that makes up a compute cluster is also called a compute node.

[0023] The distributed storage system according to this embodiment will be described in detail. A distributed storage system is composed of multiple storage nodes 220A to 220C (collectively referred to as Node 220) connected to each other by a network 203. The hardware configuration of each storage node is not particularly limited, but for example, Node 220A may have a processor (Central Processing Unit) 221, memory 222, network interface 224, drive interface 223, and storage device 225. These are connected by an internal network. Node 220A connects to the network 203 via the network interface 224 and communicates with other storage nodes (Nodes 220B to 220C). In a distributed storage system, depending on the configuration of the network 203, Node 220 may be composed of Node 220 located in geographically sufficiently distant locations.

[0024] Furthermore, although in this embodiment all of the nodes 220A to 220C constituting the distributed storage system are exemplified as storage nodes, the nodes constituting the distributed storage system are not limited to storage nodes, and may include some nodes that function as compute nodes.

[0025] A distributed storage system consists of storage nodes, each equipped with an operating system (OS) for managing and controlling them. Storage software with the functionality of a storage system runs on top of this OS, thus forming the distributed storage system. The storage software can also be run in the form of containers on the OS to create a distributed storage system. A container is a mechanism for packaging one or more software programs and configuration information. Furthermore, a hypervisor can be installed on the storage nodes, allowing the OS and software to run as virtual machines (VMs) to create a distributed storage system.

[0026] Furthermore, the present invention is also applicable to HCI. HCI is a system that enables multiple processes to be performed on a single node by running applications, middleware, management software, and containers, in addition to storage software, on top of an OS or hypervisor installed on each node.

[0027] A distributed storage system provides a host with storage pools and logical volumes (also simply called volumes) that virtualize the capacity of storage devices on multiple storage nodes.

[0028] In another embodiment, the information processing system 1 may include a storage management system. The storage management system may be part of the primary or secondary site components, or it may be a dedicated management appliance or management terminal. The management device or management terminal may be connected to the network 10.

[0029] For example, this storage management system is a computer system (one or more computers) that manages the configuration of the storage areas of storage device 130 and storage device 225, and can receive instructions from a user (or administrator) regarding the settings of storage device 130 and storage device 225. Alternatively, the user can receive instructions regarding the settings of storage device 130 and storage device 225 via a management terminal. The storage management system may also consist of separate devices for the primary site storage system and the secondary site storage system.

[0030] Administrators of distributed storage systems can perform operations such as creating, deleting, and moving volumes by issuing management commands to the distributed storage system over the network. Furthermore, the distributed storage system can notify administrators and management tools of its status, such as drive usage and processor usage, by providing information transmitted over the network.

[0031] For example, a storage management system can manage resource configuration information and application information for the primary site 100 and instruct the secondary site 200 to run the same VM (VM252) and application (Application 253) as the primary site 100, thereby creating a disaster recovery (DR) environment.

[0032] Resource configuration information and application information for the primary site 100 may be stored in the memory 112 or storage device 130 of the server system 110, or a storage management system may retrieve the information from there.

[0033] With the above configuration, data from storage device 130 can be copied and stored to storage device 225 in the distributed storage system, and the copied data can be used to restore the volume in storage device 130 within the distributed storage system. This enables the realization of a disaster recovery (DR) configuration information processing system 1 within the primary site 100, which includes storage device 130, and the secondary site 200, which includes the distributed storage system.

[0034] When establishing a disaster recovery (DR) relationship by creating a volume (destination volume) on one of the nodes of the distributed storage system, secondary site 200, for a volume (source volume) located in the storage system of primary site 100, there are several possible cases in which the destination volume is created distributed across multiple nodes of the distributed storage system. Examples of this case are shown below (1) to (4).

[0035] (1) When a destination volume is created for a source volume and assigned to a Consistency Time Group (CTG), and then a destination volume is created for another source volume in the same manner and added to the CTG, if the destination volume is created on a different node than the node where the previously created destination volume was created, the destination volumes within the same CTG will be distributed across multiple nodes.

[0036] (2) When creating destination volumes for multiple source volumes belonging to a CTG in a single batch, if the available space on the selected node is insufficient compared to the total capacity of all volumes, the destination volumes will be distributed across multiple nodes because it will be necessary to create destination volumes on other nodes due to insufficient available space.

[0037] (3) When creating destination volumes for multiple source volumes belonging to a CTG in a single batch, the destination volumes are distributed across multiple nodes to prevent the load on specific nodes from being placed on them during the initial full volume data copy process.

[0038] (4) When using an existing volume already created within a distributed storage system as a destination volume, if the node of the existing volume is different, the destination volumes within the same CTG will be distributed across multiple nodes.

[0039] A CTG (Computer-to-Data Group) is a group of volumes that perform data-consistent copying across multiple volumes. Operations can be separated by host application groups, or volumes related to the same business process can be managed under the same CTG.

[0040] In a disaster recovery (DR) environment, the journal manages the write order (update order) from the host during copy operations. A journal group is a collection of volumes consisting of one or more data volumes and journal volumes, and journal processing is performed on a journal group (JNLG) basis. First, a JNLG must be created by registering the volumes.

[0041] Since volumes belonging to the same journal volume are guaranteed to be consistent, JNLG belongs to a single CTG. Since each node in a distributed storage system operates independently, a JNL (Network Name Logistics) is provided for each node. When the nodes of the destination volumes within a CTG are distributed, a JNLG is created for each node, increasing the number of JNLGs within the CTG. If the number of nodes in the destination distributed storage system is N, and the destination volumes are distributed across N nodes, the number of JNLGs can increase by up to N times in the worst case. As the number of JNLGs increases, the JNL processing, which was previously performed collectively within each JNLG, is now performed for each JNLG, leading to decreased processing efficiency and reduced performance.

[0042] Therefore, when the destination volume is distributed across multiple nodes in the distributed storage system, the information processing system 1 implements a control to reduce the number of nodes to which it is distributed, thereby reducing the overhead of JNL processing. Details of this will be described later.

[0043] The following is an example of the operation image of Information Processing System 1. Within the primary site 100, one or more virtual machines (VMs) are created on a server (server group) consisting of one or more server systems 110. On each VM, an application specified by the user is executed. An application specified by the user is an application that provides a service to the user, and the application corresponding to the service that the user uses is specified. An operating system (OS) for managing and controlling storage devices (in a distributed storage system, an OS for managing and controlling storage nodes) is installed, and storage software with the functionality of a storage system runs on top of it. Data is stored in the storage area within storage device 130 via a storage pool that virtualizes the capacity of the storage device.

[0044] Node 220 runs a host OS for controlling the hardware, and on top of that, a hypervisor runs to run one or more guest OSs as VMs.

[0045] On top of each guest OS, a container runtime runs to operate one or more containers, and storage and computing software runs on top of that. Note that in the above software stack, if the hypervisor includes functions for controlling hardware, the host OS can be omitted. Also, if it is not necessary to run each software on a VM, the hypervisor and guest OS can be omitted, in which case the container runtime can run on the host OS. Furthermore, if the storage or computing software does not run as a container, the container runtime can be omitted, in which case the storage or computing software can run directly on the guest OS or host OS.

[0046] Figure 2 is a schematic diagram showing the relationship between the software (or control programs, modules, functions, and components) in the storage system at primary site 100 and the distributed storage system at secondary site 200. The software includes storage control 310, DR control 320, data migration control 330, and monitor 340. Each software can communicate with each other and send and receive information. Each software module is executed on the storage controller 120 of the storage system at primary site 100. In the distributed storage system, it may be on the same node as the storage node 220, or on a different node or other hardware component such as a device, terminal, or circuit, as long as it is a location where the distributed storage system can be accessed via the network 10. Furthermore, not all software needs to be implemented on the same storage controller 120 and node 220. The form in which each software is executed can be any method, such as a process or a container.

[0047] The storage control 310 controls storage devices 130 and 225. For example, when an I / O request is issued to storage devices 130 and 225, it accesses the location where the data specified in the I / O request is stored in storage device 130 and provides the data to the host. In the case of distributed storage devices at secondary site 200, when a host issues an I / O request to any storage node, the distributed storage system provides the host with access to the data by forwarding the I / O request to the storage node that holds the data specified in the I / O request.

[0048] The DR control 320 has functions such as issuing instructions for building a DR environment between the primary site 100 and the secondary site 200, forming primary and secondary systems for VMs and applications in the DR environment (determining source and destination volumes, managing CTGs (Consistency Time Groups) and journal groups (JNLG)), controlling data transfer of data stored on the storage device 130 at the primary site 100 to the secondary site 200, and restoring volumes within the secondary site 200 from data copied within the secondary site 200.

[0049] The data migration control 330 has the function of scheduling the movement of the destination volume to another node and determining whether to move the data.

[0050] The monitor 340 monitors the load on each hardware element of the distributed storage system. The data migration control 330 refers to the monitoring information from the monitor 340, determines the destination of the volume, and moves the volume.

[0051] Figure 3 shows the control information stored in the memory 122 of the storage controller and the memory 222 of node 220. Specifically, the memory stores volume information 410, CTG information 420, DR management information 430, node information 440, and storage information 460. This information is accessed as needed when each software program shown in Figure 2 performs processing, and is referenced, read, generated, written, or updated. This information may be held by the management device.

[0052] Figure 4 shows an example of volume information 410. Volume information 410 is information about the volumes stored in storage devices 130 and 225, respectively. Volume information 410 includes volume ID 411, DR presence / absence 412, CTGID 413, JNLGID 414, node ID 415, capacity 416, and usage 417.

[0053] Volume ID 411 stores the identifier of the volume provided to the host. DR presence / absence 412 indicates whether volume 411 has a copy destination volume in a DR configuration. In the case of volume information for a distributed storage system, it indicates whether there is a source volume. CTGID 413 indicates the identifier of the CTG (Consistency Time Group) to which volume 411 belongs, and JNLGID 414 indicates the identifier of the JNLG to which volume 411 belongs. Node ID 415 is provided only in the case of volume information for a distributed storage system and indicates the identifier of the node to which volume 411 is provided. Capacity 416 indicates the storage capacity of volume 411, and usage 417 indicates the amount of the storage capacity used, such as when data is stored.

[0054] Figure 5 shows an example of CTG information 420. The CTG information 420 shown in Figure 5 consists of the following items: CTGID 421, JNLG number 422, JNLGID 423, volume number 424, and volume ID 425.

[0055] CTGID421 is the identifier of the CTG configured in the DR configuration of Information Processing System 1. The number of JNLGs, 422, is the number of JNLGs belonging to CTGID421 and is the total number of JNLGs listed in JNLGID423. JNLGID423 indicates the identifier of the JNLG belonging to CTGID421. The number of volumes, 424, is the number of volumes belonging to JNLG volume ID425 and is the total number of volumes listed in volume ID.

[0056] A CTG (Continuous Data Group) is a group designed to guarantee update order across multiple volumes. It consists of multiple volumes configured on the primary and secondary storage systems of a disaster recovery (DR) configuration. By specifying a CTG, volumes can be operated collectively on a CTG basis. For example, multiple volumes used for the same task can be configured to belong to the same CTG (Computer Group).

[0057] A CTG may be created using some of the nodes in a distributed storage system. For example, if a distributed storage system consists of three nodes, from node 1 to node 3, then CTG1 may consist of nodes 1 and 2, and CTG2 may consist of nodes 2 and 3.

[0058] Figure 6 shows an example of DR management information 430. The DR management information 430 includes information indicating the corresponding DR relationship in the DR environment within the information processing system 1. The DR management information 430 exemplified in Figure 6 consists of the following items: primary 431, primary volume ID 432, secondary 433, and secondary volume ID 434.

[0059] Primary ID 431 stores an identifier, such as the serial number, that indicates the storage system constituting the primary site 100 of the DR. Primary volume ID 432 stores the identifier of the volume at the primary site. Secondary ID 433 stores an identifier, such as the serial number, that indicates the storage system constituting the secondary site 200 of the DR, corresponding to Primary ID 431. Secondary volume ID 434 stores the identifier of the volume at the secondary site.

[0060] In a disaster recovery (DR) configuration, the volume located at the primary site is called the source volume, and the volume located at the secondary site is called the destination volume. The source volume is sometimes denoted as PVOL and the destination volume as SVOL.

[0061] Figure 7 shows an example of node information 440. Node information 440 is data that shows the configuration of each node 220 that makes up the distributed storage system. In Figure 7, node information 440 has data items for node ID 441, volume 442, capacity 443, usage 444, availability 445, and IO frequency 446. Capacity 443 is the total value of the drive capacity installed in the target node 441. Usage 444 is the total capacity allocated to the volumes provided to the host from the storage devices 225, which are physical drives within node 441. Availability 445 indicates the utilization rate of the processor within the node. IO frequency 446 indicates the ratio of read / write requests to the node per unit time.

[0062] Figure 8 shows an example where the destination volumes for a disaster recovery (DR) within a CTG are distributed across multiple nodes. Volume 513 exists on storage system 510, which is built at the on-premises primary site 100. Here, volume 513 consists of two volumes, 513A and 513B. In this example, the distributed storage system consists of two nodes, 520A and 520B. The two volumes 513A and 513B are copied as the source volumes for the DR to destination volumes 523A and 523B, respectively. Volumes 513A and 513B belong to the same CTG 511. The destination volumes 523A and 523B also belong to the same CTG 521.

[0063] Volume 513A replicates data to Volume 523A located on Node 520A, and Volume 513B replicates data to Volume 523B located on Node 520B. SVOL1 is the destination volume for PVOL1, and SVOL2 is the destination volume for PVOL2. In a disaster recovery environment, the copy process from an on-premises storage system to a distributed storage system is performed using, for example, conventional asynchronous remote copy technology.

[0064] Journal groups (JNLGs) are used to manage differential copies of data between source and destination volumes. A JNLG is a collection of volumes consisting of one or more data volumes and journal volumes. Data volumes are the source and destination volumes, while journal volumes store information representing the history of write data from the host and updates to that data. When the primary site's storage system receives a write request, it writes data to the data volume and journal data to the journal volume, and returns a response to the server system. The secondary site's destination storage system reads journal data from the source storage system's journal volume asynchronously with respect to the write request and stores it in its own journal volume. The destination storage system then restores the copied data to the destination data volume based on the stored journal data. In this way, the write order is managed to be restored. The above journaling process is performed on a JNLG basis.

[0065] In Figure 8, JNLG512 or 522 includes volume 513 or 523 and the journal volume. Since node 520 of the distributed storage system operates on its own OS, JNLG is also configured for each node. Storage system 510 configures the JNLG corresponding to the destination node, so it also provides JNLG for the number of destination nodes that are distributed. Volume 513A is copied to volume 523A in node 520A, so volume 513A belongs to JNLG512A. Volume 513B is copied to volume 523B in node 520B, so volume 513B belongs to JNLG512B. If, across all configured CTGs, the destination volumes are distributed to all nodes of the distributed storage system, then the number of JNLGs required will be equal to the number of nodes for each CTG. Thus, when distributing data across multiple nodes of the destination volume within the same CTG, reducing the number of nodes to which the data is distributed is the solution to the problem. In other words, reducing the number of JNLGs within the CTG prevents a deterioration in processing efficiency and improves performance.

[0066] The procedure for creating the configuration shown in Figure 8 is as follows, and is performed by the DR control 320. For example, a CTG1 is created and registered in the CTG information 420 for the DR configuration of the DR management information 430 in Figure 6.

[0067] If the DR destination is set to nodes 520A and 520B, journal volumes JVOL3 (524A) and JVOL4 (524B) are created on nodes 520A and 520B. JVOL3 is registered to JNLG1 on node 520A, and JVOL4 is registered to JNLG2 on node 520B. Journal volumes JVOL1 (514A) and JVOL2 (514B) are created on the DR source storage system 510, and registered to JNLG1 and JNLG2 respectively. JNLG1 and JNLG2 are registered to CTG1.

[0068] If there are other CTG2s, process them similarly. For example, create a journal volume JVOL7 on node 520A and a journal volume JVOL8 on node 520B. Register JVOL7 to JNLG3 and JVOL8 to JNLG4. Create journal volumes JVOL5 and JVOL6 on the DR source storage system 510 and register them to JNLG1 and JNLG2 respectively. Register JNLG3 and JNLG4 to CTG2.

[0069] Figure 9 shows a flowchart of the process for selecting the destination volume. First, if the source storage system 510 is a volume targeted for DR, it registers PVOL1 (513A) with CTG1 (S1010). Then, storage system 510 requests the distributed storage system to create SVOL1, which is the destination volume for the DR copy of PVOL1 (S1020).

[0070] The compute node of the distributed storage system receives the request and determines which node to create SVOL1 on (S1030), and sends a request to create SVOL12 to the storage node 522, which is the result of that determination (S1040). The method for selecting the node to send the request to is to select a node with a large amount of free space in its storage capacity, or a node with a low utilization rate of its processors. Alternatively, a node with an already created but unused volume may be selected, and the unused volume may be designated as SVOL1.

[0071] Upon receiving the aforementioned request, node 1 (520A) creates SVOL1 (S1050) and registers it as the copy destination of PVOL1. SVOL1 is registered in JNLG1 of CTG1 within the node (S1060). Node 1 reports to storage system 510 via compute node that SVOL1 has been created in JNLG1 (S1070).

[0072] The storage system 510 receives the JNLGID created by SVOL1 and registers the PVOL with the received JNLG1 among the JNLGs in CTG1 (S1080).

[0073] Similarly, for PVOL2 (513B), which is the DR target volume, PVOL2 is registered with CTG1. A request is made to the distributed storage system to create SVOL, which is the destination volume for copying PVOL2. In this example, node 2 (520B) receives the request, creates SVOL, and registers it as the destination for copying PVOL2. SVOL is registered with JNLG12 of CTG1 within the node. Storage system 510 receives the JNLGID and registers PVOL with the received JNLG2 among the JNLGs in CTG1.

[0074] Here, we identified which JNLG the PVOL should register with by receiving the JNLGID from the DR copy destination, but the goal is for the PVOL and SVOL to be registered with the same JNLG, so other methods would also be acceptable. For example, if you create and register a journal volume JVOL1 on storage system 510's CTG1, and then create a copy destination volume, you can generate a new JNLG2 when SVOL is created on a different node than the previous one.

[0075] Let's explain using the examples in Figures 10 to 12. In a DR configuration of a distributed storage system consisting of a storage system 510 and a node 520, the storage system 510 manages two CTGs, CTG511 and CTG515, while the node 520 of the distributed storage system manages two CTGs, CTG521 and CTG525. CTG511 and CTG521 are managed as the same CTG. CTG515 and CTG525 are also the same CTG.

[0076] The source volume 513 belongs to CTG511, and volumes 513A (PVOL1~PVOL3 in the diagram) are copied to destination volume 523A (SVOL1~SVOL3 in the diagram) on node 520A, and volumes 513B (PVOL4~PVOL10 in the diagram) are copied to destination volume 523B (SVOL4~SVOL10) on node 520B. Since the destination volumes are distributed across two nodes, and JNLG is installed for each node, source volume 513A belongs to JNLG512A, destination volume 523A belongs to JNLG522A, source volume 513B belongs to JNLG512B, and destination volume 523B belongs to JNLG522B. JNLG512A and JNLG522A are managed as the same JNLG. JNLG512B and JNLG522B are also managed as the same JNLG. In other words, a single CTG511 can contain multiple JNLGs (JNLG512A and JNLG512A).

[0077] The same applies to CTG515. The source volume 517 belongs to CTG515, and volumes 517A (PVOL21~PVOL24 in the diagram) are copied to destination volume 527A (SVOL21~SVOL24 in the diagram) on node 520A, and volumes 517B (PVOL25~PVOL29 in the diagram) are copied to destination volume 527B (SVOL25~SVOL29) on node 520B. Since the destination volumes are distributed across two nodes and JNLG is installed for each node, source volume 517A belongs to JNLG516A, destination volume 527A belongs to JNLG526A, source volume 517B belongs to JNLG516B, and destination volume 527B belongs to JNLG526B. JNLG516A and JNLG526A are managed as the same JNLG. JNLG516B and JNLG526B are also managed as the same JNLG. In other words, multiple JNLGs (JNLG512A and JNLG512A) exist within a single CTG511.

[0078] When creating a destination volume, the system selects a node and creates the volume, taking into account factors such as the available storage space and processor utilization of that node at the time of creation. However, the node load changes over time. By referring to hardware load monitoring information, the system moves and rearranges the destination volumes between nodes to equalize the load on the JNL process. The rearrangement of volumes between nodes for equalization may be the optimal solution, or it may be designed to minimize the number of nodes where the destination volumes are distributed. Reducing the number of distributed nodes improves the efficiency of the JNL process and enhances system performance. System performance can be improved by "placing JNLGs belonging to the same CTG on as few nodes as possible to avoid distribution" and "changing the placement of multiple CTGs to equalize the processing load."

[0079] Figures 11 and 12 show an example of moving a destination volume from the configuration in Figure 10. For CTG521, all destination volumes 523A located in node 520A are moved to node 520B. Once volume 523A is moved to node 520B, it belongs to JNLG522B, and JNLG522A becomes unnecessary. Along with the movement of the destination volume, JNLG512A of CTG511 in storage system 510 is merged with JNLG512B, and JNLG512A becomes unnecessary. Unnecessary JNLGs are deleted. In this way, the number of JNLGs for CTG511 and CTG521 can be reduced. For such volume movements to be possible, it is assumed that node 520B has sufficient free space and low processor load. Similarly, with CTG515 and CTG525, the number of nodes on the destination volume can be reduced.

[0080] Furthermore, as shown in Figure 12, it is also possible to create a volume move plan that simultaneously moves volumes between CTG511 and CTG521, and between CTG515 and CTG525. In a plan to move volumes belonging to multiple CTGs simultaneously, depending on the processing method, such as moving one volume at a time, it becomes possible to move volumes even when the node has limited free capacity, and the number of JNLGs consolidated on the node increases, improving performance. The volume arrangement may be as shown in Figure 12, via Figure 11, or as shown in Figure 10, via Figure 12.

[0081] Figure 13 shows a flowchart of the process of reducing the number of nodes to which the copy destination volume within the CTG is distributed, i.e., reducing the number of JNLGs. This process is performed by the data migration control 330. The data migration control 330 may be a functional unit implemented by a program executed by one of the CPUs shown in Figure 1, or it may be a function executed by a management system, management device, or management terminal as appropriate.

[0082] First, the data migration control 330 selects the CTG to be moved (S1410). Specifically, the data migration control 330 searches for a CTG where the nodes of the destination volume are distributed. As an example, the data migration control 330 refers to the CTG information 420 and selects CTGID 421, which has a large number of JNLGs (422).

[0083] Next, the data migration control 330 selects a JNLG to be moved from among the JNLGs belonging to the selected CTG (S1420). Specifically, the data migration control 330 selects a JNLGID 423 to be moved to another node from among the JNLGID 423 of the selected CTGID 421. For example, the JNLGID 423 to be moved is selected as the JNLGID 423 with the fewest volumes 424 belonging to the JNLG.

[0084] Next, the data migration control 330 selects the destination JNLG for the JNLG to be moved (S1430). Specifically, the data migration control 330 selects the JNLGID 423 from among the JNLGID 423 of the selected CTGID 421, for example, the JNLGID 423 with the largest number of volumes (424) belonging to the JNLG.

[0085] The data migration control 330 determines whether it is possible to move all destination volumes belonging to the selected source JNLG to the node of the selected destination JNLG. First, the data migration control 330 refers to the CTG information 420 and calculates the total capacity of all destination volumes belonging to volume ID 425 of source JNLGID 423. It refers to the volume information 410, searches for volume ID 411 corresponding to volume ID 425, and obtains the capacity 416. It compares this with the free space in the storage area of ​​the destination JNLG node, and if the total capacity is less than the free space, it determines that the target JNLG can be moved (S1440). If the total capacity is greater than the free space, it determines that it cannot be moved and returns to S1410 to select another CTG. Figure 13 shows the process of returning to S1410, but as an alternative process, it may return to S1430 and re-select the destination JNLG to another JNLG, or return to S1420 and re-select the target JNLG to another JNLG.

[0086] After determining that the target JNLG is movable, the data migration control 330 then checks the hardware operating status and processor utilization of the destination node to determine whether the utilization of the destination falls within an acceptable range (S1450). For example, if the utilization is capped at 90%, the system simulates whether the increased load from the volume move will keep the utilization within 90%. If it determines that it will fall within that range, the move is permitted. If it does not, the system determines that the move is not possible and returns to S1410 to select another CTG. Figure 13 shows the process of returning to S1410, but as an alternative, the system may return to S1430 and re-select the destination JNLG to another JNLG, or return to S1420 and re-select the target JNLG to another JNLG. Here, the aforementioned 90% is just an example and should be a variable value. The utilization rate can also be predicted from the node's access frequency.

[0087] If permission to move is granted, the data migration control 330 sequentially moves the destination volumes of the source JNLG to the destination JNLG (S1460). As another example, if there are no more volumes belonging to JNLG, you can delete JVOL and then delete JNLG. There is another embodiment. In a DR configuration, the destination volume, which is a copy of the source volume, is created when the destination storage system receives a request to create the destination volume, creates the volume, and then establishes the relationship between the source volume and the destination volume. Due to this processing procedure, when a distributed storage system is configured with a DR configuration, the distributed storage system creates the destination volume on a node with available capacity. However, by specifying a node from the source storage system to create the destination volume, it is possible to create a destination volume that is not distributed across multiple nodes.

[0088] The data migration control 330 determines when to execute the process of reducing the number of nodes to which the copy destination volumes within the CTG in Figure 13 are distributed, i.e., reducing the number of JNLGs. Monitor 340 performs a process to monitor the load on each hardware element of the distributed storage system. The timing of the monitoring process is time-periodized. The time period can be set arbitrarily and may be constant or irregular. The information obtained from the monitoring process is stored in the node information 440, specifically in the availability 445 and I / O frequency 446. The monitoring process is initiated by the management device or storage control 310.

[0089] The data migration control 330 refers to the load of the hardware elements as measured by the monitor 340, determines the destination of the volume, and moves the volume. In another example, predictive resource utilization could be calculated by monitoring the load on hardware elements, and this calculated predictive resource utilization could be used to determine where to move the volume.

[0090] The timing for executing data migration is determined by setting thresholds for the load on the aforementioned hardware elements, such as an utilization rate of 445 and an I / O frequency of 446. Data migration is performed when the load is less than the threshold, i.e., when resource utilization is deemed low. Furthermore, data migration is also performed when the load is greater than the threshold, i.e., when resource utilization is deemed to be skewed towards certain nodes. Other data migration timings are determined by setting a threshold for the node's free capacity; data migration is performed if the free capacity falls below this threshold.

[0091] For example, the process shown in the flowchart in Figure 14 is performed at a specified time interval. The data migration control 330 determines whether there are any nodes 220 of the distributed storage system whose resource utilization exceeds a preset threshold (S1510). If there are nodes exceeding the threshold, the data migration process (S1540) is performed. If there are no nodes exceeding the threshold, the data migration control 330 determines whether there are distributed JNLG nodes (S1520). If there are distributed JNLG nodes, the data migration control 330 determines whether the distributed storage system is idle (S1530). In other words, it determines whether the system can operate even if the load of inter-node volume transfer processing due to the data migration process is placed on the hardware resources. If it is idle, the data migration process (S1540) is performed. If there are no distributed JNLG nodes in S1520, or if there are distributed JNLG nodes in S1520 but the node status is not idle, the process is terminated without starting the data migration process.

[0092] Thus, in Example 1, when the destination volume is distributed across multiple nodes in a distributed storage system, the overhead of JNL processing can be reduced and performance improved by minimizing the number of nodes to which it is distributed. [Examples]

[0093] As another example, Figure 15 shows a flowchart illustrating how to specify the node for creating the SVOL from the PVOL side. First, the source storage system 510 registers PVOL2 with CTG1 (S1110). Next, the storage system 510 determines which node of the distributed storage system to create SVOL2, the destination volume for PVOL2, on by checking if an SVOL already exists in the same CTG (S1120). If an SVOL already exists in the same CTG1, the storage system 510 obtains the JNLGID to which that SVOL was created (S1130). For example, if PVOL1 (513A) is the DR target volume, PVOL1 is registered in CTG1, and SVOL1 has already been created in JNLG1 of node 1, then PVOL2 belongs to CTG1, and it is determined that an SVOL already exists.

[0094] If SVOL has already been created, storage system 510 hands over the acquired JNLGID and requests the distributed storage system to create SVOL2, which is the destination volume for the DR copy of PVOL2 (S1140). If SVOL has not been created, it performs the processes S1020 to S1080 and terminates the process.

[0095] The compute node of the distributed storage system receives the request and transmits the request to the storage node 522 corresponding to the JNLGID (S1150).

[0096] Upon receiving the aforementioned request, node 1 (520A) creates SVOL2 (S1160) and registers it as the copy destination of PVOL2. SVOL2 is then registered in JNLG1 within the node (S1170). Node 1 reports to storage system 510 via the compute node that SVOL2 has been created in JNLG1 (S1180).

[0097] The storage system 510 receives the JNLGID created by SVOL2 and registers PVOL2 with JNLG1 (S1190). Here, the JNLGID information is shared between the storage system and the distributed storage system to determine which node will create the SVOL to avoid distribution across nodes. However, this is just one example; other information that indicates which node PVOL1 created the SVOL would also work. Alternatively, S1150 could provide information that allows the compute nodes of the distributed storage nodes to search for the node that created the SVOL corresponding to PVOL1. In this case, the compute nodes would determine which node to create the SVOL based on the information received from the storage system.

[0098] Figure 15 shows a case where PVOL1 belonging to CTG1 has already created SVOL1, and SVOL2 of PVOL2, also belonging to CTG1, is created on the same node. However, as another embodiment, when creating SVOLs for PVOL1 and PVOL2 belonging to CTG1 at the same time, it is also possible to create the SVOLs by selecting a node that has sufficient free capacity to create both SVOLs. In this case, the storage system 510 issues a request to the compute node to create SVOLs for multiple PVOLs on the same node, and either the compute node determines whether it can secure free capacity to determine the node on which to create the SVOLs, or the primary site storage system 510 determines whether it can secure free capacity to determine the node. Information for determining whether the primary site storage system 510 can secure free capacity can be obtained from the distributed storage system or from a management device, etc.

[0099] In another embodiment, if a particular node in a distributed storage system is under heavy load—for example, if the available capacity is below a predetermined threshold, or if the hardware utilization is above a predetermined threshold—it is possible to improve the overall system performance by moving volumes from that node to other nodes. Moving volumes may also distribute the destination volumes of a particular CTG across multiple nodes, potentially increasing the number of JNLGs.

[0100] Thus, in Example 2, the user can determine the destination volume in the DR environment and execute the copy operation without being aware of the nodes of the distributed storage system. This improves the user's flexibility in DR operations.

[0101] As described above, the information processing system 1 disclosed in the embodiment comprises a first storage system (primary site 100) having a node that provides a primary volume to a host, and a second storage system (secondary site 200) having a plurality of nodes that hold secondary volumes which are copies of the primary volume. In this information processing system, there are a plurality of pairs of primary volumes and secondary volumes, and a consistency holding group (CTG) is defined to replicate the data in the plurality of primary volumes, which is written to the plurality of primary volumes with consistency, to the plurality of secondary volumes while ensuring consistency. The replication is performed by a first journal volume provided on the same node as the primary volume and a second journal volume provided on the same node as the secondary volume. Journal processing is performed using journal volumes, and within the consistency-maintaining group, a first journal group (JNLG) including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set up, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group (JNLG), the second journal group is set up for each node where the secondary volume of the consistency-maintaining group is located, and the secondary volume is set up on each node of the second storage system based on the resource usage of each node of the second storage system and the number of nodes where the secondary volume is located in the consistency-maintaining group. Specifically, the multiple secondary volumes are arranged on the multiple nodes such that the number of second journal groups is reduced. This configuration and operation can improve performance by suppressing the distribution of destination volumes.

[0102] Furthermore, by providing the first journal group in correspondence with the second journal group, one or more first journal groups are arranged at one node. Furthermore, the second journal volume is provided for each node where the secondary volumes of the integrity maintenance group are located, and is used for journaling of one or more secondary volumes located on that node, and the first journal volume is provided in correspondence with the second journal volume, so that one or more first journal volumes are located on one node. With this configuration and operation, the number of journal groups is determined by the number of nodes on which secondary volumes are located; therefore, the number of journal groups can be reduced and performance improved by relocating secondary volumes.

[0103] Specifically, the number of journal groups and journal volumes is reduced by moving the secondary volume to a node where other secondary volumes of the same integrity group are located. Furthermore, the movement of the secondary volume between nodes reduces the number of the second journal group and the second journal volume, and this reduction in the second journal group and the second journal volume reduces the corresponding number of the first journal group and the first journal volume. This configuration and operation allows for improved performance when rearranging secondary volumes within a consistency-preserving group.

[0104] Furthermore, the information processing system 1 has multiple integrity maintenance groups, and multiple secondary volumes of the multiple integrity maintenance groups are arranged on multiple nodes of the second storage system based on the resources of each node of the second storage system. Then, for each of the multiple consistency-maintaining groups, the number of journal groups and journal volumes is reduced by moving multiple secondary volumes to the same node. Furthermore, based on the resources of the node, it is determined whether the secondary volume can be moved to a node where other secondary volumes of the same integrity group are located, and if it is determined that it can be moved, the secondary volume is moved. This configuration and operation allows for the appropriate selection of the destination node, taking into account the state of each node.

[0105] Furthermore, when the information processing system 1 creates the primary volume and secondary volume to belong to the integrity maintenance group, it creates the secondary volume on the node where the secondary volume of the same integrity maintenance group is located. This configuration and operation prevent situations where journal groups become dispersed.

[0106] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. For example, the embodiments described above are explained in detail to make the present invention easier to understand, and are not necessarily limited to those having all the configurations described. Furthermore, it is possible to replace or add configurations, not just delete them. [Explanation of Symbols]

[0107] 1. Information Processing System 10 Networks 100 Primary Sites 110 Server Systems 111,121,221 processors 112,122,222 memory 113,124,224 Network Interfaces (I / F) 120 Storage Controllers 123 Storage Interface (I / F) 130 Storage Devices 200 Secondary Sites 220 storage devices 310 Storage Control 320 DR control 330 Data Migration Control 340 monitors

Claims

1. A first storage system having a node that provides a primary volume to a host, An information processing system comprising: a second storage system having a plurality of nodes that hold secondary volumes which are copies of the primary volume; There are multiple sets of primary volumes and secondary volumes, and a consistency maintenance group is configured to replicate the data in the multiple primary volumes, which is written to the multiple primary volumes in a consistent manner, to the multiple secondary volumes while ensuring consistency. The replication is performed by journaling using a first journal volume located on the same node as the primary volume and a second journal volume located on the same node as the secondary volume. Within the consistency-maintaining group, a first journal group including the primary volume and the first journal volume, and a second journal group including the secondary volume and the second journal volume are established, and journal processing is performed while ensuring the consistency of the copies between the first journal group and the second journal group. The second journal group is located at each node where the secondary volume of the integrity-maintaining group is located. When a consistency-maintaining group exists in which multiple second journal groups are distributed across multiple nodes, the system determines whether the secondary volume of the second journal group on the source node can be moved to a destination node where other second journal groups are located. If it determines that it can be moved, the system moves the secondary volume of the second journal group on the source node to the second journal group on the source node, and deletes the second journal group and its journal volume on the source node. An information processing system characterized by the following:

2. The information processing system according to claim 1, The first journal group is provided in correspondence with the second journal group, so that one or more first journal groups are located at one node. An information processing system characterized by the following:

3. An information processing system according to Claim 2, The second journal volume is located at each node where the secondary volumes of the integrity maintenance group are located, and is used for journaling to one or more secondary volumes located at that node. The first journal volume is provided in correspondence with the second journal volume, so that one or more first journal volumes are arranged at one node. An information processing system characterized by the following:

4. The information processing system according to claim 3, The number of journal groups and journal volumes is reduced by moving the secondary volume to a node where other secondary volumes of the same integrity group are located. An information processing system characterized by the following:

5. The information processing system according to claim 4, The secondary volume moves between nodes, thereby reducing the number of the second journal group and the second journal volume. The reduction in the second journal group and the second journal volume reduces the number of corresponding first journal groups and first journal volumes. An information processing system characterized by the following:

6. The information processing system according to claim 2, It has multiple integrity-maintaining groups, Multiple secondary volumes of the multiple consistency-maintaining groups are arranged on multiple nodes of the second storage system so as to equalize the resource usage of each node of the second storage system. An information processing system characterized by the following:

7. An information processing system according to claim 6, For each of the multiple consistency-maintaining groups, the number of journal groups and journal volumes is reduced by moving multiple secondary volumes to the same node. An information processing system characterized by the following:

8. The information processing system according to claim 1, When creating the primary and secondary volumes so that they belong to the integrity maintenance group, the secondary volume to be created is created on the node where the secondary volume of the same integrity maintenance group is located. An information processing system characterized by the following:

9. A first storage system having a node that provides a primary volume to a host, A control method for an information processing system comprising: a second storage system having a plurality of nodes that hold secondary volumes which are copies of the primary volume; There are multiple sets of primary volumes and secondary volumes, and a consistency maintenance group is configured to replicate the data in the multiple primary volumes, which is written to the multiple primary volumes in a consistent manner, to the multiple secondary volumes while ensuring consistency. The replication is performed by journaling using a first journal volume located on the same node as the primary volume and a second journal volume located on the same node as the secondary volume. Within the consistency-maintaining group, a first journal group including the primary volume and the first journal volume, and a second journal group including the secondary volume and the second journal volume are established, and journal processing is performed while ensuring the consistency of the copies between the first journal group and the second journal group. The second journal group is located at each node where the secondary volume of the integrity-maintaining group is located. When a consistency-maintaining group exists in which multiple second journal groups are distributed across multiple nodes, the system determines whether the secondary volume of the second journal group on the source node can be moved to a destination node where other second journal groups are located. If it determines that it can be moved, the system moves the secondary volume of the second journal group on the source node to the second journal group on the source node, and deletes the second journal group and its journal volume on the source node. A control method for an information processing system characterized by the following:

Citation Information

Patent Citations

  • Distributed storage system and rebalancing method

    JP2021197010A

  • Forming a consistency group comprised of volumes maintained by one or more storage controllers

    US20200081806A1

  • Avoiding out-of-space conditions in asynchronous data replication environments

    US20200174919A1

  • Computer system and management method for computer system

    WO2016194096A1

  • Backup system, method therefor, and program

    WO2020158016A1