Information processing system and method for controlling information processing system
By forming consistency maintenance groups with replicated volumes and journal volumes within the same node, the distribution of copy destination volumes is controlled, enhancing performance and reducing journal processing overhead in DR environments.
Patent Information
- Application Number
- JP2024105359
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-16
AI Technical Summary
Conventional disaster recovery (DR) technologies do not effectively manage the distribution of copy destination volumes between an on-premise storage system and a distributed storage system, leading to increased journal groups and reduced processing efficiency and performance.
Implement a consistency maintenance group with primary and secondary volumes replicated within the same node, using first and second journal volumes to ensure consistency, and place secondary volumes based on node resource usage to minimize node distribution.
This approach reduces the dispersion of copy destination volumes, thereby improving performance by minimizing journal processing overhead and maintaining consistency across nodes.
Smart Images

Figure 2026006406000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system and a control method for an information processing system. [Background technology]
[0002] Disaster recovery (DR) technology is known for storing data in a remote site (secondary site) in preparation for data loss at the primary site in the event of a large-scale disaster such as an earthquake or fire. As storage operations in hybrid cloud environments progress, there are an increasing number of cases where DR environments for data from on-premise storage systems are being built in the cloud.
[0003] Distributed storage systems such as SDS (Software Defined Storage) are also used on the cloud. A DR environment is built on this distributed storage system. A distributed storage system consists of many nodes, and when creating a volume, a node with available capacity is automatically selected to create the volume. Patent Document 1 describes a technology that balances the load on nodes by taking into account the IO load on each volume and the node specifications. Patent document 1 states, "In a distributed storage system 1, a volume classifier 300 classifies multiple volumes into multiple groups based on the load fluctuation period of each volume. A processor (resource classifier 400) calculates the total load by adding up the loads of multiple volumes on the same node in the group by time, and calculates the group load based on the peak of the total load. A processor (rebalancer 500) of one of the nodes calculates the group load of the destination node when a volume that is a candidate for movement in a rebalancing operation that moves volumes between nodes is moved from the source node to the destination node, and determines the volume to be moved and the destination volume in the rebalancing operation based on the calculated group load of the destination node, and executes the rebalancing." [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-197010 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the conventional technology disclosed in Patent Document 1 does not take into consideration the case where a DR environment is constructed between an on-premise storage and a distributed storage system using DR technology. In a DR environment between an on-premise storage system and a distributed storage system, a DR relationship is established by creating a volume (copy destination volume) in one of the nodes of the distributed storage system for a volume (copy source volume) in the on-premise storage system. In this case, if the copy destination volume is created in a distributed manner across multiple nodes of the distributed storage system, the number of journal groups for managing the update order increases, which reduces processing efficiency and performance.
[0006] Therefore, an object of the present invention is to suppress the distribution of copy destination volumes to reduce the number of nodes, reduce the overhead of journal processing, and improve performance. [Means for solving the problem]
[0007] In order to achieve the above object, a representative information processing system of the present invention comprises a first storage system having a node that provides a primary volume to a host, and a second storage system having a plurality of nodes that hold secondary volumes that are copies of the primary volume, and a consistency maintenance group is formed in which there are a plurality of pairs of the primary volume and the secondary volume, and data in the plurality of primary volumes that is written to the plurality of primary volumes with consistency is replicated to the plurality of secondary volumes with consistency maintained, and the replication is carried out by a first journal volume provided in the same node as the primary volume and a second journal volume provided in the same node as the secondary volume. and wherein within the consistency maintenance group, a first journal group including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group, and the second journal group is placed for each node on which the secondary volume of the consistency maintenance group is placed, and the secondary volume is placed on each node of the second storage system based on the resource usage of each node of the second storage system and the number of nodes on which the secondary volume is placed in the consistency maintenance group. Furthermore, one representative control method for an information processing system of the present invention is a control method for an information processing system comprising a first storage system having a node that provides a primary volume to a host, and a second storage system having a plurality of nodes that hold secondary volumes that are copies of the primary volume, wherein there are a plurality of pairs of the primary volume and the secondary volume, and consistency maintenance groups are formed so that data in the plurality of primary volumes that is written with consistency to the plurality of primary volumes is replicated to the plurality of secondary volumes while ensuring consistency, and the replication is carried out by a first journal volume provided in the same node as the primary volume and a second journal volume provided in the same node as the secondary volume. and a journal processing using a consistency maintenance group, wherein within the consistency maintenance group, a first journal group including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group, and the second journal group is placed for each node on which the secondary volume of the consistency maintenance group is placed, and the secondary volume is placed on each node of the second storage system based on the resource usage of each node of the second storage system and the number of nodes on which the secondary volume is placed in the consistency maintenance group. [Effects of the Invention]
[0008] According to the present invention, it is possible to improve performance by suppressing the dispersion of copy destination volumes. Problems, configurations, and effects other than those described above will become clear from the following description of the embodiment. [Brief explanation of the drawings]
[0009] [Figure 1]1 is a block diagram showing an example of a hardware configuration of an information processing system 1 according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a software configuration. [Figure 3] FIG. 2 is a diagram showing control information stored in a memory 222. [Figure 4] FIG. 4 is a diagram showing an example of volume information 410. [Figure 5] FIG. 4 is a diagram showing an example of CTG information 420. [Figure 6] FIG. 10 is a diagram showing an example of DR management information 430. [Figure 7] FIG. 4 is a diagram illustrating an example of node information 440. [Figure 8] FIG. 10 is a diagram illustrating an example in which a copy destination volume is distributed among a plurality of nodes. [Figure 9] 10 is a flowchart illustrating an example of a processing procedure for selecting a copy destination volume. [Figure 10] FIG. 10 is a diagram showing an example in which copy destination volumes are distributed among a plurality of nodes (part 1); [Figure 11] FIG. 10 is a diagram showing an example in which copy destination volumes are distributed among a plurality of nodes (part 2); [Figure 12] FIG. 10 is a diagram showing an example in which copy destination volumes are distributed among a plurality of nodes (part 3). [Figure 13] 10 is a flowchart illustrating an example of a processing procedure for moving a copy destination volume. [Figure 14] 10 is a flowchart illustrating an example of a processing procedure for determining the timing to execute data migration. [Figure 15] 10 is a flowchart illustrating an example of a processing procedure for another process for selecting a copy destination volume. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the following embodiments are merely examples for explaining the present invention, and the present invention is not limited to these embodiments. All application examples that conform to the concept of the present invention are included in the technical scope of the present invention. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be plural or singular.
[0011] In the following explanation, the information of the present invention will be described using the expression "information," but this information may be expressed as a data structure such as a "table," "list," or "DB (database)," or in other ways. Furthermore, when describing the content of each piece of information, the expressions "identification information," "identifier," "name," "name," and "ID" can be used, and these are interchangeable.
[0012] Furthermore, although the following description may describe processing performed by executing a program, the program is executed by at least one processor (e.g., a CPU) to perform a predetermined process using storage resources (e.g., memory) and / or interface devices (e.g., communication ports) as appropriate, and therefore the processor may be the entity performing the processing. Similarly, the entity performing the processing by executing a program may be a controller, device, system, computer, node, storage system, storage device, server, management computer, client, or host having a processor. The entity performing the processing by executing a program (e.g., a processor) may include a hardware circuit that performs part or all of the processing, and may be modularized. For example, the entity performing the processing by executing a program may include a hardware circuit that performs encryption and decryption or compression and decompression. Various programs may be installed on each computer via a program distribution server or storage media. The processor operates as a functional unit that realizes a predetermined function by operating in accordance with the program. Devices and systems including a processor are devices and systems that include these functional units. Furthermore, a "read / write process" may be referred to as a "read / write process" or an "update process," etc.
[0013] In addition, the same reference numerals are used for common components in each drawing. When describing elements of the same type in each drawing without distinguishing between them, the reference numerals or common numbers in the reference numerals are used, and when describing elements of the same type with distinction between them, the reference numerals of the elements are used or, instead of the reference numerals, IDs assigned to the elements are used. [Example]
[0014] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing system 1 according to a first embodiment of the present invention. As shown in FIG. 1, the information processing system 1 is a disaster recovery system that provides a disaster recovery configuration. The information processing system 1 includes a primary site 100 and a secondary site 200, which are connected to each other via a network 10 (typically an Internet Protocol (IP) network). In this embodiment, a disaster recovery configuration is described in which an on-premise-based storage system (on-premise storage) is used for the primary site 100 and a distributed storage system is used for the secondary site 200. However, the essence of the present invention remains the same as long as at least the secondary site 200 is a distributed storage system. For ease of description, "disaster recovery" may be abbreviated as DR (Disaster Recovery). The distributed storage system may also be a cloud-based storage system (cloud storage) located in the cloud.
[0015] The primary site 100 is a storage system that provides application services to users (customers) under normal conditions, and is a so-called on-premise storage system, specifically comprising a server system 110, a storage controller 120, and a storage device 130.
[0016] The server system 110 has a processor 111, a memory 112, and a network interface (I / F) 113, and is connected to the network 10 via the I / F 113. The storage controller 120 has a memory 122, a front-end network interface (I / F) 124, a back-end storage interface (I / F) 123, and a processor 121 connected thereto. The storage controller 120 is connected to the server system 110 via the I / F 124, and to a storage device 130 via the I / F 123. The storage device 130 is a storage device that physically stores data. The server system 110 and the storage controller 120 each have redundant memories and processors.
[0017] The memory 122 stores information and one or more programs. The processor 121 executes the one or more programs to provide a storage area (here, described as a logical volume (e.g., a volume created from a storage pool that virtualizes the capacity of the storage device 130), but the essence of the invention does not depend on this) to the server system 110 and process I / O (Input / Output) requests such as write requests and read requests from the server system 110. For example, the server system 110 receives a write request or read request that specifies a volume from a higher-level device (host) used by a user (customer), and transmits it to the storage controller 120. The storage controller 120 then reads or writes data from or to the corresponding volume in the storage device 130 in response to the write request or read request.
[0018] The storage controller 120 and the storage device 130 may be configured as a single storage system, such as a high-end storage system using RAID (Redundant Array of Independent (or Inexpensive) Disks) technology or a storage system using flash memory.
[0019] The storage device 130 may be configured as, for example, a distributed storage system. Furthermore, the storage device 130 may be an HCI (Hyper-Converged Infrastructure) storage system, for example, a system having a function as a host system that issues I / O requests (for example, an execution entity of an application that issues I / O requests (for example, a virtual machine or a container)) and a function as a storage system that processes the I / O requests (for example, an execution entity of storage software (for example, a virtual machine or a container)). This is an example of the configuration of the storage device 130, and is not limited to this.
[0020] The secondary site 200 is a disaster recovery site (DR site) that stores data from the primary site 100 in preparation for data loss at the primary site 100 in the event of a large-scale disaster such as an earthquake or fire, and uses the stored data to restore the data and services of the primary site 100 in the event of a failure or other problem at the primary site 100. In the information processing system 1 according to this embodiment, the secondary site 200 is typically a distributed storage system. The secondary site 200 may also be a distributed storage system that belongs to a public cloud. The secondary site 200 may also be a cloud storage system that belongs to a public cloud and serves as the basis for a cloud storage service provided by a cloud vendor. Examples of cloud storage services include Amazon Web Services (AWS) (registered trademark), Azure (registered trademark), and Google Cloud Platform (registered trademark). The cloud storage system used by the secondary site 200 may be a storage system that belongs to another type of cloud (e.g., a private cloud) instead of a public cloud.
[0021] The secondary site 200 is configured as a distributed storage system and is connected to the network 10 . 1, data transfer between the storage device 130 at the primary site 100 and the distributed storage system at the secondary site 200 is performed by direct transfer via the network 10, but this is not limited to this and data may be transferred using a different network path or line. Furthermore, any type of network or line is not essential to the present invention.
[0022] A distributed storage system is configured by connecting multiple storage computers, each of which includes a storage device and a processor, via a network. Each computer is also called a node in the network. Each computer that makes up a distributed storage system is also called a storage node, and each computer that makes up a compute cluster is also called a compute node.
[0023] The distributed storage system according to this embodiment will be described in detail. The distributed storage system is configured by connecting multiple storage nodes 220A to 220C (collectively referred to as nodes 220) to each other via a network 203. The hardware configuration of each storage node is not particularly limited, but for example, node 220A has a processor (Central Processing Unit) 221, memory 222, a network interface 224, a drive interface 223, and a storage device 225. These are connected via an internal network. Node 220A connects to network 203 via network interface 224 and communicates with other storage nodes (nodes 220B to 220C). Depending on the configuration of network 203, the distributed storage system may be configured with nodes 220 located at sufficiently distant geographical locations.
[0024] In addition, in this embodiment, all of the nodes 220A to 220C that make up the distributed storage system are illustrated as storage nodes, but the nodes that make up the distributed storage system are not limited to storage nodes, and may be configured to include some nodes that function as compute nodes.
[0025] An operating system (OS) for managing and controlling the storage nodes that make up a distributed storage system is installed on the storage nodes, and the distributed storage system is configured by running storage software with storage system functions on top of that. A distributed storage system can also be configured by running storage software in the form of a container on the OS. A container is a mechanism for packaging one or more pieces of software and configuration information. A distributed storage system can also be configured by installing a hypervisor on the storage nodes and running the OS and software as a virtual machine (VM).
[0026] The present invention can also be applied to HCI, a system that enables multiple processes to be performed on a single node by running applications, middleware, management software, and containers in addition to storage software on the OS or hypervisor installed on each node.
[0027] A distributed storage system provides a host with a storage pool and logical volumes (also simply called volumes) that are virtualized versions of the capacities of storage devices on multiple storage nodes.
[0028] In another embodiment, the information processing system 1 may include a storage management system. The storage management system may be part of the components of the primary site or secondary site, or may reside in a dedicated management appliance or management terminal. The management appliance or management terminal may be connected to the network 10.
[0029] For example, this storage management system is a computer system (one or more computers) that manages the configuration of the storage areas of storage device 130 and storage device 225, and a user (or administrator) can instruct settings related to storage device 130 and storage device 225. Alternatively, a user can instruct settings related to storage device 130 and storage device 225 via a management terminal. There may be cases where the storage management system is a separate device for the storage system at the primary site and the storage system at the secondary site.
[0030] An administrator of the distributed storage system can perform processes such as creating, deleting, and moving volumes by issuing management commands to the distributed storage via a network. In addition, the distributed storage system can notify an administrator or a management tool of the status of the distributed storage system, such as the usage status of the drives and processors of the distributed storage system, by providing information transmitted by the distributed storage system via the network.
[0031] For example, the storage management system can manage resource configuration information and application information of the primary site 100 and, by issuing instructions to the secondary site 200, create a DR environment in which the secondary site 200 runs the same VMs (VM252) and applications (applications 253) as the primary site 100.
[0032] The resource configuration information and application information of the primary site 100 may be stored in the memory 112 or storage device 130 of the server system 110, or the storage management system may retrieve the information therefrom.
[0033] With the above configuration, data from storage device 130 can be copied and stored in storage device 225 within the distributed storage system, and the copied data can be used to restore the volume within storage device 130 within the distributed storage system, thereby realizing an information processing system 1 with a DR configuration within primary site 100 including storage device 130 and secondary site 200 including the distributed storage system.
[0034] When creating a DR relationship by creating a volume (copy destination volume) in one of the nodes of the secondary site 200, which is a distributed storage system, for a volume (copy source volume) in the storage system of the primary site 100, there are several cases where the copy destination volume is created distributed among multiple nodes of the distributed storage system. The following (1) to (4) are examples of these cases.
[0035] (1) After creating a destination volume for a source volume and assigning it to a CTG (Consistency Time Group), if a destination volume is created for another source volume in the same manner as described above and then added to the CTG, if the destination volume is created on a node different from the node on which the previously created destination volume was created, the destination volumes within the same CTG will be distributed across multiple nodes.
[0036] (2) When creating destination volumes for multiple source volumes belonging to a CTG at once, if the free space on the selected node is less than the total capacity of all volumes, the destination volumes will need to be created on other nodes due to the lack of free space, and the destination volumes will be distributed across multiple nodes.
[0037] (3) When creating destination volumes for multiple source volumes belonging to a CTG at once, the destination volumes are distributed across multiple nodes to prevent a load from being placed on a specific node during the initial volume data full copy process, etc.
[0038] (4) When an existing volume that has already been created in a distributed storage system is used as a copy destination volume, if the node of the existing volume is different, the copy destination volume within the same CTG will be distributed to multiple nodes.
[0039] A CTG is a group of volumes that copies data across multiple volumes while maintaining data integrity. Operations can be divided by host application group, and volumes for the same business can be managed by belonging to the same CTG.
[0040] In a DR environment, when performing copy processing, the write order (update order) from the host is managed using a journal. A journal group is a collection of volumes consisting of one or more data volumes and journal volumes, and journal processing is performed on a journal group (JNLG) basis. First, a JNLG must be created by registering a volume.
[0041] Since consistency is guaranteed for volumes belonging to the same journal volume, a JNLG belongs to one CTG. Since each node in a distributed storage system operates independently, a JNL is provided for each node. When the nodes of the copy destination volumes within a CTG are distributed, the number of JNLGs within the CTG increases because a JNLG is set up for each node. If the number of nodes in the distributed storage system of the copy destination is N, the number of JNLGs will increase by N times in the worst case scenario if the copy destination volumes are distributed across N nodes. As the number of JNLGs increases, JNL processing that was previously performed collectively within a JNLG is now performed for each JNLG, resulting in a decrease in processing efficiency and performance.
[0042] Therefore, in the information processing system 1, when the copy destination volume is distributed to a plurality of nodes in the distributed storage system, control is executed to reduce the number of nodes to which the volume is distributed, thereby reducing the overhead of the JNL processing. Details of this will be described later.
[0043] An example of the operation image of the information processing system 1 is shown below. Within the primary site 100, one or more virtual machines (VMs) are generated on a server (server group) consisting of one or more server systems 110. An app designated by a user is executed on each VM. An app designated by a user is an application that provides a service to the user, and an app corresponding to the service is designated by specifying the service that the user uses. An operating system (OS) for managing and controlling the storage devices (in a distributed storage system, an OS for managing and controlling the storage nodes) is installed, and storage software with the functions of the storage system runs on top of it. Data is stored in the storage area of the storage device 130 via a storage pool that virtualizes the capacity of the storage device.
[0044] In the node 220, a host OS for controlling the hardware runs, and a hypervisor for running one or more guest OSs as VMs runs on top of the host OS.
[0045] A container runtime for running one or more containers runs on each guest OS, and storage software and computing software run on top of that. In the above software stack, if the hypervisor includes a function for controlling hardware, the host OS can be omitted. Also, if there is no need to run each piece of software on a VM, the hypervisor and guest OS can be omitted, and in this case, the container runtime can be run on the host OS. Also, if the storage software and computing software do not run as containers, the container runtime can be omitted, and in this case, the storage software and computing software can be run directly on the guest OS or host OS.
[0046] FIG. 2 is a schematic diagram showing the relationship between software (or control programs, modules, functions, and components) in the storage system at the primary site 100 and the distributed storage system at the secondary site 200. The software includes storage control 310, DR control 320, data migration control 330, and monitor 340. Each piece of software can communicate with each other and send and receive information. At the primary site 100, each software module runs on the storage controller 120 of the storage system. In a distributed storage system, the software may run on the same node as the storage node 220, or on a separate node or separate hardware component such as a device, terminal, or circuit, as long as the distributed storage system is accessible via the network 10. Furthermore, all software does not need to be implemented on the same storage controller 120 and node 220. Each piece of software may be executed in any manner, such as a process or a container.
[0047] Storage control 310 controls storage device 130 and storage device 225. For example, when an I / O request is issued to storage device 130 or storage device 225, the location where the data specified by the I / O request is stored in storage device 130 is accessed and the data is provided to the host. In the case of a distributed storage device at secondary site 200, when a host issues an I / O request to one of the storage nodes, the distributed storage system provides the host with access to the data by forwarding the I / O request to the storage node that holds the data specified by the I / O request.
[0048] The DR control 320 has functions such as instructing the creation of a DR environment between the primary site 100 and the secondary site 200, creating corresponding primary and secondary volumes in the DR environment for VMs and applications (determining source and destination volumes, and managing CTGs (Consistency Time Groups) and journal groups (JNLGs)), controlling the transfer of data stored in the storage device 130 of the primary site 100 to the secondary site 200, and restoring volumes in the secondary site 200 from data copied to the secondary site 200.
[0049] The data migration control 330 has a function for scheduling the migration of the copy destination volume to another node and determining the data migration.
[0050] The monitor 340 monitors the load of each hardware element in the distributed storage system. The data migration control 330 refers to the monitoring information from the monitor 340, determines the migration destination of the volume, and migrates the volume.
[0051] FIG. 3 shows the control information stored in the memory 122 of the storage controller and the memory 222 of the node 220. Specifically, the memory stores volume information 410, CTG information 420, DR management information 430, node information 440, and storage information 460. This information is accessed as appropriate when processing is executed by each piece of software shown in Fig. 2, and reference, reading, generating, writing, updating, or the like is performed. This information may be stored in the management device.
[0052] 4 is a diagram showing an example of volume information 410. The volume information 410 is information about volumes stored in each of the storage devices 130 and 225. The volume information 410 includes a volume ID 411, DR availability 412, a CTG ID 413, a JNLG ID 414, a node ID 415, a capacity 416, and a usage amount 417.
[0053] Volume ID 411 stores the identifier of the volume provided to the host. DR presence / absence 412 indicates whether or not volume 411 has a copy destination volume in a DR configuration. In the case of volume information for a distributed storage system, it indicates whether or not there is a copy source volume. CTGID 413 indicates the identifier of the CTG (Consistency Time Group) to which volume 411 belongs, and JNLGID 414 indicates the identifier of the JNLG to which volume 411 belongs. Node ID 415 is provided only in the case of volume information for a distributed storage system, and indicates the identifier of the node to which volume 411 is provided. Capacity 416 indicates the storage capacity of volume 411, and usage 417 indicates the amount of said storage capacity that has been used for storing data, etc.
[0054] 5 is a diagram showing an example of the CTG information 420. The CTG information 420 shown in FIG.
[0055] CTGID 421 is the identifier of a CTG configured in the DR configuration of the information processing system 1. JNLG count 422 is the number of JNLGs belonging to CTGID 421, and is the total number of JNLGs listed in JNLGID 423. JNLGID 423 indicates the identifier of a JNLG belonging to CTGID 421. Volume count 424 is the number of volumes belonging to JNLG volume ID 425, and is the total number of volumes listed in volume ID.
[0056] A CTG is a group that ensures the order of updates across multiple volumes and is configured in the storage systems of the primary and secondary sites in a DR configuration. By specifying a CTG, volumes can be operated collectively as a CTG. For example, multiple volumes used for the same business can be configured to belong to the same CTG.
[0057] A CTG may be created with some of the nodes in a distributed storage system. For example, if a distributed storage system consists of three nodes, node 1 to node 3, CTG1 may consist of node 1 and node 2, and CTG2 may consist of node 2 and node 3.
[0058] Fig. 6 is a diagram showing an example of the DR management information 430. The DR management information 430 includes information indicating corresponding DR relationships in the DR environment within the information processing system 1. The DR management information 430 shown in Fig. 6 is configured to have the following items: primary 431, primary volume ID 432, secondary 433, and secondary volume ID 434.
[0059] The primary 431 stores an identifier, such as a serial number, that indicates the storage system that constitutes the DR primary site 100. The primary volume ID 432 stores an identifier for the volume at the primary site. The secondary 433 stores an identifier, such as a serial number, that indicates the storage system that constitutes the DR secondary site 200 in a form that corresponds to the primary 431. The secondary volume ID 434 stores an identifier for the volume at the secondary site.
[0060] In a DR configuration, the volume at the primary site is called the source volume, and the volume at the secondary site is called the destination volume. The source volume is sometimes referred to as PVOL, and the destination volume is sometimes referred to as SVOL.
[0061] FIG. 7 is a diagram showing an example of node information 440. The node information 440 is data indicating the configuration of each node 220 that constitutes the distributed storage system. In the case of FIG. 7, the node information 440 has data items of node ID 441, volume 442, capacity 443, usage 444, availability 445, and IO frequency 446. The capacity 443 indicates the total value of the drive capacity installed in the target node 441. Furthermore, the usage 444 indicates the total capacity allocated to the volumes provided to the host from the storage devices 225, which are physical drives within the node 441. The availability 445 indicates the usage rate of the processor within the node. The IO 446 frequency indicates the number of read / write requests to the node per unit time as a percentage.
[0062] Figure 8 shows an example in which DR destination volumes within a CTG are distributed across multiple nodes. A storage system 510 constructed at an on-premise primary site 100 has a volume 513. Here, there are two volumes 513, 513A and 513B. In this example, the distributed storage system is configured from two nodes 520A and 520B. The two volumes 513A and 513B are used as DR source volumes and are copied to destination volumes 523A and 523B, respectively. Volumes 513A and 513B belong to the same CTG 511. Destination volumes 523A and 523B also belong to the same CTG 521.
[0063] Volume 513A replicates data to volume 523A in node 520A, and volume 513B replicates data to volume 523B in node 520B. The copy destination volume for PVOL1 is SVOL1, and the copy destination volume for PVOL2 is SVOL2. Copying from the on-premise storage system to the distributed storage system in a DR environment is performed using, for example, conventional asynchronous remote copy technology.
[0064] A journal group (JNLG) is used to manage differential copying of data between a source volume and a destination volume. A JNLG is a collection of volumes consisting of one or more data volumes and journal volumes. A data volume is a source volume or a destination volume, and a journal volume is a volume that stores write data from a host and information representing the history of updates to that data. When a primary site storage system receives a write request, it writes the data to the data volume and the journal data to the journal volume, and returns a response to the server system. The secondary site destination storage system reads the journal data from the journal volume of the source storage system asynchronously with the write request and stores it in its own journal volume. The destination storage system then restores the copied data to the destination data volume based on the stored journal data. In this way, management is performed so that the write order can be restored. The journal processing described above is performed in units of JNLG.
[0065] In Figure 8, volume 513 or 523 and a journal volume belong to JNLG 512 or 522. Since node 520 in the distributed storage system runs on its own OS, JNLG is also set for each node. Since storage system 510 sets JNLG corresponding to the copy destination node, JNLGs for the number of distributed copy destination nodes are also provided in storage system 510. Volume 513A is copied to volume 523A in node 520A, so volume 513A belongs to JNLG 512A. Volume 513B is copied to volume 523B in node 520B, so volume 513B belongs to JNLG 512B. If the copy destination volumes for all configured CTGs are distributed to all nodes in the distributed storage system, the number of JNLGs required for one CTG will be equal to the number of nodes. In this way, when the copy destination volumes within the same CTG are distributed to multiple nodes, the problem can be solved by reducing the number of nodes to which they are distributed. In other words, by reducing the number of JNLGs within the CTG, it is possible to prevent a deterioration in processing efficiency and improve performance.
[0066] The procedure for creating the configuration of FIG. 8 is as follows, and is executed by the DR control 320. For example, a CTG1 is created for the DR configuration of the DR management information 430 in FIG.
[0067] Here, if the DR destinations are nodes 520A and 520B, journal volumes JVOL3 (524A) and JVOL4 (524B) are created on nodes 520A and 520B. JVOL3 is registered in JNLG1 on node 520A, and JVOL4 is registered in JNLG2 on node 520B. Journal volumes JVOL1 (514A) and JVOL2 (514B) are created on the DR source storage system 510, and registered in JNLG1 and JNLG2, respectively. JNLG1 and JNGL2 are registered in CTG1.
[0068] If there are other CTG2s, process them in the same way. For example, create a journal volume JVOL7 on node 520A, and create a journal volume JVOL8 on node 520B. JVOL7 is registered in JNLG3, and JVOL8 is registered in JNLG4. Create journal volumes JVOL5 and JVOL6 on the DR source storage system 510, and register them in JNLG1 and JNLG2, respectively. Register JNLG3 and JNGL4 in CTG2.
[0069] FIG. 9 shows a flowchart of the process for selecting a copy destination volume. First, if PVOL1 (513A) is a DR target volume, the copy source storage system 510 registers PVOL1 in CTG1 (S1010). Then, the storage system 510 requests the distributed storage system to create SVOL1, which is a copy destination volume for DR of PVOL1 (S1020).
[0070] The compute node of the distributed storage system receives the request, determines in which node SVOL1 should be created (S1030), and sends a request to create SVOL12 to the storage node 522 that is the result of the determination (S1040). The node to send the request to is selected by selecting a node with a large amount of free storage capacity or a node with a low processor utilization rate. A node that has already been created and has an unused volume may be selected, and the unused volume may be set as SVOL1.
[0071] Node 1 (520A) that received the request creates SVOL1 (S1050) and registers it as the copy destination of PVOL1. SVOL1 is registered in JNLG1 of CTG1 within the node (S1060). Node 1 reports to storage system 510 via the compute node that SVOL1 has been created in JNLG1 (S1070).
[0072] The storage system 510 receives the JNLGID in which SVOL1 was created, and registers PVOL in the received JNLG1 among the JNLGs in CTG1 (S1080).
[0073] Similarly, for PVOL2 (513B), the DR target volume, PVOL2 is registered in CTG1. A request is made to the distributed storage system to create an SVOL, which is a copy destination volume for PVOL2. In this example, node 2 (520B) receives the request, creates an SVOL, and registers it as the copy destination for PVOL2. The SVOL is registered in JNLG12 of CTG1 within the node. The storage system 510 receives the JNLGID and registers PVOL in the received JNLG2 from among the JNLGs within CTG1.
[0074] Here, the JNLG to which the PVOL is to be registered is identified by receiving the JNLGID from the DR copy destination, but other methods are also possible as long as the purpose is to register the PVOL and SVOL in the same JNLG. For example, when a journal volume JVOL1 is created and registered in CTG1 of the storage system 510 and a copy destination volume is created, JNLG2 may be newly generated at the timing when an SVOL is created in a node different from the previous time.
[0075] An explanation will be given using the examples of Figures 10 to 12. In a DR configuration of a distributed storage system consisting of a storage system 510 and a node 520, two CTGs, CTG 511 and CTG 515, are managed in the storage system 510, and CTG 521 and CTG 525, are managed in the node 520 of the distributed storage system. CTG 511 and CTG 521 are managed as the same CTG. CTG 515 and CTG 525 are also the same CTG.
[0076] Source volume 513 belongs to CTG 511, and volume 513A (PVOL1 to PVOL3 in the figure) is copied to destination volume 523A (SVOL1 to SVOL3 in the figure) of node 520A, and volume 513B (PVOL4 to PVOL10 in the figure) is copied to destination volume 523B (SVOL4 to SVOL10) of node 520B. Since the destination volumes are distributed across two nodes and a JNLG is installed for each node, source volume 513A belongs to JNLG 512A, destination volume 523A belongs to JNLG 522A, source volume 513B belongs to JNLG 512B, and destination volume 523B belongs to JNLG 522B. JNLG 512A and JNLG 522A are managed as the same JNLG. JNLG 512B and JNLG 522B are also managed as the same JNLG. That is, one CTG511 contains multiple JNLGs (JNLG512A and JNLG512A).
[0077] The same is true for CTG 515. Source volume 517 belongs to CTG 515, and volume 517A (PVOL21 to PVOL24 in the figure) is copied to destination volume 527A (SVOL21 to SVOL24 in the figure) of node 520A, and volume 517B (PVOL25 to PVOL29 in the figure) is copied to destination volume 527B (SVOL25 to SVOL29) of node 520B. Since the destination volumes are distributed across two nodes and a JNLG is installed for each node, source volume 517A belongs to JNLG 516A, destination volume 527A belongs to JNLG 526A, source volume 517B belongs to JNLG 516B, and destination volume 527B belongs to JNLG 526B. JNLG 516A and JNLG 526A are managed as the same JNLG. JNLG516B and JNLG526B are also managed as the same JNLG. In other words, multiple JNLGs (JNLG512A and JNLG512B) exist in one CTG 511.
[0078] A node is selected to create a destination volume, taking into consideration the node's free space (in storage space) and processor utilization rate at the time the destination volume was created, but the node load changes over time. Hardware load monitoring information is referenced, and the destination volume is moved and relocated between nodes to balance the JNL processing load. Relocating volumes between nodes for balancing purposes can be an optimal solution, or it can be one that minimizes the number of nodes across which the destination volume is distributed. Reducing the number of distributed nodes improves the efficiency of JNL processing and improves system performance. System performance can be improved by "avoiding dispersion when placing JNLGs that belong to the same CTG and placing them on as few nodes as possible" and "changing the placement of multiple CTGs to level out the processing load."
[0079] 11 and 12 show an example of a case where a copy destination volume is moved from the configuration of FIG. 10. For CTG 521, all copy destination volumes 523A in node 520A are moved to node 520B. When volume 523A is moved to node 520B, it belongs to JNLG 522B, and JNLG 522A becomes unnecessary. With the movement of the copy destination volume, JNLG 512A of CTG 511 in storage system 510 is integrated with JNLG 512B, and JNLG 512A becomes unnecessary. The unnecessary JNLG is deleted. In this way, the number of JNLGs in CTG 511 and CTG 521 can be reduced. Such volume movement is possible only if there is sufficient free capacity and a low processor load on the node 520B side. For CTG515 and CTG525, the number of nodes in the copy destination volume can be reduced in the same way.
[0080] It is also possible to create a volume migration plan that simultaneously migrates volumes between CTG511 and CTG521, and between CTG515 and CTG525, as shown in Figure 12. When migrating volumes that belong to multiple CTGs at the same time, it is possible to migrate volumes even when there is little free space on a node, depending on the processing method, such as moving one volume at a time, and the number of JNLGs that can be aggregated on a node increases, improving performance. The volume arrangement in FIG. 12 may be achieved by going from FIG. 10 to FIG. 11, or from FIG. 10 to FIG. 12.
[0081] 13 shows a flowchart of the process for reducing the number of nodes to which copy destination volumes in a CTG are distributed when they are distributed across multiple nodes, i.e., for reducing the number of JNLGs. This process is executed by the data migration control 330. The data migration control 330 may be a functional unit implemented by a program executed by one of the CPUs shown in FIG. 1, or may be a function executed by an appropriately provided management system, management device, or management terminal.
[0082] First, the data migration control 330 selects a CTG to be migrated (S1410). Specifically, the data migration control 330 searches for a CTG to which the nodes of the copy destination volume are distributed. As an example, the data migration control 330 references the CTG information 420 and selects a CTG ID 421 with a large number of JNLGs 422.
[0083] Next, the data migration control 330 selects a JNLG to be moved from among the JNLGs belonging to the selected CTG (S1420). Specifically, the data migration control 330 selects a JNLG ID 423 to be moved to another node from among the JNLG IDs 423 of the selected CTG ID 421. For example, the JNLG ID 423 to be moved is selected to be a JNLG ID 423 with a small number of volumes 424 belonging to the JNLG.
[0084] Next, the data migration control 330 selects a destination JNLG for the JNLG to be moved (S1430). Specifically, the data migration control 330 selects, as the destination JNLG, a JNLG ID 423 with a large number of volumes 424 belonging to the JNLG from among the JNLG IDs 423 of the selected CTG ID 421.
[0085] The data migration control 330 judges whether it is possible to migrate all copy destination volumes belonging to the selected migration source JNLG to the nodes of the selected migration destination JNLG. First, the data migration control 330 refers to the CTG information 420 and calculates the total capacity of all copy destination volumes belonging to the volume ID 425 of the migration source JNLG ID 423. It refers to the volume information 410, searches for the volume ID 411 corresponding to the volume ID 425, and acquires the capacity 416. It compares this with the free capacity of the storage area of the node of the migration destination JNLG, and if the total capacity is less than the free capacity, it determines that the target JNLG can be migrated (S1440). If the total capacity is greater than the free capacity, it determines that migration is not possible, and returns to S1410 to select another CTG. Figure 13 shows the process returning to S1410, but as an alternative processing method, it is also possible to return to S1430 and reselect another JNLG as the migration destination JNLG, or to return to S1420 and reselect another JNLG as the migration target JNLG.
[0086] After determining that the target JNLG can be migrated, the data migration control 330 next refers to the operating status of the hardware of the migration destination node and the operating rate of the processor to determine whether the operating rate of the migration destination falls within an acceptable range (S1450). For example, if the operating rate is previously set at an upper limit of 90%, a simulation is performed to determine whether the operating rate will fall within 90% due to the increase in load caused by the volume migration. If it is determined that it will fall within that range, the migration is permitted. If it will not fall within that range, it is determined that the migration is not possible, and the process returns to S1410 and another CTG is selected. In Figure 13, the process returns to S1410, but as an alternative processing method, it is also possible to return to S1430 and reselect another JNLG as the migration destination JNLG, or to return to S1420 and reselect another JNLG as the migration target JNLG. Here, the above-mentioned 90% is an example and can be set to a variable value. The operating rate may be predicted from the access frequency of the node.
[0087] If migration is permitted, the data migration control 330 sequentially migrates the copy destination volumes of the migration source JNLG to the migration destination JNLG (S1460). As another example of the above, if there are no more volumes belonging to a JNLG, the JVOL may be deleted and the JNLG may be deleted. There is also another embodiment. In a DR configuration, a destination volume that creates a replica of a source volume is created when the destination storage system receives a request to create the destination volume, creates the volume, and then creates a relationship between the source volume and the destination volume. Due to this processing procedure, when a distributed storage system is combined with a DR configuration, the distributed storage system creates the destination volume in a node with available capacity. By specifying a node from the source storage system to create the destination volume, it is possible to create a destination volume that is not distributed across multiple nodes.
[0088] When the copy destination volumes in the CTG in FIG. 13 are distributed to a plurality of nodes, the number of nodes to be distributed is reduced, that is, the timing for executing the process of reducing the number of JNLGs is determined by the data migration control 330. The monitor 340 performs a process of monitoring the load of each hardware element of the distributed storage system. The timing for executing the monitoring process is a time period. The time period can be set arbitrarily, and may or may not be a fixed period. The information obtained from the monitoring process is stored in the availability 445 and IO frequency 446 of the node information 440. The monitoring process is initiated by the management device or the storage control 310 .
[0089] The data migration control 330 refers to the load of the hardware elements as determined by the monitor 340, determines the migration destination of the volume, and migrates the volume. In another example, predicted resource utilization may be calculated by monitoring the load on hardware elements, and the calculated predicted resource utilization may be used to determine where to move a volume.
[0090] The timing for executing data migration is determined by setting thresholds for the loads of the hardware elements, such as the utilization rate 445 and IO frequency 446, and executing data migration when it is determined that the utilization rate is lower than the threshold, i.e., the resource utilization rate is low. Furthermore, data migration is executed when it is determined that the utilization rate is higher than the threshold, i.e., the resource utilization rate is biased towards the node. As for the timing of executing other data migration, a threshold is set for the free capacity of the node, and data migration is executed when the free capacity is smaller than the threshold.
[0091] For example, the process shown in the flowchart of FIG. 14 is performed at a time interval that can be specified. The data migration control 330 determines whether any of the nodes 220 in the distributed storage system have a resource usage rate exceeding a preset threshold (S1510). If there are any nodes exceeding the threshold, data migration processing (S1540) is performed. If there are no nodes exceeding the threshold, the data migration control 330 determines whether there are any nodes to which JNLG is distributed (S1520). If there are any nodes to which JNLG is distributed, the data migration control 330 determines whether the distributed storage system is idle (S1530). In other words, it determines whether the hardware resources can operate even if the load of the inter-node volume migration processing required for the data migration processing is placed on the hardware resources. If they are idle, data migration processing (S1540) is performed. If there are no nodes to which JNLG is distributed at S1520, or if there are nodes to which JNLG is distributed but the operating status of the nodes is not idle at S1520, the data migration processing is not initiated and the processing is terminated.
[0092] In this way, in the first embodiment, when the copy destination volume is distributed to a plurality of nodes in the distributed storage system, the number of nodes to which it is distributed is reduced as much as possible, thereby reducing the overhead of the JNL processing and improving performance. [Example]
[0093] As another embodiment, a method for specifying the node that creates the SVOL from the PVOL side is shown in the flowchart of FIG. First, the copy source storage system 510 registers PVOL2 in CTG1 (S1110). Next, the storage system 510 determines whether an SVOL has already been created in the same CTG (S1120) to determine in which node of the distributed storage system SVOL2, the copy destination volume of PVOL2, should be created. If an SVOL has already been created in the same CTG1, the storage system 510 acquires the JNLGID in which that SVOL was created (S1130). For example, if PVOL1 (513A) is the DR target volume, PVOL1 is registered in CTG1, and SVOL1 has already been created in JNLG1 of node 1, PVOL2 belongs to CTG1, so it is determined that an SVOL already exists.
[0094] If the SVOL has already been created, the storage system 510 hands over the acquired JNLGID and requests the distributed storage system to create SVOL2, which is the DR copy destination volume of PVOL2 (S1140). If the SVOL has not been created, it performs the processes of S1020 to S1080 and ends the process.
[0095] The compute node of the distributed storage system receives the request and sends the request to the storage node 522 corresponding to the JNLGID (S1150).
[0096] Node 1 (520A) that received the request creates SVOL2 (S1160) and registers it as the copy destination of PVOL2. SVOL2 is registered in JNLG1 within the node (S1170). Node 1 reports to storage system 510 via the compute node that SVOL2 has been created in JNLG1 (S1180).
[0097] The storage system 510 receives the JNLGID in which SVOL2 was created, and registers PVOL2 in JNLG1 (S1190). Here, the JNLGID information is shared between the storage system and the distributed storage system to determine the node where the SVOL will be created so that it is not distributed to nodes, but this is just one example, and other information may be used as long as it indicates on which node PVOL1 created the SVOL. Alternatively, the information may be information that allows the compute node of the distributed storage node to search for the node that created the SVOL corresponding to PVOL1 in S1150. In this case, the compute node determines the node where the SVOL will be created based on the information received from the storage system.
[0098] 15 illustrates a case where PVOL1 belonging to CTG1 has already created SVOL1, and SVOL2 for PVOL2 also belonging to CTG1 is created on the same node. However, as another example, when SVOLs for PVOL1 and PVOL2 belonging to CTG1 are created at the same time, the SVOL can be created by selecting a node that can secure free capacity for creating two SVOLs. In this case, the storage system 510 issues a request to a compute node to create SVOLs for multiple PVOLs on the same node, and the compute node determines whether free capacity can be secured and then decides on the node where the SVOL will be created, or the primary site storage system 510 determines whether free capacity can be secured and then decides on the node. Information for determining whether free capacity can be secured on the primary site storage system 510 side can be obtained from the distributed storage system side or from a management device, etc.
[0099] In another embodiment, if the load on a particular node in a distributed storage system is high, for example, if the free space is less than a preset threshold or the hardware utilization rate is higher than a preset threshold, it is possible to move volumes from that particular node to other nodes to improve the performance of the entire system. By moving volumes, the copy destination volumes of a certain CTG may be distributed across multiple nodes, and the number of JNLGs may increase.
[0100] In this way, in the second embodiment, the copy destination volume in the DR environment can be determined and the copy operation can be executed without the user being aware of the nodes of the distributed storage system, which improves the user's flexibility in DR operation.
[0101] As described above, the information processing system 1 disclosed in the embodiment is an information processing system comprising a first storage system (primary site 100) having a node that provides a primary volume to a host, and a second storage system (secondary site 200) having a plurality of nodes that hold secondary volumes that are copies of the primary volume, in which there are a plurality of pairs of the primary volume and the secondary volume, and a consistency maintenance group (CTG) is formed in which data in the plurality of primary volumes that is written with consistency to the plurality of primary volumes is replicated to the plurality of secondary volumes while ensuring consistency, and the replication is performed by a first journal volume provided in the same node as the primary volume and a second journal volume provided in the same node as the secondary volume. and a journal volume, wherein within the consistency maintenance group, a first journal group (JNLG) including the primary volume and the first journal volume and a second journal group including the secondary volume and the second journal volume are set, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group (JNLG), and the second journal group is placed for each node on which the secondary volume of the consistency maintenance group is placed, and the secondary volume is placed on each node of the second storage system based on the resource usage status of each node of the second storage system and the number of nodes on which the secondary volume is placed in the consistency maintenance group. Specifically, the plurality of secondary volumes are arranged on the plurality of nodes so as to reduce the number of second journal groups. This configuration and operation makes it possible to suppress dispersion of copy destination volumes and improve performance.
[0102] Furthermore, the first journal group is provided in correspondence with the second journal group, so that one or more first journal groups are arranged in one node. In addition, the second journal volume is arranged for each node on which the secondary volume of the consistency maintenance group is arranged, and is used for journal processing to one or more secondary volumes arranged on that node, and the first journal volume is arranged corresponding to the second journal volume, so that one or more first journal volumes are arranged on one node. According to this configuration and operation, the number of journal groups is determined according to the number of nodes on which secondary volumes are located, so the number of journal groups can be reduced by relocating secondary volumes, thereby improving performance.
[0103] Specifically, the secondary volume is moved to a node where another secondary volume of the same consistency maintenance group is located, thereby reducing the number of journal groups and journal volumes. Furthermore, by moving the secondary volume between nodes, the number of the second journal groups and the second journal volumes is reduced, and the reduction in the second journal groups and the second journal volumes reduces the corresponding number of the first journal groups and the first journal volumes. This configuration and operation allows performance to be improved by rearranging secondary volumes within a consistency-maintaining group.
[0104] The information processing system 1 also has a plurality of consistency maintenance groups, and a plurality of secondary volumes of the plurality of consistency maintenance groups are arranged on a plurality of nodes of the second storage system based on the resources of each node of the second storage system. Then, for each of the plurality of consistency maintenance groups, the plurality of secondary volumes are moved so as to be aggregated to the same node, thereby reducing the number of journal groups and journal volumes. Furthermore, based on the resources of the node, it is determined whether the secondary volume can be moved to a node where another secondary volume of the same consistency maintenance group is located, and if it is determined that this is possible, the secondary volume is moved. According to this configuration and operation, the destination node can be appropriately selected taking into consideration the state of each node.
[0105] Furthermore, when the information processing system 1 creates the primary volume and secondary volume so that they belong to the consistency maintenance group, the secondary volume to be created is created in the node where the secondary volume of the same consistency maintenance group is located. This configuration and operation can prevent the journal group from becoming dispersed.
[0106] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, not only can the configurations be deleted, but also replacements and additions of configurations are possible. [Explanation of symbols]
[0107] 1. Information Processing Systems 10 Network 100 primary sites 110 Server System 111,121,221 processors 112,122,222 memory 113,124,224 Network Interface (I / F) 120 Storage Controller 123 Storage Interface (I / F) 130 Storage Devices 200 secondary sites 220 Storage Devices 310 Storage Control 320 DR control 330 Data Migration Control 340 monitor
Claims
1. a first storage system having a node that provides a primary volume to a host; a second storage system having a plurality of nodes that hold a secondary volume that is a copy of the primary volume, a consistency maintenance group is configured in which there are a plurality of pairs of the primary volume and the secondary volume, and data in the plurality of primary volumes that is written to the plurality of primary volumes with consistency is replicated to the plurality of secondary volumes with consistency maintained; the replication is performed by journal processing using a first journal volume provided in the same node as the primary volume and a second journal volume provided in the same node as the secondary volume; within the consistency maintenance group, a first journal group including the primary volume and the first journal volume, and a second journal group including the secondary volume and the second journal volume are set, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group; the second journal group is arranged for each node on which the secondary volume of the consistency maintenance group is arranged, The secondary volume is allocated to each node of the second storage system based on the resource usage status of each node of the second storage system and the number of nodes in the consistency maintenance group on which the secondary volume is allocated. An information processing system comprising:
2. 2. The information processing system according to claim 1, The plurality of secondary volumes are arranged on the plurality of nodes so that the number of the second journal groups is reduced. An information processing system comprising:
3. 2. The information processing system according to claim 1, The first journal group is provided corresponding to the second journal group, so that one or more first journal groups are arranged in one node. An information processing system comprising:
4. 4. The information processing system according to claim 3, the second journal volume is arranged for each node on which the secondary volume of the consistency maintenance group is arranged, and is used for journal processing of one or more secondary volumes arranged in the node; The first journal volume is provided in correspondence with the second journal volume, so that one or more first journal volumes are arranged in one node. An information processing system comprising:
5. 5. The information processing system according to claim 4, The secondary volume is moved to a node where another secondary volume of the same consistency maintenance group is located, thereby reducing the number of journal groups and journal volumes. An information processing system comprising:
6. 6. The information processing system according to claim 5, the secondary volume is moved between nodes to reduce the number of the second journal groups and the second journal volumes; By reducing the number of the second journal groups and the second journal volumes, the number of the corresponding first journal groups and the first journal volumes is reduced. An information processing system comprising:
7. 4. The information processing system according to claim 3, It has multiple consistency groups, A plurality of secondary volumes of the plurality of consistency-maintaining groups are arranged on a plurality of nodes of the second storage system based on the resources of each node of the second storage system. An information processing system comprising:
8. 8. The information processing system according to claim 7, For each of the plurality of consistency maintenance groups, the plurality of secondary volumes are moved to be aggregated to the same node, thereby reducing the number of the journal groups and the journal volumes. An information processing system comprising:
9. 8. The information processing system according to claim 7, Based on the resources of the node, it is determined whether the secondary volume can be moved to a node where another secondary volume of the same consistency maintenance group is located, and if it is determined that the secondary volume can be moved, the secondary volume is moved. An information processing system comprising:
10. 2. The information processing system according to claim 1, When the primary volume and secondary volume are created so as to belong to the consistency maintenance group, the secondary volume is created in a node where a secondary volume of the same consistency maintenance group is located. An information processing system comprising:
11. a first storage system having a node that provides a primary volume to a host; a second storage system having a plurality of nodes that hold a secondary volume that is a copy of the primary volume, a consistency maintenance group is configured in which there are a plurality of pairs of the primary volume and the secondary volume, and data in the plurality of primary volumes that is written to the plurality of primary volumes with consistency is replicated to the plurality of secondary volumes with consistency maintained; the replication is performed by journal processing using a first journal volume provided in the same node as the primary volume and a second journal volume provided in the same node as the secondary volume; within the consistency maintenance group, a first journal group including the primary volume and the first journal volume, and a second journal group including the secondary volume and the second journal volume are set, and the journal processing is performed while ensuring the consistency of the replication between the first journal group and the second journal group; the second journal group is arranged for each node on which the secondary volume of the consistency maintenance group is arranged, The secondary volume is allocated to each node of the second storage system based on the resource usage status of each node of the second storage system and the number of nodes in the consistency maintenance group on which the secondary volume is allocated.
2. A method for controlling an information processing system comprising:
Citation Information
Patent Citations
Distributed storage system and rebalancing method
JP2021197010A
Forming a consistency group comprised of volumes maintained by one or more storage controllers
US20200081806A1
Avoiding out-of-space conditions in asynchronous data replication environments
US20200174919A1
Computer system and management method for computer system
WO2016194096A1
Backup system, method therefor, and program
WO2020158016A1