Storage system performing remote copy and path selection method for remote copy
By dynamically selecting paths based on initiator port availability and performance in remote copy operations, the system addresses performance degradation issues and enhances fault detection, leading to improved remote copy performance.
Patent Information
- Application Number
- JP2023193609
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2043-11-14
AI Technical Summary
In remote copy operations between storage systems, the performance can degrade due to the selection of paths in a round-robin manner, leading to uneven load distribution and potential delays in detecting port failures.
The system selects paths based on the availability and performance of initiator ports, allowing nodes to transmit commands via paths connected to their initiator ports as long as no abnormalities are detected, thereby avoiding command transfer between nodes.
This approach helps suppress performance degradation in remote copy operations by ensuring efficient path selection and reducing the need for command transfer between nodes, thus enhancing overall system performance and fault detection capabilities.
Smart Images

Figure 2025080456000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to remote copy between storage systems.
Background Art
[0002] For example, Patent Document 1 discloses a technique related to remote copy.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In remote copy from a PVOL (primary VOL) in a primary storage system (storage system at the primary site) to an SVOL (secondary VOL) in a secondary storage system (storage system at the secondary site), data passes through a selected path among a plurality of paths between the PVOL and the SVOL. Remote copy includes push-type remote copy (remote copy performed in response to a write from the primary storage system to the secondary storage system) and pull-type remote copy (remote copy performed in response to a read from the secondary storage system to the primary storage system). In push-type remote copy, the primary storage system can select a path, and in pull-type remote copy, the secondary storage system can select a path. Hereinafter, the storage system that selects a path is referred to as the "first storage system", and the storage system that communicates with the first storage system for remote copy is referred to as the "second storage system".
[0005] It may be considered to select paths in a round-robin manner. Since each path is selected evenly, load distribution is expected, and / or when a port failure occurs, it is expected to detect the port failure quickly.
[0006] By the way, as a first storage system, a storage system composed of a plurality of storage nodes based on SDS (Software Defined Storage) can be adopted. Those storage nodes are, for example, in an on-premises environment or a cloud environment. A storage node (hereinafter referred to as a node) is, for example, a general-purpose computer and has a VOL (logical volume).
[0007] A certain node in the first storage system will select a path in a round-robin manner. As a result, depending on the selected path, in remote copy, data transfer between the certain node and another node may be required. Therefore, the performance of remote copy may decrease.
[0008] Such problems may exist regardless of whether the first storage system is a secondary storage system or a primary storage system. Also, such problems may exist regardless of whether the remote copy is synchronous remote copy (remote copy in which a completion response to a write request is returned when data written to PVOL accompanying a write request is written to SVOL) or asynchronous remote copy (remote copy in which a completion response to a write request is returned even if data written to PVOL accompanying a write request is not written to SVOL).
Means for Solving the Problems
[0009] There are multiple paths between the first and second storage systems. For each of the multiple paths, the path communicably connects any initiator port among the multiple initiator ports of the multiple nodes constituting the first storage system and any target port among the one or more target ports of the second storage system. For each of one or two or more of the multiple nodes, the node has a first VOL that forms a remote copy pair with a second VOL among the one or more VOLs of the second storage system. When any node sends a command for remote copy in the remote copy pair to the second storage system, the node selects the path connected to the initiator port as long as no abnormality regarding the initiator port of the node is detected, and sends the command via the selected path.
Advantages of the Invention
[0010] According to the present invention, it is possible to suppress a performance degradation of remote copy performed between a second storage system and a first storage system constituted by a plurality of nodes.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0012] In the following description, the "interface device" may be one or more communication interface devices. The one or more communication interface devices may be one or more of the same type of communication interface devices (for example, one or more NICs (Network Interface Cards)), or two or more different types of communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).
[0013] Also, in the following description, the "memory" is one or more memory devices which are an example of one or more storage devices, and typically may be a main memory device. At least one of the memory devices in the memory may be a volatile memory device or a non-volatile memory device.
[0014] Also, in the following description, a "persistent storage device" may be one or more persistent storage devices, which are examples of one or more storage devices. Persistent storage devices are typically non-volatile storage devices (e.g., auxiliary storage devices), and specifically may be, for example, HDD (Hard Disk Drive), SSD (Solid State Drive), or NVMe (Non-Volatile Memory Express) drives.
[0015] Also, in the following description, a "processor" may be one or more processor devices. At least one processor device is typically a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may also be a processor device in a broad sense such as a hardware circuit (e.g., FPGA (Field-Programmable Gate Array), CPLD (Complex Programmable Logic Device), or ASIC (Application Specific Integrated Circuit)) that performs part or all of the processing.
[0016] Also, in the following description, expressions such as "xxx table" may be used to describe information from which an output is obtained for an input. However, the information may be data of any structure (e.g., structured data or unstructured data), or may also be a learning model such as a neural network, genetic algorithm, or random forest that generates an output for an input. Therefore, "xxx table" can be referred to as "xxx information". Also, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.
[0017] Also, in the following description, although the "program" may be used as the subject to describe the processing, the program is executed by the processor to perform the defined processing while appropriately using a storage device and / or an interface device, etc. Therefore, the subject of the processing may be the processor (or a device such as a controller having the processor). The program may be installed from a program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0018] Also, in the following description, when describing without distinguishing between elements of the same kind, the common part of the reference signs is used, and when distinguishing between elements of the same kind, reference signs or element identifiers may be used. For example, for PVOL, it may be distinguished from other PVOLs using an identifier such as "PVOL1".
[0019] Hereinafter, several embodiments will be described. <First Embodiment>
[0020] FIG. 1 is a schematic diagram showing an overview of the first embodiment of the present invention.
[0021] There is a positive storage system 100P at the positive site 201P. The positive storage system 100P may be a so-called disk array system, but in this embodiment, it is a system composed of a plurality of positive nodes 210P.
[0022] There is a secondary storage system 100S at the secondary site 201S. The secondary storage system 100S is a system composed of a plurality of secondary nodes 210S.
[0023] Node 210 is typically a general-purpose computer, but it may also be a device other than a general-purpose computer. Node 210 has a port 215, a VOL (Logical Volume) 102, and an SCS (Storage Control Software) 730. VOL 102 is based on a persistent storage device inside or outside Node 210.
[0024] The plurality of positive nodes 210P includes, for example, positive nodes 210P1 to 210P3. As VOL 102 in the positive storage system 100P, there are PJVOL1, PJVOL1(S), PVOL1, and PVOL1(S). "PJVOL" is the JVOL at the positive site 201P. "JVOL" is the VOL in which the JNL (Journal) is stored. "JNL" includes the data to be copied and its metadata. The metadata in JNL is a sequence number (SEQ#), which is a value for specifying the order in which the data to be copied is written, and the write destination address of the data. "PJVOL1(S)" corresponds to the standby PJVOL1 in which the redundant data (e.g., replicated JNL) of the JNL stored in PJVOL1 is stored. "PVOL" is the positive VOL. "PVOL1(S)" corresponds to the standby PVOL1 in which the redundant data (e.g., replicated data) of the data stored in PVOL1 is stored.
[0025] Also, each of the plurality of positive nodes 210P has a target port as the port 215. "Tm" (m is a natural number) means the target port m. The positive node 210P has one or more ports 215.
[0026] The plurality of secondary nodes 210S includes, for example, secondary nodes 210S1 to 210S3. As VOL102, there are SJVOL1, SJVOL1(S), SVOL1, and SVOL1(S). "SJVOL" is the JVOL at the secondary site 201S. "SJVOL1(S)" corresponds to the standby SJVOL1 in which redundant data (e.g., replicated JNL) of the JNL stored in SJVOL1 is stored. "SVOL" is the secondary VOL that forms a pair with PVOL. "SVOL1(S)" corresponds to the standby SVOL1 in which redundant data (e.g., replicated data) of the data stored in SVOL1 is stored.
[0027] Also, each of the plurality of secondary nodes 210S has an initiator port as port 215. "In" (n is a natural number) means initiator port n. The secondary node 210S has one or more ports 215.
[0028] There are a plurality of paths 60 between the primary storage system 100P and the secondary storage system 100S. For each of the plurality of paths 60, the path 60 is a path that communicably connects any initiator port and any target port. In FIG. 1, "P-ID" is the path ID. Also, a "path" is a communication path for data and commands. For example, in the case of the iSCSI protocol, the path may be an iSCSI session.
[0029] Also, there are a plurality of path groups. A path group is composed of two or more (or one) paths 60 and is associated with a remote copy pair. In FIG. 1, "PG-ID" is the path group ID. In the present embodiment, two or more paths 60 with different path groups may be connected between the same initiator port and the same target port. As an example of "two or more paths 60 with different path groups", there are a path (PG-ID: 1, P-ID: 12) and a path (PG-ID: 2, P-ID: 21). These paths are all paths connecting I2 and T2. In other words, if the path groups are different, paths connected between the same initiator port and the same target port may be treated as separate paths. Hereinafter, the path group of "PG-ID: p" may be referred to as "path group p", and the path 60 of "P-ID: q" may be referred to as "path q".
[0030] For each of two or more nodes 210, a processor of the node 210 executes a storage control software (SCS) 730 that controls I / O (Input / Output) to / from VOL102. In FIG. 1, although the SCS of the primary node 210P is not shown for convenience of illustration, the SCS 730 also exists in the primary node 210P in the same manner as in the secondary node 210S.
[0031] For each of the primary storage system 100P and the secondary storage system 100S, there are multiple (or one) SCS groups. For each of the multiple SCS groups, the SCS group is composed of an active SCS 730, i.e., SCS(A), and one or more standby SCS 730s, i.e., SCS(S). One SCS(A) and one or more SCS(S) in the SCS group are arranged on two or more different nodes 210. When a failure occurs in the node 210 having SCS(A), a failover is performed such that any SCS(S) in the same SCS group becomes SCS(A) in place of SCS(A). For example, in the secondary storage system 100S, SCSx(A) (x is a natural number) and SCSx(S) constitute SCS group x. When a failure occurs in the node 210 having SCSx(A), a failover from SCSx(A) in the node 210 to any SCSx(S) in another node 210 is performed.
[0032] In this embodiment, JNL is used, and the port 215 of the secondary storage system 100S is the initiator port. That is, the remote copy performed in this embodiment is a pull-type asynchronous remote copy. The remote copy is performed for each remote copy pair (VOL pair). The remote copy pair y is composed of PVOLy (y is a natural number) and SVOLy. According to the example shown in FIG. 1, the remote copy pair 1 is composed of PVOL1 and SVOL1.
[0033] Since it is a pull-type asynchronous remote copy, an RDJNL command (JNL read command) is transmitted from the secondary storage system 100S for remote copy in the remote copy pair. In the RDJNL command, for example, the SEQ# of the latest JNL among the un-received JNLs is specified. In response to the RDJNL command, JNL is received from the primary storage system 100P, the received JNL is stored in SJVOL1, and the JNL stored in SJVOL1 is reflected in SVOL1 (the data in the JNL is written to SVOL1).
[0034] Since the sub-storage system 100S is configured to transmit the RDJNL command, in this embodiment, the sub-storage system 100S executes the path selection program 720, and the path selection program 720 selects a path 60 to be used for transmitting the RDJNL command. The path selection program 720 is provided in each sub-node 210S. The path selection program 720 in the sub-node 210Sz (z is a natural number) may be referred to as the "path selection program z".
[0035] When the sub-node 210S transmits an RDJNL command for remote copy in a remote copy pair to the main storage system 100P, it selects the path 60 connected to the initiator port of the sub-node 210S as long as no abnormality is detected in the initiator port of the sub-node 210S, and transmits the RDJNL command via the selected path 60. Specifically, for example, the sub-node 210S that selects the path 60 is the node 210S having the SCS(A) that transmits the RDJNL command. When the SCS1(A) of the sub-node 210S1 transmits a command for remote copy in the remote copy pair 1 to the main storage system 100P, as long as no abnormality is detected in the initiator port I1 of the sub-node 210S1, the path selection program 1 selects the path 11 connected to the initiator port I1 from the path group 1 associated with the remote copy pair 1, and the SCS1(A) transmits the RDJNL command via the selected path 11.
[0036] Hereinafter, this embodiment will be described in detail.
[0037] FIG. 2 is a diagram showing a physical configuration example of the storage system 101.
[0038] There are multiple sites 201. Each site 201 is communicably connected via a network 202. The network 202 is, for example, a WAN (Wide Area Network), but is not limited to a WAN. The site 201 is, for example, a data center or the like and includes a plurality (or one) of nodes 210.
[0039] The node 210 may be a general-purpose computer. The node 210 includes, for example, one or more processor packages 213 including a processor 211 and a memory 212, one or more drives 214, and one or more ports 215. Each of these components is connected via an internal bus 216. The drive 214 is an example of a permanent storage device.
[0040] The processor 211 is, for example, a CPU (Central Processing Unit) and performs various processes.
[0041] The memory 212 is typically a volatile memory and stores control information necessary for realizing the functions of the node 210 or stores data. Also, the memory 212 stores, for example, a program executed by the processor 211. The drive 214 stores various data, programs, and the like.
[0042] The port 215 is connected to a network 220 within the site 201 and communicably connects its own node to other nodes 210 within the site 201 via the network 220. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.
[0043] Note that the physical configuration of the system is not limited to the configuration described above. For example, the networks 202 and / or 220 may be redundant. Also, for example, the network 220 may be separated into a management network and a storage network, and the connection standard may be Ethernet (registered trademark), Infiniband, or wireless, and the connection topology is not limited to the configuration shown in FIG. 2. Also, for example, the drive 214 may have a configuration independent of the node 210.
[0044] FIG. 3 is a schematic diagram showing an overview of the remote copy configuration.
[0045] A plurality of remote copy pairs are constructed between the primary site 201P and the secondary site 201S. Specifically, for example, two consistency groups 401a and 401b are constructed between the primary site 201P and the secondary site 201S. The consistency group 401 is composed of VOL102 of a plurality (or one) of remote copy pairs. In the consistency group 401, a plurality of PVOLs are copied to the SVOL while maintaining consistency. More specifically, for example, in the consistency group 401, the update difference data up to the same time for a plurality of PVOL102s is copied to a plurality of SVOLs. Also, the control (consistency control) of the consistency group 401 is managed by the PJVOL. In the PJVOL, the update difference data of a plurality (or one) of PVOLs is stored together with metadata such as its write time. When the primary site 201P transfers the data of the PVOL to the secondary site 201S, the update difference data up to the same time among the update difference data written to the PJVOL is transferred to the secondary site 201S. Thereby, data can be copied to the SVOL while maintaining the consistency of the update times among a plurality of PVOLs.
[0046] For example, according to consistency group 401a, data is copied to SVOL1 and 2 at secondary node 210S1 via PJVOL1 and SJVOL1 while maintaining the consistency of PVOL1 and 2 at primary node 210P1. According to consistency group 401b, data is copied to SVOL3 at secondary node 210S2 and SVOL4 at secondary node 210S3 via PJVOL2 at primary node 210P2, PJVOL3 at primary node 210P3, SJVOL2 at secondary node 210S2, and SJVOL3 at secondary node 210S3 while maintaining the consistency of PVOL3 at primary node 210P2 and PVOL4 at primary node 210P3. PJVOL and SJVOL do not necessarily have a 1:1 correspondence (for example, they may be 1:multiple, multiple:1, or multiple:multiple), and PJVOL may be a region on memory 212.
[0047] As can be seen from the above-described specific configuration, consistency group 401 may be composed of VOL102 in a specific node 210 within site 201, or may be composed of VOL102 in a plurality of nodes 210 within site 201.
[0048] FIG. 4 is a schematic diagram showing an overview of I / O request processing.
[0049] First, application 502 operating on host 51 issues a write request specifying PVOL1 to primary node 210P1. Primary node 210P1 that has received the write request writes data A and B associated with the write request to PVOL1, and further writes a JNL containing data A and B as updated differential data to PJVOL1.
[0050] Next, the primary node 210P1 transfers the JNL (update differential data) written to PJVOL1 to SJVOL1 and SJVOL1(S) of the secondary site 201S. At this time, if multiple paths are established between the primary site 201P and the secondary site 201S, the JNL may be transferred using any of the paths. Normally, the primary node 210P1 transfers the JNL to the secondary node 210S1 that has the ownership of SVOL1 paired with PVOL1. However, when a failure occurs in the path with the ownership, the primary node 210P1 may transfer the JNL to the secondary node 210S2 that does not have the ownership. For example, when the primary node 210P1 transfers the JNL to the secondary node 210S2 that does not have the ownership, the secondary node 210S2 transfers the received JNL to the secondary node 210S1 with the ownership, and the secondary node 210S1 writes the JNL to SJVOL1.
[0051] Next, the secondary node 210S1 writes the data A and B in the JNL written to SJVOL1 to SVOL1. Then, the data A and B written to SVOL1 are written to the drive 214a via the storage pool 504a. When the configuration of the drive 214a is DAS (Direct Attached Storage) where the node 210 and the drive 214 are connected one-to-one, the JNL is written to the drive 214a mounted on the secondary node 210S1. By writing all the data copied to SVOL1 to the drive 214a of the secondary node 210S1 with the ownership of SVOL1, when reading data from SVOL1 later, it is not necessary to read data from another node. As a result, the node-to-node transfer process can be eliminated, and a high-speed read process can be realized.
[0052] Note that the storage pool 504 may be an area based on one or more drives 214. Storage functions such as thin-provisioning, compression, or deduplication are provided, and the processing of the storage functions required for the data written to the storage pool 504 is executed.
[0053] When the secondary node 210S1 writes data to the drive 214a, in order to protect the data from node failures, it also writes redundant data of the data to be written to the drive 214b of the secondary node 210S2. Regarding the writing of redundant data, when the data protection policy is replication, a replica of the data is written to the drive 214b as redundant data. On the other hand, when the data protection policy is Erasure Coding, parity is calculated from the data, and the calculated parity is written to the drive 214b as redundant data.
[0054] Although not shown in the figure, the primary node 210P1 transfers (redundantizes) the data to be written to PVOL1 to the primary node 210P2, and the primary node 210P2 receives the data and writes it to PVOL1(S). Also, the primary node 210P2 writes the JNL to PJVOL1(S). The JNL written to PJVOL1(S) may be the JNL transferred from the primary node 210P1 or the JNL generated based on the data written to PVOL1(S). In this way, PVOL1(S) as a replica of PVOL1 and PJVOL1(S) as a replica of PJVOL1 are maintained (see Figure 1).
[0055] Figure 5 is a schematic diagram showing an overview of the recovery process from node failures.
[0056] In the secondary nodes 210S1, 210S2, and 210S3, the SCS730 is operating. The secondary node 210S has an SCS(A) and an SCS(S) corresponding to the SCS(A) in another secondary node 210S (the SCS(S) in the SCS group including the SCS(A) in another secondary node 210S). For example, the secondary node 210S1 has an SCS1(A) and an SCS3(S), the secondary node 210S2 has an SCS2(A) and an SCS1(S), and the secondary node 210S3 has an SCS3(A) and an SCS2(S). The SCSx(A) and SCSx(S) belong to the SCS group x (the redundancy group of SCSx), and there may be more than one SCSx(S).
[0057] Using the specific example shown in FIG. 5, the recovery process from a node failure will be described.
[0058] In order to inherit the remote copy pair information of the secondary node 210S1, the secondary node 210S2 replicates and holds the configuration information of SVOL1 and SJVOL1 that the secondary node 210S1 has. Also, the secondary node 210S2 stores the redundant data of the data written to the drive 214a of the secondary node 210S1 in the drive 214d. Furthermore, the secondary node 210S2 has established a path (communication path) with the primary node 210P1.
[0059] And, for example, when the secondary node 210S1 stops due to a failure, the secondary node 210S2 that has detected the failure of the secondary node 210S1 takes over the processing of SCS1(A) of the secondary node 210S1, and SCS1(S) changes to SCS1(A). The secondary node 210S2 communicates with the primary node 210P1 to continue the remote copy process between PVOL1 and SVOL1. That is, a failover from SCS1(A) of the secondary node 210S1 to SCS1(S) of the secondary node 210S2 is performed. Thereby, even if a node failure occurs in any of the secondary nodes 210S in the secondary site 201S, another secondary node 210S in the secondary site 201S can continue the remote copy from the primary site 201P.
[0060] Now, as described above, in this embodiment, the path selection program 720 in the secondary node 210S having SCS(A) that transmits the RDJNL command selects the path used for the transmission of the RDJNL command. Hereinafter, an example of path selection will be described with reference to FIGS. 1 and 6. For example, for the remote copy pair 1 (the pair of PVOL1 and SVOL1), at least one of the following (1) to (8) path selections is performed.
[0061] (1) When transmitting the RDJNL command, the path selection program 720 in the secondary node 210S having the SCS(A) that transmits the RDJNL command selects the path connected to the initiator port of the secondary node 210S as long as no abnormality regarding the initiator port of the secondary node 210S is detected, and transmits the RDJNL command via the selected path. That is, the path connected to the initiator port of the secondary node 210S having the SCS(A) that transmits the RDJNL command is preferentially selected. Specifically, for example, when SCS1(A) transmits the RDJNL command for the remote copy pair 1, the path selection program 1 preferentially selects the path 11 connected to the initiator port I1 of the secondary node 210S1 from the path group 1 associated with the remote copy pair 1.
[0062] When an abnormality regarding I1 is detected, for example, when there is a failure of the secondary node 210S1 having I1, a failure of I1, the performance of I1 is below the threshold, or the performance of the path 11 connected to I1 is below the threshold, the path connected to the initiator port of another secondary node 210S is selected, and the RDJNL command is transmitted via the path. For example, for the remote copy pair 1, at least one of the following path selections (2) to (5) is performed.
[0063] (2) At the time of failure of the secondary node 210S1 (see FIG. 6), the path 12 connected to I2 of the secondary node 210S2 where SCS1(S) (the SCS1 that becomes active instead of SCS1(A) of the secondary node 210S1) operates is selected by the path selection program 2 of the secondary node 210S2 from the path group 1 associated with the remote copy pair 1 (see the thick solid line arrow in FIG. 6).
[0064] (3) At the time of port failure of I1, the path 12 connected to I2 of the secondary node 210S2 where SCS1(S) operates is selected by the path selection program 1 of the secondary node 210S1 from the path group 1 associated with the remote copy pair 1 (see the dashed-dotted line arrow in FIG. 1).
[0065] (4) In (2) and / or (3), when there is a port failure of I2, the path 13 connected to I3 of another secondary node 210S3 without SCS1(S) is selected from the path group 1 associated with the remote copy pair 1 by the path selection program 1 or the path selection program 2.
[0066] (5) When the performance of I1 is exhausted, the path connected to the initiator port of another secondary node 210S among the path group 1 is selected by the path selection program 1 of the secondary node 210S1 (see the dashed arrow in FIG. 1). Note that "when the performance of I1 is exhausted" means, for example, when the performance of I2 is below the threshold value or when the performance of the path 11 connected to I1 is below the threshold value. Also, the "another secondary node 210S" mentioned in this paragraph may be a secondary node 210S with a resource utilization rate below the threshold value.
[0067] (6) In at least one of (1) to (5), when there is no selected path to the initiator port of the secondary node 210S, the path selection program 720 of the secondary node 210S having the initiator port dynamically generates a path connected to the initiator port and selects the generated path. For example, if there is no path connected to I1, the path selection program 1 dynamically generates a path connected to I1. Also, for example, if there is no path connected to I2, the path selection program 1 or the path selection program 2 dynamically generates a path connected to I2. The dynamically generated path may be deleted by the path selection program 720 that generated the path when the abnormality regarding the initiator port that caused the generation of the path is resolved. That is, the dynamically generated path may belong to the path group 1 temporarily (or permanently) (and may be managed as a temporary (or permanent) path available for remote copy regarding the remote copy pair 1).
[0068] (7) When a plurality of paths in path group 1 are connected to I1, path selection program 1 selects a path in a round-robin manner from among the plurality of paths. For example, for remote copy pair 1, path 11 may be selected for the transmission of a certain RDJNL command, and path 14 may be selected for the transmission of the next RDJNL command. In response to the RDJNL command via path 14, JNL may be transmitted from PJVOL1(S) to secondary node 210S1 via path 14 by primary node 210P2 (see FIG. 1).
[0069] (8) When a plurality of paths in path group 1 are connected to I1, path selection program 1 identifies the statistically most excellent path (for example, the path with the best performance or the fewest error counts) among the plurality of paths based on statistics (for example, performance or error count statistics) for each of the plurality of paths, and selects the identified path.
[0070] The above path selection makes use of the arrangement of SCS(A) and SCS(S) in secondary storage system 100S. Specifically, for example, according to at least one of (2) to (5), the initiator ports of secondary nodes 210S other than secondary node 210S1 are preferentially used for the transmission of RDJNL commands of SCS(A) arranged in the secondary node 210S for remote copy pairs other than remote copy pair 1. Therefore, if an abnormality occurs in the initiator port, the abnormality is detected by the secondary node 210S. In other words, if the abnormality in the initiator port has not been detected by the secondary node 210S, it means that the path connected to the initiator port is usable. That is, in the transmission of the RDJNL command of SCS1(A), even if there is no selection of a path via another secondary node 210S by round-robin, it is possible to know whether the path via the other secondary node 210S is available before path selection. Thus, in this embodiment, it is possible to perform path selection that effectively utilizes the arrangement of SCS(A) and SCS(S) in secondary storage system 100S.
[0071] FIG. 7 is a diagram showing an example of data and programs held in the memory 212.
[0072] Information is read from the drive 214 into the memory 212. For example, various tables, the SCS 730, and the path selection program 720 included in the control information table 710 are expanded on the memory 212 during the execution of the processes for which they are used, but are stored in a non-volatile storage area such as the drive 214 in case of a power failure or the like when not in use. The control information table 710 includes a system configuration management table 711, a pair configuration management table 712, and a path management table 713.
[0073] FIG. 8 is a diagram showing an example of the system configuration management table 711.
[0074] The system configuration management table 711 includes a node configuration management table 810, a drive configuration management table 820, and a port configuration management table 830. For each site 201, there is a node configuration management table 810 for a plurality of nodes 210 existing in each site 201, and the node 210 has a drive configuration management table 820 and a port configuration management table 830 regarding the drive 214 within its own node 210.
[0075] The node configuration management table 810 is provided for each site 201 and stores information indicating the configuration (such as the relationship between the node 210 and the drive 214) regarding the node 210 provided in the site 201. More specifically, the node configuration management table 810 stores information such as a node ID 811, a status 812, a CPU usage rate 815, a memory usage rate 816, a drive ID list 813, and a port ID list 814 for each node 210.
[0076] The node ID 811 is the ID of node 210. The status 812 indicates the status of node 210 (e.g., "Normal", "Warning", or "Failure", etc.). The CPU usage rate 815 represents the CPU usage rate of node 210. The memory usage rate 816 represents the memory usage rate of node 210. The drive ID list 813 is a list of the IDs of the drives 214 provided in node 210. The port ID list 814 is a list of the IDs of the ports 215 provided in node 210.
[0077] The drive configuration management table 820 is provided for each node 210 and stores information indicating the configuration related to the drive 214 provided in node 210. More specifically, the drive configuration management table 820 stores information such as the drive ID 821, the status 822, the BE bandwidth usage rate 824, the drive usage rate 825, and the size 823 for each drive 214.
[0078] The drive ID 821 is the ID of the drive 214. The status 822 indicates the status of the drive 214. The BE bandwidth usage rate 824 represents the usage rate of the communication bandwidth (backend bandwidth) between the processor 211 and the drive 214. The drive usage rate 825 represents the ratio of the used capacity to the capacity of the drive 214. The size 823 indicates the capacity of the drive 214.
[0079] The port configuration management table 830 is provided for each node 210 and stores information indicating the configuration related to the port 215 provided in node 210. More specifically, the port configuration management table 830 stores information such as the port ID 831, the status 832, the NW bandwidth usage rate 834, and the address 833 for each port.
[0080] The port ID 831 is the ID of port 215. The status 832 indicates the status of port 215. The NW bandwidth utilization rate 834 represents the utilization rate of the bandwidth of the network connected to port 215 (the value of the NW bandwidth utilization rate 834 may be an example of the performance of port 215). The address 833 indicates the address on the network assigned to port 215. The form of the address may be an IP (Internet Protocol), or a WWN (World Wide Name), a MAC (Media Access Control) address, etc.
[0081] Figure 9 is a diagram showing an example of the pair configuration management table 712.
[0082] The pair configuration management table 712 is composed of a VOL management table 910, a pair management table 920, and a JNL management table 930.
[0083] The VOL management table 910 stores information indicating the configuration related to VOL102. More specifically, the VOL management table 910 stores information such as a VOL ID 911, an owner node ID 912, a fallback destination node ID 913, a size 914, and an attribute 915 for each VOL102.
[0084] The VOL ID 911 is the ID of VOL102. The owner node ID 912 is the ID of node 210 that has the ownership of VOL102. The fallback destination node ID 913 is the ID of node 210 that takes over the process in case of a failure of node 210 that has the ownership of the SVOL. The size 914 indicates the capacity of VOL102.
[0085] The attribute 915 indicates the attribute of VOL102. "NML_VOL" means a normal VOL not used in remote copy. "PVOL" means a positive VOL. "PJVOL" means a JVOL that stores the updated differential data of the PVOL. Although not shown, the attributes also include SVOL and SJVOL.
[0086] The pair management table 920 stores information indicating the configuration related to the remote copy pair. More specifically, for each consistency group, the pair management table 920 stores information such as a pair group ID 921, a PJVOL ID 922, a PVOL ID 923, an SJVOL ID 924, an SVOL ID 925, a PG-ID 927, and a status 926.
[0087] The pair group ID 921 is the ID of the consistency group. The PJVOL ID 922 is a list of the IDs of the PJVOLs belonging to the consistency group. The PVOL ID 923 is a list of the IDs of the PVOLs belonging to the consistency group. The SJVOL ID 924 is a list of the IDs of the SJVOL102JSs belonging to the consistency group. The SVOL ID 925 is a list of the IDs of the SVOL102Ss belonging to the consistency group. The PG-ID 927 is the ID of the path group for each remote copy pair belonging to the consistency group. The status 926 indicates the status of each remote copy pair in the consistency group (e.g., "PAIR", "COPY", "SUSPEND", etc.). "PAIR" is the state in which the write to the PVOL102P is periodically reflected in the SVOL102S. "COPY" is the state during the initial copy. "SUSPEND" is the pair interruption state (the state in which synchronization between the PVOL102P and the SVOL102S is not performed).
[0088] The JNL management table 930 stores information related to the JNL. More specifically, for each JNL, the JNL management table 930 stores information such as a pair group ID 931, a JNL ID 932, a P / SVOL ID 933, a P / SVOL address 934, a size 935, and a cache segment ID 936.
[0089] The pair group ID 931 is the ID of the consistency group to which the JNL belongs. The JNL ID 932 is the ID of the JNL. The JNL ID corresponds to the SEQ#, and in the consistency group, it is, for example, a sequential number. That is, the JNL ID represents the writing order, and the data within the JNL is stored in SVOL102S in the consistency group in the order of the JNL IDs.
[0090] The P / SVOL ID 933 includes the ID of PVOL102P where the data within the JNL is written and the ID of SVOL102S where the data within the JNL is written. The P / SVOL address 934 includes the storage destination address of the data within the PVOL102P where the data within the JNL is written and the storage destination address of the data within the SVOL102S where the data within the JNL is written.
[0091] The size 935 represents the size of the JNL. For example, one JNL contains one or more pieces of data. The cache segment ID 936 is the ID of the cache segment (a region in the cache provided in the memory 212) where the data within the JNL is written.
[0092] Figure 10 is a diagram showing an example of the path management table 713.
[0093] The path management table 713 is a table regarding the path 60. The path management table 713 stores information such as the PG-ID 1031, P-ID 1034, protocol information 1035, status 1036, average response 1037, Local Address 1032, and Destination Address 1033 for each path 60.
[0094] The PG-ID 1031 is the ID of the path group to which the path 60 belongs. The P-ID 1034 is the ID of the path 60.
[0095] The protocol information 1035 is information indicating the communication protocol of path 60. The communication protocol of 60 may be iSCSI, or may be a vendor-specific proprietary protocol in addition to FC (Fibre Channel) and NVMe-oF (NVMe over Fabrics).
[0096] The state 1036 represents the state of path 60. A predetermined value other than "Normal" may mean an abnormality regarding path 60. For example, "Failure" means a path failure.
[0097] The average response 1037 represents the average response time in communication via path 60. The value of the average response 1037 may be an example of the performance of path 60.
[0098] Local Address 1032 represents the address of the initiator port of path 60. Destination Address 1033 represents the address of the target port of path 60.
[0099] Hereinafter, an example of the process performed in this embodiment will be described. In the following description, remote copy pair 1 will be taken as an example. The secondary node 210S having the ownership of SVOL1 in remote copy pair 1 is secondary node 210S1.
[0100] FIG. 11 is a diagram showing the flow of the remote copy process.
[0101] SCS1(A) issues an RDJNL command (S1101). Path selection processing is performed (S1102).
[0102] SCS1(A) transmits the RDJNL command to the primary storage system 100P via the path 60 selected in the path selection process of S1102 (S1103). In response to the RDJNL command, via the path 60, SCS1(A) receives the JNL from the primary storage system 100P (S1104). SCS1(A) stores the received JNL in SJVOL1 (S1105).
[0103] SCS1(A) reflects the unreflected JNL in the order of SEQ# to SVOL1 (S1106). That is, the data in the unreflected JNL is stored in SVOL1.
[0104] Although details are omitted, in the remote copy process, the status 926 of the pair group including the remote copy pair 1 and the values in various other tables are appropriately updated.
[0105] FIG. 12 is a diagram showing the flow of the path selection process. In the description with reference to FIGS. 12 to 14, the remote copy pair 1 and the path selection program 1 are taken as examples. However, when a failover occurs from SCS1(A) of the secondary node 210S1 to SCS1(S) of the secondary node 210S2 due to a node failure of the secondary node 210S1, the path selection process may be performed by the path selection program 2. Also, for the purpose of taking the remote copy pair 1 as an example, the connection destination in the positive storage system 100P of the dynamically generated path is the target port (that is, T1 or T2) of the positive node 210P having PVOL1 or PVOL1(S). The address of the target port can be specified from the path management table 713.
[0106] The path selection program 1 specifies the path group ID: 1 corresponding to the remote copy pair 1 from the pair management table 920 (S1201). The path selection program 1 refers to the path management table 713 using the path group ID: 1 as a key and determines which of the following (A) to (C) applies (S1202). (A) There is no normal path (a path with status 1036 being "Normal") in the secondary node 210S1. (B) The performance of the path (the path connected to the initiator port I1) of the secondary node 210S1 is insufficient. That is, the performance of the path that is the same node as SCS1(A) and should be preferentially selected is insufficient. (C) It does not fall under either (A) or (B).
[0107] The determination of whether (A) applies may be based on whether the states 1036 of path 11 and path 14 (the path connected to I1 of sub-node 210S1) are both other than "Normal". If neither path 11 nor path 14 is "Normal", then (A) applies. When (A) applies, path selection program 1 performs path selection processing in case of priority path failure (S1203).
[0108] The determination of whether (B) applies may be based on whether the average response 1037 of path 11 and path 14 (the path connected to I1 of sub-node 210S1) is both equal to or greater than the threshold value (and / or whether the number of RDJNL commands transmitted per unit time exceeds the threshold value). If the average response 1037 of path 11 and path 14 (the path connected to I1 of sub-node 210S1) is both equal to or greater than the threshold value (and / or the number of RDJNL commands transmitted per unit time exceeds the threshold value), then (B) applies. When (B) applies, path selection program 1 performs path selection processing in case of insufficient priority path performance (S1205).
[0109] When (C) applies, path selection program 1 selects the normal path 11 or normal path 14 connected to I1 (S1204). In S1204, the normal path may be selected by round-robin, or the normal path with the smallest value of average response 1037 may be selected.
[0110] FIG. 13 is a diagram showing the flow of path selection processing in case of priority path failure.
[0111] Path selection program 1 determines whether the sub-node 210S2 having SCS1(S) has a normal path (S1301). This determination may be made by path selection program 1 collecting the necessary information from path selection program 2, or may be based on the path management table 713 of all sub-nodes 210S being shared among each sub-node 210S and using the shared table.
[0112] When the determination result of S1301 is true (S1301: YES), the path selection program 1 selects one normal path from the normal paths connected to the initiator port I2 of the secondary node 210S2 having SCS1(S) (S1302). Note that in S1302, path 12 in path group 1 is selected based on the path management table 713.
[0113] When the determination result of S1301 is false (S1301: NO), the path selection program 1 selects one normal path from the normal paths connected to the initiator ports of other secondary nodes 210S (S1303). Note that the "other secondary nodes 210S" mentioned in this paragraph refer to any secondary node 210S other than secondary node 210S1 and secondary node 210S2. In S1303, path 13 in path group 1 is selected based on the path management table 713.
[0114] In S1301 to S1303, when there is no normal path, the path selection program 1 may dynamically generate a path. The connection destination of the dynamically generated path may be determined based on the Destination Address 1033 in the path management table 713. Specifically, for the remote copy pair 1, the connection destination of the dynamically generated path may be T1 of the primary node 210P1 having PVOL1, or T2 of the primary node 210P2 having PVOL1(S).
[0115] Also, before the determination of S1301, a self-node port determination may be performed to determine whether the state 832 of I1 of the secondary node 210S1 is "Normal" (normal). S1301 may be performed when the result of the self-node port determination is false. When the result of the self-node port determination is true, the path selection program 1 may dynamically regenerate the path connected from I1 to T1 or T2 for the remote copy pair 1.
[0116] FIG. 14 is a diagram showing the flow of path selection processing when the priority path performance is insufficient.
[0117] The path selection program 1 determines whether the resource utilization rate of the secondary node 210S1 is below the threshold (S1401). For the remote copy pair 1, the "resource utilization rate" mentioned in this paragraph may be at least one of the CPU utilization rate 815 of the secondary node 210S1, the memory utilization rate 816 of the secondary node 210S1, the statistical value (e.g., average value) of the BE bandwidth utilization rate 824 for the drive 214 of the secondary node 210S1, and the NW bandwidth utilization rate 834 of I1.
[0118] If the determination result of S1401 is true (S1401: YES), the path selection program 1 dynamically generates a path connected to I1 and selects the path (S1402). That is, although the performance of path 11 and path 14 is insufficient, if the resource utilization rate of the secondary node 210S1 is below the threshold, a path connected to I1 is dynamically generated.
[0119] If the determination result of S1401 is false (S1401: NO), the path selection program 1 determines whether the resource utilization rate of other secondary nodes 210S is below the threshold (S1403). The "other secondary nodes 210S" mentioned in this paragraph may be each secondary node 210S other than the secondary node 210S1. Also, in this paragraph, the "resource utilization rate" of other secondary nodes 210S may be the same as the "resource utilization rate" of the secondary node 210S1.
[0120] If the determination result of S1403 is false (S1403: NO), the path selection program 1 selects path 11 or path 14 from I1 (S1404). That is, although the performance of path 11 and path 14 is insufficient, if the resource utilization rate of other secondary nodes 210S exceeds the threshold, there may be a low advantage in selecting a path connected to other secondary nodes 210S, so path 11 or path 14 is selected.
[0121] If the determination result of S1403 is true (S1403: YES), the path selection program 1 determines, based on the path management table 713, whether a normal path belonging to path group 1 is connected to other secondary nodes 210S (S1405).
[0122] When the determination result of S1405 is false (S1405: NO), the path selection program 1 causes the path selection program 720 of another sub-node 210S to dynamically generate a path connected to the initiator port of the other sub-node 210S (S1406), and selects the path (S1407). That is, if the performance of path 11 and path 14 is insufficient and the resource utilization rate of another sub-node 210S is below the threshold, but no path belonging to path group 1 is connected to the other sub-node 210S, a path connected to the initiator port of the other sub-node 210S is dynamically generated.
[0123] When the determination result of S1405 is true (S1405: YES), the path selection program 1 selects a normal path 12 or path 13 connected to another sub-node 210S from path group 1 (S1407). <Second Embodiment>
[0124] The second embodiment of the present invention will be described. At this time, the differences from the first embodiment will be mainly described, and the description of the common points with the first embodiment will be omitted or simplified.
[0125] FIG. 15 is a schematic diagram showing the outline of the second embodiment.
[0126] In the first embodiment, pull-type asynchronous remote copy is adopted, but in the second embodiment, push-type synchronous remote copy is adopted. That is, the positive node 210P has a path selection program 720, and the path selection program 720 of the positive node 210P having SCS(A) for a remote copy pair (for example, a pair of PVOL1 and SVOL1) performs path selection processing. For the remote copy pair, the port of the positive node 210P is the initiator port, and the port of the sub-node 210S is the target port. The WRJNL command (JNL write command) is transmitted from the SCS(A) of the positive node 210P to the secondary storage system 100S via the path selected by the path selection processing. Thereby, it is expected to suppress a decrease in performance related to the processing of push-type remote copy.
[0127] The above has described an embodiment of the present invention. However, this is an exemplification for the description of the present invention and is not intended to limit the scope of the present invention to this embodiment. The present invention can be implemented in various other forms.
[0128] Also, the above description can be summarized as follows. The following summary may include supplementary explanations and explanations of modified examples of the above description. According to the first and second embodiments, the present invention can be applied regardless of whether the remote copy is of the pull type or the push type, and regardless of whether it is a synchronous remote copy or an asynchronous remote copy. Therefore, in the following summary, the first storage system may be either the secondary storage system 100S or the primary storage system 100P, and the second storage system may be the primary storage system 100P when the first storage system is the secondary storage system 100S, and may be the secondary storage system 100S when the first storage system is the primary storage system 100P.
[0129] A storage system as a first storage system that communicates with a second storage system includes a plurality of nodes each having an initiator port, a memory, a processor, and a VOL. There are a plurality of paths between the first storage system and the second storage system. For each of the plurality of paths, the path communicably connects any one of the plurality of initiator ports of the plurality of nodes to any one of the one or more target ports of the second storage system. For each of one or two or more of the plurality of nodes, the node has a first VOL that forms a remote copy pair with a second VOL among the one or more VOLs of the second storage system, and when the node transmits a command for remote copy in the remote copy pair to the second storage system, the path connected to the initiator port is selected unless an abnormality related to the initiator port of the node is detected, and the command is transmitted via the selected path. As a result, command transfer between nodes is not required in command transmission, and thus it can be expected to suppress a performance degradation of remote copy. Also, "one or two or more of the plurality of nodes" may mean that the plurality of nodes may include so-called spare nodes other than "one or two or more nodes".
[0130] For each of two or more nodes, a processor of the node may execute an SCS that controls I / O (Input / Output) to / from VOL. There are one or more SCS groups, and for each of the one or more SCS groups, the SCS group may be composed of an SCS (A) that is one active SCS and one or more SCSs (S) that are standby SCSs. The one SCS (A) and the one or more SCSs (S) may be arranged on two or more different nodes among a plurality of nodes. A failover may be performed such that when a failure occurs in the node having the SCS (A), any one of the SCSs (S) in the SCS group becomes the SCS (A) in place of the SCS (A) in the SCS group. The node that selects a path for a remote copy pair for which remote copy is performed may be the node having the SCS (A) that transmits a command for the remote copy. Therefore, for each of two or more nodes, an initiator port of the node is used for command transmission from the SCS (A) in the node, and thus, it is possible to determine whether there is an abnormality regarding the initiator port of each node. As a result, it is possible to know whether a path connected to another node can be used before path selection without the node having the SCS (A) selecting a path connected to the node and a path connected to another node in a round-robin manner.
[0131] For example, any one of two or more nodes may be a first node (e.g., secondary node 210S1). Any SCS group other than the first SCS group including the first SCS (A) (e.g., SCS1(A)) owned by the first node may be a second SCS group. A node having a second SCS (A) (e.g., SCS2(A) or SCS3(A)) within the second SCS group may be a second node (e.g., secondary node 210S2 or secondary node 210S3). A remote copy pair for which remote copy is performed by the first SCS (A) may be a first remote copy pair (e.g., remote copy pair 1). When an abnormality related to the initiator port (e.g., I1) of the first node is detected, the node having the first SCS (A) selects a second path (e.g., path 12 or path 13) connected to the initiator port, as long as an abnormality related to the initiator port (e.g., I2 or I3) of the second node has not been detected by the second node, and may transmit a command for remote copy in the first remote copy pair via the second path. Thereby, it is possible to avoid a situation where there is an abnormality when selecting a path connected to the second node, and thus it can be expected to suppress a performance degradation of remote copy.
[0132] Note that the second node may be a node (e.g., secondary node 210S2) having a first SCS (S) (e.g., SCS1(S)) within the first SCS group. Thereby, it can be expected that the first SCS (S) takes over the process in place of the first SCS (A) and continues the remote copy without command transfer between the second node and other nodes. Specifically, for example, the second node may have a standby first VOL which is a VOL in which redundant data of data stored in the first VOL constituting the first remote copy pair is stored. A command for remote copy in the first remote copy pair and transmitted via the second path may be a command for remote copy between the standby first VOL and a second VOL in the first remote copy pair.
[0133] Any SCS group other than the first SCS group and the second SCS group may be a third SCS group. A node having a third SCS (A) (for example, SCS3(A)) within the third SCS group may be a third node (for example, secondary node 210S3). When an abnormality regarding the initiator port of the first node is detected, if an abnormality regarding the initiator port of the node having the first SCS(A) is detected by the second node, and unless an abnormality regarding the initiator port of the third node is detected by the third node, a third path connected to the initiator port may be selected, and a command for remote copy in the first remote copy pair may be transmitted via the third path. Thereby, it is possible to avoid a situation where there is an abnormality when a path connected to the third node is selected, and thus it can be expected to suppress a decrease in the performance of remote copy.
[0134] The abnormality regarding the initiator port may be a failure of the node having the initiator port, a failure of the initiator port, a failure of the path connected to the initiator port, the performance of the initiator port being below a threshold value, or the performance of the path connected to the initiator port being below a threshold value.
[0135] The abnormality regarding the initiator port of the first node may be that the performance of the initiator port is below a threshold value, or the performance of the path connected to the initiator port is below a threshold value. In that case, the second node may be a node with a resource utilization rate below a threshold value. That is, when an abnormality regarding the initiator port of the first node is detected and there is a second node as another node with a resource utilization rate below a threshold value, a path connected to the initiator port of the second node may be selected. Since there is a margin in the resources regarding the second node to which the selected path is connected, it can be expected to suppress a decrease in the performance of remote copy.
[0136] For each of one or more nodes, an abnormality related to the initiator port of the node may be that the performance of the path connected to the initiator port is below a threshold value. When the usage rate of the resources related to the node is below the threshold value, the node may dynamically generate a path connected to the initiator port and select the generated path. Thereby, when there is a margin in the resources related to the node even though the performance of the path connected to the node is low, a path connected to the node is dynamically generated and the path is selected, so that it can be expected to suppress a decrease in the performance of remote copy.
[0137] When an abnormality related to the initiator port of the node or another node is detected and it is necessary to select a path connected to the initiator port of the node (further, when there is a shortage of paths (for example, normal and sufficiently performing paths) connected to the initiator port of the node), the path selected by the node may be a path dynamically generated by the node. Thereby, it can be expected to suppress a decrease in the performance of remote copy. When the abnormality related to the initiator port is resolved, the dynamically generated path may be deleted by the node that dynamically generated the path. Thereby, it can be expected to avoid the management burden or resource excess associated with leaving unnecessary paths.
[0138] When a plurality of paths are connected to the initiator port of a node, when the node transmits a command to a second storage system, the node may select a path from the plurality of paths in a round-robin manner. Thereby, since it is possible to avoid load concentration on the path while eliminating the need for command transfer between nodes, it is expected to further suppress a decrease in the performance of remote copy.
[0139] When a plurality of paths are connected to the initiator port of a node, when the node transmits the command to the second storage system, the node may identify the statistically most excellent path among the plurality of paths based on the statistics for each of the plurality of paths, and select the identified path. Thereby, it is possible to maintain the selection of the most excellent path while eliminating the need for command transfer between nodes, and it is expected to further suppress the performance degradation of remote copy.
[0140] When a node transmits a command for remote copy in a remote copy pair to a second storage system, as long as an abnormality regarding the initiator port of the node is not detected, the path connected to the initiator port may be selected from a path group associated with the remote copy pair and including the path connected to the initiator port. Different path groups can be associated with different remote copy pairs, whereby paths between the same initiator port and target port can be managed as different paths. For example, as a path connecting I2 and T2, management such as path 12 belonging to path group 1 associated with remote copy pair 1 and path 21 belonging to path group 2 associated with another remote copy pair (the pair of PVOL5 and SVOL7) is possible. Also, for each remote copy pair, the paths as options are narrowed down. As a result, it is expected to suppress the performance degradation of remote copy.
[0141] Also, in the above description, whether each of the first storage system and the second storage system is a primary storage system or a secondary storage system may differ depending on the remote copy pair. That is, the storage system having PVOL may be the primary storage system, and the storage system having SVOL may be the secondary storage system.
Explanation of Signs
[0142] 201 Site 210 Node
Claims
1. A storage system as a first storage system that communicates with a second storage system, comprising a plurality of nodes each having an initiator port, a memory, a processor, and a VOL (Logical Volume), there are a plurality of paths, and for each of the plurality of paths, the path communicably connects any one of the plurality of initiator ports of the plurality of nodes to any one of the one or more target ports of the second storage system, for each of one or two or more of the plurality of nodes, the node has a first VOL that forms a remote copy pair with a second VOL among the one or more VOLs of the second storage system, when the node transmits a command for remote copy in the remote copy pair to the second storage system, unless an abnormality related to the initiator port of the node is detected, it selects the path connected to the initiator port and transmits the command via the selected path, A storage system.
2. For each of the two or more nodes, the processor of the node executes SCS (Storage Control Software) that controls I / O (Input / Output) to the VOL, there are one or more SCS groups, and for each of the one or more SCS groups, the SCS group is composed of an SCS (A) that is one active SCS and one or more SCSs (S) that are standby SCSs, and the one SCS (A) and the one or more SCSs (S) are arranged on two or more different nodes among the plurality of nodes, and when a failure occurs in the node having the SCS (A), any one of the SCSs (S) in the SCS group becomes SCS (A) instead of SCS (A) in the SCS group, and a failover is performed, The node that selects a path for the remote copy pair for which remote copy is performed is the node having the SCS (A) that transmits a command for the remote copy, The storage system according to claim 1.
3. Any one of the two or more nodes is a first node, Any SCS group other than the first SCS group including the first SCS (A) possessed by the first node is the second SCS group, The node having the second SCS (A) within the second SCS group is the second node, The remote copy pair for which remote copy is performed by the first SCS (A) is the first remote copy pair, When an abnormality related to the initiator port of the first node is detected, the node having the first SCS (A) selects a second path connected to the initiator port as long as the second node has not detected an abnormality related to the initiator port of the second node, and transmits a command for remote copy in the first remote copy pair via the second path. The storage system according to claim 2.
4. The second node is a node having the first SCS (S) within the first SCS group. The storage system according to claim 3.
5. The second node has a standby first VOL which is a VOL in which redundant data of data stored in the first VOL constituting the first remote copy pair is stored. The command for remote copy in the first remote copy pair and transmitted via the second path is a command for remote copy between the standby first VOL and the second VOL in the first remote copy pair. The storage system according to claim 4.
6. Any SCS group other than the first SCS group and the second SCS group is the third SCS group, The node having the third SCS (A) within the third SCS group is the third node, When an abnormality related to the initiator port of the first node is detected, the node having the first SCS (A) selects a third path connected to the initiator port as long as the third node has not detected an abnormality related to the initiator port of the third node if the second node has detected an abnormality related to the initiator port of the second node, and transmits a command for remote copy in the first remote copy pair via the third path. The storage system according to claim 4.
7. An abnormality related to an initiator port means a failure of a node having the initiator port, a failure of the initiator port, a failure of a path connected to the initiator port, the performance of the initiator port being below a threshold value, or the performance of a path connected to the initiator port being below a threshold value. The storage system according to claim 1.
8. An abnormality related to the initiator port of the first node means that the performance of the initiator port is below a threshold value, or the performance of a path connected to the initiator port is below a threshold value. The second node is a node whose resource utilization rate is below a threshold value. The storage system according to claim 3.
9. For each of the one or more nodes, an abnormality related to the initiator port of the node means that the performance of a path connected to the initiator port is below a threshold value. when the resource utilization rate related to the node is below a threshold value, the node dynamically generates a path connected to the initiator port and selects the generated path. The storage system according to claim 1.
10. When an abnormality related to the initiator port of a node or another node is detected and a path connected to the initiator port of the node needs to be selected, the path selected by the node is a path dynamically generated by the node. The storage system according to claim 1.
11. When the abnormality related to the initiator port is resolved, the dynamically generated path is deleted by the node that dynamically generated the path. The storage system according to claim 10.
12. When a plurality of paths are connected to the initiator port of a node, when the node transmits the command to the second storage system, the node selects a path in a round-robin manner from the plurality of paths. The storage system according to claim 1.
13. When a plurality of paths are connected to the initiator port of a node, when the node transmits the command to the second storage system, the node identifies the statistically most excellent path among the plurality of paths based on the statistics for each of the plurality of paths and selects the identified path. The storage system according to claim 1.
14. When a node sends a command for remote copy in a remote copy pair to the second storage system, unless an abnormality related to the initiator port of the node is detected, a path connected to the initiator port is selected from a path group associated with the remote copy pair and including the path connected to the initiator port. The storage system according to claim 1.
15. A path selection method performed by a storage system as a first storage system communicating with a second storage system and configured by a plurality of nodes, when any one or two or more nodes among the plurality of nodes send a command for remote copy in a remote copy pair to the second storage system, select a path connected to the initiator port by the node unless an abnormality related to the initiator port of the node is detected, send the command by the node via the selected path, there are a plurality of paths, and for each of the plurality of paths, the path is a path that communicably connects any one of the plurality of initiator ports of the plurality of nodes and any one of the one or more target ports of the second storage system, for each of the one or two or more nodes, the node has a first VOL, the remote copy pair is a pair composed of the first VOL and a second VOL among the one or more VOLs of the second storage system, Path selection method.
Citation Information
Patent Citations
Storage device system and data transfer method
JP2003296290A
Storage device system and data replication method
JP2004145855A
Storage control apparatus, storage system, storage control method, and storage control program
JP2015162091A
Storage system and data processing method
JP2022191854A
Computer system and management method for computer system
WO2016194096A1