Storage system for remote replication and path selection method in remote replication
By establishing multiple paths between storage systems and detecting exceptions before sending remote copy commands, and selecting an exception-free path for data transmission, the problem of performance degradation during remote copying is solved, and more stable data transmission is achieved.
Patent Information
- Application Number
- CN202411247038.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-09-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During remote replication, the path-selected storage system results in performance degradation, especially in the event of a node failure or a port failure.
By establishing multiple paths between the first storage system and the second storage system, each path is connected to a target port by an initiator port, the node detects an exception at the initiator port before sending a remote copy command, and if there is no exception, the path connected to the port is selected for data transmission.
It effectively suppresses the degradation of remote replication performance, ensuring that the node failure or port failure can be quickly detected and selected alternate paths, and maintains the stability of data transmission.
Smart Images

Figure CN120010756A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to remote replication between storage systems. Background Art
[0002] For example, Patent Document 1 discloses a technique related to remote copying.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: WO2016 / 194096 Summary of the invention
[0006] Problems to be solved by the invention
[0007] In remote replication from a PVOL (primary VOL) in a primary storage system (storage system in a primary site) to an SVOL (secondary VOL) in a secondary storage system (storage system in a secondary site), data is transferred via a selected path from among a plurality of paths between the PVOL and the SVOL. In remote replication, there are push-type remote replication (remote replication performed in response to writing from a primary storage system to a secondary storage system) and pull-type remote replication (remote replication performed in response to reading from a secondary storage system to a primary storage system), but in push-type remote replication, the primary storage system can select a path, and in pull-type remote replication, the secondary storage system can select a path. Hereinafter, the storage system that selects a path is referred to as a "first storage system", and the storage system that communicates with the first storage system for remote replication is referred to as a "second storage system".
[0008] It is conceivable to select the paths in round-robin scheduling. Since each path is selected equally, it is expected that the load can be distributed and / or the failure of a port can be detected quickly when the failure of the port occurs.
[0009] In addition, as the first storage system, a storage system composed of a plurality of storage nodes that is the basis of SDS (Software Defined Storage) can be used. These storage nodes are located in, for example, an on-premises environment or a cloud environment. The storage node (hereinafter referred to as a node) is, for example, a general-purpose computer and has a VOL (logical volume).
[0010] A certain node in the first storage system selects a path in the round-robin scheduling, but as a result, data transmission between the certain node and another node is required in remote replication according to the selected path, so the performance of remote replication may be reduced.
[0011] Such a problem may exist regardless of whether the first storage system is a secondary storage system or a primary storage system. In addition, such a problem may exist regardless of whether the remote replication is synchronous remote replication (a remote replication that returns a completion response to a write request when the data written to the PVOL accompanying the write request is written to the SVOL) or asynchronous remote replication (a remote replication that returns a completion response to a write request even if the data written to the PVOL accompanying the write request is not written to the SVOL).
[0012] Means for solving problems
[0013] There are multiple paths between the first storage system and the second storage system. For each of the multiple paths, the path is a path that connects one of the multiple initiator ports of the multiple nodes constituting the first storage system to one of the one or more target ports of the second storage system in a communicative manner. For each of one or more than two of the multiple nodes, the node has a first VOL that forms a remote replication pair with a second VOL of the one or more VOLs of the second storage system. When a node sends a command for remote replication under a remote replication pair to the second storage system, the node selects a path connected to the initiator port and sends the command via the selected path unless an abnormality related to the initiator port of the node is detected.
[0014] Effects of the Invention
[0015] According to the present invention, it is possible to suppress a decrease in the performance of remote replication performed between a second storage system and a first storage system composed of a plurality of nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram showing the outline of the first embodiment of the present invention.
[0017] Figure 2 This is a diagram showing an example of the physical configuration of the entire system.
[0018] Figure 3 This is a schematic diagram showing an overview of the remote copy configuration.
[0019] Figure 4 This is a schematic diagram showing an overview of I / O request processing.
[0020] Figure 5 This is a schematic diagram showing an overview of recovery processing from a node failure.
[0021] Figure 6 This is a schematic diagram showing an overview of path selection when a node fails.
[0022] Figure 7 This is a diagram showing an example of data and programs stored in the memory.
[0023] Figure 8 This is a diagram showing an example of a system configuration management table.
[0024] Fig. 9 This is a diagram showing an example of a pairing configuration management table.
[0025] Fig.10 This is a diagram showing an example of a path management table.
[0026] Fig.11 This is a diagram showing the flow of remote copy processing.
[0027] Fig.12 This is a diagram showing the flow of the path selection process.
[0028] Fig.13 This is a diagram showing the flow of path selection processing when a priority path fails.
[0029] Fig.14 This is a diagram showing the flow of path selection processing when the performance of the priority path is insufficient.
[0030] Fig.15 It is a schematic diagram showing the outline of the second embodiment of the present invention.
[0031] Description of Reference Numerals
[0032] 201 Site
[0033] 210 nodes DETAILED DESCRIPTION
[0034] In the following description, an "interface device" may be one or more communication interface devices. The one or more communication interface devices may be one or more communication interface devices of the same type (e.g., one or more NICs (Network Interface Cards)), or two or more communication interface devices of different types (e.g., a NIC and an HBA (Host Bus Adapter)).
[0035] In the following description, "memory" is an example of one or more storage devices, that is, one or more storage devices, typically a main storage device. At least one storage device in the memory may be a volatile storage device or a non-volatile storage device.
[0036] In addition, in the following description, "permanent storage device" may be an example of one or more storage devices, i.e., one or more permanent storage devices. Permanent storage devices may typically be non-volatile storage devices (e.g., auxiliary storage devices), and specifically, may be, for example, HDD (Hard Disk Drive), SSD (Solid State Drive), or NVMe (Non-Volatile Memory Express) drives.
[0037] In addition, in the following description, "processor" can be more than one processor device. At least one processor device can typically be a microprocessor device such as a CPU (Central Processing Unit), or it can be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device can be a single core or a multi-core. The at least one processor device can be a processor core. At least one processor device can also be a generalized processor device such as a hardware circuit that performs part or all of the processing (for example, an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device) or an ASIC (Application Specific Integrated Circuit)).
[0038] In addition, in the following description, sometimes the information that is output for input is expressed as "xxx table", but the information can be data of any structure (for example, it can be structured data or unstructured data), or it can be a learning model represented by a neural network, genetic algorithm, or random forest that generates output for input. Therefore, "xxx table" can be called "xxx information". In addition, in the following description, the composition of each table is an example, and a table can be divided into two or more tables, or all or part of two or more tables can be one table.
[0039] In addition, in the following description, "program" is sometimes used as the subject to explain the processing, but the program is executed by the processor, and the specified processing is performed appropriately using the storage device and / or interface device, so the subject of the processing can also be the processor (or a device such as a controller having the processor). The program can also be installed from a program source to a device such as a computer. The program source can be, for example, a program distribution server or a computer-readable (e.g., non-temporary) recording medium. In addition, in the following description, two or more programs can be implemented as one program, and one program can be implemented as two or more programs.
[0040] In addition, in the following description, when the same elements are not distinguished for description, the common parts of the reference numerals are used, and when the same elements are distinguished for description, the reference numerals or element identifiers are sometimes used. For example, for PVOL, an identifier such as "PVOL1" is sometimes used to distinguish it from other PVOLs.
[0041] Several embodiments are described below.
[0042] <First Embodiment>
[0043] Figure 1 It is a schematic diagram showing the outline of the first embodiment of the present invention.
[0044] The primary site 201P has a primary storage system 100P. The primary storage system 100P may be a so-called disk array system, but in the present embodiment, it is a system composed of a plurality of primary nodes 210P.
[0045] The secondary site 201S has a secondary storage system 100S. The secondary storage system 100S is a system composed of a plurality of secondary nodes 210S.
[0046] The node 210 is typically a general-purpose computer, but may be a device other than a general-purpose computer. The node 210 has a port 215, a VOL (logical volume) 102, and an SCS (Storage Control Software) 730. The VOL 102 is based on a permanent storage device in the node 210 or external to the node 210.
[0047] The plurality of master nodes 210P include, for example, master nodes 210P1 to 210P3. As VOL102 in the primary storage system 100P, there are PJVOL1, PJVOL1(S), PVOL1, and PVOL1(S). "PJVOL" is a JVOL located at the primary site 201P. "JVOL" is a VOL storing a JNL (journal). "JNL" includes data of a replication object and metadata thereof. The metadata in the JNL includes a sequence number (SEQ#) or a write destination address of the data, which is a value for determining the order in which the data of the replication object is written. "PJVOL1(S)" is equivalent to a standby PJVOL1 in which redundant data (e.g., a replicated JNL) of the JNL stored in PJVOL1 is stored. "PVOL" is a primary VOL. "PVOL1(S)" is equivalent to a standby PVOL1 in which redundant data (e.g., a replicated data) of the data stored in PVOL1 is stored.
[0048] In addition, each of the plurality of master nodes 210P has a target port as the port 215 . “Tm” (m is a natural number) means the target port m. The master node 210P has one or more ports 215 .
[0049] The plurality of sub-nodes 210S include, for example, sub-nodes 210S1 to 210S3. As VOL102, there are SJVOL1, SJVOL1(S), SVOL1, and SVOL1(S). "SJVOL" is a JVOL located in the sub-site 201S. "SJVOL1(S)" is equivalent to a standby SJVOL1 in which redundant data (e.g., a copy of JNL) of the JNL stored in SJVOL1 is stored. "SVOL" is a sub-VOL paired with PVOL. "SVOL1(S)" is equivalent to a standby SVOL1 in which redundant data (e.g., a copy of data) of the data stored in SVOL1 is stored.
[0050] In addition, each of the plurality of slave nodes 210S has an initiator port as the port 215 . “In” (n is a natural number) means the initiator port n. The slave node 210S has one or more ports 215 .
[0051] There are a plurality of paths 60 between the primary storage system 100P and the secondary storage system 100S. Each of the plurality of paths 60 is a path that connects a certain initiator port to a certain target port in a communicative manner. Figure 1 In the example, "P-ID" is a path ID. In addition, a "path" is a communication path for data or commands. For example, in the case of the iSCSI protocol, a path may be an iSCSI session.
[0052] In addition, there are multiple path groups. A path group is composed of two or more (or one) paths 60 and is associated with a remote replication pair. Figure 1 In the figure, "PG-ID" is the path group ID. In the present embodiment, two or more paths 60 with different path groups can be connected between the same initiator port and the same target port. Examples of "two or more paths 60 with different path groups" include path (PG-ID: 1, P-ID: 12) and path (PG-ID: 2, P-ID: 21). These paths are paths connecting I2 and T2. In other words, if the path groups are different, the paths connected between the same initiator port and the same target port can be processed as different paths. Hereinafter, the path group of "PG-ID: p" is sometimes referred to as "path group p", and the path 60 of "P-ID: q" is sometimes referred to as "path q".
[0053] For each of the two or more nodes 210, the processor of the node 210 executes SCS (Storage Control Software) 730 that controls I / O (Input / Output) with respect to VOL 102. Figure 1 In the figure, for the convenience of drawing, the SCS of the master node 210P is not shown, but the SCS 730 exists in the master node 210P and the slave node 210S in the same way.
[0054] For each of the primary storage system 100P and the secondary storage system 100S, there are multiple (or one) SCS groups. For each of the multiple SCS groups, the SCS group is composed of an activated SCS730, namely SCS(A) and one or more standby SCS730, namely SCS(S). One SCS(A) and one or more SCS(S) in the SCS group are configured on two or more different nodes 210. In the event of a failure of the node 210 having the SCS(A), a failover is performed in which one of the SCS(S) in the same SCS group replaces the SCS(A) and becomes the SCS(A). For example, in the secondary storage system 100S, SCSx(A) (x is a natural number) and SCSx(S) constitute an SCS group x, and in the event of a failure of the node 210 having the SCSx(A), a failover is performed from the SCSx(A) in the node 210 to one of the SCSx(S) in another node 210.
[0055] In this embodiment, using JNL, the port 215 of the secondary storage system 100S is the initiator port. That is, the remote replication performed in this embodiment is a pull-type asynchronous remote replication. Remote replication is performed on each remote replication pair (VOL pair). PVOLy (y is a natural number) and SVOLy constitute a remote replication pair y. According to Figure 1 In the example shown, remote replication pair 1 is formed by PVOL1 and SVOL1.
[0056] Since it is a pull-type asynchronous remote copy, the secondary storage system 100S sends a RDJNL command (a command to read JNL) for remote copy in a remote copy pair. In the RDJNL command, for example, the SEQ# of the latest JNL among the unreceived JNLs is specified, and in response to the RDJNL command, the JNL is received from the primary storage system 100P, the received JNL is stored in SJVOL1, and the JNL stored in SJVOL1 is reflected in SVOL1 (the data in the JNL is written into SVOL1).
[0057] Since the secondary storage system 100S sends the RDJNL command, in the present embodiment, the secondary storage system 100S executes the path selection program 720, and the path selection program 720 selects the path 60 used in sending the RDJNL command. The path selection program 720 is configured in each secondary node 210S. Sometimes, the path selection program 720 in the secondary node 210Sz (z is a natural number) is called "path selection program z".
[0058] When the slave node 210S sends the RDJNL command for remote replication under the remote replication pairing to the primary storage system 100P, as long as no abnormality related to the initiator port of the slave node 210S is detected, the path 60 connected to the initiator port is selected, and the RDJNL command is sent via the selected path 60. Specifically, for example, the slave node 210S that selects the path 60 is the node 210S having the SCS(A) that sends the RDJNL command. When the SCS1(A) of the slave node 210S1 sends the command for remote replication under the remote replication pairing 1 to the primary storage system 100P, as long as no abnormality related to the initiator port I1 of the slave node 210S1 is detected, the path selection program 1 selects the path 11 connected to the initiator port I1 from the path group 1 associated with the remote replication pairing 1, and the SCS1(A) sends the RDJNL command via the selected path 11.
[0059] Hereinafter, this embodiment will be described in detail.
[0060] Figure 2 It is a diagram showing an example of the physical configuration of the storage system 101 .
[0061] There are a plurality of sites 201. The sites 201 are connected to each other via a network 202 so as to be able to communicate. The network 202 is, for example, a WAN (Wide Area Network), but is not limited to a WAN. The site 201 is a data center or the like, and includes a plurality of (or one) nodes 210.
[0062] The node 210 may be a general-purpose computer. The node 210 includes, for example, one or more processor packages 213 including a processor 211 and a memory 212, one or more drivers 214, and one or more ports 215. These components are connected via an internal bus 216. The driver 214 is an example of a permanent storage device.
[0063] The processor 211 is, for example, a CPU (Central Processing Unit), and performs various processes.
[0064] The memory 212 is typically a volatile memory, and stores control information or data required to realize the functions of the node 210. The memory 212 stores, for example, a program executed by the processor 211. The driver 214 stores various data, programs, and the like.
[0065] The port 215 is connected to a network 220 in the site 201, and the node is communicatively connected to other nodes 210 in the site 201 via the network 220. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.
[0066] Note that the physical configuration of the system is not limited to the above configuration. For example, the network 202 and / or 220 may be redundant. In addition, for example, the network 220 may be separated into a management network and a storage network, and the connection standard may be Ethernet (registered trademark), Infiniband or wireless, and the connection topology is not limited to Figure 2 In addition, for example, the driver 214 may be a structure independent of the node 210.
[0067] Figure 3 This is a schematic diagram showing an overview of the remote copy configuration.
[0068] Multiple remote replication pairs are constructed between the primary site 201P and the secondary site 201S. Specifically, for example, two consistency groups 401a and 401b are constructed between the primary site 201P and the secondary site 201S. The consistency group 401 is composed of multiple (or one) remote replication paired VOLs 102, and in the consistency group 401, multiple PVOLs are copied to SVOLs in a consistent state. More specifically, for example, in the consistency group 401, the updated differential data up to the same moment for multiple PVOLs 102 are copied to multiple SVOLs. In addition, the control (consistency control) of the consistency group 401 is managed by the PJVOL. In the PJVOL, the updated differential data of multiple (or one) PVOLs are stored together with metadata such as the time of writing. When the primary site 201P transmits the data of the PVOL to the secondary site 201S, the updated differential data up to that moment in the updated differential data written to the PJVOL is transmitted to the secondary site 201S. This makes it possible to copy data to the SVOL while maintaining the consistency of update timings among a plurality of PVOLs.
[0069] For example, according to the consistency group 401a, while maintaining the matching of PVOL1 and 2 located at the primary node 210P1, data is copied to SVOL1 and 2 located at the secondary node 210S1 via PJVOL1 and SJVOL1. According to the consistency group 401b, while maintaining the matching of PVOL3 located at the primary node 210P2 and PVOL4 located at the primary node 210P3, data is copied to SVOL3 located at the secondary node 210S2 and SVOL4 located at the secondary node 210S3 via PJVOL2 located at the primary node 210P2, PJVOL3 located at the primary node 210P3, SJVOL2 located at the secondary node 210S2, and SVOVOL3 located at the secondary node 210S3. PJVOL and SJVOL may not necessarily correspond to 1:1 (for example, 1:many, many:1, or many:many), and PJVOL may be an area on the memory 212.
[0070] It can be seen from the above specific structure that the consistency group 401 can be composed of the VOL 102 in a specific node 210 in the site 201, and can also be composed of the VOL 102 in multiple nodes 210 located in the site 201.
[0071] Figure 4 This is a schematic diagram showing an overview of I / O request processing.
[0072] First, the application 502 running on the host 51 issues a write request to the master node 210P1 specifying PVOL1. The master node 210P1 receiving the write request writes the data A and B associated with the write request to PVOL1, and further writes JNL including the data A and B as updated differential data to PJVOL1.
[0073] Next, the master node 210P1 transmits the JNL (updated differential data) written to PJVOL1 to SJVOL1 and SJVOL1(S) of the slave site 201S. At this time, when multiple paths are established between the master site 201P and the slave site 201S, any path can be used to transmit the JNL. Normally, the master node 210P1 transmits the JNL to the slave node 210S1 that owns SVOL1 that is paired with PVOL1. However, when a failure occurs in the path with ownership, the master node 210P1 can also transmit the JNL to the slave node 210S2 that does not own the ownership. For example, when the master node 210P1 transmits the JNL to the slave node 210S2 that does not own the ownership, the slave node 210S2 transmits the received JNL to the slave node 210S1 that has ownership, and the slave node 210S1 writes the JNL to SJVOL1.
[0074] Next, the secondary node 210S1 writes the data A and B written to the JNL of SJVOL1 to SVOL1. Then, the data A and B written to SVOL1 are written to the drive 214a via the storage pool 504a. In the case where the drive 214a is configured as a DAS (Direct Attached Storage) in which the node 210 and the drive 214 are connected in a one-to-one manner, the JNL is written to the drive 214a mounted on the secondary node 210S1. By writing all the data copied to SVOL1 to the drive 214a of the secondary node 210S1 that owns SVOL1 in this way, when reading data from SVOL1 later, there is no need to read data from other nodes. As a result, inter-node transfer processing can be eliminated, and high-speed read processing can be achieved.
[0075] In addition, the storage pool 504 may be an area based on more than one drive 214. Storage functions such as thin provisioning, compression, or deduplication are provided, and processing of storage functions required for data written to the storage pool 504 is performed.
[0076] In order to protect the data from node failure when writing data to the driver 214a, the secondary node 210S1 also writes redundant data of the written data to the driver 214b of the secondary node 210S2. For the writing of redundant data, when the data protection strategy is replication, a copy of the data is written to the driver 214b as redundant data. On the other hand, when the data protection strategy is erasure coding, parity is calculated based on the data, and the calculated parity is written to the driver 214b as redundant data.
[0077] In addition, although not shown in the figure, the master node 210P1 transmits the data of the write object written to PVOL1 to the master node 210P2 (redundancy), and the master node 210P2 receives the data and writes it to PVOL1(S). In addition, the master node 210P2 writes the JNL to PJVOL1(S). The JNL written to PJVOL1(S) may be the JNL transmitted from the master node 210P1 or the JNL generated based on the data written to PVOL1(S). In this way, PVOL1(S) which is a copy of PVOL1 and PJVOL1(S) which is a copy of PJVOL1 are maintained (refer to Figure 1 ).
[0078] Figure 5 This is a schematic diagram showing an overview of recovery processing from a node failure.
[0079] SCS730 operates in slave nodes 210S1, 210S2, and 210S3. The slave node 210S has an SCS(A) and an SCS(S) corresponding to the SCS(A) in another slave node 210S (an SCS(S) in an SCS group including the SCS(A) in another slave node 210S). For example, the slave node 210S1 has SCS1(A) and SCS3(S), the slave node 210S2 has SCS2(A) and SCS1(S), and the slave node 210S3 has SCS3(A) and SCS2(S). SCSx(A) and SCSx(S) belong to SCS group x (a redundant group of SCSx), and SCSx(S) is not limited to one, but may exist in multiple numbers.
[0080] use Figure 5 The specific example shown is used to illustrate the recovery process from a node failure.
[0081] In order to inherit the remote replication pairing information of the secondary node 210S1, the secondary node 210S2 copies and holds the configuration information of SVOL1 and SJVOL1 of the secondary node 210S1. In addition, the secondary node 210S2 stores the redundant data of the data written to the drive 214a of the secondary node 210S1 in the drive 214d. In addition, a path (communication path) is established between the secondary node 210S2 and the primary node 210P1.
[0082] For example, when the secondary node 210S1 stops due to a failure, the secondary node 210S2 that detects the failure of the secondary node 210S1 transfers the processing of SCS1(A) of the secondary node 210S1, and SCS1(S) becomes SCS1(A). The secondary node 210S2 communicates with the primary node 210P1 and continues the remote replication processing between PVOL1 and SVOL1. That is, failover is performed from the SCS1(A) of the secondary node 210S1 to the SCS1(S) of the secondary node 210S2. As a result, even if a node failure occurs in one of the secondary nodes 210S of the secondary site 201S, the other secondary node 210S of the secondary site 201S can continue the remote replication from the primary site 201P.
[0083] As described above, in this embodiment, the path selection program 720 in the slave node 210S having the SCS (A) that sends the RDJNL command selects the path for sending the RDJNL command. Figure 1 as well as Figure 6 , an example of path selection is described. For example, for remote copy pair 1 (pair of PVOL1 and SVOL1), at least one of the following path selections (1) to (8) is performed.
[0084] (1) In order to prevent communication between slave nodes 210S during the transmission of the RDJNL command, the path selection program 720 in the slave node 210S having the SCS (A) that transmits the RDJNL command selects a path connected to the initiator port and transmits the RDJNL command via the selected path, unless an abnormality related to the initiator port of the slave node 210S is detected. That is, the path connected to the initiator port of the slave node 210S having the SCS (A) that transmits the RDJNL command is preferentially selected. Specifically, for example, when SCS1 (A) transmits the RDJNL command to the remote replication pair 1, the path selection program 1 preferentially selects the path 11 connected to the initiator port I1 of the slave node 210S1 from the path group 1 associated with the remote replication pair 1.
[0085] When an abnormality related to I1 is detected, such as a failure of the slave node 210S1 having I1, a failure of I1, a performance of I1 being below a threshold, or a performance of the path 11 connected to I1 being below a threshold, a path connected to the initiator port of another slave node 210S is selected, and an RDJNL command is sent via the path. For example, at least one of the following (2) to (5) is selected for the remote replication pair 1.
[0086] (2) When the slave node 210S1 fails (refer to Figure 6 ), through the path selection program 2 of the secondary node 210S2, a path 12 (refer to Figure 6 thick solid arrow).
[0087] (3) When port I1 fails, path selection program 1 of slave node 210S1 selects path 12 connected to I2 of slave node 210S2 operating SCS1(S) from path group 1 associated with remote copy pair 1 (see Figure 1 single-dotted arrow).
[0088] (4) In (2) and / or (3), when the port of I2 fails, path 13 connected to I3 of another secondary node 210S3 that does not have SCS1(S) is selected from path group 1 associated with remote replication pair 1 through path selection program 1 or path selection program 2.
[0089] (5) Regarding I1, in the case of performance exhaustion, the path selection program 1 of the slave node 210S1 selects a path in the path group 1 that is connected to the initiator port of another slave node 210S (see Figure 1 In addition, “regarding I1, in the case of performance exhaustion” means, for example, that the performance of I2 is below the threshold or the performance of the path 11 connected to I1 is below the threshold. In addition, the “another secondary node 210S” mentioned in this paragraph may be a secondary node 210S whose resource utilization rate is below the threshold.
[0090] (6) In at least one of (1) to (5), when there is no selected path for the initiator port of the slave node 210S, the path selection program 720 of the slave node 210S having the initiator port dynamically generates a path connected to the initiator port and selects the generated path. For example, if there is no path connected to I1, path selection program 1 dynamically generates a path connected to I1. In addition, for example, if there is no path connected to I2, path selection program 1 or path selection program 2 dynamically generates a path connected to I2. When the abnormality related to the initiator port that caused the generation of the path is eliminated, the dynamically generated path can be deleted by the path selection program 720 that generated the path. That is, the dynamically generated path can temporarily (or permanently) belong to path group 1 (can be managed as a temporary (or permanent) path that can be used for remote replication of remote replication pair 1).
[0091] (7) When multiple paths in path group 1 are connected, path selection program 1 selects a path from the multiple paths by round-robin scheduling. For example, for remote replication pair 1, path 11 may be selected when a certain RDJNL command is sent, and path 14 may be selected when the next RDJNL command is sent. In response to the RDJNL command via path 14, JNL may be sent from PJVOL1(S) to slave node 210S1 via path 14 through master node 210P2 (see Figure 1 ).
[0092] (8) I1 When multiple paths in path group 1 are connected, path selection program 1 determines the path with the best statistics among the multiple paths (for example, the path with the best performance or the least number of errors) based on statistics for each of the multiple paths (for example, statistics on performance and number of errors), and selects the determined path.
[0093] The above path selection utilizes the configuration of SCS(A) and SCS(S) in the secondary storage system 100S. Specifically, for example, according to at least one of (2) to (5), the initiator port of the secondary node 210S other than the secondary node 210S1 is preferentially used for the RDJNL command transmission of the SCS(A) configured in the secondary node 210S for remote replication pairings other than remote replication pairing 1, so that when an abnormality related to the initiator port occurs, the abnormality is detected by the secondary node 210S. In other words, if the abnormality related to the initiator port is not detected by the secondary node 210S, the path connected to the initiator port can be used. That is, in the RDJNL command transmission of SCS1(A), even if the path via another secondary node 210S is not selected by polling scheduling, it is possible to know whether the path via the other secondary node 210S can be used before path selection. In this way, in the present embodiment, path selection can be performed to activate the configuration using the SCS (A) and SCS (S) in the secondary storage system 100S.
[0094] Figure 7 This is a diagram showing an example of data and programs stored in the memory 212 .
[0095] Information is read from the driver 214 to the memory 212. For example, the various tables, SCS 730, and path selection program 720 included in the control information table 710 are expanded on the memory 212 during the execution of the processing used by each, but otherwise, they are stored in a non-volatile storage area such as the driver 214 in order to prepare for power failures, etc. The control information table 710 includes a system configuration management table 711, a pair configuration management table 712, and a path management table 713.
[0096] Figure 8 This is a diagram showing an example of the system configuration management table 711.
[0097] The system configuration management table 711 includes a node configuration management table 810, a driver configuration management table 820, and a port configuration management table 830. In addition, each site 201 has a node configuration management table 810 for a plurality of nodes 210 existing in each site 201, and a node 210 has a driver configuration management table 820 and a port configuration management table 830 for a driver 214 in the node 210.
[0098] The node configuration management table 810 is provided for each site 201, and stores information indicating the configuration of the node 210 provided at the site 201 (such as the relationship between the node 210 and the driver 214). More specifically, the node configuration management table 810 stores information such as the node ID 811, the status 812, the CPU usage 815, the memory usage 816, the driver ID list 813, and the port ID list 814 for each node 210.
[0099] The node ID 811 is the ID of the node 210. The state 812 indicates the state of the node 210 (for example, "Normal", "Warning", or "Failure", etc.). The CPU usage 815 indicates the CPU usage of the node 210. The memory usage 816 indicates the memory usage of the node 210. The driver ID list 813 is a list of the IDs of the drivers 214 set in the node 210. The port ID list 814 is a list of the IDs of the ports 215 set in the node 210.
[0100] The driver configuration management table 820 is provided for each node 210, and stores information indicating the configuration of the driver 214 provided in the node 210. More specifically, the driver configuration management table 820 stores information such as a driver ID 821, a status 822, a BE band usage rate 824, a driver usage rate 825, and a size 823 for each driver 214.
[0101] The driver ID 821 is the ID of the driver 214. The status 822 indicates the status of the driver 214. The BE band usage rate 824 indicates the usage rate of the communication band (backend band) between the processor 211 and the driver 214. The driver usage rate 825 indicates the ratio of the used capacity to the capacity of the driver 214. The size 823 indicates the capacity of the driver 214.
[0102] The port configuration management table 830 is provided for each node 210, and stores information indicating the configuration of the port 215 provided in the node 210. More specifically, the port configuration management table 830 stores information such as a port ID 831, a state 832, a NW band usage rate 834, and an address 833 for each port.
[0103] Port ID 831 is the ID of port 215. Status 832 indicates the status of port 215. NW band usage rate 834 indicates the usage rate of the band of the network connected to port 215 (the value of NW band usage rate 834 may be an example of the performance of port 215). Address 833 indicates the address on the network assigned to port 215. The address may be in the form of IP (Internet Protocol), WWN (World Wide Name), MAC (Media Access Control) address, etc.
[0104] Fig. 9 It is a diagram showing an example of the pairing configuration management table 712 .
[0105] The pair configuration management table 712 is configured to include a VOL management table 910 , a pair management table 920 , and a JNL management table 930 .
[0106] The VOL management table 910 stores information indicating the configuration of the VOL 102. More specifically, the VOL management table 910 stores information such as a VOL ID 911, an owner node ID 912, a fallback destination node ID 913, a size 914, and an attribute 915 for each VOL 102.
[0107] VOL ID 911 is the ID of VOL 102. Owner node ID 912 is the ID of node 210 having ownership of VOL 102. Fallback destination node ID 913 is the ID of node 210 that inherits processing when node 210 having ownership of SVOL fails. Size 914 indicates the capacity of VOL 102.
[0108] Attribute 915 indicates the attribute of VOL 102. "NML_VOL" is a normal VOL not used in remote replication. "PVOL" is a primary VOL. "PJVOL" is a JVOL storing updated differential data of PVOL. Although not shown, SVOL and SJVOL are also included as attributes.
[0109] The pair management table 920 stores information indicating the configuration of remote copy pairs. More specifically, the pair management table 920 stores information such as a pair group ID 921, a PJVOL ID 922, a PVOL ID 923, a SVOL ID 924, a SJVOL ID 925, a PG-ID 927, and a status 926 for each consistency group.
[0110] Pairing group ID 921 is the ID of the consistency group. PJVOL ID 922 is a list of IDs of PJVOLs belonging to the consistency group. PVOL ID 923 is a list of IDs of PVOLs belonging to the consistency group. SJVOL ID 924 is a list of IDs of SJVOL 102JS belonging to the consistency group. SVOL ID 925 is a list of IDs of SVOL 102S belonging to the consistency group. PG-ID 927 is the ID of the path group of each remote command pairing belonging to the consistency group. Status 926 indicates the status of each remote copy pairing in the consistency group (for example, "PAIR", "COPY", "SUSPEND", etc.). "Pairing" is a state in which writes to PVOL 102P are periodically reflected to SVOL 102S. "Copy" is a state in initial replication. "Suspend" is a pairing interruption state (a state in which synchronization between PVOL 102P and SVOL 102S is not performed).
[0111] The JNL management table 930 stores information on the JNL. More specifically, the JNL management table 930 stores information such as a pairing group ID 931, a JNL ID 932, a P / SVOL ID 933, a P / SVOL address 934, a size 935, and a cache segment ID 936 for each JNL.
[0112] Pairing group ID 931 is the ID of the consistency group to which JNL belongs. JNL ID 932 is the ID of JNL. The ID of JNL is equivalent to SEQ#, which is, for example, a sequence number in a consistency group. That is, the ID of JNL indicates the order of writing, and the data in JNL is stored in SVOL 102S in the consistency group according to the order of the ID of JNL.
[0113] The P / SVOL ID 933 includes the ID of the PVOL 102P and the ID of the SVOL 102S in which the data is written in the JNL. The P / SVOL address 934 includes the storage destination address of the data in the PVOL 102P and the storage destination address of the data in the SVOL 102S in which the data is written in the JNL.
[0114] The size 935 indicates the size of the JNL. For example, one JNL includes one or more data. The cache segment ID 936 is the ID of the cache segment (an area set in the cache of the memory 212) where the data in the JNL is written.
[0115] Fig.10 It is a diagram showing an example of the path management table 713.
[0116] The path management table 713 is a table related to the path 60. The path management table 713 stores information such as PG-ID 1031, P-ID 1034, protocol information 1035, status 1036, average response 1037, local address (Local Address) 1032, and destination address (Destination Address) 1033 for each path 60.
[0117] The PG-ID 1031 is the ID of the path group to which the path 60 belongs. The P-ID 1034 is the ID of the path 60 .
[0118] The protocol information 1035 is information indicating the communication protocol of the path 60. The communication protocol of the path 60 may be iSCSI, FC (Fibre Channel), NVMe-oF (NVMe over Fabrics), or a vendor-specific original protocol.
[0119] The status 1036 indicates the status of the path 60. A specified value other than "normal" may indicate an abnormality associated with the path 60. For example, "failed" indicates a path failure.
[0120] The average response 1037 indicates the average response time in the communication via the path 60. The value of the average response 1037 may be an example of the performance of the path 60.
[0121] The local address 1032 indicates the address of the initiator port of the path 60. The destination address 1033 indicates the address of the target port of the path 60.
[0122] An example of processing performed in the present embodiment will be described below. In the following description, remote copy pair 1 is taken as an example. The slave node 210S that has ownership of SVOL1 in remote copy pair 1 is the slave node 210S1.
[0123] Fig.11 This is a diagram showing the flow of remote copy processing.
[0124] SCS1(A) issues a RDJNL command (S1101) and performs a path selection process (S1102).
[0125] SCS1(A) sends an RDJNL command to the primary storage system 100P via the path 60 selected in the path selection process of S1102 (S1103). In response to the RDJNL command, SCS1(A) receives JNL from the primary storage system 100P via the path 60 (S1104). SCS1(A) stores the received JNL in SJVOL1 (S1105).
[0126] SCS1(A) reflects the unreflected JNL in SVOL1 in the order of SEQ# (S1106). That is, the data in the unreflected JNL is stored in SVOL1.
[0127] Although details are omitted, in the remote copy process, the status 926 of the pair group including the remote copy pair 1 and values in other various tables are updated as appropriate.
[0128] Fig.12 is a diagram showing the flow of path selection processing. Figure 12 to Figure 14 In the description, remote replication pair 1 and path selection program 1 are taken as an example, but in the case where a failover occurs from SCS1(A) of the secondary node 210S1 to SCS1(S) of the secondary node 210S2 due to a node failure of the secondary node 210S1, the path selection process can be performed by path selection program 2. In addition, in order to take remote replication pair 1 as an example, the connection destination of the dynamically generated path in the primary storage system 100P is the target port (i.e., T1 or T2) of the primary node 210P having PVOL1 or PVOL1(S). The address of the target port can be determined based on the path management table 713.
[0129] The path selection program 1 identifies the path group ID: 1 corresponding to the remote copy pair 1 from the pair management table 920 (S1201). The path selection program 1 refers to the path management table 713 using the path group ID: 1 as a key, and determines which of the following (A) to (C) is applicable (S1202).
[0130] (A) There is no normal path (a path whose status 1036 is "normal") in the secondary node 210S1.
[0131] (B) The performance of the path of the slave node 210S1 (the path connected to the initiator port I1) is insufficient. In other words, the performance of the path to be preferentially selected is insufficient at the same node as SCS1 (A).
[0132] (C) Neither (A) nor (B) meets the requirements.
[0133] The determination of whether (A) is met may be other than the determination of whether the status 1036 of path 11 and path 14 (the path connected to I1 of the secondary node 210S1) are both "normal". If both path 11 and path 14 are not "normal", (A) is met. If (A) is met, the path selection program 1 performs the path selection process when the priority path fails (S1203).
[0134] Whether (B) is met can be determined by determining whether the average response 1037 of path 11 and path 14 (path connected to I1 of secondary node 210S1) is above the threshold (and / or whether the number of RDJNL commands sent per unit time exceeds the threshold). If the average response 1037 of path 11 and path 14 (path connected to I1 of secondary node 210S1) is above the threshold (and / or the number of RDJNL commands sent per unit time exceeds the threshold), (B) is met. If (B) is met, path selection program 1 performs path selection processing when priority path performance is insufficient (S1205).
[0135] If (C) is met, the path selection program 1 selects the normal path 11 or the normal path 14 connected to I1 (S1204). In S1204, the normal path may be selected from the polling scheduling library or the normal path with the smallest average response 1037 value may be selected.
[0136] Fig.13 This is a diagram showing the flow of path selection processing when a priority path fails.
[0137] The path selection program 1 determines whether the secondary node 210S2 having SCS1(S) has a normal path (S1301). This determination can be performed by the path selection program 1 collecting necessary information from the path selection program 2, or the path management table 713 of all the secondary nodes 210S is shared by each secondary node 210S and performed based on the shared table.
[0138] If the result of S1301 is true (S1301: Yes), the path selection program 1 selects one normal path from the normal paths connected to the initiator port I2 of the slave node 210S2 having SCS1(S) (S1302). In step S1302, the path 12 in the path group 1 is selected based on the path management table 713.
[0139] When the result of the determination in S1301 is false (S1301: No), the path selection program 1 selects a normal path from the normal paths connected to the initiator ports of the other slave nodes 210S (S1303). In addition, the "other slave nodes 210S" mentioned in this paragraph refers to a slave node 210S other than the slave node 210S1 and the slave node 210S2. In S1303, the path 13 in the path group 1 is selected based on the path management table 713.
[0140] In S1301 to S1303, in the case where there is no normal path, the path selection program 1 may also dynamically generate a path. The connection destination of the dynamically generated path may be determined based on the destination address 1033 of the path management table 713. Specifically, for the remote replication pair 1, the connection destination of the dynamically generated path may be T1 of the master node 210P1 having PVOL1, or T2 of the master node 210P2 having PVOL1(S).
[0141] In addition, before the determination of S1301, the determination of the node port may be performed to determine whether the state 832 of I1 as the secondary node 210S1 is "normal" (normal). S1301 may also be performed when the result of the node port determination is false. When the result of the node port determination is true, the path selection program 1 may dynamically regenerate a path from I1 to T1 or T2 for the remote replication pair 1.
[0142] Fig.14 This is a diagram showing the flow of path selection processing when the performance of the priority path is insufficient.
[0143] The path selection program 1 determines whether the resource utilization rate of the secondary node 210S1 is below the threshold value (S1401). With respect to the remote replication pair 1, the "resource utilization rate" mentioned in this paragraph may be at least one of the CPU utilization rate 815 of the secondary node 210S1, the memory utilization rate 816 of the secondary node 210S1, the statistical value (e.g., average value) of the BE band utilization rate 824 of the driver 214 of the secondary node 210S1, and the NW band utilization rate 834 of I1.
[0144] If the result of S1401 is true (S1401: Yes), the path selection program 1 dynamically generates a path connected to I1 and selects the path (S1402). That is, if the performance of paths 11 and 14 is insufficient, but the resource utilization rate of the secondary node 210S1 is below the threshold, a path connected to I1 is dynamically generated.
[0145] When the determination result of S1401 is false (S1401: No), the path selection program 1 determines whether the resource utilization rate of other subnodes 210S is below the threshold value (S1403). The "other subnodes 210S" mentioned in this paragraph may be each subnode 210S other than the subnode 210S1. In addition, in this paragraph, the "resource utilization rate" of other subnodes 210S may be the same as the "resource utilization rate" of the subnode 210S1.
[0146] If the result of S1403 is false (S1403: No), the path selection program 1 selects path 11 or path 14 from I1 (S1404). That is, the performance of path 11 and path 14 is insufficient, but if the resource utilization rate of other slave nodes 210S exceeds the threshold, it is possible that the advantage of selecting the path connected to other slave nodes 210S is low, so path 11 or path 14 is selected.
[0147] When the determination result of S1403 is true (S1403: Yes), the path selection program 1 determines whether a normal path belonging to the path group 1 is connected to another secondary node 210S based on the path management table 713 (S1405).
[0148] If the result of the determination in S1405 is false (S1405: No), the path selection program 1 causes the path selection program 720 of the other slave node 210S to dynamically generate a path connected to the initiator port of the other slave node 210S (S1406), and selects the path (S1407). That is, if the performance of the path 11 and the path 14 is insufficient and the resource utilization rate of the other slave node 210S is below the threshold, but if the path belonging to the path group 1 is not connected to the other slave node 210S, a path connected to the initiator port of the other slave node 210S is dynamically generated.
[0149] When the determination result of S1405 is true (S1405: Yes), the path selection program 1 selects the normal path 12 or the path 13 connected to the other secondary node 210S in the path group 1 (S1407).
[0150] <Second Embodiment>
[0151] Here, the second embodiment of the present invention will be described. At this time, the differences from the first embodiment will be mainly described, and the description of the same points as the first embodiment will be omitted or simplified.
[0152] Fig.15 It is a schematic diagram showing the outline of the second embodiment.
[0153] In the first embodiment, pull-type asynchronous remote replication is adopted, but in the second embodiment, push-type synchronous remote replication is adopted. That is, the master node 210P has a path selection program 720, and the path selection program 720 of the master node 210P having SCS (A) performs path selection processing with respect to the remote replication pairing (for example, the pairing of PVOL1 and SVOL1). With respect to the remote replication pairing, the port of the master node 210P is the initiator port, and the port of the secondary node 210S is the target port. The WRJNL command (JNL write command) is sent from the SCS (A) of the master node 210P to the secondary storage system 100S via the path selected by the path selection processing. As a result, it is expected that the reduction in performance involved in the processing of push-type remote replication can be suppressed.
[0154] Although one embodiment of the present invention has been described above, this is an example for explaining the present invention and is not intended to limit the scope of the present invention to this embodiment. The present invention can also be implemented in various other forms.
[0155] In addition, the above description can be summarized as follows. The following summary may include supplementary descriptions of the above descriptions and descriptions of modified examples. In addition, according to the first and second embodiments, the present invention can be applied regardless of whether the remote replication is a pull type or a push type and whether it is a synchronous remote replication or an asynchronous remote replication. Therefore, in the following summary, the first storage system can be one of the secondary storage system 100S and the primary storage system 100P, and the second storage system can be the primary storage system 100P when the first storage system is the secondary storage system 100S, and can be the secondary storage system 100S when the first storage system is the primary storage system 100P.
[0156] The storage system as the first storage system communicating with the second storage system each has a plurality of nodes having an initiator port, a memory, a processor, and a VOL. There are a plurality of paths between the first storage system and the second storage system. For each of the plurality of paths, the path is a path that connects one of the plurality of initiator ports of the plurality of nodes to one of the one or more target ports of the second storage system in a communicative manner. For each of one or more than two of the plurality of nodes, the node has a first VOL that forms a remote replication pair with a second VOL of the one or more VOLs of the second storage system. When the node sends a command for remote replication under the remote replication pair to the second storage system, the node selects a path connected to the initiator port and sends the command via the selected path unless an abnormality related to the initiator port of the node is detected. As a result, command transmission between nodes is not required in command transmission, so it is expected that the performance reduction of remote replication can be suppressed. In addition, "one or more than two of the plurality of nodes" may mean that the plurality of nodes may include so-called spare nodes other than "one or more than two nodes".
[0157] It may be that, for each of two or more nodes, the processor of the node executes the SCS that controls the I / O (Input / Output) with respect to the VOL. It may be that there are one or more SCS groups, and for each of the one or more SCS groups, the SCS group is composed of one active SCS, namely SCS (A) and one or more standby SCS, namely SCS (S). It may be that the one SCS (A) and the one or more SCS (S) are configured on two or more different nodes among the plurality of nodes. It may be that, in the event of a failure of a node having the SCS (A), a failover is performed in which the SCS (A) in the SCS group is replaced and a certain SCS (S) in the SCS group becomes the SCS (A). It may be that the node that selects the path for the remote replication pair for performing remote replication is the node having the SCS (A) that sends the command for the remote replication. Therefore, for each of the two or more nodes, the initiator port of the node is used in the command sent from the SCS (A) in the node, so the presence or absence of abnormality related to the initiator port of each node can be known. Therefore, when a node having SCS (A) does not select a path connected to the node and a path connected to other nodes by polling scheduling, it can be known whether the path connected to other nodes can be used before path selection.
[0158] For example, one of the two or more nodes may be the first node (e.g., the secondary node 210S1). One of the SCS groups other than the first SCS group including the first SCS (A) having the first node (e.g., SCS1 (A)) may be the second SCS group. A node having the second SCS (A) in the second SCS group (e.g., SCS2 (A) or SCS3 (A)) may be the second node (e.g., the secondary node 210S2 or the secondary node 210S3). A remote replication pairing for performing remote replication based on the first SCS (A) may be the first remote replication pairing (e.g., remote replication pairing 1). When an abnormality related to the initiator port (e.g., I1) of the first node is detected, the node having the first SCS (A) may select a second path (e.g., path 12 or path 13) connected to the initiator port, and send a command for remote replication under the first remote replication pairing via the second path, as long as an abnormality related to the initiator port (e.g., I2 or I3) of the second node is not detected by the second node. This can avoid an abnormality when a path to connect to the second node is selected, and can expect to suppress a decrease in the performance of remote copying.
[0159] It should be noted that the second node may be a node (e.g., a secondary node 210S2) having the first SCS(S) (e.g., SCS1(S)) in the first SCS group. Thus, it can be expected that the first SCS(S) will inherit the processing instead of the first SCS(A) and continue the remote replication without the transmission of commands between the second node and other nodes. Specifically, for example, the second node may have a VOL in which redundant data of data stored in the first VOL constituting the first remote replication pairing is stored, that is, the standby first VOL. The command for remote replication under the first remote replication pairing and sent via the second path may be a command for remote replication between the standby first VOL and the second VOL in the first remote replication pairing.
[0160] It may be that one of the SCS groups other than the first SCS group and the second SCS group is the third SCS group. It may be that a node having a third SCS(A) (for example, SCS3(A)) in the third SCS group is a third node (for example, sub-node 210S3). It may be that, in the case where an abnormality related to the initiator port of the first node is detected, if the second node detects an abnormality related to the initiator port of the second node in the node having the first SCS(A), then as long as the abnormality related to the initiator port of the third node is not detected by the third node, a third path connected to the initiator port is selected, and a command for remote replication under the first remote replication pairing is sent via the third path. Thus, it is possible to avoid the presence of an abnormality when the path connected to the third node is selected, and therefore, it is possible to expect to suppress the performance degradation of remote replication.
[0161] The abnormality related to the initiator port may be a failure of a node having the initiator port, a failure of the initiator port, a failure of a path connected to the initiator port, a performance of the initiator port being below a threshold, or a performance of a path connected to the initiator port being below a threshold.
[0162] The abnormality related to the initiator port of the first node may be that the performance of the initiator port is below a threshold or the performance of the path connected to the initiator port is below a threshold. In this case, the second node may be a node whose resource usage rate is below the threshold. That is, when an abnormality related to the initiator port of the first node is detected and there is a second node as another node whose resource usage rate is below the threshold, the path connected to the initiator port of the second node is selected. Since there is surplus resource related to the second node connected to the selected path, it is expected that the performance degradation of remote replication can be suppressed.
[0163] For each of one or more nodes, the abnormality associated with the initiator port of the node is that the performance of the path connected to the initiator port is below a threshold value. When the utilization rate of the resource associated with the node is below the threshold value, the node can dynamically generate a path connected to the initiator port and select the generated path. Thus, even if the performance of the path connected to the node is reduced, when there is surplus resource associated with the node, a path connected to the node is dynamically generated and selected, so it is expected that the performance reduction of remote replication can be suppressed.
[0164] It is possible that, in the case where an abnormality related to the initiator port of a node or other nodes is detected and a path connected to the initiator port of the node needs to be selected (and further, in the case where the path connected to the initiator port of the node (for example, a normal and sufficiently performing path) is insufficient), the path selected by the node is a path dynamically generated by the node. Thus, it can be expected that the performance degradation of remote replication can be suppressed. In addition, when the abnormality related to the initiator port is eliminated, the dynamically generated path can be deleted by the node that dynamically generated the path. Thus, it can be expected that the management burden or excess resources associated with the remaining unnecessary paths can be avoided.
[0165] When a node's initiator port is connected to multiple paths, when the node sends a command to the second storage system, a path can be selected from the multiple paths through polling scheduling. This avoids the load concentration of the path while eliminating the need for command transmission between nodes, and thus can further suppress the performance degradation of remote replication.
[0166] It may be that when a node's initiator port is connected to multiple paths, the node determines the path with the best statistics among the multiple paths based on statistics for each of the multiple paths when sending the command to the second storage system, and selects the determined path. In this way, the selection of the best path can be maintained without requiring command transmission between nodes, and it is expected that the performance degradation of remote replication can be further suppressed.
[0167] It may be that, when a node sends a command for remote replication under a remote replication pairing to a second storage system, as long as an abnormality related to the initiator port of the node is not detected, a path connected to the initiator port is selected from a path group associated with the remote replication pairing and including a path connected to the initiator port. Different path groups can be associated with different remote replication pairings, so that the paths between the same initiator port and the target port can be managed as different paths. For example, as paths connecting I2 and T2, path 12 belonging to path group 1 associated with remote replication pairing 1 and path 21 belonging to path group 2 associated with another remote replication pairing (pairing of PVOL5 and SVOL7) can be managed. In addition, the paths that are options are narrowed down for each remote command pairing. As a result, it can be expected that the performance degradation of remote replication can be suppressed.
[0168] In the above description, whether the first storage system and the second storage system are primary storage systems or secondary storage systems may differ according to the remote replication pairing. That is, the storage system with PVOL may be the primary storage system, and the storage system with SVOL may be the secondary storage system.
Claims
1. A storage system is a storage system that is a first storage system communicating with a second storage system, The storage system includes a plurality of nodes each having an initiator port, a memory, a processor, and a VOL (logical volume). There are a plurality of paths, and for each of the plurality of paths, the path is a path that connects one of the plurality of initiator ports of the plurality of nodes to one of the one or more target ports of the second storage system in a communicative manner, For each of one or more than two nodes of the plurality of nodes, The node has a first VOL that forms a remote replication pair with a second VOL of the one or more VOLs of the second storage system. When the node sends a command for remote replication in the remote replication pairing to the second storage system, the node selects a path connected to the initiator port and sends the command via the selected path unless an abnormality related to the initiator port of the node is detected.
2. The storage system according to claim 1, For each of the two or more nodes, a processor of the node executes an SCS for controlling I / O for the VOL, wherein I / O refers to input / output and SCS refers to storage control software. The storage system has one or more SCS groups. For each of the one or more SCS groups, the SCS group is composed of an activated SCS, namely SCS (A), and one or more standby SCSs, namely SCS (S). The one SCS (A) and the one or more SCS (S) are configured at two or more different nodes among the plurality of nodes. When a node having the SCS (A) fails, a certain SCS (S) in the SCS group replaces the SCS (A) in the SCS group and becomes a failover of the SCS (A). The node that selects a path for the remote copy pair that performs remote copy is the node having the SCS (A) that sends the command for the remote copy.
3. The storage system according to claim 2, One of the two or more nodes is a first node, The first SCS group includes the first SCS (A) of the first node, and an SCS group other than the first SCS group is a second SCS group. The node having the second SCS (A) within the second SCS group is the second node, The remote replication pair that performs remote replication based on the first SCS (A) is a first remote replication pair. In the event of detecting an abnormality associated with the launcher port of the first node, the node having the first SCS (A) selects a second path connected to the launcher port and sends a command for remote replication under the first remote replication pairing via the second path, as long as the abnormality associated with the launcher port of the second node is not detected through the second node.
4. The storage system according to claim 3, The second node is a node having a first SCS(S) within the first SCS group.
5. The storage system according to claim 4, The second node has a VOL in which redundant data of data stored in the first VOL constituting the first remote replication pair is stored, that is, a standby first VOL. The command for remote replication in the first remote replication pair and sent via the second path is a command for remote replication between the standby first VOL and the second VOL in the first remote replication pair.
6. The storage system according to claim 4, One SCS group other than the first SCS group and the second SCS group is a third SCS group, A node having a third SCS (A) within said third SCS group is a third node, In the case where an abnormality related to the launcher port of the first node is detected, if the abnormality related to the launcher port of the second node is detected through the second node, the node having the first SCS (A) selects a third path connected to the launcher port and sends a command for remote replication under the first remote replication pairing via the third path, as long as the abnormality related to the launcher port of the third node is not detected through the third node.
7. The storage system according to claim 1, The abnormality related to the initiator port is a failure of a node having the initiator port, a failure of the initiator port, a failure of a path connected to the initiator port, a performance of the initiator port being below a threshold, or a performance of a path connected to the initiator port being below a threshold.
8. The storage system according to claim 3, The abnormality related to the initiator port of the first node is that the performance of the initiator port is below a threshold or the performance of the path connected to the initiator port is below a threshold, The second node is a node whose resource usage rate is below a threshold.
9. The storage system according to claim 1, For each of the one or more nodes, The abnormality related to the initiator port of the node is that the performance of the path connected to the initiator port is below the threshold. When the usage rate of the resource related to the node is below a threshold, the node dynamically generates a path connected to the initiator port and selects the generated path.
10. The storage system according to claim 1, When an abnormality related to the initiator port of a node or another node is detected and a path connected to the initiator port of the node needs to be selected, the path selected by the node is a path dynamically generated by the node.
11. The storage system according to claim 10, When the abnormality related to the initiator port is resolved, the dynamically generated path is deleted by the node that dynamically generated the path.
12. The storage system according to claim 1, When a plurality of paths are connected to the initiator port of the node, the node selects a path from the plurality of paths by round-robin scheduling when sending the command to the second storage system.
13. The storage system according to claim 1, When a plurality of paths are connected to the initiator port of a node, the node determines a path with the best statistics among the plurality of paths based on statistics for each of the plurality of paths when sending the command to the second storage system, and selects the determined path.
14. The storage system according to claim 1, When the node sends a command for remote replication under a remote replication pairing to the second storage system, as long as no abnormality related to the initiator port of the node is detected, the node selects a path connected to the initiator port from a path group associated with the remote replication pairing and including the path connected to the initiator port.
15. A path selection method, which is a path selection method performed by a storage system, wherein the storage system is a storage system that is a first storage system communicating with a second storage system and is a storage system composed of a plurality of nodes, When one of the plurality of nodes or one of the two or more nodes sends a command for remote replication in a remote replication pair to the second storage system, As long as no anomaly is detected with respect to the initiator port of the node, a path connected to the initiator port is selected through the node. The command is sent by the node via the selected path, There are a plurality of paths, and for each of the plurality of paths, the path is a path that connects one of the plurality of initiator ports of the plurality of nodes to one of the one or more target ports of the second storage system in a communicative manner, For each of the one or more nodes, the node has a first VOL, The remote copy pair is a pair consisting of the first VOL and a second VOL among one or more VOLs included in the second storage system.
Citation Information
Patent Citations
Computer system and management method for computer system
WO2016194096A1