Storage system and storage control method

By dynamically determining the need for redundancy and non-volatilization based on data type in each node of the storage system, the performance degradation associated with high write frequencies in asynchronous remote copy systems is mitigated, ensuring efficient and reliable data management.

JP2025088562AActive Publication Date: 2025-06-11HITACHI VANTARA LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023203339
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-06-11
Estimated Expiration
2043-11-30

AI Technical Summary

Technical Problem

In asynchronous remote copy systems, the high frequency of writing from cache to non-volatile media leads to performance degradation, and this issue is also present in other storage control systems with multiple nodes that perform data duplication.

Method used

Each node in the storage system determines, based on the type of data, whether redundancy and non-volatilization are required for the data in the cache segment, and controls the performance of these operations accordingly.

Benefits of technology

This approach reduces the writing frequency from cache to non-volatile media, thereby improving performance and maintaining data integrity even if data is lost due to node failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088562000001_ABST
    Figure 2025088562000001_ABST
Patent Text Reader

Abstract

To appropriately reduce the frequency of writing from a cache to a nonvolatile medium in a storage system composed of a plurality of nodes.SOLUTION: In each of a plurality of nodes constituting a storage system, when there is data to be stored in a segment allocated from a cache, the node determines, for the segment, whether redundancy (redundant data of data in the segment is transferred to another node) is required and whether non-volatilization (data in the segment is stored in a nonvolatile area) is required according to a type of the data, and controls, on the basis of a result of determination, whether to perform the redundancy of data in the segment and whether to perform the non-volatilization of data in the segment.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to memory control.

Background Art

[0002] As an example of memory control, there is remote copy. Regarding remote copy, for example, the technology disclosed in Patent Document 1 is known.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] As at least the secondary storage system among the primary storage system (the storage system at the primary site) and the secondary storage system (the storage system at the secondary site), SDS (Software Defined Storage) can be adopted. SDS is based on one or a plurality (typically a plurality) of storage nodes. Those storage nodes are, for example, in an on-premises environment or a cloud environment. The storage nodes (hereinafter referred to as nodes) are, for example, general-purpose computers and have a cache and a VOL (logical volume). The cache is typically in volatile memory, and the VOL is typically based on a persistent storage device.

[0005] The nodes that form the basis of SDS generally do not have a battery. For this reason, when a power failure of the node occurs, the data in the cache (typically volatile memory) of the node may be lost. To prevent such data loss, the node performs a data protection process of writing data from the cache to the VOL.

[0006] Specifically, for example, in asynchronous remote copy, as data stored in the cache, there are a JNL (Journal) and data to be copied from PVOL (Primary VOL) to SVOL (Secondary VOL). The JNL includes data to be copied (replication) and metadata of the data.

[0007] Assume that the secondary storage system includes a first and a second node. Assume that the first node has a first cache, a first JVOL (JNL VOL), a first memory evacuation area, and a first SVOL. Assume that the first JVOL, the first memory evacuation area, and the first SVOL are areas based on a persistent storage device inside or outside the first node. Assume that the second node has a second cache, a second JVOL, a second memory evacuation area, and a second SVOL. Assume that the second JVOL, the second memory evacuation area, and the second SVOL are areas based on a persistent storage device inside or outside the second node. Assume that the second SVOL is a mirror VOL of the first SVOL.

[0008] When the first node receives a JNL from the primary storage system, for example, the following processing is performed. · The first node writes the received JNL into the first cache, copies the data in the JNL into the first cache, and writes the JNL and the data into the first memory evacuation area. Also, for data redundancy, the first node transfers the JNL and the data stored in the first cache to the second storage node. The second node writes the JNL and the data into the second cache and writes the JNL and the data into the second memory evacuation area. This increases the possibility of restoring the JNL and the data even if the JNL and the data disappear from the first cache due to a power failure of the first node. · The first node writes the log to the first JVOL and then writes the JNL to the first JVOL. The first node writes the log to the first SVOL and then writes the data in the JNL to the first SVOL. Similarly, the second node writes the log to the first JVOL and then writes the JNL to the first JVOL. The second node writes the log to the second SVOL and then writes the data in the JNL to the second SVOL. As a result, the data is duplicated in the first SVOL and the second SVOL, and even if a failure occurs in one of the first and second nodes, the data can be restored from the other node.

[0009] However, in this process, the frequency of writing from the cache to the VOL is high in asynchronous remote copy, so there is concern about performance degradation of asynchronous remote copy.

[0010] The high frequency of writing from the cache to the non-volatile medium (for example, the non-volatile medium that is the basis of the VOL) may also be an issue for storage control other than asynchronous remote copy in a storage system composed of multiple nodes (a storage system in which data duplication is performed between nodes).

Means for Solving the Problem

[0011] In each of the multiple nodes constituting the storage system, when there is data to be stored in the segment secured from the cache, the node determines, according to the type of the data, whether redundancy (transferring redundant data of the data in the segment to another node) is required for the segment and whether non-volatilization (storing the data in the segment in the non-volatile area) is required for the segment, and controls whether to perform redundancy of the data in the segment and whether to perform non-volatilization of the data in the segment based on the result of the determination.

Effect of the Invention

[0012] According to the present invention, in a storage system composed of a plurality of nodes, the writing frequency from the cache to the non-volatile medium can be appropriately reduced.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14A

Figure 14B

Figure 15A

Figure 15B

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0014] In the following description, the "interface device" may be one or more communication interface devices. The one or more communication interface devices may be one or more of the same type of communication interface devices (for example, one or more NICs (Network Interface Cards)), or may be two or more different types of communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).

[0015] Also, in the following description, the "memory" is one or more memory devices which are an example of one or more storage devices, and typically may be a main memory device. At least one of the memory devices in the memory may be a volatile memory device or a non-volatile memory device.

[0016] Also, in the following description, the "persistent storage device" may be one or more persistent storage devices which are an example of one or more storage devices. The persistent storage device is typically a non-volatile storage device (for example, an auxiliary storage device), and specifically may be, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an NVMe (Non-Volatile Memory Express) drive.

[0017] Also, in the following description, the "processor" may be one or more processor devices. At least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may also be a processor device in a broad sense that performs part or all of the processing, such as a hardware circuit (e.g., FPGA (Field-Programmable Gate Array), CPLD (Complex Programmable Logic Device), or ASIC (Application Specific Integrated Circuit)).

[0018] Also, in the following description, in expressions such as "xxx table", information from which an output can be obtained for an input may be described. However, the information may be data of any structure (e.g., structured data or unstructured data), or may also be a learning model such as a neural network that generates an output for an input, a genetic algorithm, or a random forest. Therefore, "xxx table" can be referred to as "xxx information". Also, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.

[0019] In the following description, the "program" may be used as the subject to describe the processing. However, since the program performs the defined processing by being executed by a processor, appropriately using a storage device and / or an interface device, etc., the subject of the processing may be the processor (or a device such as a controller having the processor). The program may be installed from a program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0020] Also, in the following description, when describing without distinguishing between elements of the same kind, a common part of the reference signs is used, and when distinguishing between elements of the same kind, a reference sign or an identifier of the element may be used. For example, for PVOL, the reference sign may be used as in "PVOL102P1", or the identifier may be used as in "PVOL1".

[0021] FIG. 1 is a schematic diagram showing an overview of an embodiment of the present invention.

[0022] There are a host 51 and a primary storage system 100P at the primary site 201P. The host 51 may be a physical computer or a logical computer (e.g., a virtual machine). The primary storage system 100P may be a so-called disk array system, but in this embodiment, it is a system composed of a plurality of primary nodes 210P.

[0023] There is a secondary storage system 100S at the secondary site 201S. The secondary storage system 100S is a system composed of a plurality of secondary nodes 210S.

[0024] Node 210 is typically a general-purpose computer, but it may also be a device other than a general-purpose computer. Node 210 typically does not have a battery. Node 210 has a cache 55, a VOL (Logical Volume) 102, and a non-volatile area 103. The cache 55 is typically volatile memory. VOL 102 and the non-volatile area 103 are based on a persistent storage device inside or outside Node 210.

[0025] A plurality of positive nodes 210P include, for example, a first positive node 210P1 and a second positive node 210P2. As VOL 102, there are PJVOL102JP and PVOL102P. "PJVOL" is the JVOL at the positive site 201P. "JVOL" is the VOL in which the JNL (Journal) is stored. "JNL" includes the data to be copied and its metadata. The metadata in the JNL is a sequence number (SEQ#) which is a value for specifying the order in which the data to be copied is written, and the write destination address of the data. "PVOL" is the positive VOL.

[0026] A plurality of secondary nodes 210S include, for example, a first secondary node 210S1 and a second secondary node 210S2. As VOL 102, there are SJVOL102JS and SVOL102S. "SJVOL" is the JVOL at the secondary site 201S. "SVOL" is the secondary VOL that forms a pair with the PVOL.

[0027] In this embodiment, in each of the positive storage system 100P and the secondary storage system 100S, a redundancy flag 2 and a non-volatility flag 3 are provided for each segment (an example of a cache area) in the cache 55.

[0028] The redundancy flag 2 indicates the necessity of redundancy. When the value of the redundancy flag 2 is "on", redundancy (copying to other nodes) of the data in the segment in the cache 55 is performed. When the value of the redundancy flag 2 is "off", redundancy of the data in the segment in the cache 55 is not performed.

[0029] The non-volatile flag 3 indicates the necessity of non-volatilization. When the value of the non-volatile flag 3 is "on", non-volatilization of the data within the segment in the cache 55 (writing to the non-volatile area 103) is performed. When the value of the redundancy flag 2 is "off", non-volatilization of the data within the segment in the cache 55 is not performed.

[0030] According to the type of data in the segment, the values of the corresponding redundancy flag 2 and non-volatile flag 3 for that segment are respectively controlled.

[0031] For example, in this embodiment, the following processing is performed. In FIG. 1, PJVOL102JP1S in the second positive node 210P2 is the standby VOL (mirror VOL) of PJVOL102JP1A in the first positive node 210P1. PVOL102P1S in the second positive node 210P2 is the standby VOL of PVOL102P1A in the first positive node 210P1. SJVOL102JS1S in the second secondary node 210S2 is the standby VOL of SJVOL102JS1A in the first secondary node 210S1. SVOL102S2S in the second secondary node 210S2 is the standby VOL of SVOL102S1A in the first secondary node 210S1.

[0032] The first positive node 210P1 receives a write request specifying PVOL102P1A and writes the data to be written, which is associated with the write request, to the first segment of the cache 55P1. The first positive node 210P1 sets the values of the redundancy flag 2P1 and the non-volatility flag 3P1 of the first segment to "on", respectively. For this reason, the first positive node 210P1 transfers the data in the first segment to the second positive node 210P2 (redundancy) and writes it to the non-volatile area 103P1 (non-volatility). The second positive node 210P2 receives the data, writes it to the segment of the cache 55P2, and writes it to the non-volatile area 103P2. Although not shown, for the segment of the cache 55P2 (the segment to which the data from the first positive node 210P1 is written), the redundancy flag may be set to "off" and the non-volatility flag may be set to "on". Thereby, the redundancy of the data is skipped and the non-volatility of the data (writing the data to the non-volatile area 103P2) is performed.

[0033] The first positive node 210P1 generates a JNL including the data in the first segment of the cache 55P1 and writes the JNL to the second segment of the cache 55P1. The first positive node 210P1 sets the values of the redundancy flag 2P2 and the non-volatility flag 3P2 of the second segment to "on", respectively. For this reason, the first positive node 210P1 transfers the JNL in the second segment to the second positive node 210P2 (redundancy) and writes it to the non-volatile area 103P1 (non-volatility). The second positive node 210P2 receives the data, writes it to the segment of the cache 55P2, and writes it to the non-volatile area 103P2. Although not shown, for the segment of the cache 55P2 (the segment to which the JNL from the first positive node 210P1 is written), the redundancy flag may also be set to "off" and the non-volatility flag may be set to "on". Thereby, the redundancy of the JNL is skipped and the non-volatility of the JNL is performed.

[0034] Although not shown, the first positive node 210P1 writes the data in the cache 55P1 to PVOL102P1A and writes the JNL in the cache 55P1 to PJVOL102JP1A. Similarly, the second positive node 210P2 writes the data in the cache 55P2 to PVOL102P1S and writes the JNL in the cache 55P2 to PJVOL102JP1S.

[0035] The first secondary node 210S1 receives the JNL from the first positive node 210P1. In this embodiment, the reception of the JNL is performed in response to a JNL read request (for example, a read request specifying the SEQ# of the JNL to be read) from the first secondary node 210S1 to the first positive node 210P1. However, the reception of the JNL may also be the reception of a JNL write request (for example, a write request associated with the JNL to be written) from the first positive node 210P1 to the first secondary node 210S1.

[0036] The first secondary node 210S1 writes the received JNL to the first segment of the cache 55S1. The first secondary node 210S1 sets the values of the redundancy flag 2S1 and the non-volatility flag 3S1 of the first segment to "off", respectively. Therefore, both the redundancy and non-volatility of the JNL are skipped.

[0037] The first secondary node 210S1 writes the data (data to be copied) in the JNL in the first segment of the cache 55S1 to the second segment of the cache 55S1. The first secondary node 210S1 sets the value of the redundancy flag 2S2 in the second segment to "on" and the value of the non-volatile flag 3S2 to "off". Therefore, data redundancy is performed, but data non-volatilization is skipped. That is, the first secondary node 210S1 writes the data in the second segment of the cache 55S1 to the SVOL102S1A and transfers the data to the second primary node 210P2. The second primary node 210P2 receives the data, writes it to the segment of the cache 55S2, and writes it to the SVOL102S1S. Although not shown, for the segment of the cache 55S2 (the segment where the data (data in the JNL) from the first secondary node 210S1 is written), both the redundancy flag and the non-volatile flag may be set to "off". Thereby, both the redundancy and non-volatilization of the data in the JNL are skipped.

[0038] According to the above processing, the redundancy and non-volatilization of the JNL in the first segment of the cache 55S1 are skipped, and the non-volatilization of the data in the second segment of the cache 55S1 is skipped. Even if redundancy or non-volatilization is skipped, if the JNL or data disappears from the cache 55S1 due to a power failure or the like of the first secondary node 210S1, the JNL or data can be restored from the primary site 201P.

[0039] Specifically, in PJVOL102JP1A (and 102JP1S), JNL containing the data is stored until the data is written to the first and second SVOL102S1A (and 102S1S). That is, when data is written to SVOL102S1A (and 102S1S), the first secondary node 210S1 notifies the first primary node 210P1 of the SEQ# of the JNL containing the data that has been copied to SVOL102S1A (and 102S1S). When the first primary node 210P1 receives a notification of the SEQ# of the JNL for which data copying (reflection to SVOL102S) has been completed, the JNL having the SEQ# is purged from JVOL102JP1A (the second primary node 210P2 also purges the JNL having the SEQ# from JVOL102JP1S). In this way, until data is written to SVOL102S1A (and 102S1S), JNL containing the data is stored at the primary site 201P. The first secondary node 210S1 sends a JNL read request specifying the SEQ# of the JNL to be restored to the first primary node 210P1 for restoring JNL and data that have disappeared from the cache 55S1 before being written to JVOL102JS1A or SVOL102S1A. Thereby, the first secondary node 210S1 can obtain the JNL having the SEQ# from the first primary node 210P1 and can also obtain the data within the JNL.

[0040] Hereinafter, this embodiment will be described in detail.

[0041] FIG. 2 is a diagram showing a physical configuration example of the storage system 101.

[0042] There are a plurality of sites 201. Each site 201 is communicably connected via a network 202. The network 202 is, for example, a WAN (Wide Area Network), but is not limited to a WAN. The site 201 is a data center or the like and includes a plurality (or one) of nodes 210.

[0043] Node 210 may be a general-purpose computer. Node 210 includes, for example, one or more processor packages 213 including a processor 211 and a memory 212, one or more drives 214, and one or more ports 215. Each of these components is connected via an internal bus 216. The drive 214 is an example of a persistent storage device.

[0044] The processor 211 is, for example, a CPU (Central Processing Unit) and performs various processes.

[0045] The memory 212 is typically a volatile memory and stores control information necessary for realizing the functions of the node 210 and stores data. Also, the memory 212 stores, for example, a program executed by the processor 211. The drive 214 stores various data, programs, and the like.

[0046] The port 215 is connected to the network 220 within the site 201 and communicatively connects its own node to other nodes 210 within the site 201 via the network 220. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.

[0047] Note that the physical configuration of the system is not limited to the configuration described above. For example, the network 202 and / or 220 may be redundant. Also, for example, the network 220 may be separated into a management network and a storage network, the connection standard may be Ethernet (registered trademark), Infiniband, or wireless, and the connection topology is not limited to the configuration shown in FIG. 2. Also, for example, the drive 214 may have a configuration independent of the node 210.

[0048] FIG. 3 is a diagram showing a configuration example of the software platform of the site 201.

[0049] For example, the secondary site 201S can adopt a software platform having the configuration illustrated in FIG. 3. At site 201, there is a network storage service 30 that provides a plurality of persistent stores 32 to a plurality of nodes 210 via a network 220. The persistent store 32 is a storage area based on one or more drives 214.

[0050] The node 210 has an instance store 65, a hypervisor 64, and a virtual machine 61.

[0051] The instance store 65 provides block-level temporary storage for instances. This storage may be on a drive 214 physically attached to the node 210.

[0052] The hypervisor 64 dynamically generates and deletes virtual machines 61.

[0053] The virtual machine 61 manages one or more virtual drives 63 and executes storage control software (SCS) 720.

[0054] The SCS 720 controls I / O (Input / Output) to the virtual drive 63. The storage control program 720 is redundant among the nodes 210. That is, when a failure occurs in a node 210, the SCS (Standby) 720 of another node 210 changes from Standby to Active to replace the SCS (Active) 720 of the node 210.

[0055] The virtual drive 63 is a storage area to which the instance store 65 or the persistent store 32 is allocated. The virtual drive 63 may be treated as VOL102.

[0056] In this way, at site 201, the instance store 65 by DAS (Direct Attached Storage) and the storage via the network 220 such as iSCSI (network storage service 30) are used. The hypervisor 64 may not be present. For example, DAS and the network storage service 30 may be configured in bare metal.

[0057] FIG. 4 is a schematic diagram showing an overview of the remote copy configuration.

[0058] Remote copy pairs are constructed between the primary site 201P and the secondary site 201S among a plurality of VOL102. Specifically, for example, two consistency groups 401a and 401b are constructed between the primary site 201P and the secondary site 201S. The consistency group 401 is composed of VOL102 of a plurality (or one) of remote copy pairs. In the consistency group 401, a plurality of PVOLs are copied to the SVOL while maintaining consistency. More specifically, for example, in the consistency group 401, the update difference data up to the same time for a plurality of PVOL102 is copied to a plurality of SVOLs. Also, the control (consistency control) of the consistency group 401 is managed by the PJVOL. In the PJVOL, the update difference data of a plurality (or one) of PVOL102P is stored together with metadata such as its write time. When the primary site 201P transfers the data of the PVOL to the secondary site 201S, among the update difference data written to the PJVOL, the update difference data up to the same time is transferred to the secondary site 201S. Thereby, data can be copied to the SVOL while maintaining the consistency of the update times among a plurality of PVOLs.

[0059] For example, according to consistency group 401a, data is copied to SVOL1 and 2 at secondary node 210S1 via PJVOL1 and SJVOL1 while maintaining the consistency of PVOL1 and 2 at primary node 210P1. According to consistency group 401b, data is copied to SVOL3 at secondary node 210S2 and SVOL4 at secondary node 210S3 via PJVOL2 at primary node 210P2, PJVOL3 at primary node 210P3, SJVOL2 at secondary node 210S2, and SJVOL3 at secondary node 210S3 while maintaining the consistency of PVOL3 at primary node 210P2 and PVOL4 at primary node 210P3. PJVOL and SJVOL do not necessarily have to correspond on a 1:1 basis (for example, they may be 1:many, many:1, or many:many), and PJVOL may be an area on memory 212.

[0060] As can be seen from the specific configuration described above, consistency group 401 may be composed of VOL102 within a specific node 210 within site 201, or may be composed of VOL102 at a plurality of nodes 210 within site 201.

[0061] FIG. 5 is a schematic diagram showing an overview of I / O request processing.

[0062] First, application 502 operating on host 51 issues a write request specifying PVOL1 to primary node 210P1. Primary node 210P1 that has received the write request writes data A and B associated with the write request to PVOL1, and further writes a JNL containing data A and B as updated differential data to PJVOL1.

[0063] Next, the primary node 210P1 transfers the JNL (update differential data) written to PJVOL1 to SJVOL1 and SJVOL1 (standby) of the secondary site 201S. At this time, if multiple communication paths are established between the primary site 201P and the secondary site 201S, the JNL may be transferred using any of the communication paths. Normally, the primary node 210P1 transfers the JNL to the secondary node 210S1 that has the ownership of SVOL1 paired with PVOL1. However, when a failure occurs in the communication path with the ownership, the primary node 210P1 may transfer the JNL to the secondary node 210S2 or the like that does not have the ownership. For example, when the primary node 210P1 transfers the JNL to the secondary node 210S2 that does not have the ownership, the secondary node 210S2 transfers the received JNL to the secondary node 210S1 with the ownership, and the secondary node 210S1 writes the JNL to SJVOL1.

[0064] Next, the secondary node 210S1 writes the JNL written to SJVOL1 to SVOL1. Then, the data A and B written to SVOL1 are written to the drive 214a via the storage pool 504a. When the configuration of the drive 214a is DAS (Direct Attached Storage) in which the node 210 and the drive 214 are connected one-to-one, the JNL is written to the drive 214a mounted on the secondary node 210S1. By writing all the data copied to SVOL1 to the drive 214a of the secondary node 210S1 with the ownership of SVOL1, when reading data from SVOL1 later, there is no need to read data from another node. As a result, the inter-node transfer process can be eliminated, and a high-speed read process can be realized.

[0065] Note that the storage pool 504 may be an area based on one or more drives 214. Storage functions such as thin-provisioning, compression, or deduplication are provided, and the processing of the storage functions required for the data written to the storage pool 504 is executed.

[0066] When the secondary node 210S1 writes data to the drive 214a, in order to protect the data from node failures, it also writes the redundant data of the data to be written to the drive 214b of the secondary node 210S2. Regarding the writing of redundant data, when the data protection policy is replication, the replica of the data is written to the drive 214b as redundant data. On the other hand, when the data protection policy is Erasure Coding, parity is calculated from the data, and the calculated parity is written to the drive 214b as redundant data.

[0067] FIG. 6 is a schematic diagram showing an overview of the recovery process from node failures.

[0068] In the secondary nodes 210S1, 210S2, and 210S3, the SCS720 is operating. The secondary node 210S has an operating SCS (Active) and a standby SCS (Standby) corresponding to the SCS (Active) in another secondary node 210S. For example, the secondary node 210S1 has SCS1 (Active) and SCS3 (Standby), the secondary node 210S2 has SCS2 (Active) and SCS1 (Standby), and the secondary node 210S3 has SCS3 (Active) and SCS2 (Standby). SCSx (Active) and SCSx (Standby) belong to the redundancy group of SCSx, and there may be more than one SCSx (Standby) (x is a natural number).

[0069] Using the specific example shown in FIG. 6, the recovery process from node failures will be described.

[0070] The secondary node 210S2 replicates and holds the configuration information of SVOL1 and SJVOL1 that the secondary node 210S1 has in order to inherit the remote copy pair information of the secondary node 210S1. In addition, the secondary node 210S2 stores the redundant data of the data written to the drive 214a of the secondary node 210S1 in the drive 214d. Furthermore, the secondary node 210S2 has established a communication path with the primary node 210P1.

[0071] And, for example, when the secondary node 210S1 stops due to a failure, the secondary node 210S2 that has detected the failure of the secondary node 210S1 takes over the processing of SCS1 (Active) of the secondary node 210S1, and SCS1 (Standby) changes to SCS1 (Active). The secondary node 210S2 communicates with the primary node 210P1 to continue the remote copy process between PVOL1 and SVOL1. That is, a failover from SCS1 of the secondary node 210S1 to SCS1 of the secondary node 210S2 is performed. Thereby, even if a node failure occurs at any of the secondary sites 201S, another secondary site 201S can continue the remote copy from the primary site 201P.

[0072] FIG. 7 is a diagram showing an example of data and programs held in the memory 212.

[0073] Information is read from the drive 214 into the memory 212. For example, various tables included in the control information table 710 and various programs included in the SCS 720 are expanded on the memory 212 during the execution of the processes for which they are used, but otherwise are stored in a non-volatile storage area such as the drive 214 in preparation for a power outage or the like.

[0074] The control information table 710 includes a system configuration management table 711, a pair configuration management table 712, and a cache management table 713.

[0075] The SCS 720 includes a pair formation processing program 721, an initial JNL creation processing program 722, a restore processing program 723, a JNL read processing program 724, a JNL purge processing program 725, a cache storage processing program 726, a destage processing program 727, an updated JNL creation processing program 728, and a pair recovery processing program 729.

[0076] FIG. 8 is a diagram showing an example of the system configuration management table 711.

[0077] The system configuration management table 711 includes a node configuration management table 810, a drive configuration management table 820, and a port configuration management table 830. For each site 201, there is a node configuration management table 810 for a plurality of nodes 210 existing in each site 201, and the node 210 has a drive configuration management table 820 and a port configuration management table 830 regarding the drives 214 within its own node 210.

[0078] The node configuration management table 810 is provided for each site 201 and stores information indicating the configuration (such as the relationship between the node 210 and the drive 214) related to the nodes 210 provided in the site 201. More specifically, the node configuration management table 810 stores information such as a node ID 811, a status 812, a drive ID list 813, and a port ID list 814 for each node 210.

[0079] The node ID 811 is the ID of the node 210. The status 812 indicates the status of the node 210 (for example, "Normal", "Warning", or "Failure", etc.). The drive ID list 813 is a list of the IDs of the drives 214 provided in the node 210. The port ID list 814 is a list of the IDs of the ports 215 provided in the node 210.

[0080] The drive configuration management table 820 is provided for each node 210 and stores information indicating the configuration related to the drives 214 provided in the node 210. More specifically, the drive configuration management table 820 stores information such as a drive ID 821, a status 822, and a size 823 for each drive 214.

[0081] The drive ID 821 is the ID of the drive 214. The status 822 indicates the status of the drive 214. The size 823 indicates the capacity of the drive 214.

[0082] The port configuration management table 830 is provided for each node 210 and stores information indicating the configuration related to the port 215 provided in the node 210. More specifically, the port configuration management table 830 stores information such as a port ID 831, a status 832, and an address 833 for each port.

[0083] The port ID 831 is the ID of the port 215. The status 832 indicates the status of the port 215. The address 833 indicates the address on the network assigned to the port 215. The form of the address may be an IP (Internet Protocol), or a WWN (World Wide Name), a MAC (Media Access Control) address, etc.

[0084] FIG. 9 is a diagram showing an example of the pair configuration management table 712.

[0085] The pair configuration management table 712 is configured to include a VOL management table 910, a pair management table 920, and a JNL management table 930.

[0086] The VOL management table 910 stores information indicating the configuration related to the VOL102. More specifically, the VOL management table 910 stores information such as a VOL ID 911, an owner node ID 912, a fallback destination node ID 913, a size 914, and an attribute 915 for each VOL102.

[0087] The VOL ID 911 is the ID of the VOL102. The owner node ID 912 is the ID of the node 210 having the ownership of the VOL102. The fallback destination node ID 913 is the ID of the node 210 that takes over the process in case of a failure of the node 210 having the ownership of the SVOL. The size 914 indicates the capacity of the VOL102.

[0088] Attribute 915 indicates the attribute of VOL102. "NML_VOL" means a normal VOL that does not belong to a consistency group. "PAIR_VOL" means a PVOL or SVOL that belongs to a consistency group. "JNL_VOL" means JVOL.

[0089] The pair management table 920 stores information indicating the configuration related to the remote copy pair. More specifically, the pair management table 920 stores information such as a pair group ID 921, a PJVOL ID 922, a PVOL ID 923, an SJVOL ID 924, an SVOL ID 925, and a status 926 for each consistency group.

[0090] The pair group ID 921 is the ID of the consistency group. The PJVOL ID 922 is a list of the IDs of PJVOL102JP that belong to the consistency group. The PVOL ID 923 is a list of the IDs of PVOL102P that belong to the consistency group. The SJVOL ID 924 is a list of the IDs of SJVOL102JS that belong to the consistency group. The SVOL ID 925 is a list of the IDs of SVOL102S that belong to the consistency group. The status 926 indicates the status of each remote copy pair in the consistency group (e.g., "PAIR", "COPY", "SUSPEND", etc.). "PAIR" is a state in which the write to PVOL102P is periodically reflected in SVOL102S. "COPY" is a state during initial copy. "SUSPEND" is a pair interruption state (a state in which synchronization between PVOL102P and SVOL102S is not performed).

[0091] The JNL management table 930 stores information related to JNL. More specifically, the JNL management table 930 stores information such as a pair group ID 931, a JNL ID 932, a P / SVOL ID 933, a P / SVOL address 934, a size 935, and a cache segment ID 936 for each JNL.

[0092] The pair group ID 931 is the ID of the consistency group to which JNL belongs. The JNL ID 932 is the ID of JNL. The JNL ID corresponds to SEQ#, and in the consistency group, it is, for example, a sequential number. That is, the JNL ID represents the writing order, and the data in JNL will be stored in SVOL102S in the consistency group in the order of the JNL IDs.

[0093] The P / SVOL ID 933 includes the ID of PVOL102P where the data in JNL is written and the ID of SVOL102S where the data in JNL is written. The P / SVOL address 934 includes the storage destination address of the data in PVOL102P where the data in JNL is written and the storage destination address of the data in SVOL102S where the data in JNL is written.

[0094] The size 935 represents the size of JNL. For example, one JNL contains one or more pieces of data. The cache segment ID 936 is the ID of the cache segment where the data in JNL is written.

[0095] Figure 10 is a diagram showing an example of the cache management table 713.

[0096] The cache management table 713 is a table regarding the cache 55. The cache management table 713 includes a dirty queue 1001, a clean queue 1002, a free queue 1003, and a cache segment management table 1004. In this embodiment, there are a plurality of different segment sizes as the segment size, and segments of a plurality of different segment sizes are prepared in advance. However, the segment size may be a variable size according to the cache allocation request.

[0097] The dirty queue 1001 is a queue of the IDs (addresses) of the dirty segments in which the dirty data to be written by the drive 214 is stored for each drive 214. A "dirty segment" is a segment in which dirty data is stored. "Dirty data" is data that has not been written to the drive 214.

[0098] The clean queue 1002 is a queue of the IDs (addresses) of the clean segments having the segment size for each segment size. A "clean segment" is a segment in which clean data is stored. "Clean data" is data that has been written to the drive 214.

[0099] The free queue 1003 is a queue of the IDs (addresses) of the free segments having the segment size for each segment size. A "free segment" is a segment to which new data may be written. A segment of a desired size is allocated from the free queue 1003, and data is written to the allocated segment.

[0100] The cache segment management table 1004 stores information regarding cache segments. More specifically, the cache segment management table 1004 stores information such as a segment ID 1041, a memory address 1042, a size 1043, a VOL ID 1044, a VOL address 1045, and a redundancy - non - volatility flag 1046 for each cache segment.

[0101] The segment ID 1041 is the ID of the cache segment. The memory address 1042 is the address of the cache segment (the address in the cache 55). The size 1043 is the size of the cache segment.

[0102] VOL ID1044 is the ID of VOL102 where data in the cache segment is to be written. VOL address 1045 is the address in the destination VOL102 (the address where data is to be written).

[0103] The redundancy - non - volatility flag 1046 includes a redundancy flag 2 and a non - volatility flag 3 corresponding to the cache segment. "1" means "on" and "0" means "off".

[0104] Hereinafter, an example of the processing performed in this embodiment will be described. In the description with reference to FIGS. 11 to 17, to avoid confusion, "P" is appended to the end of the reference numerals of the programs at the primary site 201P, and "S" is appended to the end of the reference numerals of the programs at the secondary site 201S. Also, in the following description, "initial JNL" refers to the JNL generated in the initial copy, and "updated JNL" means the JNL generated in response to the update of PVOL102P after the initial copy.

[0105] FIG. 11 is a diagram showing the flow of the pair formation process.

[0106] According to the pair formation process, a remote copy pair is formed through communication between the primary site 201P and the secondary site 201S. In the description of FIG. 11, the pair formation process program 721P is a program in the primary node 210P having a PVOL candidate. The pair formation process program 721S is a program in the secondary node 210S having an SVOL candidate.

[0107] The pair formation process program 721P sends a pre - check request to the pair formation process program 721S (S1101). The pair formation process program 721S receives the request (S1102) and performs a predetermined pre - check such as whether there is a VOL102 to be paired and whether the information of the partner device is correct. The pair formation process program 721S returns a response to the pre - check request (S1103). The response represents the result of the pre - check. The pair formation process program 721P receives the response (S1104).

[0108] If the response is a predetermined response, the pair formation processing program 721P transmits a pair formation request to set the PVOL candidate as PVOL102P and the SVOL candidate as SVOL102S to the pair formation processing program 721S (S1105). The pair formation processing program 721S receives the request (S1106), forms a VOL pair (registers information in the pair management table 920), and sets the state 926 of the pair to "COPY" (S1107). The pair formation processing program 721S activates the restore process (Figure 13) (S1108) and returns a response to the pair formation request (S1109). The pair formation processing program 721S waits for the completion of the initial copy (S1110). When the initial copy is completed (S1111: Yes), specifically, when the synchronization of the PVOL data and the SVOL data is completed by the restore processing program 723S, the pair formation processing program 721S sets the state 926 of the pair to "PAIR" (S1112).

[0109] The pair formation processing program 721P receives the response transmitted in S1109 (S1113). If the response is a predetermined response, the pair formation processing program 721P forms a VOL pair (registers information in the pair management table 920) and sets the state 926 of the pair to "COPY" (S1114). If there is a resynchronization option for the pair (S1115: Yes), the pair formation processing program 721P sets the resynchronization option (S1116).

[0110] The pair formation processing program 721P activates the initial JNL creation process (S1117). The pair formation processing program 721P waits for the completion of the initial copy (S1118). When all the initial JNLs are purged (S1119: Yes), specifically, when the synchronization of the PVOL data and the SVOL data is completed by the restore processing program 723S, the pair formation processing program 721P sets the state 926 of the pair to "PAIR" (S1120).

[0111] Figure 12 is a diagram showing the flow of the initial JNL creation process at the primary site 201P.

[0112] When the initial JNL creation process is started, the process shown in FIG. 12 is performed. In this process, the data of PVOL102P is generated as the initial JNL. The JNL is basically placed on the cache 55P. In the case of cache full, it is destaged to PJVOL102JP, and the destaged part is purged from the cache 55P. There are full copy and differential copy for the initial copy. When the state 926 of the pair is "SUSPEND", a JNL is created only for the update difference during pair interruption by differential copy. Redundancy flags and non-volatility flags are set for the cache segment, and writing is controlled according to these flags.

[0113] If the resynchronization option is not set (S1201: No), the initial JNL creation process program 722P targets the data in the entire area of PVOL102P for creating the initial JNL (S1202). If the resynchronization option is set (S1201: Yes), the initial JNL creation process program 722P refers to a difference management table (not shown) representing the difference between PVOL102P and SVOL102S, and targets the data in the difference area for creating the initial JNL (S1203).

[0114] The initial JNL creation process program 722P secures a cache segment for the initial JNL from the free queue 1003 (S1204), and acquires the VOL address 1045 corresponding to the segment (the PVOL address of the PVOL area where the data to be included in the initial JNL is located) (S1205). The initial JNL creation process program 722P reads data from the address (S1206), creates the metadata of the initial JNL (S1207), creates an initial JNL including the data read in S1206 and the metadata created in S1207, and sets the initial JNL as the storage target (S1208). The metadata includes, for example, JNL ID, LBA (VOL address), transfer length, pair ID, PVOL ID, and SVOL ID.

[0115] The initial JNL creation processing program 722P sets both the redundancy flag and the non-volatile flag in the redundancy-non-volatile flag 1046 to "1" for the segments secured in S1204 (S1209). The initial JNL creation processing program 722P activates the cache storage processing (Figure 15A) (S1210).

[0116] When the cache usage rate (for example, the ratio of the total capacity of dirty segments and clean segments to the total capacity of the cache 55P) exceeds a predetermined value (S1211: Yes), the initial JNL creation processing program 722P designates the initial JNL in the cache segment as a destage target (S1212) and activates the destage processing (Figure 15B) (S1213). The initial JNL creation processing program 722P releases the cache segment having the initial JNL destaged to PJVOL102JP, that is, makes it a free segment (S1214).

[0117] If there is an initial JNL that has not been created (S1215: No), the process returns to S1204. When all initial JNLs have been created (S1215: Yes), the initial JNL creation process ends.

[0118] Figure 13 is a diagram showing the flow of the restore process at the secondary site 201S.

[0119] When the restore process is activated, the process shown in Figure 13 is performed. In this process, the JNL from the primary site 201P is reflected in SVOL102S. By setting the redundancy flag and the non-volatile flag of the cache segment where the JNL from the primary site 201P is written to "0" respectively, the processing load is reduced. If the JNL disappears during a failure of the secondary node, the JNL is retransferred from the primary site 201P and the synchronization state between PVOL and SVOL is restored. After the JNL is reflected in SVOL102S, a purge notification is issued to the primary site 102P to discard the JNL at the primary site 102P.

[0120] The restore processing program 723S waits for a certain period of time (S1301) and sends a JNL read request (S1302). In this request, the ID (for example, SEQ#) of the JNL to be read may be specified. This request is sent to the positive node 210P having the PVOL represented by the metadata of the JNL. In response to this request, S1401 in FIG. 14A is performed.

[0121] The restore processing program 723S secures the cache segment of the JNL to be read from the free queue 1003 (S1303).

[0122] When S1412 in FIG. 14A is performed, the restore processing program 723S receives a response to the JNL read request sent in S1302 (S1304). The restore processing program 723S stores the JNL included in the response in the buffer (S1305).

[0123] For the segment secured in S1303, the restore processing program 723S sets both the redundancy flag and the non-volatility flag in the redundancy - non-volatility flag 1046 to "0" (S1306). The restore processing program 723S designates the JNL stored in S1305 as a cache storage target (S1307) and starts the cache storage process (FIG. 15A) (S1308).

[0124] When the cache usage rate exceeds a predetermined value (S1309: Yes), the restore processing program 723S designates the JNL in the cache segment as a destage target (S1310) and starts the destage process (FIG. 15B) (S1311). The restore processing program 723S releases the cache segment having the JNL destaged to the SJVOL102JS (S1312).

[0125] When the determination result of S1309 is false (S1309: No), or after S1312, the restore processing program 723S secures a cache segment as the copy destination of the data in the JNL stored in the segment at S1308 from the free queue 1003 (S1313). The restore processing program 723S sets both the redundancy flag and the non-volatility flag to "0" for the segment secured at S1313 (S1314). The restore processing program 723S designates the data in the JNL stored in the segment at S1308 as the cache storage target (S1315) and starts the cache storage process (Figure 15A) (S1316).

[0126] The restore processing program 723S sets the redundancy flag to "1" and the non-volatility flag to "0" for the segment secured at S1313 (S1317). The restore processing program 723S designates the data in the segment secured at S1313 as the destage target (S1318) and starts the destage process (Figure 15B) (S1319). The restore processing program 723S releases the cache segment having the data destaged to the SVOL102S (S1320). The restore processing program 723S updates the ID (SEQ#) of the restored (reflected) JNL (S1321) and sends a purge notification including the updated JNL ID to the primary site 201P (S1322). In response to this notification, S1451 in Figure 14B is performed. When S1453 in Figure 14B is performed, the restore processing program 723S receives a response to the notification sent at S1322 (S1323). The process returns to S1301.

[0127] Figure 14A is a diagram showing the flow of the JNL read process at the primary site 201P.

[0128] The JNL read processing program 724P receives a JNL read request (S1401) and determines whether there is an untransferred JNL at the secondary site 201S (S1402). For example, it may be determined whether the JNL ID specified in the request is the ID of an untransferred JNL.

[0129] When the determination result of S1402 is true (S1402: Yes), the JNL read processing program 724P determines whether the untransferred JNL is cached (S1403).

[0130] When the determination result of S1403 is false (S1403: No), the JNL read processing program 724P secures the cache segment of the JNL (S1404), reads the JNL from PJVOL102JP into the buffer (S1405). The JNL read processing program 724P sets both the redundancy flag and the non-volatility flag to "0" for the segment secured in S1404 (S1406). The JNL read processing program 724P designates the JNL read in S1405 as an object for cache storage (S1407), and activates the cache storage process (Figure 15A) (S1408). The JNL read processing program 724P includes the JNL in the response to the JNL read request received in S1401 (S1410), and returns the response (S1412).

[0131] When the determination result of S1403 is true (S1403: Yes), the JNL read processing program 724P acquires the JNL to be transferred from the cache 55P (S1409), includes the JNL in the response to the JNL read request (S1410), and returns the response (S1412). The "JNL to be transferred" may be the JNL with the youngest JNL ID among the untransferred JNLs.

[0132] When the determination result of S1402 is false (S1402: No), the JNL read processing program 724P includes a value indicating no JNL in the response to the JNL read request (S1411), and returns the response (S1412).

[0133] Figure 14B is a diagram showing the flow of JNL purge processing at the positive site 201P.

[0134] The JNL purge processing program 725P receives a purge notification including a JNL ID (S1451), purges all JNLs having JNL IDs up to the JNL ID represented by the notification from PJVOL102JP, and returns a response indicating completion (S1453). Note that the purge notification may be included as a parameter in a JNL read request.

[0135] FIG. 15A is a diagram showing the flow of cache storage processing at the secondary site 201S.

[0136] When the cache storage processing is activated, the processing shown in FIG. 15A is performed. In this processing, the JNL or data to be stored in the cache is stored in the secured cache segment, and depending on the flag, the redundancy or non-volatility of the JNL or data is controlled whether to be implemented.

[0137] The cache storage processing program 726S stores the JNL or data in the secured cache segment (S1501).

[0138] When the redundancy flag corresponding to the segment is "1" (S1502: Yes), the cache storage processing program 726S transfers the JNL or data in the segment to another secondary node 210S (S1503). In other words, when the redundancy flag corresponding to the segment is "0" (S1502: No), the cache storage processing program 726S skips transferring the JNL or data in the segment to another secondary node 210S.

[0139] When the non-volatility flag corresponding to the segment is "1" (S1504: Yes), the cache storage processing program 726S writes the JNL or data in the segment to the non-volatile area 103S (S1505). In other words, when the non-volatility flag corresponding to the segment is "0" (S1504: No), the cache storage processing program 726S skips writing the JNL or data in the segment to the non-volatile area 103S.

[0140] FIG. 15B is a diagram showing the flow of the destage process at the secondary site 201S.

[0141] When the destage process is activated, the process shown in FIG. 15B is performed. In this process, a JNL is written from a segment of the cache 55S to SJVOL102JS, or data is written from the segment to SVOL102S. At that time, if the redundancy flag corresponding to the segment is "1", the JNL or data is redundant to another secondary node. When the redundancy flag is "0", the JNL or data may be written to the instance store of the secondary node.

[0142] The destage process program 727S selects a JNL or data to be destaged (S1551). Specifically, a dirty segment is selected from the dirty queue 1001.

[0143] When the redundancy flag corresponding to the segment having the selected JNL or data is "1" (S1552: Yes), the destage process program 727S transfers the JNL or data to another secondary node 210S (S1553), and writes the JNL or data to SJVOL102JS or SVOL102S of its own secondary node 210S (S1555).

[0144] When the non-volatile flag corresponding to the segment having the selected JNL or data is "0" (S1552: No, S1554: Yes), the destage process program 727S writes the JNL or data to the instance store of its own secondary node 210S (S1556). Note that S1555 may be performed instead of S1556.

[0145] FIG. 16 is a diagram showing the flow of the updated JNL creation process at the primary site 201P.

[0146] In this process, an updated JNL is created in response to a write request from host 51. The updated JNL is transferred to the secondary site 201S by the restore process of the secondary site 201S and reflected in the SVOL102S in the same way as the initial JNL. If the updated JNL at the primary site 201P is lost due to a node failure, the updated JNL cannot be recovered. For this reason, the data and JNL on the cache are made redundant and non-volatile.

[0147] The updated JNL creation processing program 728P receives a write request (S1601) and stores the write data (data to be written) associated with the write request in a buffer (S1602). The updated JNL creation processing program 728P secures a cache segment for the write data from the free queue 1003 (S1603), sets the redundancy flag of the segment to "1", and sets the write data as the storage target (S1605).

[0148] The updated JNL creation processing program 728P determines whether the attribute 915 of VOL102 specified in the write request is "PAIR_VOL", that is, whether the VOL102 is PVOL102P (S1606).

[0149] If the determination result of S1606 is true (S1606: Yes), the updated JNL creation processing program 728P determines whether the state 926 of the pair including the VOL is "SUSPEND" (S1607). If the determination result of S1607 is true (S1607: Yes), the updated JNL creation processing program 728P updates the difference management table representing the difference between the PVOL102P and the SVOL102S according to the write destination address according to the write request (S1608).

[0150] If the determination result of S1607 is false (S1607: No), the updated JNL creation processing program 728P secures a cache segment for the updated JNL (S1609). The updated JNL creation processing program 728P creates the metadata of the updated JNL (S1610), sets the updated JNL as the cache storage target (S1611), and sets the redundancy flag of the segment to "1" (S1612).

[0151] When the determination result of S1606 is false (S1606: No), after S1608, or after S1612, the update JNL creation processing program 728P activates the cache storage process (Figure 15A) (S1613) and returns a response to the write request to the host 51 (S1614).

[0152] The update JNL creation processing program 728P targets the write data for destaging (S1615) and determines whether the cache usage rate exceeds a predetermined value (S1616). When the determination result of S1616 is true (S1616: Yes), the update JNL creation processing program 728P targets the update JNL for destaging (S1617).

[0153] When the determination result of S1616 is false (S1616: No), or after S1617, the update JNL creation processing program 728P activates the destaging process (Figure 15B) (S1618), releases the cache segment where the destaged update JNL exists (S1619: Yes, S1620), and releases the cache segment where the destaged write data exists (S1621).

[0154] In addition, in the process shown in Figure 16, for both the cache segment of the write data and the cache segment of the update JNL, the update JNL creation processing program 728P may set the non-volatile flag to "1" and write the write data and update JNL in those segments to the non-volatile area 103P.

[0155] Figure 17 is a diagram showing the flow of the pair recovery process.

[0156] In this process, the secondary site 201S detects a failure, transitions the pair state to "SUSPEND", and notifies the primary site 201P of the failure. In response to this notification, the JNL is retransferred from the primary site 201P to the secondary site 201S. The pair is recovered using the retransferred JNL. For example, when the network 202 becomes temporarily unavailable and the pair state becomes "SUSPEND", and then the network 202 recovers later, the pair recovery process may be performed. In this embodiment, pair recovery is automatically performed when the pair state transitions to "SUSPEND". The "failure" mentioned in this paragraph may include, in addition to network failures, node failures, power failures, or drive failures. In response to a failure at the primary site 201P, the pair state at the primary site 201P may be set to "SUSPEND", and that state may be notified to the secondary site 201S, and pair recovery may be performed in response to that notification.

[0157] The pair recovery processing program 729S detects a failure at the secondary site 201S (S1701), and sets the state 926 of the VOL pair affected by the failure to "SUSPEND" (S1702). The pair recovery processing program 729S determines whether the number of executions of the series of processes from S1704 to S1706 has exceeded the retry threshold (S1703).

[0158] If the determination result in S1703 is false (S1703: No), the pair recovery processing program 729S performs the series of processes from S1704 to S1706. That is, the pair recovery processing program 729S notifies the primary node 210P having the PVOL102P belonging to the pair, whose state 926 was set to "SUSPEND" in S1702, of the failure of the pair (S1704), monitors the recovery of the pair (S1705), and determines whether the pair state has recovered normally (S1706). If the determination result in S1705 is true (S1705: Yes), the processing of the pair recovery processing program 729S ends. If the determination result in S1705 is false (S1705: No), the processing returns to S1703.

[0159] The pair recovery processing program 729P receives a pair failure notification from the pair recovery processing program 729S (S1751), sets a resynchronization option (S1752), and starts a pair formation process (Figure 11) (S1753).

[0160] As described above, although one embodiment of the present invention has been described, this is an exemplification for the purpose of explaining the present invention, and is not intended to limit the scope of the present invention to this embodiment. The present invention can be implemented in various other forms.

[0161] Also, the above description can be summarized as follows. The following summary may include supplementary explanations and explanations of modified examples of the above description. In the following description, the subject of the process is the processor 211. Specifically, for example, the subject of the process may be the SCS720. Also, in the following description, for reference, the reference numerals of the elements are mainly described with reference to FIG. 1.

[0162] A plurality of nodes 210 including a memory 212 provided with a cache 55 and a processor 211 connected to the memory 212 are provided. In each of the plurality of nodes 210, when there is data to be stored in a segment secured from the cache 55, the processor 211 determines, according to the type of the data, whether or not to perform redundancy (transferring redundant data of the data in the segment to another node) and whether or not to perform non-volatilization (storing the data in the segment in the non-volatile area 103) for the segment, and controls whether or not to perform redundancy of the data in the segment and whether or not to perform non-volatilization of the data in the segment based on the determination. For each of the plurality of nodes 210, the non-volatile area 103 is an area based on one or more non-volatile media inside or outside the node 210 (the non-volatile media may be called a permanent storage device, and an example of the non-volatile media is the drive 214). Thereby, the writing frequency from the cache 55 to the non-volatile media in the storage system 100 configured by the plurality of nodes 210 can be appropriately reduced.

[0163] The following description takes, as an example, determining the value of the redundancy flag 2 as the determination of the necessity of redundancy. However, in the present invention, the determination of the necessity of redundancy is not limited to determining the value of the redundancy flag 2. Similarly, the following description takes, as an example, determining the value of the non-volatility flag 3 as the determination of the necessity of non-volatility. However, in the present invention, the determination of the necessity of redundancy is not limited to determining the value of the non-volatility flag 3 (for example, information in a form other than a flag may be adopted). Further, in the following description, "redundant data" may be replicated data (for example, replication of data included in JNL or JNL), or parity data (parity data of data included in JNL or JNL).

[0164] The plurality of nodes 210 may be a plurality of secondary nodes 210S that constitute the secondary storage system 100S. Among the plurality of secondary nodes 210, the first secondary node 210S1 may have an SJVOL102JS1A which is a VOL102 where the JNL is stored, an SVOL102S1A that forms a pair with a PVOL102P1A in the primary storage system 100P, a first secondary cache 55S1 which is a cache 55 in the secondary node 210S1, and a first secondary processor (hereinafter, for convenience, referred to as "211S1") which is a processor 211 in the secondary node 210S1. The JNL may include data written to the PVOL102P1A in the primary storage system 100P and metadata including a SEQ# (sequence number) which is the order in which the data is written. In the above-described embodiment, an example of the SEQ# is the JNL ID. The first secondary processor 211S1 may secure a segment in which the JNL is stored from the first secondary cache 55S1. When the data stored in the secured segment is the JNL, for the segment, both the redundancy flag 2 and the non-volatility flag 3 may be set to "off". When both the redundancy flag 2 and the non-volatility flag 3 are "off" for the segment, the first secondary processor 211S1 may write the JNL in the segment to the SJVOL102JS1A and transfer the redundant data of the JNL in the segment to a second secondary node 210S2 having an SVOL102S1S which is a mirror VOL of the SVOL102S1A, or may not perform either the redundancy which is storing the JNL in the segment to the non-volatile area 103S1 or the non-volatility. Thereby, in the asynchronous remote copy using the JNL, the writing frequency from the cache 55S to the non-volatile medium can be appropriately reduced. Also, even if such a writing frequency is reduced, when the JNL or data is lost in the secondary node 210S1 or 210S1, the JNL or data can be restored using the JNL from the primary storage system 100P. Note that for each of the plurality of secondary nodes 210S, the VOL102 (SJVOL102JS or SVOL102S) may be an area based on one or more non-volatile media inside or outside the secondary node 210S.

[0165] The first sub-processor 211S1 secures a segment in which data in the JNL is stored from the first sub-cache 55S1. If the data stored in the secured segment is the data in the JNL, for the segment, the redundancy flag 2 may be set to "on" and the non-volatility flag 3 may be set to "off". When the redundancy flag 2 is "on" and the non-volatility flag 3 is "off" for the segment, the first sub-processor 211S1 writes the data in the segment to SVOL102S1A, performs redundancy by transferring the redundant data of the data in the segment to the second sub-node 210S2, and may not perform non-volatilization by storing the data in the segment in the non-volatile area 103S1. Thereby, the writing frequency from the cache 55S to the non-volatile medium in the asynchronous remote copy can be appropriately reduced. Note that the second sub-processor (hereinafter, for convenience, referred to as "211S2") which is the processor 211 in the second sub-node 210S2 secures a segment from the second sub-cache 55S2 which is the cache 55 in the sub-node 210S2, stores the redundant data from the first sub-node 210S1 in the segment, and may write the data in the segment to SVOL102S1S. Since the data in the segment is the redundant data from another sub-node 210S for the segment, the second sub-processor 211S2 may set both the redundancy flag 2 and the non-volatility flag 3 to "off".

[0166] The positive storage system 100P may be composed of a plurality of positive nodes 210P1 which are a plurality of nodes 210. Among the plurality of positive nodes 210P1, the first positive node 210P1 may include a PJVOL102PJ1A which is a VOL102 in which a JNL is stored, the PVOL102P1A, a first positive cache 55P1 which is a cache 55 in the positive node 210P1, and a first positive processor which is a processor 211 in the positive node 210P1 (hereinafter, for convenience, referred to as "211P1"). The first positive processor 211P1 secures a segment in which a JNL is stored from the first positive cache 55P1. When the data stored in the secured segment is a JNL, for the segment, both the redundancy flag 2 and the non-volatility flag 3 may be set to "on". When both the redundancy flag 2 and the non-volatility flag 3 are "on" for the segment, the first positive processor 211P1 writes the JNL in the segment to the PJVOL102PJ1A, and transfers the redundant data of the JNL in the segment to a second positive node 210P2 having a PVOL102JP1S which is a mirror VOL of the PVOL102P1A, which is the redundancy process, or stores the JNL in the segment in the non-volatile area 103P1, which is the non-volatility process. Thereby, the certainty that the JNL required for restoration in the secondary storage system 100S exists in the primary storage system 100P can be enhanced. For each of the plurality of positive nodes 210P1, the VOL102 (PJVOL102JP or PVOL102P) may be an area based on one or more non-volatile media inside or outside the positive node 210P1. Note that the second positive processor which is a processor 211 in the second positive node 210P2 (hereinafter, for convenience, referred to as "211P2") secures a segment from the second positive cache 55P2 which is a cache 55 in the positive node 210P2, stores the redundant data (redundant data of the JNL) from the first positive node 210P1 in the segment, and may write the data in the segment to the PJVOL102JP1S.For the second main processor 211P2, for this segment, since the data within this segment is redundant data from another main node 210P, the redundancy flag 2 may be set to "off", but the non-volatile flag 3 may be set to "on". Therefore, the second main processor 211P2 may store the data within this segment in the non-volatile area 103P2.

[0167] When the first main processor 211P1 secures a segment from the first main cache 55P1, which is the data included in the JNL and where the data from the host 51 is stored, and the data stored in the secured segment is the data from the host 51 (the data written to PVOL102JP1A), for this segment, both the redundancy flag 2 and the non-volatile flag 3 may be set to "on". When both the redundancy flag 2 and the non-volatile flag 3 are "on" for this segment, the first main processor 211P1 may write the data within this segment to PVOL102P1A and transfer the redundant data of the data within this segment to the second main node 210P2. That is, both the redundancy process and the non-volatile process of storing the data within this segment in the non-volatile area 103P1 may be performed. Thereby, the possibility that the JNL required for restoration in the secondary storage system 100S can be regenerated again in the primary storage system 100P can be increased. Note that the second main processor 211P2 may secure a segment from the second main cache 55P2, store the redundant data from the first main node 210P1 in this segment, and write the data within this segment to PVOL102P1S. For the second main processor 211P2, for this segment, since the data within this segment is redundant data from another main node 210P, the redundancy flag 2 may be set to "off", but the non-volatile flag 3 may be set to "on". Therefore, the second main processor 211P2 may store the data within this segment in the non-volatile area 103P2.

[0168] When the processor 211 (either the secondary processor 211S or the primary processor 211P, for example) determines that the non-volatilization flag 3 for a segment is "off", the processor 211 may write the data within the segment to the instance store 65 in the node 210 that has the processor 211. The instance store 65 may be a region based on a volatile storage medium.

[0169] After the first secondary processor 211S1 writes data to SVOL102S1A (and SVOL102S1S), the first secondary processor 211S1 may update the SEQ# of the JNL that has been reflected and notify the updated SEQ# to the primary storage system 100P. The primary storage system 100P may purge the JNL having SEQ# up to the SEQ# notified from the first secondary processor 211S1 from the primary storage system 100P. In this way, the JNL can be stored in the primary storage system 100P until the JNL is reflected in SVOL102S1A (and SVOL102S1S).

[0170] When the first secondary processor 211S1 detects that a pair has failed, the first secondary processor 211S1 may notify the primary storage system 100P of the failure of the pair. The first secondary node 210S1 receives, from the primary storage system 100P that has received the notification, a JNL including data as a difference between SVOL102S1A and PVOL102P1A, and the first secondary processor 211S1 may secure a segment from the first secondary cache 55S1 and store the JNL in the segment. As a result, the above-described processing is performed, and thus, recovery from the failure is automatically performed. For example, the formation of the pair between SVOL102S1A and PVOL102P1A may be performed in response to a pair formation request from the primary storage system 100P. In response to the above-described failure notification, the first secondary node 210S1 receives a pair formation request from the primary storage system 100P, and in the process performed in response to the pair formation request, the first secondary node 210S1 may receive, from the primary storage system 100P, a JNL including data as a difference. That is, when a pair failure is detected, recovery from the failure can be automatically performed by running the pair formation process again.

Explanation of Symbols

[0171] 201 Site 210 Node

Claims

1. A storage system comprising a plurality of nodes each including a memory provided with a cache and a processor connected to the memory, wherein in each of the plurality of nodes, when there is data to be stored in a segment secured from the cache, the processor determines, according to the type of the data, for the segment, the necessity of redundancy by transferring redundant data of the data in the segment to another node and the necessity of non-volatility by storing the data in the segment in a non-volatile area, controls, based on the determination, whether to perform redundancy of the data in the segment and whether to perform non-volatility of the data in the segment, and for each of the plurality of nodes, the non-volatile area is an area based on one or more non-volatile media inside or outside the node. Storage system.

2. The plurality of nodes are a plurality of secondary nodes constituting a secondary storage system, and among the plurality of secondary nodes, a first secondary node has a secondary journal volume in which a journal including metadata including data written to a primary volume in a primary storage system and a sequence number which is the order in which the data is written is stored, a secondary volume forming a pair with the primary volume in the primary storage system, a first secondary cache which is the cache in the secondary node, and a first secondary processor which is the processor in the secondary node, and the first secondary processor secures a segment in which the journal is stored from the first secondary cache, when the data stored in the secured segment is a journal, determines that redundancy is unnecessary and non-volatility is also unnecessary for the segment, when redundancy and non-volatility are both unnecessary for the segment, writes the journal in the segment to the secondary journal volume and transfers redundant data of the journal in the segment to a second secondary node having a secondary mirror volume which is a mirror volume of the secondary volume, and does not perform redundancy which is to transfer redundant data of the journal in the segment to another node and does not perform non-volatility which is to store the journal in the segment in a non-volatile area, and for each of the plurality of secondary nodes, the volume is an area based on one or more non-volatile media inside or outside the secondary node. The storage system according to claim 1.

3. The first secondary processor Secure a segment in the first secondary cache for storing data in the journal, if the data stored in the secured segment is the data in the journal, determine that redundancy is required for the segment and determine that non-volatilization is not required for the segment, if redundancy is required for the segment but non-volatilization is not required, perform redundancy by writing the data in the segment to the secondary volume and transferring redundant data of the data in the segment to the second secondary node, and do not perform non-volatilization by storing the data in the segment in the non-volatile area, The storage system according to claim 2.

4. The positive storage system is composed of a plurality of positive nodes which are a plurality of nodes, among the plurality of positive nodes, the first positive node has a positive journal volume which is a volume storing a journal, the positive volume, a first positive cache which is a cache in the positive node, and a first positive processor which is a processor in the positive node, the first positive processor, secures a segment storing a journal from the first positive cache, if the data stored in the secured segment is a journal, determine that redundancy is required for the segment and determine that non-volatilization is also required for the segment, if both redundancy and non-volatilization are required for the segment, write the journal in the segment to the positive journal volume, transfer redundant data of the journal in the segment to a second positive node having a positive mirror volume which is a mirror volume of the positive volume, and perform both redundancy which is to transfer redundant data of the journal in the segment to a second positive node having a positive mirror volume which is a mirror volume of the positive volume and non-volatilization which is to store the journal in the segment in the non-volatile area, For each of the plurality of positive nodes, the volume is an area based on one or more non-volatile media inside or outside the secondary node, The storage system according to claim 2.

5. The first positive processor, secures a segment storing data included in the journal and storing data from the host from the first positive cache, if the data stored in the secured segment is the data from the host, determine that redundancy is required for the segment and determine that non-volatilization is also required for the segment, When redundancy and non-volatility are required for the segment, the data in the segment is written to the positive volume, and redundancy, which is to transfer the redundant data of the data in the segment to the second positive node, and non-volatility, which is to store the data in the segment in the non-volatile area, are both performed. The storage system according to claim 4.

6. When non-volatility is not required for the segment, the processor writes the data in the segment to the instance store in the node having the processor. The storage system according to claim 1.

7. The redundant data is duplicate data or parity data. The storage system according to claim 1.

8. After data is written to the secondary volume, the first secondary processor updates the sequence number of the journal that has been reflected, and notifies the positive storage system of the updated sequence number. The positive storage system purges from the positive storage system the journal having the sequence numbers up to the sequence number notified from the first secondary processor. The storage system according to claim 2.

9. When the first secondary processor detects that the pair has failed, it notifies the positive storage system of the failure of the pair. The first secondary node receives from the positive storage system, in response to the notification, a journal including data as the difference between the secondary volume and the positive volume. The first secondary processor secures a segment from the first secondary cache and stores the journal in the segment. The storage system according to claim 2.

10. The formation of the pair between the secondary volume and the positive volume is performed in response to a pair formation request from the positive storage system. In response to the notification, the first secondary node receives a pair formation request from the positive storage system. In the process performed in response to the pair formation request, the first secondary node receives from the positive storage system a journal including data as the difference. The storage system according to claim 9.

11. In each of the plurality of nodes constituting the storage system, when there is data to be stored in a segment secured from the cache, the node Based on the type of the data, for the segment, make a redundancy-non-volatile determination which is a determination of the necessity of redundancy that transfers redundant data of the data in the segment to another node, and the necessity of non-volatilization that stores the data in the segment in a non-volatile area, Based on the determination, control whether to perform redundancy of the data in the segment and whether to perform non-volatilization of the data in the segment, A storage control method.

Citation Information

Patent Citations

  • Remote copy system

    JP2005018736A

  • Information processing apparatus, method and program for managing memory

    JP2011159101A

  • Backup system and method

    JP2023011448A

  • Storage system and storage control method

    JP2023152247A