Memory system and memory control method

By setting redundant and nonvolatile flags in a multi-node storage system, the writing frequency and storage method of data are controlled according to the data type, the performance reduction and data recovery difficulties caused by high asynchronous remote replication frequency are solved, and more efficient storage performance and better data recovery capabilities are achieved.

CN120066391AInactive Publication Date: 2025-05-30HITACHI VANDALA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411247037.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-09-06
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In multi-node storage systems, the frequency of asynchronous remote replication is high, resulting in reduced performance and difficulty in data recovery in the event of node failure.

Method used

By setting redundant flags and nonvolatile flags in each node, whether to perform redundant and nonvolatile processing is determined according to the type of data, thereby controlling the writing frequency and storage method of data.

Benefits of technology

Effectively reduces the frequency of writing from cache to nonvolatile media, improves the performance of the storage system, and improves the possibility of data recovery in the event of node failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066391A_ABST
    Figure CN120066391A_ABST
Patent Text Reader

Abstract

The invention provides a storage system and a storage control method, and the problem to be solved is to appropriately reduce the frequency of writing from a cache to a non-volatile medium in a storage system composed of a plurality of nodes. In each of a plurality of nodes constituting a storage system, if there is data to be stored in a segment secured from a cache, the node stores the data to be stored in the segment in accordance with the type of the data. It is determined whether or not redundancy is required (redundant data of the data in the segment is transferred to another node) and whether or not non-volatile is required (the data in the segment is stored in the non-volatile region), and on the basis of the result of the determination, the data in the segment is stored in the non-volatile region. And controls the presence or absence of the implementation of redundancy of the data in the segment and the presence or absence of non-volatile processing of the data in the segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to storage control. Background Art

[0002] As an example of storage control, there is remote replication. Regarding remote replication, for example, the technology disclosed in Patent Document 1 is known.

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2005-18736 Summary of the Invention

[0006] Problems to be Solved by the Invention

[0007] As at least the secondary storage system among the primary storage system (the storage system in the primary site) and the secondary storage system (the storage system in the secondary site), SDS (Software Defined Storage) can be adopted. SDS is based on one or more (typically multiple) storage nodes. These storage nodes are, for example, in a local deployment environment or a cloud environment. The storage nodes (hereinafter referred to as nodes) are, for example, general-purpose computers, having a cache and a VOL (logical volume). The cache is typically a volatile memory, and the VOL is typically based on a permanent storage device.

[0008] The nodes that are the basis of SDS generally do not have a battery. Therefore, when a power failure of a node occurs, the data in the cache (typically a volatile memory) of the node may be lost. To prevent such data loss, the node performs a data protection process of writing data from the cache to the VOL.

[0009] Specifically, for example, in asynchronous remote replication, as the data stored in the cache, there are JNL (journal) and the data to be replicated from the PVOL (primary VOL) to the SVOL (secondary VOL). The JNL includes the data to be replicated (replication) and the metadata of the data.

[0010] Assume that the secondary storage system includes a first node and a second node. Assume that the first node has a first cache, a first JVOL (JNL VOL), a first memory protection area, and a first SVOL. The first JVOL, the first memory protection area, and the first SVOL are areas based on a permanent storage device inside or outside the first node. Assume that the second node has a second cache, a second JVOL, a second memory protection area, and a second SVOL. The second JVOL, the second memory protection area, and the second SVOL are areas based on a permanent storage device inside or outside the second node. The second SVOL is a mirror VOL of the first SVOL.

[0011] When the first node receives a JNL from the primary storage system, for example, the following processing is performed.

[0012] · The first node writes the received JNL to the first cache, copies the data in the JNL to the first cache, and writes the JNL and the data to the first memory protection area. In addition, for data redundancy, the first node transmits the JNL and the data stored in the first cache to the second storage node. The second node writes the JNL and the data to the second cache, and writes the JNL and the data to the second memory protection area. Thus, even if the JNL and the data disappear from the first cache due to a power failure of the first node, the possibility of recovering the JNL and the data is increased.

[0013] · The first node writes a log to the first JVOL, and then writes the JNL to the first JVOL. The first node writes a log to the first SVOL, and then writes the data in the JNL to the first SVOL. Similarly, the second node writes a log to the first JVOL, and then writes the JNL to the first JVOL. The second node writes a log to the second SVOL, and then writes the data in the JNL to the second SVOL. Thus, the data is redundant in the first SVOL and the second SVOL, and even if a failure occurs in one of the first and second nodes, the data can be recovered from the other node.

[0014] However, in this processing, since the frequency of writing from the cache to the VOL is high in asynchronous remote replication, there is a concern about a decrease in the performance of asynchronous remote replication.

[0015] Regarding a storage system composed of multiple nodes (a storage system that performs data redundancy between nodes), there is also a problem that the frequency of writing from the cache to a non-volatile medium (for example, a non-volatile medium that forms the basis of the VOL) is high for storage control other than asynchronous remote replication.

[0016] Means for Solving the Problem

[0017] In each of the multiple nodes constituting the storage system, when there is data of a storage object for a section ensured from the cache, the node determines, based on the type of the data, whether redundancy is required (transferring redundant data of the data in the section to another node) and whether non-volatilization is required (storing the data in the section in a non-volatile area) for the section, and controls the presence or absence of the implementation of redundancy of the data in the section and the presence or absence of non-volatilization of the data in the section based on the result of the determination.

[0018] Advantageous Effects of the Invention

[0019] According to the present invention, it is possible to appropriately reduce the write frequency from the cache to the non-volatile medium in a storage system constituted by multiple nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram showing an outline of an embodiment of the present invention.

[0021] Figure 2 It is a diagram showing an example of the physical configuration of the entire system.

[0022] Figure 3 It is a diagram showing an example of the configuration of the software platform of the site.

[0023] Figure 4 It is a schematic diagram showing an outline of the remote copy configuration.

[0024] Figure 5 It is a schematic diagram showing an outline of the I / O request processing.

[0025] Figure 6 It is a schematic diagram showing an outline of the recovery processing from a node failure.

[0026] Figure 7 It is a diagram showing an example of the data and programs held in the memory.

[0027] Figure 8 It is a diagram showing an example of the system configuration management table.

[0028] Figure 9 It is a diagram showing an example of the pairing configuration management table.

[0029] Figure 10 It is a diagram showing an example of the cache management table.

[0030] Figure 11 It is a diagram showing the flow of the pairing formation process.

[0031] Figure 12 It is a diagram showing the flow of the initial JNL creation process in the primary site.

[0032] Figure 13 This is a diagram showing the process of restoration processing in the secondary site.

[0033] Figure 14A This is a diagram showing the process of JNL reading processing in the primary site.

[0034] Figure 14B This is a diagram showing the process of JNL clearing processing in the primary site.

[0035] Figure 15A This is a diagram showing the process of cache storage processing in the secondary site.

[0036] Figure 15B This is a diagram showing the process of destage processing in the secondary site.

[0037] Figure 16 This is a diagram showing the process of updated JNL production processing in the primary site.

[0038] Figure 17 This is a diagram showing the process of pairing restoration processing.

[0039] Explanation of reference numerals

[0040] Site 201

[0041] Node 210 Detailed implementation mode

[0042] In the following description, the "interface device" can be one or more communication interface devices. One or more communication interface devices can be one or more of the same type of communication interface devices (for example, one or more NICs (Network Interface Cards)), or can be two or more different types of communication interface devices (for example, NICs and HBAs (Host Bus Adapters)).

[0043] In addition, in the following description, the "memory" is an example of one or more storage devices, that is, one or more storage devices, typically the main storage device. At least one of the storage devices in the memory can be a volatile storage device or a non-volatile storage device.

[0044] In addition, in the following description, the "permanent storage device" may be an example of one or more storage devices, that is, one or more permanent storage devices. A permanent storage device can typically be a non-volatile storage device (such as an auxiliary storage device). Specifically, for example, it can be an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an NVMe (Non-Volatile Memory Express) drive.

[0045] In addition, in the following description, the "processor" may be one or more processor devices. At least one processor device can typically be a microprocessor device such as a CPU (Central Processing Unit), or can also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device can be single-core or multi-core. The at least one processor device can be a processor core. At least one processor device can also be a generalized processor device such as a hardware circuit (such as an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)) that performs part or all of the processing.

[0046] In addition, in the following description, sometimes the information that obtains an output from an input is described as "xxx table", but this information can be data of any structure (for example, it can be either structured data or unstructured data), or can also be a learning model represented by a neural network, a genetic algorithm, or a random forest that generates an output for an input. Therefore, the "xxx table" can be referred to as "xxx information". In addition, in the following description, the composition of each table is an example. One table can be divided into two or more tables, and all or part of two or more tables can also be one table.

[0047] In addition, in the following description, the "program" is sometimes used as the subject to describe the processing. However, since the program is executed by a processor and appropriately uses a storage device and / or an interface device, etc. to perform the specified processing, the subject of the processing can also be the processor (or a device such as a controller having the processor). The program can also be installed from a program source into a device such as a computer. The program source can be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In addition, in the following description, two or more programs can be implemented as one program, and one program can be implemented as two or more programs.

[0048] In addition, in the following description, when the same kind of elements are not distinguished for description, the common parts in the reference numerals are used. When the same kind of elements are distinguished for description, sometimes the reference numerals or the identifiers of the elements are used. For example, for PVOL, there are cases where the reference numeral is used as in "PVOL102P1", and cases where the identifier is used as in "PVOL1".

[0049] Figure 1 It is a schematic diagram showing an outline of an embodiment of the present invention.

[0050] The main site 201P includes a host 51 and a main storage system 100P. The host 51 can be a physical computer or a logical computer (e.g., a virtual machine). The main storage system 100P can be a so-called disk array system, but in this embodiment, it is a system composed of a plurality of main nodes 210P.

[0051] The secondary site 201S has a secondary storage system 100S. The secondary storage system 100S is a system composed of a plurality of secondary nodes 210S.

[0052] The node 210 is typically a general-purpose computer, but can also be a device other than a general-purpose computer. The node 210 typically does not have a battery. The node 210 has a cache 55, a VOL (logical volume) 102, and a non-volatile area 103. The cache 55 is typically a volatile memory. The VOL102 and the non-volatile area 103 are based on a permanent storage device in or outside the node 210.

[0053] Multiple primary nodes 210P include, for example, a first primary node 210P1 and a second primary node 210P2. As VOL102, there are PJVOL102JP and PVOL102P. "PJVOL" is the JVOL located at the primary site 201P. "JVOL" is a VOL that stores a JNL (journal). The "JNL" contains the data of the replication object and its metadata. The metadata in the JNL includes a sequence number (SEQ#) that is a value for determining the order in which the data of the replication object is written or the write destination address of the data. "PVOL" is the primary VOL.

[0054] Multiple secondary nodes 210S include, for example, a first secondary node 210S1 and a second secondary node 210S2. As VOL102, there are SJVOL102JS and SVOL102S. "SJVOL" is the JVOL located at the secondary site 201S. "SVOL" is a secondary VOL that forms a pair with the PVOL.

[0055] In the present embodiment, in each of the primary storage system 100P and the secondary storage system 100S, a redundancy flag 2 and a non-volatile flag 3 are set for each section (an example of a cache area) in the cache 55.

[0056] The redundancy flag 2 indicates whether redundancy is required. When the value of the redundancy flag 2 is "on", redundancy of the data in the section in the cache 55 (replication to other nodes) is performed. When the value of the redundancy flag 2 is "off", redundancy of the data in the section in the cache 55 is not performed.

[0057] The non-volatile flag 3 indicates whether non-volatility is required. When the value of the non-volatile flag 3 is "on", non-volatility of the data in the section in the cache 55 (writing to the non-volatile area 103) is performed. When the value of the redundancy flag 2 is "off", non-volatility of the data in the section in the cache 55 is not performed.

[0058] The values of the redundancy flag 2 and the non-volatile flag 3 corresponding to the section are controlled according to the type of data in the section.

[0059] For example, in the present embodiment, the following processing is performed. In addition, in Figure 1Among them, PJVOL102JP1S in the second primary node 210P2 is the standby VOL (mirror VOL) of PJVOL102JP1A in the first primary node 210P1. PVOL102P1S in the second primary node 210P2 is the standby VOL of PVOL102P1A in the first primary node 210P1. SJVOL102JS1S in the second secondary node 210S2 is the standby VOL of SJVOL102JS1A in the first secondary node 210S1. SVOL102S2S in the second secondary node 210S2 is the standby VOL of SVOL102S1A in the first secondary node 210S1.

[0060] The first primary node 210P1 receives a write request specifying PVOL102P1A, and writes the data to be written associated with the write request into the first section of the cache 55P1. The first primary node 210P1 sets the values of the redundancy flag 2P1 and the non-volatile flag 3P1 of the first section to "on" respectively. Therefore, the first primary node 210P1 transfers the data in the first section to the second primary node 210P2 (for redundancy), and writes it into the non-volatile area 103P1 (for non-volatilization). The second primary node 210P2 receives this data, writes it into the section of the cache 55P2, and writes it into the non-volatile area 103P2. In addition, although not shown, for the section of the cache 55P2 (the section where data from the first primary node 210P1 is written), the redundancy flag can be set to "off", and the non-volatile flag can be set to "on". Thus, the redundancy of the data is skipped, and the non-volatilization of the data (writing the data into the non-volatile area 103P2) is implemented.

[0061] The first primary node 210P1 generates a JNL containing the data in the first section of the cache 55P1, and writes the JNL into the second section of the cache 55P1. The first primary node 210P1 sets the values of the redundancy flag 2P2 and the non-volatile flag 3P2 of the second section to "on" respectively. Therefore, the first primary node 210P1 transfers the JNL in the second section to the second primary node 210P2 (for redundancy), and writes it into the non-volatile area 103P1 (for non-volatilization). The second primary node 210P2 receives this data, writes it into the section of the cache 55P2, and writes it into the non-volatile area 103P2. In addition, although not shown, for the section of the cache 55P2 (the section where the JNL is written from the first primary node 210P1), the redundancy flag can also be set to "off", and the non-volatile flag can be set to "on". Thus, the redundancy of the JNL is skipped, and the non-volatilization of the JNL is implemented.

[0062] Although not shown, the first master node 210P1 writes the data in cache 55P1 to PVOL102P1A and writes the JNL in cache 55P1 to PJVOL102JP1A. Similarly, the second master node 210P2 writes the data in cache 55P2 to PVOL102P1S and writes the JNL in cache 55P2 to PJVOL102JP1S.

[0063] The first slave node 210S1 receives the JNL from the first master node 210P1. In this embodiment, the reception of the JNL is performed in response to a JNL read request (e.g., a read request specifying the SEQ# of the JNL to be read) from the first slave node 210S1 to the first master node 210P1. However, the reception of the JNL may also be a reception of a JNL write request (e.g., a write request associated with the JNL to be written) from the first master node 210P1 to the first slave node 210S1.

[0064] The first slave node 210S1 writes the received JNL to the first section of cache 55S1. The first slave node 210S1 sets the values of the redundancy flag 2S1 and the non-volatile flag 3S1 of the first section to "off". Therefore, both the redundancy and non-volatility of the JNL are skipped.

[0065] The first slave node 210S1 writes the data (the data to be copied) in the JNL in the first section of cache 55S1 to the second section of cache 55S1. The first slave node 210S1 sets the value of the redundancy flag 2S2 of the second section to "on" and the value of the non-volatile flag 3S2 to "off". Therefore, the redundancy of the data is implemented, but the non-volatility of the data is skipped. That is, the first slave node 210S1 writes the data in the second section of cache 55S1 to SVOL102S1A and transfers the data to the second master node 210P2. The second master node 210P2 receives the data, writes it to the section of cache 55S2, and writes it to SVOL102S1S. In addition, although not shown, for the section of cache 55S2 (the section written with the data (the data in the JNL) from the first slave node 210S1), both the redundancy flag and the non-volatile flag may be "off". Thus, the non-volatility of the redundancy of the data in the JNL is also skipped.

[0066] According to the above processing, the redundancy and non-volatility of the JNL in the first section of cache 55S1 are skipped, and the non-volatility of the data in the second section of cache 55S1 is skipped. Even if the redundancy and non-volatility are skipped, in the case where the JNL and data disappear from cache 55S1 due to a power failure or the like of the first slave node 210S1, the JNL and data can be restored from the main site 201P.

[0067] Specifically, in PJVOL102JP1A (and 102JP1S), the JNL containing the data is saved until the data is written to the first and second SVOL102S1A (and 102S1S). That is, when the data is written to SVOL102S1A (and 102S1S), the first secondary node 210S1 notifies the first primary node 210P1 of the SEQ# of the JNL containing the data that has been completely replicated to SVOL102S1A (and 102S1S). When the first primary node 210P1 receives the notification of the SEQ# of the JNL indicating the completion of data replication (completion of reflection to SVOL102S), the JNL with that SEQ# is cleared from JVOL102JP1A (the second primary node 210P2 also clears the JNL with that SEQ# from JVOL102JP1S). In this way, the JNL containing the data is saved in the primary site 201P until the data is written to SVOL102S1A (and 102S1S). The first secondary node 210S1 sends a JNL read request specifying the SEQ# of the JNL of the restoration object to the first primary node 210P1 in order to restore the JNL or data that has disappeared from the cache 55S1 before writing to JVOL102JS1A or SVOL102S1A. As a result, the first secondary node 210S1 can obtain the JNL with that SEQ# from the first primary node 210P1, and can also obtain the data within that JNL.

[0068] Hereinafter, this embodiment will be described in detail.

[0069] Figure 2 FIG. is a diagram showing an example of the physical configuration of the storage system 101.

[0070] It has a plurality of sites 201. Each site 201 is communicably connected via a network 202. The network 202 is, for example, a WAN (Wide Area Network), but is not limited to a WAN. The site 201 is a data center or the like and includes a plurality of (or one) nodes 210.

[0071] The node 210 may be a general-purpose computer. The node 210 includes, for example, one or more processor packages 213 including a processor 211 and a memory 212, one or more drivers 214, and one or more ports 215. These respective components are connected via an internal bus 216. The driver 214 is an example of a permanent storage device.

[0072] The processor 211 is, for example, a CPU (Central Processing Unit) and performs various processes.

[0073] The memory 212 is typically a volatile memory that stores control information required to implement the functions of the implementation node 210 or stores data. Additionally, the memory 212 stores, for example, a program executed by the processor 211. The driver 214 stores various data, programs, etc.

[0074] The port 215 is connected to the network 220 within the site 201 and connects this node to other nodes 210 within the site 201 via the network 220 in a communicable manner. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.

[0075] Note that the physical configuration of the system is not limited to the above configuration. For example, the network 202 and / or 220 can be made redundant. Additionally, for example, the network 220 can also be separated into a management network and a storage network, and the connection standard can be Ethernet (registered trademark), Infiniband, or wireless, and the connection topology is not limited to Figure 2 the configuration shown. Additionally, for example, the driver 214 can also be a configuration independent of the node 210.

[0076] Figure 3 It is a diagram showing an example configuration of the software platform of the site 201.

[0077] For example, a software platform having the Figure 3 illustrated configuration can be adopted at the secondary site 201S. The site 201 includes a network storage service 30 that provides multiple permanent memories 32 to multiple nodes 210 via the network 220. The permanent memory 32 is a storage area based on one or more drivers 214.

[0078] The node 210 has an instance store 65, a hypervisor 64, and a virtual machine 61.

[0079] The instance store 65 provides block-level temporary storage for instances. This storage can be located on the driver 214 physically attached to the node 210.

[0080] The hypervisor 64 dynamically generates or deletes the virtual machine 61.

[0081] The virtual machine 61 manages one or more virtual drives 63 and executes storage control software (SCS) 720.

[0082] The SCS 720 controls the I / O (Input / Output) to the virtual drive 63. The storage control program 720 is redundant among the nodes 210. That is, in the case of a failure of the node 210, the SCS (Active) 720 of another node 210 changes from standby to active, replacing the SCS (Standby) 720 of the node 210.

[0083] The virtual drive 63 is a storage area to which the instance memory 65 or the permanent memory 32 is assigned. The virtual drive 63 can be handled as VOL102.

[0084] In this way, in the site 201, the instance memory 65 based on DAS (Direct Attached Storage) and the memory via the network 220 such as iSCSI (network storage service 30) are used. The hypervisor 64 may not be present. For example, DAS and the network storage service 30 may be configured by bare metal.

[0085] Figure 4 It is a schematic diagram showing an outline of the remote replication configuration.

[0086] A remote replication pair is established between the main site 201P and the secondary site 201S among multiple VOL102s. Specifically, for example, two consistency groups 401a and 401b are established between the main site 201P and the secondary site 201S. The consistency group 401 is composed of multiple (or one) VOL102s of the remote replication pair. In the consistency group 401, multiple PVOLs are replicated to the SVOL in a state where consistency is maintained. More specifically, for example, in the consistency group 401, the update differential data up to the same time for multiple PVOL102s is replicated to multiple SVOLs. In addition, the control (consistency control) of the consistency group 401 is managed by the PJVOL. In the PJVOL, the update differential data of multiple (or one) PVOL102Ps and metadata such as its write time are stored. When the main site 201P transfers the data of the PVOL to the secondary site 201S, the update differential data up to that time written in the PJVOL is transferred to the secondary site 201S. Thereby, data can be replicated to the SVOL while maintaining the matching of the update times among multiple PVOLs.

[0087] For example, according to consistency group 401a, while maintaining the matching state of PVOL1 and 2 located at the primary node 210P1, data is replicated to SVOL1 and 2 located at the secondary node 210S1 via PJVOL1 and SJVOL1. According to consistency group 401b, while maintaining the matching state of PVOL3 located at the primary node 210P2 and PVOL4 located at the primary node 210P3, data is replicated to SVOL3 located at the secondary node 210S2 and SVOL4 located at the secondary node 210S3 via PJVOL2 located at the primary node 210P2, PJVOL3 located at the primary node 210P3, SJVOL2 located at the secondary node 210S2, and SVOVOL3 located at the secondary node 210S3. PJVOL and SJVOL do not have to correspond to 1:1 (e.g., 1:many, many:1, or many:many), and PJVOL can be a region on the memory 212.

[0088] From the above specific configuration, it can be seen that the consistency group 401 can be composed of VOL102 within a specific node 210 within the site 201, or can be composed of VOL102 of multiple nodes 210 located within the site 201.

[0089] Figure 5 It is a schematic diagram showing an overview of I / O request processing.

[0090] First, the application program 502 operating on the host 51 issues a write request specifying PVOL1 to the primary node 210P1. The primary node 210P1 that receives the write request writes data A and B associated with the write request to PVOL1, and further writes a JNL containing data A and B as updated differential data to PJVOL1.

[0091] Next, the primary node 210P1 transfers the JNL (updated differential data) written to PJVOL1 to SJVOL1 and SJVOL1 (standby) of the secondary site 201S. At this time, when multiple communication paths are established between the primary site 201P and the secondary site 201S, any communication path can be used to transfer the JNL. Usually, the primary node 210P1 transfers the JNL to the secondary node 210S1 that has the ownership of SVOL1 paired with PVOL1. However, when a failure occurs in the communication path with the ownership, the primary node 210P1 can also transfer the JNL to a secondary node 210S2 that does not have the ownership, etc. For example, when the primary node 210P1 transfers the JNL to the secondary node 210S2 that does not have the ownership, the secondary node 210S2 transfers the received JNL to the secondary node 210S1 that has the ownership, and the secondary node 210S1 writes the JNL to SJVOL1.

[0092] Next, the secondary node 210S1 writes the JNL written to SJVOL1 to SVOL1. Then, data A and B written to SVOL1 are written to the drive 214a via the storage pool 504a. In the case where the configuration of the drive 214a is DAS (Direct Attached Storage) in which the node 210 and the drive 214 are connected one-to-one, the JNL is written to the drive 214a mounted on the secondary node 210S1. By writing all the data copied to SVOL1 to the drive 214a of the secondary node 210S1 that has the ownership of SVOL1 like this, when reading data from SVOL1 later, there is no need to read data from other nodes. As a result, the inter-node transfer process can be excluded, and high-speed read processing can be achieved.

[0093] In addition, the storage pool 504 may be an area based on one or more drives 214. Provide storage functions such as Thin-provisoning, compression, or deduplication, and perform processing of the storage functions required for the data written to the storage pool 504.

[0094] In order to protect the data from node failures when writing data to the drive 214a, the secondary node 210S1 also writes redundant data of the written data to the drive 214b of the secondary node 210S2. For the writing of redundant data, when the data protection policy is Replication, a copy of the data is written as redundant data to the drive 214b. On the other hand, when the data protection policy is Erasure Coding, parity is calculated based on the data, and the calculated parity is written as redundant data to the drive 214b.

[0095] Figure 6 It is a schematic diagram showing an outline of the recovery process from a node failure.

[0096] SCS720 operates in the secondary nodes 210S1, 210S2, and 210S3. The secondary node 210S has an SCS (active) of the working system and an SCS (standby) of the standby system corresponding to the SCS (active) in another secondary node 210S. For example, the secondary node 210S1 has SCS1 (active) and SCS3 (standby), the secondary node 210S2 has SCS2 (active) and SCS1 (standby), and the secondary node 210S3 has SCS3 (active) and SCS2 (standby). SCSx (active) and SCSx (standby) belong to the redundancy group of SCSx, and there can be more than one SCSx (standby) (x is a natural number).

[0097] Use Figure 6A specific example is shown to illustrate the recovery process from a node failure.

[0098] In order to inherit the remote replication pairing information of the secondary node 210S1, the secondary node 210S2 copies and holds the configuration information of SVOL1 and SJVOL1 that the secondary node 210S1 has. In addition, the secondary node 210S2 stores the redundant data of the data written to the drive 214a of the secondary node 210S1 in the drive 214d. Furthermore, a communication path is established between the secondary node 210S2 and the primary node 210P1.

[0099] For example, when the secondary node 210S1 stops due to a failure, the secondary node 210S2 that detects the failure of the secondary node 210S1 takes over the processing of SCS1 (activated) of the secondary node 210S1, and SCS1 (standby) becomes SCS1 (activated). The secondary node 210S2 communicates with the primary node 210P1 and continues the remote replication process between PVOL1 and SVOL1. That is, failover is performed from SCS1 of the secondary node 210S1 to SCS1 of the secondary node 210S2. Thus, even if a node failure occurs in one secondary site 201S, the other secondary site 201S can continue the remote replication from the primary site 201P.

[0100] Figure 7 It is a diagram showing an example of the data and programs held in the memory 212.

[0101] Information is read from the drive 214 to the memory 212. For example, various tables included in the control information table 710 and various programs included in the SCS720 are expanded on the memory 212 during the execution of the respective processes they are used for, but when not, they are stored in a non-volatile storage area such as the drive 214 in advance to cope with power outages, etc.

[0102] The control information table 710 includes a system configuration management table 711, a pairing configuration management table 712, and a cache management table 713.

[0103] The SCS720 includes a pairing formation processing program 721, an initial JNL creation processing program 722, a restoration processing program 723, a JNL reading processing program 724, a JNL clearing processing program 725, a cache storage processing program 726, a downgrading processing program 727, an updated JNL creation processing program 728, and a pairing restoration processing program 729.

[0104] Figure 8 It is a diagram showing an example of the system configuration management table 711.

[0105] The system configuration management table 711 includes a node configuration management table 810, a drive configuration management table 820, and a port configuration management table 830. In addition, for each site 201, the node configuration management table 810 exists for multiple nodes 210 present in each site 201, and for a node 210, the drive configuration management table 820 and the port configuration management table 830 exist regarding the drives 214 within the node 210.

[0106] The node configuration management table 810 is set for each site 201 and stores information representing the configuration of the nodes 210 (such as the relationship between the nodes 210 and the drives 214) set in the site 201. More specifically, the node configuration management table 810 stores information such as a node ID 811, a status 812, a drive ID list 813, and a port ID list 814 for each node 210.

[0107] The node ID 811 is the ID of the node 210. The status 812 represents the status of the node 210 (e.g., "Normal", "Warning", or "Failure", etc.). The drive ID list 813 is a list of the IDs of the drives 214 set in the node 210. The port ID list 814 is a list of the IDs of the ports 215 set in the node 210.

[0108] The drive configuration management table 820 is set for each node 210 and stores information representing the configuration of the drives 214 set in the node 210. More specifically, the drive configuration management table 820 stores information such as a drive ID 821, a status 822, and a size 823 for each drive 214.

[0109] The drive ID 821 is the ID of the drive 214. The status 822 represents the status of the drive 214. The size 823 represents the capacity of the drive 214.

[0110] The port configuration management table 830 is set for each node 210 and stores information representing the configuration related to the ports 215 set in the node 210. More specifically, the port configuration management table 830 stores information such as a port ID 831, a status 832, and an address 833 for each port.

[0111] The port ID 831 is the ID of the port 215. The status 832 represents the status of the port 215. The address 833 represents the address on the network assigned to the port 215. The form of the address can be an IP (Internet Protocol), a WWN (World Wide Name), a MAC (Media Access Control) address, etc.

[0112] Figure 9 This is a diagram showing an example of the pairing configuration management table 712.

[0113] The pairing configuration management table 712 is configured to include a VOL management table 910, a pairing management table 920, and a JNL management table 930.

[0114] The VOL management table 910 stores information representing the configuration of VOL102. More specifically, the VOL management table 910 stores information such as VOL ID 911, owner node ID 912, fallback destination node ID 913, size 914, and attribute 915 for each VOL102.

[0115] The VOL ID 911 is the ID of VOL102. The owner node ID 912 is the ID of the node 210 that has the ownership of VOL102. The fallback destination node ID 913 is the ID of the node 210 that inherits the processing in case of a failure of the node 210 that has the ownership of the SVOL. The size 914 represents the capacity of VOL102.

[0116] The attribute 915 represents the attribute of VOL102. "NML_VOL" refers to a normal VOL that does not belong to a consistency group. "PAIR_VOL" refers to a PVOL or SVOL that belongs to a consistency group. "JNL_VOL" refers to a JVOL.

[0117] The pairing management table 920 stores information representing the configuration related to remote replication pairing. More specifically, the pairing management table 920 stores information such as pairing group ID 921, PJVOL ID 922, PVOL ID 923, SJVOL ID 924, SVOL ID 925, and status 926 for each consistency group.

[0118] The pairing group ID 921 is the ID of the consistency group. The PJVOL ID 922 is a list of the IDs of the PJVOL102JP that belongs to the consistency group. The PVOL ID 923 is a list of the IDs of the PVOL102P that belongs to the consistency group. The SJVOL ID 924 is a list of the IDs of the SVOL102JS that belongs to the consistency group. The SVOL ID 925 is a list of the IDs of the SVOL102S that belongs to the consistency group. The status 926 represents the status of each remote replication pairing in the consistency group (e.g., "PAIR", "COPY", "SUSPEND", etc.). "PAIR" is the state where the writes to PVOL102P are periodically reflected to SVOL102S. "COPY" is the state during initial replication. "SUSPEND" is the pairing interruption state (the state where synchronization between PVOL102P and SVOL102S is not performed).

[0119] The JNL management table 930 stores information about JNLs. More specifically, the JNL management table 930 stores information such as the pairing group ID 931, JNL ID 932, P / SVOL ID 933, P / SVOL address 934, size 935, and cache segment ID 936 for each JNL.

[0120] The pairing group ID 931 is the ID of the consistency group to which the JNL belongs. The JNL ID 932 is the ID of the JNL. The ID of the JNL is equivalent to the SEQ#, for example, a serial number in the consistency group. That is, the ID of the JNL represents the write order, and the data within the JNL is stored in the SVOL 102S in the consistency group in the order of the JNL ID.

[0121] The P / SVOL ID 933 contains the ID of the PVOL 102P in which the data written into the JNL is stored and the ID of the SVOL 102S in which the data written into the JNL is stored. The P / SVOL address 934 contains the storage destination address of the data in the PVOL 102P where the data written into the JNL is stored and the storage destination address of the data in the SVOL 102S where the data written into the JNL is stored.

[0122] The size 935 represents the size of the JNL. For example, one JNL contains one or more data. The cache segment ID 936 is the ID of the cache segment in which the data written into the JNL is stored.

[0123] Figure 10 It is a diagram showing an example of the cache management table 713.

[0124] The cache management table 713 is a table related to the cache 55. The cache management table 713 contains a dirty queue 1001, a clean queue 1002, a free queue 1003, and a cache segment management table 1004. In addition, in the present embodiment, there are multiple different segment sizes as the segment size, and segments of multiple different segment sizes are prepared in advance. However, the segment size can also be a variable size corresponding to the request ensured in the cache.

[0125] The dirty queue 1001 is a queue that stores the ID (address) of the dirty segment in which the dirty data for which the drive 214 becomes the write destination is stored for each drive 214. A "dirty segment" is a segment that stores dirty data. "Dirty data" is data that has not been written to the drive 214.

[0126] The clean queue 1002 is a queue that has the ID (address) of the clean segment of that segment size for each segment size. A "clean segment" refers to a segment that stores clean data. "Clean data" refers to data that has been written to the drive 214.

[0127] The free queue 1003 is a queue of IDs (addresses) of free segments each having a segment size. A "free segment" is a segment into which new data can be written. A segment of a desired size is ensured from the free queue 1003, and data is written into the ensured segment.

[0128] The cache segment management table 1004 stores information about cache segments. More specifically, the cache segment management table 1004 stores information such as segment ID 1041, memory address 1042, size 1043, VOL ID 1044, VOL address 1045, and redundancy-non-volatility flag 1046 for each cache segment.

[0129] The segment ID 1041 is the ID of the cache segment. The memory address 1042 is the address of the cache segment (the address in the buffer 55). The size 1043 is the size of the cache segment.

[0130] The VOL ID 1044 is the ID of the VOL 102 that is the write destination of the data in the cache segment. The VOL address 1045 is the address in the VOL 102 that is the write destination (the address of the data write destination).

[0131] The redundancy-non-volatility flag 1046 includes a redundancy flag 2 and a non-volatility flag 3 corresponding to the cache segment. "1" indicates "on", and "0" indicates "off".

[0132] Hereinafter, an example of the processing performed in the present embodiment will be described. In the following description, to avoid confusion, "P" is appended to the reference numerals of the programs in the primary site 201P, and "S" is appended to the reference numerals of the programs in the secondary site 201S. Also, in the following description, an "initial JNL" refers to a JNL generated in the initial replication, and an "update JNL" refers to a JNL generated based on the update of the PVOL 102P after the initial replication. Figures 11 to 17 is a diagram showing the flow of the pairing formation process.

[0133] Figure 11 According to the pairing formation process, a remote replication pair is formed through communication between the primary site 201P and the secondary site 201S. In the following description, the pairing formation process program 721P is a program in the primary node 210P having a PVOL candidate. The pairing formation process program 721S is a program in the secondary node 210S having an SVOL candidate.

[0134] Figure 11

[0135] ​​The pairing formation handler 721P sends a pre-check request to the pairing formation handler 721S (S1101). The pairing formation handler 721S receives the request (S1102), and performs a prescribed pre-check such as whether there is a VOL102 for pairing formation object or whether the information of the other device is correct. The pairing formation handler 721S returns a response to the pre-check request (S1103). This response indicates the result of the pre-check. The pairing formation handler 721P receives this response (S1104).

[0136] If this response is a prescribed response, the pairing formation handler 721P sends a pairing formation request with the PVOL candidate set to PVOL102P and the SVOL candidate set to SVOL102S to the pairing formation handler 721S (S1105). The pairing formation handler 721S receives this request (S1106), forms a VOL pairing (registers information in the pairing management table 920), and sets the status 926 of this pairing to "copy" (S1107). The pairing formation handler 721S starts the restoration process ( Figure 13 )(S1108), and returns a response to the pairing formation request (S1109). The pairing formation handler 721S waits for the completion of the initial copy (S1110). When the initial copy is completed (S1111: yes), specifically, when the synchronization of the data of PVOL and the data of SVOL is completed by the restoration handler 723S, the pairing formation handler 721S sets the status 926 of this pairing to "paired" (S1112).

[0137] The pairing formation handler 721P receives the response sent in S1109 (S1113). If this response is a prescribed response, the pairing formation handler 721P forms a VOL pairing (registers information in the pairing management table 920), and sets the status 926 of this pairing to "copy" (S1114). If there is a resynchronization option in this pairing (S1115: yes), the pairing formation handler 721P sets the resynchronization option (S1116).

[0138] The pairing formation handler 721P starts the initial JNL production process (S1117). The pairing formation handler 721P waits for the completion of the initial copy (S1118). When all the initial JNLs are cleared (S1119: yes), specifically, when the synchronization of the data of PVOL and the data of SVOL is completed by the restoration handler 723S, the pairing formation handler 721P sets the status 926 of this pairing to "paired" (S1120).

[0139] Figure 12This is a diagram showing the process of initial JNL creation processing in the main site 201P.

[0140] When the initial JNL creation processing is started, the following processing is performed. Figure 12 In this processing, the data of PVOL102P is generated as the initial JNL. The JNL is basically placed on the cache 55P, but is demoted to PJVOL102JP when the cache is full, and the demoted amount is cleared from the cache 55P. The initial replication includes full replication and differential replication. When the pairing state 926 is "paused", only the update differentials during pairing interruption are used to create the JNL by differential replication. Redundancy flags and non-volatile flags are set for the cache sections, and writing is controlled according to these flags.

[0141] If the resynchronization option is not set (S1201: No), the initial JNL creation processing program 722P uses the data in the entire area of PVOL102P as the object for creating the initial JNL (S1202). If the resynchronization option is set (S1201: Yes), the initial JNL creation processing program 722P refers to a differential management table (not shown) indicating the difference between PVOL102P and SVOL102S, and uses the data in the differential area as the object for creating the initial JNL (S1203).

[0142] The initial JNL creation processing program 722P secures a cache section for the initial JNL from the free queue 1003 (S1204), and obtains the VOL address 1045 corresponding to this section (the PVOL address of the PVOL area containing the data in the initial JNL) (S1205). The initial JNL creation processing program 722P reads the data from this address (S1206), creates the metadata of the initial JNL (S1207), creates an initial JNL containing the data read in S1206 and the metadata created in S1207, and uses this initial JNL as the storage object (S1208). The metadata includes, for example, JNL ID, LBA (VOL address), transfer length, pairing ID, PVOL ID, and SVOL ID.

[0143] The initial JNL creation processing program 722P sets both the redundancy flag and the non-volatile flag in the redundancy-non-volatile flag 1046 to "1" for the section secured in S1204 (S1209). The initial JNL creation processing program 722P starts the cache storage processing ( Figure 15A )(S1210).

[0144] When the cache usage rate (e.g., the ratio of the total capacity of the dirty and clean segments to the total capacity of cache 55P) exceeds a specified value (S1211: Yes), the initial JNL production handler 722P designates the initial JNL in the cache segment as a downgrade target (S1212) and starts the downgrade process ( Figure 15B )(S1213). The initial JNL production handler 722P releases the cache segment having the initial JNL downgraded by PJVOL102JP, which then becomes a free segment (S1214).

[0145] If there is an initial JNL that has not been produced (S1215: No), the process returns to S1204. If all initial JNLs have been produced (S1215: Yes), the initial JNL production process ends.

[0146] Figure 13 FIG. is a diagram showing the process of the restoration process in the secondary site 201S.

[0147] When the restoration process is started, the process shown in Figure 13 is performed. In this process, the JNL from the primary site 201P is reflected in SVOL102S. By setting the redundancy flag and the non-volatile flag of the cache segment in which the JNL from the primary site 201P is written to "0" respectively, the processing load is reduced. When the JNL disappears in the event of a secondary node failure, the JNL is retransmitted from the primary site 201P, and the synchronization state between PVOL and SVOL is restored. After the JNL is reflected in SVOL102S, a purge notice is sent to the primary site 102P to discard the JNL in the primary site 102P.

[0148] The restoration handler 723S stands by for a certain period of time (S1301) and sends a JNL read request (S1302). In this request, the ID (e.g., SEQ#) of the JNL to be read can be specified. This request is sent to the primary node 210P of PVOL having the metadata representation of the JNL. In response to this request, S1401 of Figure 14A is performed.

[0149] The restoration handler 723S secures the cache segment of the JNL to be read from the free queue 1003 (S1303).

[0150] In the case where S1412 of Figure 14A is performed, the restoration handler 723S receives the response to the JNL read request sent in S1302 (S1304). The restoration handler 723S stores the JNL included in this response in the buffer (S1305).

[0151] The restoration processing program 723S sets both the redundancy flag and the non-volatile flag in the redundancy-non-volatile flag 1046 to "0" for the section ensured in S1303 (S1306). The restoration processing program 723S sets the JNL stored in S1305 as an object to be cached (S1307) and starts the cache storage processing( Figure 15A )(S1308).

[0152] When the cache usage rate exceeds a specified value (S1309: Yes), the restoration processing program 723S sets the JNL in the cache section as an object to be demoted (S1310) and starts the demotion processing( Figure 15B )(S1311). The restoration processing program 723S releases the cache section having the tiered JNL in SJVOL102JS (S1312).

[0153] When the determination result in S1309 is false (S1309: No), or after S1312, the restoration processing program 723S ensures a cache section for the copy destination of the data in the JNL stored in the section in S1308 from the free queue 1003 (S1313). The restoration processing program 723S sets both the redundancy flag and the non-volatile flag to "0" for the section ensured in S1313 (S1314). The restoration processing program 723S sets the data in the JNL stored in the section in S1308 as an object to be cached (S1315) and starts the cache storage processing( Figure 15A )(S1316).

[0154] The restoration processing program 723S sets the redundancy flag to "1" and the non-volatile flag to "0" for the section ensured in S1313 (S1317). The restoration processing program 723S sets the data in the section ensured in S1313 as an object to be demoted (S1318) and starts the demotion processing( Figure 15B )(S1319). The restoration processing program 723S releases the cache section having the demoted data in SVOL102S (S1320). The restoration processing program 723S updates the ID (SEQ#) of the JNL for which restoration is completed (reflection is completed), and sends a purge notification including the updated JNL ID to the main site 201P (S1322). In response to this notification, Figure 14B of S1451 is performed. In the case where Figure 14B of S1453 is performed, the restoration processing program 723S receives a response to the notification sent in S1322 (S1323). The process returns to S1301.

[0155] Figure 14AThis is a diagram showing the process of JNL reading processing in the main site 201P.

[0156] The JNL reading processing program 724P receives a JNL reading request (S1401) and determines whether there is an untransmitted JNL in the secondary site 201S (S1402). For example, it can be determined whether the JNL ID specified by the request is the ID of an untransmitted JNL.

[0157] If the determination result in S1402 is true (S1402: Yes), the JNL reading processing program 724P determines whether the untransmitted JNL is cached (S1403).

[0158] If the determination result in S1403 is false (S1403: No), the JNL reading processing program 724P ensures a cache section for the JNL (S1404), reads the JNL from PJVOL102JP into a buffer (S1405). The JNL reading processing program 724P sets both the redundancy flag and the non-volatile flag to "0" for the section ensured in S1404 (S1406). The JNL reading processing program 724P uses the JNL read in S1405 as an object to be cached in the cache (S1407) and starts the cache storage processing ( Figure 15A )(S1408). The JNL reading processing program 724P includes the JNL in the response to the JNL reading request received in S1401 (S1410) and returns the response (S1412).

[0159] If the determination result in S1403 is true (S1403: Yes), the JNL reading processing program 724P obtains the JNL to be transmitted from the cache 55P (S1409), includes the JNL in the response to the JNL reading request (S1410), and returns the response (S1412). The "JNL to be transmitted" can be the JNL with the smallest JNL ID among the untransmitted JNLs.

[0160] If the determination result in S1402 is false (S1402: No), the JNL reading processing program 724P includes a value indicating no JNL in the response to the JNL reading request (S1411) and returns the response (S1412).

[0161] Figure 14B This is a diagram showing the process of JNL clearing processing in the main site 201P.

[0162] The JNL clearing handler 725P receives a clearing notification containing a JNL ID (S1451), clears all JNLs with JNL IDs up to the JNL ID indicated by the notification from PJVOL102JP, and returns a response indicating completion (S1453). In addition, the clearing notification may be included as a parameter in the JNL read request.

[0163] Figure 15A It is a diagram showing the process of cache storage processing in the secondary site 201S.

[0164] When the cache storage processing is started, the following Figure 15A processing is performed. In this processing, the JNL or data of the cache storage object is stored in the ensured cache section, and the implementation of redundancy and non-volatility of the JNL or data is controlled according to the flag.

[0165] The cache storage handler 726S stores the JNL or data in the ensured cache section (S1501).

[0166] When the redundancy flag corresponding to this section is "1" (S1502: Yes), the cache storage handler 726S transfers the JNL or data in this section to another secondary node 210S (S1503). In other words, when the redundancy flag corresponding to this section is "0" (S1502: No), the cache storage handler 726S skips transferring the JNL or data in this section to another secondary node 210S.

[0167] When the non-volatility flag corresponding to this section is "1" (S1504: Yes), the cache storage handler 726S writes the JNL or data in this section to the non-volatile area 103S (S1505). In other words, when the non-volatility flag corresponding to this section is "0" (S1504: No), the cache storage handler 726S skips writing the JNL or data in this section to the non-volatile area 103S.

[0168] Figure 15B It is a diagram showing the process of degradation processing in the secondary site 201S.

[0169] When the degradation processing is started, the following Figure 15B processing is performed. In this processing, the JNL is written from the section of the cache 55S to SJVOL102JS, or the data is written from the section to SVOL102S. At this time, if the redundancy flag corresponding to the section is "1", the JNL or data is redundant in another secondary node. When the redundancy flag is "0", the JNL or data may also be written in the instance memory of the secondary node.

[0170] The downgrade handler 727S selects the JNL or data (S1551) to be downgraded. Specifically, it selects a dirty section from the dirty queue 1001.

[0171] When the redundancy flag corresponding to the section with the selected JNL or data is "1" (S1552: Yes), the downgrade handler 727S transmits the JNL or data to another secondary node 210S (S1553), and writes the JNL or data to the SJVOL102JS or SVOL102S of this secondary node 210S (S1555).

[0172] When the non-volatile flag corresponding to the section with the selected JNL or data is "0" (S1552: No, S1554: Yes), the downgrade handler 727S writes the JNL or data to the instance memory of this secondary node 210S (S1556). In addition, S1555 can be performed instead of S1556.

[0173] Figure 16 It is a diagram showing the flow of the update JNL production process in the primary site 201P.

[0174] In this process, an update JNL is produced in response to a write request from the host 51. The update JNL is transmitted to the secondary site 201S through the restoration process of the secondary site 201S in the same way as the initial JNL, and is reflected in the SVOL102S. If the update JNL in the primary site 201P disappears due to a node failure, the update JNL cannot be restored. Therefore, the data and JNL on the cache are redundant and non-volatile.

[0175] The update JNL production handler 728P receives a write request (S1601), and stores the write data (the data to be written) attached to the write request in the buffer (S1602). The update JNL production handler 728P secures a cache section for the write data from the free queue 1003 (S1603), sets the redundancy flag of this section to "1", and sets the write data as the storage object (S1605).

[0176] The update JNL production handler 728P determines whether the attribute 915 of the VOL102 specified by the write request is "PAIR_VOL", that is, determines whether this VOL102 is the PVOL102P (S1606).

[0177] When the determination result in S1606 is true (S1606: Yes), update whether the status 926 of the JNL production processing program 728P determining the pairing including this VOL is "paused" (S1607). When the determination result in S1607 is true (S1607: Yes), update the JNL production processing program 728P to update the differential management table indicating the difference between this PVOL102P and SVOL102S according to the address of the write destination in accordance with the write request (S1608).

[0178] When the determination result in S1607 is false (S1607: No), update the JNL production processing program 728P to ensure the cache section for updating the JNL (S1609). Update the JNL production processing program 728P to produce the metadata for updating the JNL (S1610), store the updated JNL as a cache storage object (S1611), and set the redundancy flag of this section to "1" (S1612).

[0179] When the determination result in S1606 is false (S1606: No), after S1608 or after S1612, update the JNL production processing program 728P to start the cache storage process Figure 15A )(S1613), and return the response to the write request to the host 51 (S1614).

[0180] The JNL production processing program 728P sets the write data as the object to be degraded, and determines whether the cache usage rate exceeds the specified value (S1616). When the determination result in S1616 is true (S1616: Yes), the JNL production processing program 728P sets the updated JNL as the object to be degraded (S1617).

[0181] When the determination result in S1616 is false (S1616: No), or after S1617, update the JNL production processing program 728P to start the degradation process Figure 15B )(S1618), release the cache section where the updated JNL to be degraded exists (S1619: Yes, S1620), and release the cache section where the write data to be degraded exists (S1621).

[0182] In addition, in the Figure 16 processing shown, for either the cache section of the write data or the cache section of the updated JNL, the JNL production processing program 728P can also set the non-volatile flag to "1", and write the write data and the updated JNL in these sections to the non-volatile area 103P.

[0183] Figure 17 It is a diagram showing the process of the pairing recovery process.

[0184] In this process, the secondary site 201S detects a failure, changes the state of the pairing to "paused", and notifies the primary site 201P of the failure. In response to this notification, the JNL is retransmitted from the primary site 201P to the secondary site 201S. The retransmitted JNL is used to resume the pairing. For example, if the network 202 is temporarily unavailable and the state of the pairing becomes "paused", but later, when the network 202 is restored, the pairing restoration process can be performed. In the present embodiment, the pairing restoration is automatically performed when the state of the pairing changes to "paused". As the "failure" mentioned in this paragraph, in addition to network failures, there may also be node failures, power failures, or drive failures. It is also possible that the state of the pairing in the primary site 201P is set to "paused" in response to a failure in the primary site 201P, and this state is notified to the secondary site 201S, and the pairing restoration can also be performed in response to this notification.

[0185] The pairing restoration processing program 729S detects a failure in the secondary site 201S (S1701), and sets the state 926 of the VOL pairing affected by this failure to "paused" (S1702). The pairing restoration processing program 729S determines whether the number of executions of the series of processes from S1704 to S1706 exceeds the retry threshold (S1703).

[0186] When the determination result in S1703 is false (S1703: No), the pairing restoration processing program 729S executes the series of processes from S1704 to S1706. That is, the pairing restoration processing program 729S notifies the primary node 210P having the PVOL102P belonging to the pairing that the state 926 of the pairing set to "paused" in S1702 has failed (S1704), monitors the restoration of the pairing (S1705), and determines whether the state has returned to normal (S1706). When the determination result in S1705 is true (S1705: Yes), the processing of the pairing restoration processing program 729S ends. When the determination result in S1705 is false (S1705: No), the process returns to S1703.

[0187] The pairing restoration processing program 729P receives the notification of the failure of the pairing from the pairing restoration processing program 729S (S1751), sets the resynchronization option (S1752), and starts the pairing formation process ( Figure 11 )(S1753).

[0188] As described above, an embodiment of the present invention has been described, but this is an example for explaining the present invention and is not intended to limit the scope of the present invention to this embodiment. The present invention can also be implemented in various other ways.

[0189] In addition, the above description can be summarized as follows. The following summary may include supplementary explanations and variations of the above description. Further, in the following description, the subject of processing is the processor 211. Specifically, for example, the subject of processing may be the SCS720. Further, in the following description, for reference, the reference numerals of the elements are mainly referred to Figure 1 to describe the reference numerals of the elements.

[0190] There are a plurality of nodes 210 each including a memory 212 provided with a cache 55 and a processor 211 connected to the memory 212. In each of the plurality of nodes 210, when there is data to be stored in a section ensured from the cache 55, the processor 211 determines, based on the type of the data, whether redundancy (transferring redundant data of the data in the section to another node) is required and whether non-volatilization (storing the data in the section in the non-volatile area 103) is required for the section, and controls the presence or absence of the implementation of the redundancy of the data in the section and the presence or absence of the non-volatilization of the data in the section based on the determination. For each of the plurality of nodes 210, the non-volatile area 103 is an area based on one or more non-volatile media in or outside the node 210 (the non-volatile media may be referred to as a permanent storage device, and an example of the non-volatile media is the drive 214). Thus, in the storage system 100 constituted by the plurality of nodes 210, the write frequency from the cache 55 to the non-volatile media can be appropriately reduced.

[0191] The following description uses, as an example, determining the value of the redundancy flag 2 as the determination of whether redundancy is required. However, in the present invention, the determination of whether redundancy is required is not limited to determining the value of the redundancy flag 2. Similarly, the following description uses, as an example, determining the value of the non-volatilization flag 3 as the determination of whether non-volatilization is required. However, in the present invention, the determination of whether redundancy is required is not limited to determining the value of the non-volatilization flag 3 (for example, information in a form other than a flag may be used). Further, in the following description, the "redundant data" may be replicated data (for example, a replication of a JNL and data included in the JNL) or parity data (parity data of data included in a JNL and a JNL).

[0192] It is possible that the multiple nodes 210 are multiple secondary nodes 210S that constitute the secondary storage system 100S. It is possible that the first secondary node 210S1 among the multiple secondary nodes 210 has the SJVOL102JS1A of the VOL102 stored as the JNL, the SVOL102S1A that forms a pair with the PVOL102P1A in the primary storage system 100P, the first secondary cache 55S1 as the cache 55 in this secondary node 210S1, and the first secondary processor (hereinafter, for convenience, denoted as "211S1") of the processor 211 in this secondary node 210S1. The JNL may include the data written to the PVOL102P1A in the primary storage system 100P and the metadata including the SEQ# (sequence number) as the order of writing this data. In the above embodiment, an example of the SEQ# is the JNL ID. It is possible that the first sub-processor 211S1 secures a section for storing the JNL from the first sub-cache 55S1. When the data stored in the secured section is the JNL, the first sub-processor 211S1 may set both the redundancy flag 2 and the non-volatile flag 3 to "off". For this section, when both the redundancy flag 2 and the non-volatile flag 3 are "off", the first secondary processor 211S1 may also not perform the redundancy process of writing the JNL in this section to the SJVOL102JS1A, transmitting the redundant data of the JNL in this section to the second secondary node 210S2 having the mirror VOL of the SVOL102S1A, and the non-volatilization of storing the JNL in this section in the non-volatile area 103S1. Thus, in the asynchronous remote replication using the JNL, it is possible to appropriately reduce the writing frequency from the cache 55S to the non-volatile medium. In addition, even if such a writing frequency is reduced, when the JNL or data disappears in the secondary node 210S1 or 210S1, it is possible to restore the JNL or data using the JNL from the primary storage system 100P. In addition, for each of the multiple secondary nodes 210S, the VOL102 (SJVOL102JS or SVOL102S) may be an area based on one or more non-volatile media inside or outside this secondary node 210S.

[0193] It is possible that the first secondary processor 211S1 secures a section for storing data in the JNL from the first secondary cache 55S1. When the data stored in the secured section is the data in the JNL, for this section, the redundancy flag 2 is set to "on" and the non-volatile flag 3 is set to "off". It is possible that when the redundancy flag 2 is "on" and the non-volatile flag 3 is "off" for this section, the first secondary processor 211S1 writes the data in this section to SVOL102S1A, performs redundancy for transferring the redundant data of the data in this section to the second secondary node 210S2, and does not perform non-volatilization for storing the data in this section in the non-volatile area 103S1. Thus, it is possible to appropriately reduce the writing frequency from the cache 55S to the non-volatile medium in asynchronous remote replication. Further, it is possible that the second secondary processor (hereinafter, for convenience of explanation, labeled "211S2") of the processor 211 in the second secondary node 210S2 secures a section from the second secondary cache 55S2 which is the cache 55 in this secondary node 210S2, stores the redundant data from the first secondary node 210S1 in this section, and writes the data in this section to SVOL102S1S. For this section, since the data in this section is the redundant data from another secondary node 210S, the second secondary processor 211S2 may also set both the redundancy flag 2 and the non-volatile flag 3 to "off".

[0194] The main storage system 100P may be composed of a plurality of main nodes 210P1 that are a plurality of nodes 210. The first main node 210P1 among the plurality of main nodes 210P1 may have PJVOL102PJ1A which is VOL102 for storing JNL, the PVOL102P1A, the first main cache 55P1 which is the cache 55 in this main node 210P1, and the first main processor (hereinafter, for convenience, denoted as "211P1") which is the processor 211 in this main node 210P1. The first main processor 211P1 may secure a section for storing JNL from the first main cache 55P1, and when the data stored in the secured section is JNL, for this section, both the redundancy flag 2 and the non-volatile flag 3 are set to "on". When both the redundancy flag 2 and the non-volatile flag 3 are "on" for this section, the first main processor 211P1 performs redundancy such as writing the JNL in this section to PJVOL102PJ1A, transmitting the redundant data of the JNL in this section to the second main node 210P2 which is the mirror VOL having PVOL102P1A, i.e., PVOL102JP1S, and non-volatilization of storing the JNL in this section in the non-volatile area 103P1. Thereby, the reliability that the JNL required for restoration in the secondary storage system 100S exists in the main storage system 100P can be improved. For each of the plurality of main nodes 210P1, VOL102 (PJVOL102JP or PVOL102P) may be an area based on one or more non-volatile media in or outside this main node 210P1. In addition, the second main processor (hereinafter, for convenience, denoted as "211P2") which is the processor 211 in the second main node 210P2 may secure a section from the second main cache 55P2 which is the cache 55 in this main node 210P2, store the redundant data (the redundant data of JNL) from the first main node 210P1 in this section, and write the data in this section to PJVOL102JP1S. For this section, since the data in this section is the redundant data from another main node 210P, the second main processor 211P2 may also set the redundancy flag 2 to "off", but set the non-volatile flag 3 to "on". Therefore, the second main processor 211P2 may store the data in the section in the non-volatile area 103P2.

[0195] It is possible that the first main processor 211P1 ensures the data included in the JNL from the first main cache 55P1, that is, the section storing the data from the host 51. When the data stored in the ensured section is the data from the host 51 (the data written to PVOL102JP1A), for this section, both the redundancy flag 2 and the non-volatile flag 3 can be set to "on". It is possible that when both the redundancy flag 2 and the non-volatile flag 3 are "on" for this section, the first main processor 211P1 writes the data in this section to PVOL102P1A, transfers the redundant data of the data in this section to the second main node 210P2 for redundancy, and stores the data in this section in the non-volatile area 103P1 for non-volatilization. Thus, the possibility of regenerating in the main storage system 100P the JNL that needs to be restored in the secondary storage system 100S can be increased. In addition, the second main processor 211P2 can ensure a section from the second main cache 55P2, store the redundant data from the first main node 210P1 in this section, and write the data in this section to PVOL102P1S. For this section, since the data in this section is the redundant data from another main node 210P, the second main processor 211P2 can also set the redundancy flag 2 to "off", but set the non-volatile flag 3 to "on". Therefore, the second main processor 211P2 can store the data in the section in the non-volatile area 103P2.

[0196] It is possible that when the non-volatile flag 3 is "off" for a section in the processor 211 (for example, one of the secondary processor 211S and the main processor 211P), the data in this section is written to the instance memory 65 in the node 210 having this processor 211. The instance memory 65 can be an area based on a volatile storage medium.

[0197] It is possible that the first secondary processor 211S1 updates the SEQ# of the JNL after reflection is completed after writing data in SVOL102S1A (and SVOL102S1S), and notifies the updated SEQ# to the main storage system 100P. It is possible that the main storage system 100P clears from the main storage system 100P the JNL having the SEQ# up to the SEQ# notified by the first sub-processor 211S1. In this way, this JNL can be pre-saved to the main storage system 100P until the JNL is reflected in SVOL102S1A (and SVOL102S1S).

[0198] It is possible that when the first sub-processor 211S1 detects a failure of a pair, it notifies the main storage system 100P of the failure of the pair. It is possible that the first secondary node 210S1 receives a JNL containing data that is the difference between SVOL102S1A and PVOL102P1A from the main storage system 100P that has received the notification, and the first sub-processor 211S1 secures a section from the first secondary cache 55S1 and stores the JNL in the section. As a result, the above processing is performed, and thus the recovery from the failure is automatically performed. For example, the formation of the pair of SVOL102S1A and PVOL102P1A can be performed in response to a pair formation request from the main storage system 100P. In response to the notification of the above failure, the first secondary node 210S1 receives a pair formation request from the main storage system 100P, and in the process of performing the processing in response to the pair formation request, the first secondary node 210S1 can receive a JNL containing data that is the difference from the main storage system 100P. That is, when a failure of a pair is detected, by running the pair formation process again, the recovery from the failure can be automatically performed.

Claims

1. A storage system, A plurality of nodes each having a cache are provided, wherein the plurality of nodes include a memory and a processor connected to the memory. In each of the plurality of nodes, when there is data of a storage object of a segment secured from a cache, the processor executes: Depending on the type of data, for the segment, a decision is made as to whether the redundant data of the data in the segment needs to be transmitted to another node for redundancy, and whether the data in the segment needs to be stored in a nonvolatile area for non-volatility. Based on the determination, whether or not the data in the segment is to be made redundant and whether or not the data in the segment is to be made non-volatile are controlled. For each of the plurality of nodes, the non-volatile area is an area based on one or more non-volatile media in or outside the node.

2. The storage system according to claim 1, The plurality of nodes are a plurality of secondary nodes constituting a secondary storage system, A first secondary node among the plurality of secondary nodes has: a secondary journal volume as a volume stored as a journal containing data written to the primary volume in the primary storage system and metadata, wherein the metadata contains a sequence number as the order in which the data written to the primary volume in the primary storage system was written; A secondary volume that is paired with a primary volume in a primary storage system; A first secondary cache as a cache in the secondary node; as well as As the first secondary processor of the processor in the secondary node, The first secondary processor executes: Ensure that the log is stored in the first secondary cache, If the data stored in the secured segment is a log, it is determined that redundancy is not required for the segment, and non-volatility is not required. When redundancy and non-volatility are not required for the segment, the log in the segment is written to the secondary log volume, and the redundant data of the log in the segment is not transferred to the second secondary node having the secondary mirror volume as the mirror volume of the secondary volume for redundancy, nor is the log in the segment stored in the non-volatile area for non-volatility. For each of the plurality of slave nodes, a volume is based on an area of ​​one or more non-volatile media in or outside the slave node.

3. The storage system according to claim 2, The first secondary processor executes: secure the section where the data in the log is stored from the first secondary cache, If the data stored in the secured segment is data in a log, redundancy is determined to be necessary for the segment and non-volatility is determined not to be necessary. When the segment needs to be made redundant but not non-volatile, the data in the segment is written to the secondary volume, and the redundant data in the segment is transferred to the second secondary node for redundancy, without storing the data in the segment in the non-volatile area for non-volatility.

4. The storage system according to claim 2, The main storage system is composed of a plurality of master nodes as a plurality of nodes, A first master node among the plurality of master nodes has a master log volume as a volume in which logs are stored, the master volume, a first master cache as a cache in the master node, and a first master processor as a processor in the master node, The first main processor executes: secure the segment where the log is stored from the first primary cache, If the data stored in the secured segment is a log, redundancy is determined to be necessary for the segment, and non-volatility is determined to be necessary. When the segment needs to be made redundant and non-volatile, the log in the segment is written to the primary log volume, the redundant data of the log in the segment is transferred to the second primary node having the primary mirror volume as the mirror volume of the primary volume for redundancy, and the log in the segment is stored in the non-volatile area for non-volatile storage, For each of the plurality of primary nodes, a volume is based on an area of ​​one or more non-volatile media in or outside the secondary node.

5. The storage system according to claim 4, The first main processor executes: From the first primary cache, secure the segment where the data contained in the log, i.e. the data from the host, is stored, When the data stored in the secured segment is data from the host, it is determined that the segment needs to be made redundant and non-volatile. When the segment needs to be made redundant and non-volatile, the data in the segment is written to the primary volume, the redundant data in the segment is transferred to the second primary node for redundancy, and the data in the segment is stored in the non-volatile area for non-volatility.

6. The storage system according to claim 1, The processor writes the data in the segment to the instance memory in the node having the processor when non-volatility is not required for the segment.

7. The storage system according to claim 1, The redundant data is duplicate data or parity data.

8. The storage system according to claim 2, After writing data to the secondary volume, the first secondary processor updates the serial number of the log reflecting the completion, and notifies the primary storage system of the updated serial number. The primary storage system clears logs having sequence numbers up to the sequence number notified from the first secondary processor from the primary storage system.

9. The storage system according to claim 2, When the first sub-processor detects that the pair has failed, the first sub-processor notifies the primary storage system of the failure of the pair. The first secondary node receives a log including data that is a difference between the secondary volume and the primary volume from the primary storage system that has received the notification, and the first secondary processor secures a segment from the first secondary cache and stores the log in the segment.

10. The storage system according to claim 9, The pairing of the secondary volume with the primary volume is performed in response to a pairing formation request from the primary storage system. In response to the notification, the first slave node receives a pairing formation request from the primary storage system, and in a process performed in response to the pairing formation request, the first slave node receives a log containing data as the difference from the primary storage system.

11. A storage control method, In each of the plurality of nodes constituting the storage system, when there is data to be stored in a segment secured from the cache, the node executes: Based on the type of data, a redundancy-non-volatility decision is made for the segment, wherein the redundancy-non-volatility decision is a decision on whether the redundant data of the data in the segment needs to be transmitted to another node for redundancy, and whether the data in the segment needs to be stored in a non-volatile area for non-volatility. Based on the determination, whether or not the data in the segment is made redundant and whether or not the data in the segment is made nonvolatile are controlled.

Citation Information

Patent Citations

  • Remote copy system

    JP2005018736A