Computer system, remote replication control method, and remote replication control program product
By dispersing the generation of volumes and log volumes of the replication destination in a decentralized storage system of multiple storage nodes, the complex problems of asynchronous remote replication settings and execution are solved, and the guarantee of data writing order is achieved.
Patent Information
- Application Number
- CN202411341410.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-09-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In a distributed storage system with multiple storage nodes, the setting and execution of asynchronous remote replication are complicated, and it is difficult to ensure the data writing order of multiple volumes.
By dispersing the volumes of the replication destination among multiple storage nodes, and generating a log volume for storing data of the updated contents at each storage node, corresponding log volumes are also generated in the positive storage system to ensure the writing order of data.
It realizes simple setting and execution of asynchronous remote replication from the storage system of the replication source to the storage system of the replication destination of multiple storage nodes, ensuring the order of writing of data.
Smart Images

Figure CN120104046A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a data replication technology between storage systems. Background Art
[0002] For example, Patent Document 1 discloses a method in which, in asynchronous remote copying between multiple storage devices, the updateable location in each device is determined based on write order information, thereby ensuring the order of data update across the devices.
[0003] In addition, Patent Document 2 discloses the following method: In a distributed storage system composed of multiple storage nodes, in order to ensure data redundancy and maintain responsiveness, each node shares the responsibility of executing I / O processing, and the physical area of a certain node is preferentially allocated as the storage area processed by the node.
[0004] In addition, in a computer system, in remote replication from a primary storage system (replication source storage system) to a secondary storage system (replication destination storage system), a CTG (consistency group) is sometimes formed as a range to ensure the writing order of data to multiple volumes.
[0005] Patent Document 1: Japanese Patent Application Publication No. 2007-264946
[0006] Patent Document 2: Japanese Patent Application Publication No. 2019-101702 Summary of the invention
[0007] For example, when asynchronously remotely replicating a volume of a primary storage system in a computer system to a secondary storage system, the secondary storage system may be a distributed storage system composed of a plurality of storage nodes. In this case, volumes serving as replication destinations for a plurality of volumes belonging to the same CTG in the primary storage system may be generated in a distributed manner in the plurality of storage nodes.
[0008] In this way, when multiple volumes belonging to the same CTG in the secondary storage system are distributed to multiple storage nodes, a log volume storing data indicating the updated content of the volume needs to be generated for each storage node, and corresponding log volumes need to be generated in the primary storage system, which takes time for the user to set. In addition, when multiple volumes belonging to the same CTG in the secondary storage system are distributed to multiple storage nodes, a process must be performed to ensure the update order of the data of the multiple volumes belonging to the CTG.
[0009] On the other hand, sometimes a volume that becomes a replication destination for each of the multiple volumes belonging to the same CTG in the primary storage system is generated on one storage node. In this case, the user must make settings different from the case where the multiple volumes belonging to the same CTG in the secondary storage system are distributed to multiple storage nodes, which is a complicated process for the user.
[0010] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technology that can easily and appropriately set and execute asynchronous remote replication from a replication source storage system to a replication destination secondary storage system including a plurality of storage nodes.
[0011] In order to achieve the above-mentioned purpose, a computer system of one viewpoint is a computer system including a first storage system and a second storage system, wherein the first storage system manages multiple first volumes belonging to a consistency group that ensures the writing order of data, the second storage system has multiple storage nodes, and the second storage system generates multiple second volumes that become the replication destinations of the multiple first volumes in a dispersed manner on the multiple storage nodes, and a second log volume storing log data of the write content in the first volume that is the replication source of the second volume is respectively present in the multiple storage nodes that generate the multiple second volumes, the first storage system generates multiple first log volumes corresponding to each of the second log volumes and storing log data about the multiple first volumes, and performs write order guarantee processing that controls the processing of the multiple first volumes so that the log data representing the write content to the multiple first volumes can be stored in the multiple first log volumes in a state that can ensure the writing order of the multiple first log volumes.
[0012] According to the present invention, it is possible to easily and appropriately set and execute asynchronous remote copy from a copy source storage system to a copy destination storage system including a plurality of storage nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a diagram showing the overall structure of the computer system according to the first embodiment.
[0014] Figure 2 This is a diagram showing the structure of a storage node according to the first embodiment.
[0015] Figure 3 This is a configuration diagram of a host computer and a management terminal according to the first embodiment.
[0016] Figure 4 It is a diagram for explaining the structure of a memory of a storage node of the storage system on the positive side according to the first embodiment.
[0017] Figure 5It is a diagram for explaining the structure of a memory of a storage node of the secondary-side storage system according to the first embodiment.
[0018] Figure 6 This is a diagram showing the structure of the replication pair management table according to the first embodiment.
[0019] Figure 7 This is a diagram showing the structure of a volume management table of the storage system on the positive side of the first embodiment.
[0020] Figure 8 It is a diagram showing the structure of a write order management table of the storage system on the positive side of the first embodiment.
[0021] Fig. 9 This is a diagram showing the structure of a volume management table of the secondary storage system according to the first embodiment.
[0022] Fig.10 This is a diagram showing the structure of a write order management table of the secondary storage system according to the first embodiment.
[0023] Fig.11 This is a flowchart of the pair generation process according to the first embodiment.
[0024] Fig.12 This is a flowchart of the paired volume generation process according to the first embodiment.
[0025] Fig.13 This is a flowchart of the pairing creation preparation process according to the first embodiment.
[0026] Fig.14 This is a flowchart of the pair addition process in the same node intra-pattern according to the first embodiment.
[0027] Fig.15 This is a flowchart of the pairing addition process in the cross-node mode according to the first embodiment.
[0028] Fig.16 This is a flowchart of the host I / O process according to the first embodiment.
[0029] Fig.17 This is a flowchart of the write order management process according to the first embodiment.
[0030] Fig.18 This is a flowchart of the log data transmission process according to the first embodiment.
[0031] Fig.19 This is a flowchart of the data reflection arbitration process according to the first embodiment.
[0032] Fig. 20 This is a flowchart of the pairing generation process according to the second embodiment.
[0033] Fig.21This is a flowchart of the pair addition process in the same node mode according to the second embodiment.
[0034] Fig. 22 This is a flowchart of the pairing addition process in the cross-node mode according to the second embodiment.
[0035] Fig.23 This is a flowchart of the write order management process according to the third embodiment.
[0036] Fig.24 It is an overall structural diagram of a computer system according to the fifth embodiment.
[0037] Fig.25 It is a diagram showing the configuration of memories of a plurality of device management terminals according to the fifth embodiment.
[0038] Fig.26 It is a diagram showing the structure of the operation information table according to the fifth embodiment.
[0039] Fig. 27 It is a diagram showing the structure of the service level information table according to the fifth embodiment.
[0040] Fig.28 This is a flowchart of the optimal pair generation process according to the fifth embodiment.
[0041] Fig.29 This is a flowchart of paired volume creation processing according to the fifth embodiment. DETAILED DESCRIPTION
[0042] The embodiments will be described with reference to the drawings. The embodiments described below do not limit the inventions to the scope of the claims, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solutions of the invention.
[0043] In the following description, information is sometimes described by expressing the "AAA table", but the information can also be expressed by any data structure. In other words, in order to indicate that the information does not depend on the data structure, the "AAA table" can be referred to as "AAA information".
[0044] In the following description, the structure of each table is an example, and one table may be divided into two or more tables, and all or part of two or more tables may be one table.
[0045] In addition, in the following description, "program" is sometimes used as the main body of the action to illustrate the processing, but the program is executed by a processor (for example, CPU (Central Processing Unit)), thereby appropriately using memory or communication I / F to perform specified processing, so the main body of the processing action can also be the processor (or a controller, computer, or other device having the processor).
[0046] In addition, the program may be installed from a program source to a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-temporary) recording medium. In addition, in the following description, two or more programs may be implemented as one program, or one program may be implemented as two or more programs.
[0047] In addition, in the following description, when describing the same type of elements without distinguishing them, reference symbols are used, and when describing the same type of elements, the ID (for example, identification number) of the element is sometimes used. For example, when describing the storage node without distinguishing it in particular, it is recorded as "storage node 102", and when describing the various nodes in a distinguished manner, it is sometimes recorded as "storage node 1" or "storage node 2". In addition, in the following description, sometimes n is added to the name of the element in the storage node n (n is a natural number) to distinguish which node the element belongs to.
[0048] Hereinafter, one embodiment of the present invention will be described based on the drawings. In addition, the present invention is not limited to the embodiment described below.
[0049] [First embodiment]
[0050] Figure 1 It is a diagram showing the overall structure of the computer system according to the first embodiment.
[0051] The computer system 10 includes a storage system 100 , a primary host computer 104 , a primary management terminal 105 , a storage system 101 , a secondary host computer 107 , and a secondary management terminal 108 .
[0052] The storage system 100 and the storage system 101 are connected via an inter-storage system network 110. The inter-storage system network 110 is, for example, a wired LAN (Local Area Network), a wireless LAN, or a WAN (Wide Area Network), and may be, for example, a network using Ethernet (registered trademark) or Fibre Channel.
[0053] The storage system 100, the main host computer 104, and the main management terminal 105 are connected via a main network 106. In this embodiment, the range connected to the main network 106 is called a main system, and each configuration is treated as a main configuration. The main network 106 is, for example, a LAN.
[0054] The storage system 100 is an example of a first storage system, and has one or more storage nodes 111. The storage node 111 is an example of a computer, and provides a plurality of volumes. Each storage node 111 is provided for redundancy, and in the storage system 100, it operates as one storage node as a whole.
[0055] The host computer 104 executes applications and the like to perform various processes associated with I / O requests on the storage system 100. The management terminal 105 instructs the storage system 100 to create a volume and the like to manage the storage system 100.
[0056] The storage system 101, the secondary host computer 107, and the secondary management terminal 108 are connected via a secondary network 109. In this embodiment, the area connected to the secondary network 109 is referred to as a secondary system, and each configuration is treated as a secondary configuration. The secondary network 109 is, for example, a LAN.
[0057] The storage system 101 is an example of a second storage system, and has a plurality of storage nodes 102. In the present embodiment, the storage system 101 is a distributed storage system such as a scale-out storage system, which is composed of a plurality of storage nodes 102 communicatively connected via an inter-node network 103, and collaborates as a cluster to provide a plurality of volumes.
[0058] The secondary host computer 107 executes applications and the like, and performs various processes associated with I / O requests on the storage system 101. The secondary management terminal 108 instructs the storage system 101 to create a volume and the like, and performs processes for managing the storage system 101.
[0059] In the computer system 10, one or both of the primary-side system and the secondary-side system may be operated on the cloud.
[0060] In this embodiment, in order to cope with disasters and perform backup, it is assumed that the data written by the primary host computer 104 to the primary storage system 100 is remotely copied to the secondary storage system 101. When the primary system fails, etc., the secondary host computer 107 in the secondary system performs recovery processing based on the data saved in the storage system 101 and starts processing again.
[0061] An example of the configuration of a pair of volumes (copy pair) of a copy source and a copy destination involved in remote copying in this embodiment will be described. The copy pair is generated by a pair generation process (see Fig.11 ) and so on.
[0062] exist Figure 1In the example, the system ID of storage system 100 is set to 1, and the system ID of storage system 101 is set to 2. Storage system 101 is composed of three storage nodes 102 (storage node 1) with ID 1, storage node 102 (storage node 2) with ID 2, and storage node 102 (storage node 3) with ID 3.
[0063] In the computer system 10, a CTG (consistency group) is managed, which represents a range in which the writing order of data is guaranteed between multiple volumes in the data reflection to the secondary side system through remote replication. Each replication pair is managed as belonging to any CTG. Specifically, in the computer system 10, there are CTG112 (CTG1) and CTG113 (CTG2). The same CTG (112, 113) is assigned the same ID in the storage systems 100 and 101.
[0064] Regarding CTG1, in the primary storage system 100, there are PVOL (primary volume) 114 (PVOL1) and PVOL 115 (PVOL2) as volumes to be subjected to I / O processing by the primary host computer 104. In addition, in the storage system 100, there is JVOL 116 (JVOL5) as a JVOL (journal volume), which is a volume for storing differential data (journal data) representing written contents used for remote replication to the secondary storage system 101 asynchronously with the I / O processing on the PVOLs when a write I / O request is issued to these PVOLs 114 and 115.
[0065] In the secondary storage system 101, JVOL 117 (JVOL5) is paired with JVOL 116 (JVOL5) of the storage system 100 and is a volume that receives and temporarily stores differential data of remote replication in storage node 1. In addition, in storage node 1, SVOL (secondary volume) 118 (SVOL1) and SVOL 119 (SVOL2) are present as target volumes that are replication destinations of PVOL 114 and PVOL 115 and reflect differential data stored in JVOL 117.
[0066] In CTG1, in the secondary storage system 101, the SVOL exists in one storage node and is not distributed across multiple storage nodes (does not span storage nodes). Therefore, the computer system 10 operates in a mode that ensures the data write order by reflecting data to multiple SVOLs using only one JVOL.
[0067] Regarding CTG2, in the primary storage system 100, there are PVOL120 (PVOL3) and PVOL121 (PVOL4) as volumes to be subjected to I / O processing by the primary host computer 104. In addition, in the storage system 100, there is JVOL122 (JVOL6) as a JVOL, which is a volume for storing differential data for remote replication to the secondary storage system 101 when a write I / O request is issued to PVOL120, asynchronously with the I / O processing to PVOL120. In addition, in the storage system 100, there is JVOL123 (JVOL7) as a JVOL, which is a volume for storing differential data for remote replication to the secondary storage system 101 when a write I / O request is issued to PVOL121, asynchronously with the I / O processing to PVOL121.
[0068] In the secondary storage system 101, in the storage node 2, there is a volume JVOL124 (JVOL6) which is paired with the JVOL122 (JVOL6) of the storage system 100 and receives and temporarily stores the differential data of the remote replication. In addition, in the storage node 2, there is a SVOL126 (SVOL3) which is the replication destination of PVOL120 (PVOL3) and reflects the differential data stored in JVOL124. In addition, in the storage node 3, there is a volume JVOL125 (JVOL7) which is paired with the JVOL123 (JVOL7) of the storage system 100 and receives and temporarily stores the differential data of the remote replication. In addition, in the storage node 3, there is a SVOL127 (SVOL4) which is the replication destination of PVOL121 (PVOL4) and reflects the differential data stored in JVOL125.
[0069] The JVOL on the secondary side does not perform data reflection processing on the SVOL across the storage node 102. Therefore, in the CTG2 across the storage node 2 and the storage node 3, multiple JVOLs are prepared in the storage system 100 on the primary side and the storage system 101 on the secondary side, and the computer system 10 performs processing to ensure the data writing order across multiple JVOLs.
[0070] In the computer system 10 of the present embodiment, in an environment where an environment for guaranteeing a write order limited to a storage node such as CTG1 and an environment for guaranteeing a write order across storage nodes such as CTG2 coexist, the user's operation method can be facilitated (for example, in the same way as in a case where such environments are not mixed).
[0071] Next, the structure of the storage nodes 111 and 102 will be described.
[0072] Figure 2 1 is a diagram showing the structure of a storage node according to the first embodiment. The storage nodes 111 and 102 have the same structure, so for ease of use Figure 2 Provide explanation.
[0073] The storage node 111 (102) is an example of a computer, and includes a CPU 201 as an example of a processor, a memory 202 as an example of a storage unit, a storage device 203, and a communication (interface) I / F 204. These components 201 to 204 are connected to each other via an internal bus or the like so as to be able to communicate with each other. There may be one or more CPU 201, memory 202, storage device 203, and communication I / F 204.
[0074] The CPU 201 is responsible for overall operation control of the storage node 111 (102). The CPU 201 performs various processes based on the programs and management information stored in the memory 202. The CPU 201 may be a physical CPU of a physical computer or a virtual CPU that is virtually allocated to the physical CPU of a physical computer using the virtualization function of the cloud.
[0075] The memory 202 is a volatile semiconductor memory such as SRAM (Static RAM (Random Access Memory)) or DRAM (Dynamic RAM), and stores various programs executed by the CPU 201 and management information referenced or updated by the CPU 201. The memory 202 may be a physical memory or a virtual memory in which a physical memory is virtually allocated using a virtualization function of the cloud.
[0076] The storage device 203 is a storage device that stores user data used by the primary host computer 104, the secondary host computer 107, etc. Typically, the storage device 203 may be a non-volatile storage device. The storage device 203 may be, for example, a HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage device 203 may be a physical storage device or a virtual storage device that is a physical storage device virtually allocated using a virtualization function of the cloud.
[0077] The communication I / F 204 is an interface for communication via a network (communication between storage nodes via the inter-node network 103, communication with the primary host computer 104 and the primary management terminal 105 via the primary network 106, communication with the secondary host computer 107 and the secondary management terminal 108 via the secondary network 109, and communication between storage systems via the inter-storage system network 110), such as a NIC (Network Interface Card) or a FC (Fibre Channel) card. The communication I / F 204 may be a physical communication I / F, or may be a communication I / F that is virtually allocated using a cloud virtualization function.
[0078] Next, the configurations of the main host computer 104, the main management terminal 105, the sub-host computer 107, and the sub-management terminal 108 will be described.
[0079] Figure 3 104, the main management terminal 105, the secondary host computer 107, and the secondary management terminal 108 have the same structure, so for ease of use Figure 3 In addition, Figure 2 The same structures as those described above are denoted by the same reference numerals.
[0080] The main host computer 104 (main management terminal 105, sub-host computer 107, sub-management terminal 108) includes a CPU 201 as an example of a processor, a memory 202 as an example of a storage unit, and a communication I / F 204. These components 201, 202, 204 are connected to each other via an internal bus or the like so as to be able to communicate with each other. There may be one or more CPU 201, memory 202, and communication I / F 204, respectively.
[0081] The CPU 201 performs processing to control the primary host computer 104 (primary management terminal 105, secondary host computer 107, secondary management terminal 108) according to the program and management information stored in the memory 202. The memory 202 stores the program executed by the CPU 201 and the management information referenced or updated by the CPU 201. The communication I / F 204 is an interface for communicating with the storage system via a network (for communicating with the storage system 100 via the primary network 106 or communicating with the storage system 101 via the secondary network 109).
[0082] Next, a diagram is provided to explain the structure of the memory 202 of the storage node 111 of the storage system 100 on the main side.
[0083] Figure 4It is a diagram for explaining the structure of a memory of a storage node of the storage system on the positive side according to the first embodiment.
[0084] The memory 202 of the storage node 111 stores a replication pair management program 401 , a host I / O processing program 402 , a write order management program 403 , a log data transfer program 404 , a volume management program 405 , a replication pair management table 406 , a volume management table 407 , and a write order management table 408 .
[0085] The copy pair management program 401 is executed by the CPU 201 to perform processes such as creation, status change, and deletion of a copy pair (a pair of a primary volume and a secondary volume related to remote copy) in accordance with instructions from the management terminal 105 .
[0086] The host I / O processing program 402 is executed by the CPU 201 to perform I / O processing (read processing, write processing) according to an I / O request (read request, write request) issued from the host computer 104 .
[0087] The write order management program 403 is executed by the CPU 201 to store differential data (journal data) to which time information is added for write data issued by the primary volume (PVOL) constituting the replication pair in the journal volume (JVOL).
[0088] Here, when there are multiple log volumes belonging to the same CTG as the copy destination, the order of adding records to the table and transmitting differential data changes due to the difference in the time taken for processing and the time taken for communication, so when the data is reflected in multiple SVOLs via multiple log volumes, the writing order of the data may not be guaranteed. Therefore, the writing order management program 403 is executed by the CPU 201, and in order to make the time section of the data consistent, in the writing order guarantee of the data via the multiple log volumes, the writing process corresponding to the writing request from the main host computer 104 is suppressed, and the time information for the differential data is updated during this period.
[0089] The log data transfer program 404 is executed by the CPU 201 to transfer the differential data stored in the log volume to the secondary storage system 101 via the inter-storage system network 110 .
[0090] The volume management program 405 is executed by the CPU 201 to perform volume management processing such as creation and deletion of volumes according to instructions from the management terminal 105 .
[0091] The replication pair management table 406 stores information related to the replication pair. The volume management table 407 stores information related to the volume. The write order management table 408 stores information related to the write data corresponding to the write request issued to the volume constituting the primary side of the replication pair. The details of the replication pair management table 406, the volume management table 407, and the write order management table 408 will be described later.
[0092] Next, a diagram is provided to explain the configuration of the memory 202 of the storage node 102 of the storage system 101 on the secondary side.
[0093] Figure 5 It is a diagram for explaining the structure of a memory of a storage node of the secondary storage system according to the first embodiment.
[0094] The memory 202 of the storage node 102 stores a log data receiving program 501, a host I / O processing program 502, a data reflection arbitration program 503, a replication pair management program 504, a volume management program 505, a replication pair management table 506, a volume management table 507, and a write order management table 508. Here, the programs 401 to 405 of the memory 202 of the storage node 111 and the programs 501 to 505 of the memory 202 of the storage node 102 are examples of remote copy control programs.
[0095] The log data receiving program 501 is executed by the CPU 201 to perform processing for receiving data transferred by the log data transfer program 404 of the primary storage system 100 and storing the data in the log volume.
[0096] The host I / O processing program 502 is executed by the CPU 201 to perform I / O processing (read processing, write processing) according to an I / O request (read request, write request) from the sub-host computer 107 .
[0097] The data reflection arbitration program 503 is executed by the CPU 201 to reflect the data stored in the log volume to the secondary volume (SVOL) while ensuring the writing order in the CTG based on the time information.
[0098] The copy pair management program 504 is executed by the CPU 201 to perform processes such as creation, status change, and deletion of copy pairs in accordance with instructions from the secondary management terminal 108 .
[0099] The volume management program 505 is executed by the CPU 201 to perform volume management processing such as creation and deletion of volumes according to instructions from the secondary management terminal 108 or the storage system 100 .
[0100] The replication pair management table 506 stores information related to the replication pair. The volume management table 507 stores information related to the volume. The write order management table 508 stores information related to the write data corresponding to the write request issued to the volume constituting the primary side of the replication pair. The details of the replication pair management table 506, the volume management table 507, and the write order management table 508 will be described later.
[0101] In the storage system 101 , these programs and data may be stored in all the storage nodes 102 , or a specific storage node 102 may be used as a representative node (representative node) and the representative node may hold a table related to all the storage nodes 102 .
[0102] Next, the replication pair management tables 406 and 506 will be described in detail.
[0103] Figure 6 406 and 506 have the same structure. Figure 6 A specific example of the replication pair management table 406 is shown.
[0104] The replication pair management table 406 (506) stores a record for each replication pair. The record of the replication pair management table 406 (506) includes fields for CTG ID 601, primary system ID 602, primary log ID 603, primary volume ID 604, status 605, secondary system ID 606, secondary log ID 607, secondary volume ID 608, operation mode 609, and replication path ID 610.
[0105] CTG ID 601 stores the ID (identification information) of the CTG that is the range for guaranteeing the writing order. Since the CTGs that include the volumes that constitute the replication pair for remote replication are the same, the same ID is assigned to the CTG in the primary storage system 100 and the secondary storage system 101. Figure 6 In the example of FIG. 1 , CTG ID 601 stores 1 which is the ID of CTG 112 (CTG1) or 2 which is the ID of CTG 113 (CTG2).
[0106] The primary system ID 602 stores the ID of the storage system that becomes the primary in remote replication. Figure 6 In the example, the ID of the storage system 100, that is, 1, is stored in the main-side system ID 602.
[0107] The ID (log ID) of the log stored before the data to be written is asynchronously transmitted when writing to the primary volume (PVOL) is stored in the primary log ID 603. In one CTG, one or more primary log volumes and one or more secondary log volumes are required. When multiple secondary volumes (SVOL) of the secondary storage system span multiple storage nodes (distributed to multiple storage nodes), the primary storage system and the secondary storage system require multiple log volumes according to the number of spanned storage nodes. Figure 6 In the example, the record corresponding to CTG1 stores the volume ID 5 of JVOL116 (JVOL5), the record corresponding to CTG2 stores the volume ID 6 of JVOL122 (JVOL6), and the record corresponding to JVOL123 (JVOL7) stores the volume ID 7. Figure 6 In the example, the volume ID of the log volume is used as the log ID, but the log ID is not limited to this, and an ID different from the volume ID may be used.
[0108] The primary volume ID 604 stores the ID (identification information) of the volume that becomes the primary side in remote replication. Figure 6 In the first record, the ID of volume 114, ie, PVOL1, is stored in the front volume ID 604.
[0109] Status 605 stores the status of the replication pair. As the status of the replication pair, there are, for example, "pairing in progress", "pairing", "pairing split", "failure", etc., where "pairing in progress" is a status in which the processing of forming a pair is being executed to transfer the existing storage data of the primary volume to the secondary volume of the replication pair that becomes the storage system of the secondary side immediately after the replication pair is generated, "pairing" is a status in which the formation of the pair is completed and the processing of transferring the updated write data is started, "pairing allocation" is a status in which the processing of transferring the data to the storage system of the secondary side is interrupted and only the data that can guarantee the write order in the CTG is reflected, and "failure" is a status in which the write order of the data is lost due to a failure. In addition, as the status of the replication pair, there may be other states.
[0110] The secondary system ID 606 stores the ID of the storage system that serves as the secondary in remote replication. Figure 6 In the example of FIG. 6 , the ID 2 of the storage system 101 is stored in the secondary system ID 606 .
[0111] The secondary log ID 607 stores the ID (log ID) of the log storing the differential data received from the primary storage system. Figure 6In the example, the record corresponding to CTG1 stores the volume ID 5 of JVOL117 (JVOL5), the record corresponding to CTG2 stores the volume ID 6 of JVOL124 (JVOL6), and the record corresponding to JVOL125 (JVOL7). Figure 6 In the example, the volume ID of the log volume is used as the log ID, but the log ID is not limited to this, and an ID different from the volume ID may be used.
[0112] The information of the action mode for ensuring the update order of asynchronous remote replication is stored in the action mode 609. As the information of the action mode, there is "in-node" which indicates the action mode when the log volume of the secondary side of CTG is located in the same storage node, that is, the same node mode, and "cross-node" which indicates the action mode when the log volume of the secondary side of CTG spans multiple storage nodes, that is, the cross-node mode. In this embodiment, the write order management program 403 switches the method of ensuring the update order of data based on the information of the action mode 609.
[0113] The replication path ID 610 stores the ID of a previously constructed path (path: replication path) in the inter-storage system network 110 used for data transmission between the replication pair of the storage system 100 and the storage system 101. Multiple replication paths may be constructed between the primary and secondary storage systems.
[0114] Next, the volume management table 407 of the storage system 100 on the primary side will be described in detail.
[0115] Figure 7 This is a diagram showing the structure of a volume management table of the storage system on the positive side of the first embodiment.
[0116] The volume management table 407 manages records for each volume. The records of the volume management table 407 include fields for volume ID 701 , volume attribute 702 , maximum capacity 703 , data storage destination storage device 704 , and write I / O restriction flag 705 .
[0117] The volume ID 701 stores the ID of the volume in the storage system 100 .
[0118] The attributes of the volume corresponding to the record are stored in the volume attribute 702. As the attributes of the volume, there are "normal" indicating that it is a normal volume connected to the host computer and can perform I / O, and "journal" indicating that it is a journal volume (JVOL) that temporarily stores data written to the normal volume for asynchronous remote replication.
[0119] The maximum capacity 703 stores the maximum capacity of the volume corresponding to the record.
[0120] The name of the storage device that is the data storage destination of the volume assigned to the record is stored in the data storage destination storage device 704. As a method of managing the name of the storage device, a thin provisioning function is typically applied, but other methods may be used.
[0121] The state (ON or OFF) of the write I / O restriction flag indicating whether writing to the volume corresponding to the record can be accepted is stored in the write I / O restriction flag 705. The state of the write I / O restriction flag is an example of restriction information. When the write I / O restriction flag is on, writing to the volume corresponding to the record is suppressed. In this embodiment, the writing order of data is guaranteed by the cooperation of suppressing writing to the volume according to the write I / O restriction flag and the updating control of the time information.
[0122] For example, the first row of the volume management table 407 corresponds to the volume with volume ID 1, namely PVOL114, indicating that the volume type is normal, the maximum capacity is 100GB, the data storage destination storage device is storage device 1, and the write I / O restriction flag is off.
[0123] Next, the write order management table 408 of the primary storage system 100 will be described in detail.
[0124] Figure 8 It is a diagram showing the structure of a write order management table of the storage system on the positive side of the first embodiment.
[0125] The write order management table 408 is a table for managing the write process to the primary volume in the paired state in the storage system 100, the temporary storage destination of data, and the state until the reflection to the secondary volume of the storage system 101, and stores a record of each write process for the PVOL. The record of the write order management table 408 includes fields of ID 801, CTG ID 802, time information (generation number) 803, write target volume ID 804, write address 805, log volume ID 806, log volume address 807, and reflection status 808.
[0126] The ID 801 stores the writing order of data written by the replication pair in the storage system 100 to the volume in the pair state.
[0127] The CTG ID 802 stores the ID of the CTG to which the write process corresponding to the record is to be applied.
[0128] In the time information (generation number) 803, time information (for example, time period) related to the time when the write process corresponding to the record is performed is stored. In this embodiment, the time information is managed in a manner of incrementing at regular intervals, and based on the time information, it is possible to grasp in which time period the write process corresponding to the record is performed. In addition, in this embodiment, the time information is a real number, but it is not limited to this, and it can also be the time information provided by the NTP server or the timer information in the storage system.
[0129] The write target volume ID 804 stores the ID of a volume in a paired state, that is, the volume that is the target of a write request by the host computer 104 .
[0130] The write address 805 stores an address (LBA: Logical Block Address) to be the target of the write process corresponding to the record.
[0131] The log volume ID 806 stores the volume ID of the log volume storing the data to be written corresponding to the record.
[0132] The address 807 on the log volume stores the address (LBA) of the log volume storing the data to be the target of the write processing corresponding to the record.
[0133] The reflection status 808 stores the reflection status of the data to be written corresponding to the entry to the secondary volume. The reflection status includes "not transferred" indicating that the target data has not been transferred to the secondary storage system, and "transfer completed" indicating that the target data has been transferred to the secondary storage system but has not been reflected to the secondary volume. The reflection status may include other statuses.
[0134] For example, the record in the first row of the write order management table 408 is the record of the earliest write processing with a write order of 1. The write processing is targeted at the CTG with an ID of 1, the time information for performing the write processing is 1, the ID of the volume of the write object is 1 (that is, the volume of the write object is PVOL114 (PVOL1)), the write object address is 0 to 255, the volume ID of the log volume storing the data of the write object is 5 (that is, JVOL116 (JVOL5)), and the address of the write log volume is 0 to 255, indicating that the data of the write object has been transferred to the storage system on the secondary side.
[0135] Next, the volume management table 507 of the secondary storage system 101 will be described in detail.
[0136] Fig. 9 This is a diagram showing the structure of a volume management table of the secondary storage system according to the first embodiment.
[0137] The volume management table 507 stores entries for each volume of the secondary storage system. The records of the volume management table 507 include fields for volume ID 901 , volume attribute 902 , maximum capacity 903 , volume storage node 904 , data storage destination storage device 905 , and write I / O restriction flag 906 .
[0138] The volume ID 901 stores the ID of the volume in the storage system 101 .
[0139] The attributes of the volume corresponding to the record are stored in the volume attribute 902. As the attributes of the volume, there are "normal" indicating that it is a normal volume connected to the host computer and can perform I / O, and "journal" indicating that it is a journal volume (JVOL) that temporarily stores data written to the normal volume for asynchronous remote replication.
[0140] The maximum capacity 903 stores the maximum capacity of the volume corresponding to the record.
[0141] The volume storage node 904 stores the ID of the storage node storing the volume corresponding to the record.
[0142] The name of the storage device that is the data storage destination of the volume assigned to the record is stored in the data storage destination storage device 905. As a method of managing the name of the storage device, the function of ThinProvisioning is typically applied, but other methods may be used.
[0143] The state of the write I / O restriction flag indicating whether writing to the volume corresponding to the record can be accepted (ON or OFF) is stored in the write I / O restriction flag 906. When the write I / O restriction flag is on, writing to the volume corresponding to the record is suppressed. In this embodiment, in the storage system 101 on the secondary side, the write process is suppressed by turning on the write I / O restriction flag in the record, thereby ensuring the data at a certain time point on the primary side.
[0144] For example, the record in the first row of the volume management table 507 is the volume with volume ID 1, that is, the record corresponding to SVOL118, indicating that the volume type is normal, the maximum capacity is 100GB, the storage node of the storage volume is storage node 1, the data storage destination storage device is storage device 1, and the write I / O restriction flag is on.
[0145] Next, the write order management table 508 of the secondary storage system 101 will be described in detail.
[0146] Fig.10 This is a diagram showing the structure of a write order management table of the secondary storage system according to the first embodiment.
[0147] The write order management table 508 is a table for managing the state from the journal volume in the storage system 101 to the volume on the secondary side of the paired state, and stores a record of each write process for the SVOL that sent the data to be written to the secondary side storage system 101. The record of the write order management table 508 includes fields of ID 1001, CTG ID 1002, time information (generation number) 1003, write target volume ID 1004, write address 1005, journal volume ID 1006, address on the journal volume 1007, and reflection status 1008.
[0148] The ID 1001 stores the writing order of data transferred to the secondary volume whose replication pair is in the paired state within the storage system 101 .
[0149] The CTG ID 1002 stores the ID of the CTG to which the write process corresponding to the record is to be applied.
[0150] In the time information (generation number) 1003, time information (for example, time period) related to the time when the write process corresponding to the record is performed is stored. In this embodiment, the time information is managed in a manner of incrementing at regular intervals, and based on the time information, it is possible to grasp in which time period the write process corresponding to the record is performed. In addition, in this embodiment, the time information is a real number, but it is not limited to this, and it can also be the time information provided by the NTP server or the timer information in the storage system.
[0151] The write target volume ID 1004 stores the ID of the volume to which the corresponding record is to be written.
[0152] The write address 1005 stores an address (LBA: Logical Block Address) to be the target of the write process corresponding to the record.
[0153] The log volume ID 1006 stores the volume ID of the log volume storing the data to be written corresponding to the record.
[0154] The address 1007 on the log volume stores the address (LBA) of the log volume storing the data to be the target of the write processing corresponding to the record.
[0155] The reflection status 1008 stores the reflection status of the data to be written corresponding to the entry to the secondary volume. The reflection status includes "not reflected" indicating that the target data has not been reflected to the secondary volume, and "already reflected" indicating that the target data has been reflected to the secondary volume. The reflection status may include other statuses.
[0156] For example, the write order management table 508 is a table indicating the status that the data of the write object corresponding to the record with ID 801 of 8 in the write order management table 408 has not arrived at the storage system 101. The record in the first row of the write order management table 508 is the record of the earliest write processing with a write order of 1, the write processing targets the CTG with ID 1, the time information of the write processing is 1, the ID of the volume of the write object is 1 (that is, the volume of the write object is SVOL118 (SVOL1)), the write object address is 0 to 255, the volume ID of the log volume storing the data of the write object is 5 (that is, JVOL117 (JVOL5)), and the address of the written log volume is 0 to 255, indicating that the data of the write object has not been reflected in the SVOL.
[0157] Next, a pair creation process for creating a replication pair between the storage system 100 and the storage system 101 of the computer system 10 will be described.
[0158] Fig.11 This is a flowchart of the pair generation process according to the first embodiment.
[0159] Here, when executing the pairing generation processing, the managing terminal 105 sends a pairing generation instruction (the ID of the CTG as the guaranteed range of the write order, the ID of the primary side volume (first volume), and the system ID of the secondary side storage system 101) to the storage node 111 of the storage system 100, for example, according to the user's instructions.
[0160] The replication pair management program 401 of the active storage node 111 (strictly speaking, the CPU 201 executing the replication pair management program 401 ) receives a pair creation instruction from the active management terminal 105 ( S1100 ).
[0161] Next, the replication pair management program 401 refers to the volume management table 407 to confirm the maximum capacity of the leading volume specified in the pair creation instruction ( S1101 ).
[0162] Next, the replication pair management program 401 sends a paired volume creation instruction to the storage system 101 having the system ID specified in the paired volume creation instruction (S1102). Here, the paired volume creation instruction includes the ID of the specified CTG and the capacity of the secondary volume to be created as a pair.
[0163] The volume management program 505 of any storage node 102 of the storage system 101 executes the following paired volume creation process (see Fig.12 ): Receive a paired volume creation instruction (S1103), determine the storage node 102 configured as the volume (SVOL) to be replicated based on the ID of the CTG specified in the paired volume creation instruction, and create a log volume as needed (S1104).
[0164] The volume management program 505 sends a response including the CTG ID, the volume ID of the SVOL that becomes the replication pair, and the volume ID of the journal volume associated with the SVOL to the storage node 111 (S1105).
[0165] The replication pair management program 401 performs the following pair generation preparation processing: receiving a response from the storage system 101 (S1106), and generating a new log volume in the storage node 111 as needed based on the CTG ID and the volume ID of the log volume received as a response (see Fig.13 )(S1107).
[0166] Next, the replication pair management program 401 determines whether the secondary volumes of all replication pairs belonging to the CTG of the specified CTG ID belong to the same storage node (S1108). Specifically, the replication pair management program 401 refers to the replication pair management table 406 to determine whether the secondary log IDs of all replication pairs included in the CTG of the same CTG ID and the volume IDs of the log volumes included in the response received in step S1106 are all the same ID.
[0167] As a result, if the result is "true" (S1108: Yes), the copy pair management program 401 advances the process to step S1109, and if the result is "false" (S1108: No), the copy pair management program 401 advances the process to step S1112.
[0168] In step S1109, the replication pair management program 401 performs the following pair addition processing in the same-node mode: adds replication pair information to the replication pair management table 406 as the same-node mode, instructs the storage system 101 to add pair information to the replication pair management table 506 (see Fig.14 ).
[0169] Next, the replication pair management program 401 copies all the data of the primary volume of the newly created replication pair to the secondary volume of the pair of the storage system 101 (S1110). Here, in the process of copying all the data, for example, the replication pair management program 401 performs a process corresponding to an I / O request from the primary host computer 104, and typically allocates a bitmap to each address to indicate whether there is uncopied data and manages the same in order to copy all the data.
[0170] Next, the replication pair management program 401 updates the state 605 of the record corresponding to the replication pair generated in the replication pair management table 406 to pair, and also issues an instruction to update the state of the replication pair to the storage system 101 (S1111).
[0171] On the other hand, in step S1112, the replication pair management program 401 performs the following cross-node mode pair addition processing: adds pair information to the replication pair management table 406 as a cross-node mode, instructs the storage system 101 to add pair information to the replication pair management table 506 (see Fig.15 ).
[0172] Next, the replication pair management program 401 copies all the data of the primary volume of the newly created replication pair to the secondary volume of the pair of the storage system 101 (S1113). Here, in the process of copying all the data, for example, the replication pair management program 401 performs a process corresponding to an I / O request from the primary host computer 104, and typically allocates a bitmap to each address to indicate whether there is uncopied data and manages the same in order to copy all the data.
[0173] Next, the replication pair management program 401 updates the state 605 of the record in the replication pair management table 406 corresponding to the generated replication pair to pair, and also issues an instruction to update the state of the replication pair to the storage system 101 (S1114). As a result, the storage system 101 updates the state 605 of the record in the replication pair management table 506 corresponding to the generated replication pair to pair.
[0174] Next, the paired volume creation process ( S1104 ) in the storage system 101 will be described.
[0175] Fig.12 This is a flowchart of the paired volume generation process according to the first embodiment.
[0176] The volume management program 505 of the secondary storage system 101 refers to the replication pair management table 506 ( S1200 ), and checks whether the ID of the CTG designated by the paired volume creation instruction is registered ( S1201 ).
[0177] As a result, if the CTG ID is registered in the copy pair management table 506 (True: S1201: Yes), the volume management program 505 advances the process to step S1202, and if not registered (False: S1201: No), the process advances to step S1207.
[0178] In step S1202, the volume management program 505 refers to the volume management table 507 to confirm free IDs and on which storage node a large number of volumes are allocated.
[0179] Next, the volume management program 505 determines on which storage node to generate the volume (paired volume) to be the replication pair based on the information confirmed in step S1202 (S1203). The storage node as the destination for generating the volume can be determined based on the number of volumes (volume count) and the remaining capacity of the storage device. Typically, a storage node with fewer volumes and a storage node with more remaining capacity of the storage device can be determined as the destination for generating the volume.
[0180] Next, the volume management program 505 generates a paired volume (second volume) for the determined storage node, and adds a record corresponding to the paired volume to the volume management table 507 (S1204).
[0181] Next, the volume management program 505 determines whether or not a log volume belonging to the ID of the CTG specified by the paired volume creation instruction exists in the storage node where the paired volume is created (S1205).
[0182] As a result, if a log volume belonging to the ID of the CTG specified by the paired volume creation instruction exists in the storage node where the paired volume is created (if it is "true": S1205: Yes), the volume management program 505 ends the processing. On the other hand, if a log volume belonging to the ID of the CTG specified by the paired volume creation instruction does not exist in the storage node where the paired volume is created (if it is "false": S1205: No), the volume management program 505 creates a log volume (second log volume) in the storage node, adds a record corresponding to the log volume to the volume management table 507 (S1206), and ends the processing.
[0183] In step S1207, the volume management program 505 refers to the volume management table 507 to confirm free IDs and on which storage node a large number of volumes are allocated.
[0184] Next, the volume management program 505 determines on which storage node to generate the volume (paired volume) to be the replication pair based on the information confirmed in step S1207 (S1208). The storage node to be the destination of the volume generation can be determined based on the number of volumes and the remaining capacity of the storage device. Typically, the storage node with fewer volumes and the storage node with more remaining capacity of the storage device can be determined as the destination of the volume generation.
[0185] Next, the volume management program 505 generates a paired volume for the determined storage node, and adds a record corresponding to the paired volume to the volume management table 507 (S1209).
[0186] Next, the volume management program 505 generates a journal volume in the determined storage node, adds a record corresponding to the journal volume to the volume management table 507 (S1210), and terminates the processing.
[0187] Next, the pair creation preparation process ( S1107 ) in the storage system 100 will be described.
[0188] Fig.13 This is a flowchart of the pairing creation preparation process according to the first embodiment.
[0189] The replication pair management program 401 of the storage system 100 refers to the replication pair management table 406 to confirm the existing CTGID (S1300). Next, the replication pair management program 401 refers to the volume management table 407 to confirm the existing log volume (S1301).
[0190] Next, the replication pair management program 401 determines whether or not the CTG ID specified by the pair creation instruction exists based on the confirmation result of step S1300 ( S1302 ).
[0191] As a result, if the CTG ID specified by the pairing generation instruction exists (if it is "true": S1302: Yes), the copy pairing management program 401 causes the processing to enter step S1303. On the other hand, if the CTG ID specified by the pairing generation instruction does not exist (if it is "false": S1302: No), the copy pairing management program 401 causes the processing to enter step S1305.
[0192] In step S1303 , the replication pair management program 401 determines whether the ID of the journal volume of the storage system 101 received in S1106 exists in the replication pair management table 406 .
[0193] As a result, when the ID of the log volume of the storage system 101 exists in the replication pair management table 406 (“true”: S1303 : Yes), the replication pair management program 401 ends the processing.
[0194] On the other hand, when the ID of the log volume of the storage system 101 does not exist in the replication pairing management table 406 (the case of "false": S1303: No), the replication pairing management program 401 generates a corresponding log volume (first log volume) in the storage system 100, adds a record of the generated log volume to the volume management table 407 (S1304), and ends the processing.
[0195] In step S1305, the replication pair management program 401 generates a corresponding journal volume in the storage system 100, and adds a record of the generated journal volume to the volume management table 407. Thereafter, the replication pair management program 401 ends the processing.
[0196] Next, the pairing addition process ( S1109 ) of the same-node intra-pattern in the computer system 10 will be described.
[0197] Fig.14 This is a flowchart of the pair addition process in the same node intra-pattern according to the first embodiment.
[0198] The replication pair management program 401 adds a record corresponding to the generated replication pair to the replication pair management table 406 with the operation mode 609 being in-node (S1400). Next, the replication pair management program 401 sends a pair addition instruction to the storage system 101 serving as the secondary (S1401).
[0199] In contrast, the replication pair management program 504 of the storage system 101 receives the pair addition instruction ( S1402 ), sets the operation mode 609 to in-node, adds the record of the generated replication pair to the replication pair management table 506 ( S1403 ), and transmits a pair addition completion response to the storage system 100 ( S1404 ).
[0200] Next, the replication pair management program 401 of the storage system 100 receives the completion response ( S1405 ), and ends the pair addition process in the intra-node mode.
[0201] Next, the pairing addition process ( S1112 ) in the cross-node mode in the computer system 10 will be described.
[0202] Fig.15 This is a flowchart of the pairing addition process in the cross-node mode according to the first embodiment.
[0203] The replication pair management program 401 of the storage system 101 adds a record of the replication pair created in the replication pair management table 406 with the operation mode 609 being cross-node ( S1500 ).
[0204] Next, the replication pair management program 401 checks the replication pair management table 406, and determines whether the operation mode 609 of all the replication pairs recorded with the same CTG ID is within the node (S1501). As a result, if the operation mode 609 of the other replication pairs recorded with the same CTG ID is within the node (if it is "true": S1501: Yes), the replication pair management program 401 advances the processing to step S1502. On the other hand, if the operation mode 609 of the other replication pair records is across nodes (if it is "false": S1501: No), the processing advances to step S1503.
[0205] In step S1502, the replication pair management program 401 checks the replication pair management table 406, switches the operation mode 609 of all replication pairs with the same CTGID to cross-node, and advances the process to step S1503.
[0206] In step S1503 , the replication pair management program 401 sends a pair addition instruction to the storage system 101 .
[0207] The replication pair management program 504 of the storage system 101 receives the pair addition instruction ( S1504 ), and adds a record of the replication pair to be added in the replication pair management table 506 with the operation mode 609 being cross-node ( S1505 ).
[0208] Next, the replication pair management program 504 checks the replication pair management table 506, and determines whether the action mode 609 of the record of the other replication pair of the same CTG ID is within the node (S1506). As a result, if the action mode 609 of the record of the other replication pair of the same CTG ID is within the node (if it is "true": S1506: Yes), the replication pair management program 504 advances the processing to step S1507. On the other hand, if the action mode 609 of the record of the other replication pair is across nodes (if it is "false": S1506: No), the processing advances to step S1508.
[0209] In step S1507, the replication pair management program 504 checks the replication pair management table 506, switches the operation mode 609 of all replication pairs with the same CTGID to cross-node, and advances the process to step S1508.
[0210] In step S1508 , the replication pair management program 504 sends a pair addition completion response to the storage system 100 .
[0211] Next, the replication pair management program 401 of the storage system 100 receives the completion response ( S1509 ), and ends the pair addition process in the inter-node mode.
[0212] Next, host I / O processing in the computer system 10 will be described.
[0213] Fig.16 This is a flowchart of the host I / O process according to the first embodiment.
[0214] The host I / O processing program 402 of the storage system 100 receives an I / O request issued by the host computer 104 (S1600). Here, the I / O request includes LUN (Logical Unit Number), the type of I / O (whether it is a write process or a read process), the start address, the amount of data, etc., and if it is a write process, it also includes the actual data of the write object. In addition, the correspondence between LUN and volume ID is managed in the storage system 100 by existing technology.
[0215] Next, the host I / O processing program 402 determines whether the I / O request is a write process (S1601). As a result, if the I / O request is a write process ("True": S1601: Yes), the host I / O processing program 402 advances the process to step S1603. On the other hand, if the I / O process is a read process ("False": S1601: No), the process advances to step S1602.
[0216] In step S1602, the host I / O processing program 402 performs a read process. Typically, if past data of the address that is the target of the I / O request exists in the cache, the host I / O processing program 402 reads the cache data, and if past data does not exist, reads the data from the storage device of the storage destination, and sends the read data to the host computer 104.
[0217] In step S1603, the host I / O processing program 402 refers to the replication pair management table 406 to determine whether the volume ID corresponding to the LUN of the I / O request is recorded as the primary volume ID and whether its status is paired. As a result, if the volume ID corresponding to the LUN is recorded as the primary volume ID and its status is paired (if it is "true": S1603: Yes), the host I / O processing program 402 advances the processing to step S1606, and if the volume ID corresponding to the LUN is recorded as the primary volume ID and its status is not paired (if it is "false": S1603: No), the host I / O processing program 402 advances the processing to step S1604.
[0218] In step S1604, the host I / O processing program 402 performs a write process. Typically, the host I / O processing program 402 updates the cache data if there is past data of the address that is the target of the I / O request in the cache, and writes new data to the cache if there is no past data. Next, the host I / O processing program 402 sends a completion response of the write process to the host computer 104 (S1605).
[0219] In step S1606, the host I / O processing program 402 refers to the volume management table 407 to determine whether the write I / O restriction flag 705 of the target volume of the write processing is on. As a result, if the write I / O restriction flag 705 of the target volume is on (if it is "true": S1606: Yes), the host I / O processing program 402 advances the processing to step S1607, and on the other hand, if the write I / O restriction flag 705 of the target volume is not on (if it is "false": S1606: No), the processing advances to step S1608.
[0220] In step S1607 , the host I / O processing program 402 waits for a certain period of time to wait for the process of updating the time information of the log in order to make the time sections of the data consistent, and then advances the process to step S1606 .
[0221] In step S1608, the host I / O processing program 402 performs a write process.
[0222] Next, the host I / O processing program 402 issues a log generation instruction (S1609) by specifying the parameters specified in the write process for the data written by the write order management program 403 in step S1608. Here, the parameters include the volume ID and the start address.
[0223] The write order management program 403 receives the log generation instruction (S1610). Next, the write order management program 403 refers to the replication pair management table 406, confirms the ID of the primary log volume corresponding to the volume to be written, and stores the written data in a free area of the log volume of the confirmed ID (S1611).
[0224] Next, the write order management program 403 adds a record including the address of the data written in step S1611 to the write order management table 408 (S1612). Next, the write order management program 403 sends a completion of adding to the log to the host I / O processing program 402 (S1613).
[0225] The host I / O processing program 402 receives the completion response (S1614). Next, the host I / O processing program 402 sends a completion response of the write process to the host computer 104 (S1615).
[0226] According to the above-described host I / O processing, for a volume whose write I / O restriction flag 705 is on, it is possible to appropriately suppress execution of a write process in order to wait for a process to update the time information of a log.
[0227] Next, the write order management process in the computer system 10 will be described.
[0228] Fig.17 This is a flowchart of the write order management process according to the first embodiment.
[0229] The write order management process is started, for example, when a copy pair is generated in the computer system 10. Here, the host I / O process and the write order management process are examples of the write order guarantee process.
[0230] The write order management program 403 refers to the copy pair management table 406 to obtain information on all existing copy pairs ( S1700 ).
[0231] Next, the write order management program 403 determines whether there is a volume whose status 605 is a pair and whose operation mode 609 is a primary side across nodes based on the acquired copy pair information (S1701). As a result, if such a primary side volume exists (if it is "true": S1701: Yes), the write order management program 403 advances the process to S1702. On the other hand, if such a primary side volume does not exist (if it is "false": S1701: No), the process advances to step S1705.
[0232] In step S1702, the write order management program 403 turns on the write I / O restriction flag 705 of the volume management table 407 for all primary volumes whose status 605 is paired and whose operation mode 609 is cross-node. This inhibits the host computer from writing to these volumes.
[0233] Next, the write order management program 403 increases the setting value of the time information given when the log is added by 1 (1 generation update) (S1703). Thus, the setting value is given to the write processing thereafter.
[0234] Next, the write order management program 403 turns off the write I / O restriction flag 705 of the volume management table 407 for all primary volumes whose status 605 is paired and whose operation mode 609 is cross-node (S1704), and proceeds to step S1705. Thus, the host computer restarts the write process to these volumes.
[0235] In step S1705, the write order management program 403 waits for a certain period of time until the next time information update timing, and after waiting, the processing enters step S1700. In addition, in the above-mentioned write order management processing, steps S1702 to S1704 are executed for all volumes whose status 605 is paired and whose action mode 609 is the positive side across nodes, but the present invention is not limited to this. For example, the processing of steps S1702 to S1704 can also be performed at staggered timing for each CTG. In this way, the influence of the suppression of the write processing on the computer system 10 can be suppressed.
[0236] Next, the log data transfer process in the computer system 10 will be described.
[0237] Fig.18 This is a flowchart of the log data transmission process according to the first embodiment.
[0238] The log data transfer program 404 of the storage system 100 refers to the write order management table 408 to obtain information on all write processes (S1800). Next, the log data transfer program 404 determines whether there is any untransferred log data based on the reflection status 808 (S1801).
[0239] As a result, if there is untransmitted log data (the case of "true": S1801: yes), the log data transmission program 404 causes the processing to enter step S1802. On the other hand, if there is no untransmitted log data (the case of "false": S1801: no), the processing is terminated.
[0240] In step S1802, the log data transfer program 404 transfers the untransmitted log data to the storage system 101 as the secondary storage system. Specifically, the log data transfer program 404 refers to the ID of the log volume of the record ID 1006 of the write order management table 408 corresponding to the untransmitted log data and the address 1007 on the log volume, reads the data from the address corresponding to the log volume corresponding to the ID, refers to the copy pair management table 406, selects the copy path of the corresponding record copy path ID 610, specifies the volume ID of the recorded secondary volume ID 608 as a parameter, and transfers the log data.
[0241] When receiving the transferred log data (S1803), the log data receiving program 501 of the storage system 101 refers to the replication pair management table 506, identifies the secondary log ID corresponding to the designated secondary volume ID, and stores the log data in the identified log volume (S1804).
[0242] Next, the log data receiving program 501 registers information of the received log data and the address of the log volume where the log data is written as a record in the write order management table 508 (S1805). Next, the log data receiving program 501 sends a completion response to the storage system 100 (S1806).
[0243] The log data transfer program 404 of the storage system 100 receives the completion response (S1807). Next, the log data transfer program 404 updates the reflection status 808 of the record corresponding to the transferred log data in the write order management table 408 to transfer completed (S1808), and the process proceeds to step S1800.
[0244] In addition, Fig.18 In the log data transmission process, an example of transmitting log data from the storage system on the primary side to the storage system on the secondary side starting from the primary side is shown, but the present invention is not limited to this. The storage system on the secondary side may also periodically request the storage system on the primary side to transmit log data, and in response, the storage system on the primary side sends log data.
[0245] Next, the data reflection arbitration process in the storage system 101 will be described.
[0246] Fig.19This is a flowchart of the data reflection arbitration process according to the first embodiment.
[0247] The data reflection arbitration process is performed, for example, periodically. The data reflection arbitration program 503 of the storage system 101 refers to the write order management table 508 to obtain information on records corresponding to all log data (S1900). Next, the data reflection arbitration program 503 determines whether there is log data that is not reflected in the volume by determining whether there is a record that does not reflect the reflection status 1008 in the record (S1901). As a result, if there is unreflected log data (in the case of "true": S1901: yes), the data reflection arbitration program 503 advances the process to step S1902, and if there is no unreflected log data (in the case of "false": S1901: no), the data reflection arbitration program 503 ends the process.
[0248] In step S1902 , the data reflection arbitration program 503 refers to the replication pair management table 506 based on the volume ID of the write target corresponding to each log data, and determines whether the operation mode 609 of the record of the replication pair corresponding to the volume ID is cross-node.
[0249] As a result, when the action mode 609 is across nodes (when it is "true": S1902: Yes), the data reflects that the arbitration procedure 503 causes the processing to enter step S1903. On the other hand, when the action mode 609 is within the node (when it is "false": S1902: No), the processing enters step S1907.
[0250] In step S1903, the data reflection arbitration program 503 refers to the write order management table 508, and collects the latest time information 1003 of the log data corresponding to each secondary log ID belonging to the same CTG ID (S1903).
[0251] Here, all log data up to the generation before the latest generation reached by each secondary log belonging to the same CTG (i.e., the generation whose time information is the generation number obtained by subtracting 1 from the latest generation number) has arrived at the secondary log volume. Therefore, the data reflection arbitration program 503 reflects the unreflected log data up to the time information of the previous generation to the address of the write address 1005 of the volume (SVOL) corresponding to the volume ID 1004 that is the write target of the secondary volume (S1904).
[0252] The data reflection arbitration program 503 updates the reflection status 1008 of the record corresponding to the log data whose reflection to the secondary volume has been completed to reflection completion (S1905). At this time, the data reflection arbitration program 503 may also notify the storage system 100 that the log data has been reflected, and delete the record corresponding to the log data from the write order management table 408, 508.
[0253] After the log data is reflected, the data reflection arbitration program 503 waits for a certain period of time (S1906), and then proceeds to step S1901.
[0254] In step S1907, the data reflection arbitration program 503 collects the log data of each secondary-side log ID belonging to the same CTG ID.
[0255] Here, since the order of the log IDs of the data in the log volume is consistent with the writing order, arbitration between multiple logs is not required. Therefore, the data reflection arbitration program 503 reflects the unreflected log data in order from the smallest ID number to the address corresponding to the write address 1005 of the volume corresponding to the volume ID of the write target volume ID 1004 that becomes the secondary side volume (S1908).
[0256] Next, the data reflection arbitration program 503 updates the reflection status 1008 of the record corresponding to the log data whose reflection to the secondary volume has been completed to reflection completion (S1909), and the process proceeds to step S1901. At this time, the data reflection arbitration program 503 may also notify the storage system 100 that the log data has been reflected, and delete the record corresponding to the log data from the write order management table 408, 508.
[0257] In this embodiment, data reflection arbitration processing is performed in the secondary side storage system, but it can also be that the primary side storage system 100 refers to the reflection status 808 in the record of the write order management table 408 as the time information of the transmitted record, and regularly notifies the secondary side storage system of the previous generation that can be reflected to the latest generation number that has arrived, and the secondary side storage system that receives it reflects the log data to the secondary volume.
[0258] In addition, in the present embodiment, log data is generated during I / O processing and continuously transmitted to the storage system on the secondary side, but a snapshot (Snapshot) in a state where write I / O is suppressed may be generated periodically at a granularity such as a few minutes, hours, or days, and the difference from the previous snapshot may be transmitted to the storage system on the secondary side, thereby reflecting (copying) the data to the secondary side.
[0259] [Second embodiment]
[0260] Next, a computer system according to a second embodiment will be described. In this embodiment, the parts different from the computer system according to the first embodiment will be basically described. In addition, the same functional parts as those of the computer system according to the first embodiment will be described using the same reference numerals.
[0261] In the second embodiment, the function sharing of each storage node 102 in the storage system 101, the table holding method, and the information of the pair generation instruction from the user are different.
[0262] First, the function sharing and table holding method of each storage node 102 will be described. Each storage node 102 operates independently like the storage system 100, but has the same ID as the storage system.
[0263] The storage nodes 102 of the storage system 101 include representative nodes and other nodes (general nodes). In the representative node, all programs similar to those in the first embodiment are executed, and each table contains records of all storage nodes of the storage system 101. On the other hand, in the general node, only records related to its own volume or replication pair are included, and it executes according to instructions from the representative node, so only part of the program is executed.
[0264] Next, the information specified by the user as the pair generation instruction is described. When generating a pair, it is assumed that the path between the primary and secondary storage systems has been established in advance, and the log volume, primary and secondary volumes have also been generated by the user.
[0265] In this structure, the user specifies the ID of the CTG, the ID of the log volume, the ID of the primary volume, and the ID of the secondary volume as a pairing generation instruction. In the communication system of this embodiment, it is not possible to understand from the primary storage system that the secondary storage system is a storage system composed of multiple nodes, and it is not possible to perform direct operations on appropriate storage nodes. Therefore, in this embodiment, the operation instruction is transmitted between the storage nodes of the secondary storage system so that the storage node having the volume to be operated can be operated.
[0266] Fig. 20 is a flowchart of the pairing generation process of the second embodiment. Fig.11 The same processing steps in the pair generation process of the first embodiment shown are denoted by the same reference numerals.
[0267] Here, when the pairing generation processing is executed, the managing terminal 105 sends a pairing generation instruction (the ID of the CTG as the guaranteed range of the write order, the ID of the primary side log volume, the ID of the secondary side log volume, the ID of the primary side volume, and the ID of the secondary side volume) to the storage node 111 of the storage system 100, for example, according to the user's instructions.
[0268] The replication pair management program 401 of the active storage node 111 (strictly speaking, the CPU 201 executing the replication pair management program 401 ) receives a pair creation instruction from the active management terminal 105 ( S2000 ).
[0269] Next, the replication pair management program 401 refers to the replication pair management table 406 to determine whether the main-side log ID belonging to the CTG ID specified by the pair generation instruction is a single ID and the same ID as the main-side log ID specified by the pair generation instruction (S2001). As a result, if the main-side log ID belonging to the specified CTG ID is a single ID and the same ID as the main-side log ID specified by the pair generation instruction (in the case of "true": S2001: Yes), the replication pair management program 401 advances the processing to step S2002. On the other hand, if the main-side log ID belonging to the specified CTG ID is not a single ID and the same ID as the main-side log ID specified by the pair generation instruction (in the case of "false": S2001: No), the processing advances to step S2005.
[0270] In step S2002, the replication pair management program 401 executes pair addition processing in the same node mode (see Fig.21 ), execute data replication of all addresses (S1100) and update of the replication pair management table 406 (S1111).
[0271] In step S2005, the replication pair management program 401 executes the pair addition process in the cross-node mode (see Fig. 22 ), execute data replication of all addresses (S1113) and update of the replication pair management table 406 (S1114).
[0272] Next, the pairing addition process ( S2002 ) of the same-node intra-pattern in the computer system 10 will be described.
[0273] Fig.21 is a flowchart of the pairing addition process of the same node mode in the second embodiment. Fig.21 In, with Fig.14 The same steps in the pair addition process in the same node mode of the first embodiment are denoted by the same reference numerals. In the second embodiment, the processes of steps S1402 , S1403 , and S1404 are executed by a representative storage node (representative node) among the plurality of storage nodes 102 .
[0274] The replication pair management program 504 of the representative node refers to the volume management table 507 to confirm to which storage node 202 the log volume of the replication pair added in step S1403 belongs ( S2100 ).
[0275] Next, the replication pair management program 504 of the representative node instructs the identified storage node 202 to add a record related to the replication pair to the replication pair management table 506 having only records of volumes associated with the storage node ( S2101 ).
[0276] Next, the pairing addition process ( S2005 ) in the cross-node mode in the computer system 10 will be described.
[0277] Fig. 22 is a flowchart of the pairing addition process of the cross-node mode of the second embodiment. Fig. 22 In, with Fig.15 The same steps as those in the cross-node mode pairing addition process of the first embodiment are marked with the same reference numerals. In the second embodiment, the processes of steps S1504, S1505, S1506, S1507, and S1508 are performed by a representative storage node (representative node) among the storage nodes 102.
[0278] The replication pair management program 504 of the representative node refers to the volume management table 507 to identify all storage nodes belonging to the CTG of the specified CTG ID and the storage node to which the log volume of the specified secondary log ID belongs ( S2200 ).
[0279] Next, the replication pair management program 504 of the representative node instructs the storage node where the designated log volume exists to add a replication pair record to the replication pair management table 506 of the storage node ( S2201 ).
[0280] Furthermore, the replication pair management program 504 of the representative node issues an instruction to all storage nodes belonging to the replication pair in the same CTG to rewrite the operation mode 609 of the existing pairs in the replication pair management table 506 of these storage nodes from intra-node to inter-node (S2202).
[0281] According to the computer system of the second embodiment, in the secondary storage system, the representative storage node executes processing for notifying other storage nodes of various information, so that in the primary storage system, processing can be performed regardless of the structure of the secondary storage node.
[0282] [Third Embodiment]
[0283] Next, a computer system according to a third embodiment will be described. In this embodiment, the parts different from the computer system according to the second embodiment will be basically described. In addition, the same functional parts as those of the computer system according to the first embodiment will be described using the same reference numerals.
[0284] In the computer system of the third embodiment, the storage system 101 of the second embodiment operates as a primary storage system (first storage system), and the storage system 100 operates as a secondary storage system (second storage system). In this embodiment, the storage node 102 of the storage system 101 also stores programs required for the primary side (write order management program 403, etc.) as programs of the storage node 111, and the storage node 111 stores programs required for the secondary side (data reflection arbitration program 503, etc.) as programs of the storage node 102.
[0285] Next, the write order management process in the computer system 10 will be described.
[0286] Fig.23 3 is a flowchart of the write order management process of the third embodiment. Fig.23 In, with Fig.17 The same steps in the write order management process of the first embodiment are denoted by the same reference numerals. In the third embodiment, the processes of steps S1700 , S1701 , and S1705 are executed by a representative storage node (representative node) among the storage nodes 102 .
[0287] The representative node's write order management program 403 simultaneously issues an instruction to set the write I / O restriction flag 906 of the volume management table 507 to be enabled for all replication pairs operating in the cross-node mode (S2300). As a result, each storage node having a volume that becomes a replication pair receives the instruction, sets the write I / O restriction flag 906 of the entry corresponding to the corresponding volume in the volume management table 507 to be enabled, and sends a notification of the completion of the setting to the representative node.
[0288] The representative node write order management program 403 waits until all storage nodes 102 that issued the instruction have completed receiving the instruction (S2301). When receiving the instruction from all storage nodes 102, the representative node increases the setting value of the time information assigned when the log is appended by 1 (1 generation update), and instructs each storage node to set the write I / O restriction flag 906 of the record corresponding to the volume that becomes the replication pair in the volume management table 507 to off (S2302). Thus, in the future, each storage node assigns the setting value to the write process.
[0289] [Fourth Embodiment]
[0290] Next, a computer system according to a fourth embodiment will be described. In this embodiment, the parts different from the computer system according to the third embodiment will be basically described. In addition, the same functional parts as those of the computer system according to the first embodiment will be described using the same reference numerals.
[0291] The computer system of the fourth embodiment is an example in which the secondary management terminal 108 is responsible for a part of the processing when the copy pair operation is performed on the storage system 101 in the computer system of the third embodiment.
[0292] In the present embodiment, when generating a pair, the secondary management terminal 108 stores information equivalent to the copy pair management table 506 in the memory 202 and performs the same operation as described above. Fig.23 The write order management process shown is the same process.
[0293] Here, the storage system 101 is composed of a plurality of storage nodes 102, but the ID of the storage system is one, so it is processed as one node from the secondary management terminal 108. Therefore, the secondary management terminal 108 cannot perform the processing of being aware of the plurality of storage nodes as performed in step S2300 and send an instruction to one storage node.
[0294] Therefore, in this embodiment, when a storage node receives an instruction from the secondary management terminal 108, it transmits the received instruction to the representative node, and the representative node determines the location of the storage node to which the volume belongs, transmits the instruction to the storage node with reference to the information in the replication pair management table 506 and the volume management table 507, and collects the completion responses obtained from each storage node into one. Then, the representative node instructs the storage node that received the instruction from the secondary management terminal 108 to respond to the secondary management terminal 108 with a completion response.
[0295] By such processing, the secondary management terminal 108 can perform similar control regardless of whether the primary storage system is the storage system 100 or the storage system 101 composed of a plurality of nodes.
[0296] [Fifth Embodiment]
[0297] Next, a computer system according to a fifth embodiment will be described. In this embodiment, the parts different from those of the computer system according to the first embodiment will be basically described. In addition, the same functional parts as those of the computer system according to the first embodiment will be described using the same reference numerals.
[0298] In this embodiment, multiple device management terminals are connected to both the primary storage system and the secondary storage system and have the function of managing multiple storage systems. The multiple device management terminals extract and effectively utilize operation information from multiple storage systems to perform processing such as determining volume configuration for efficiently utilizing resources such as CPU and capacity while maintaining the RPO (Recovery Point Objective) specified by the user.
[0299] For example, when writing order control is performed across multiple storage nodes, log data at a time point that is common to all logs and has not yet arrived is not reflected in the secondary volume, so RPO increases. In this embodiment, volume configuration is performed that can maintain the RPO specified by the user and improve resource utilization efficiency.
[0300] Fig.24 It is an overall structural diagram of a computer system according to the fifth embodiment.
[0301] The computer system 10A further includes a plurality of device management terminals 2400 in the computer system 10. The plurality of device management terminals 2400 are connected to the storage system 100 via a network 2401 and are connected to the storage system 101 via a network 2402. The plurality of device management terminals 2400 are an example of a management device and may be Figure 3 The hardware structure shown can also be a virtual machine on the cloud.
[0302] Fig.25 It is a diagram showing the configuration of memories of a plurality of device management terminals according to the fifth embodiment.
[0303] The memory 202 of the plurality of device management terminals 2400 includes an operation information collection program 2500 , a service level maintenance program 2501 , an optimal pair generation program 2502 , a replication pair management table 2503 , a volume management table 2504 , a service level information table 2505 , and an operation information table 2506 .
[0304] The operating information collection program 2500 is executed by the CPU 201 to periodically collect operating information such as the CPU operating rate in each system from the storage system 100 and the storage system 101 , and performs processing to store in the operating information table 2506 .
[0305] The service level maintenance program 2501 is executed by the CPU 201 to store the indexes such as RPO specified by the user in the service level information table 2505 , and to periodically check whether these indexes are satisfied, and to notify the user of an alarm if they are not satisfied.
[0306] The optimal pair generation program 2502 is executed by the CPU 201 to perform the following processing: accept an instruction to generate a replication pair from the user, determine the volume configuration of the remote replication pair so that resources such as the CPU can be used to the maximum extent while maintaining the specified service level, and instruct the storage system 100 and the storage system 101 to generate volumes and replication pairs.
[0307] The replication pair management table 2503 holds information related to the replication pairs of the storage systems 100 and 101. The replication pair management table 2503 has the same records as the replication pair management table 406 and the replication pair management table 506.
[0308] The volume management table 2504 holds information related to the volumes of the storage systems 100 and 101. The volume management table 2504 has the same records as the volume management table 407 and the volume management table 507.
[0309] The service level information table 2505 stores service level indicators such as RPO specified by the user. The operation information table 2506 stores operation information such as the CPU operation rate of the storage systems 100 and 101. The service level information table 2505 and the operation information table 2506 will be described in detail later.
[0310] Next, the operation information table 2506 will be described in detail.
[0311] Fig.26 It is a diagram showing the structure of the operation information table according to the fifth embodiment.
[0312] The operation information table 2506 is a table for managing operation information and stores a record for each storage node. The record of the operation information table 2506 includes fields for system ID 2600 , node ID 2601 , time 2602 , CPU operation rate 2603 , and available capacity 2604 .
[0313] The ID of the storage system to which the storage node corresponding to the entry belongs is stored in the system ID 2600. The node ID of the storage node corresponding to the entry is stored in the node ID 2601. In addition, since the storage system 100 has multiple storage nodes but is a node for redundancy, in the storage system 100, multiple nodes work as one storage node, and therefore no node ID is set in the node ID 2601.
[0314] The time at which the operation information of the storage node corresponding to the entry is acquired is stored in the time 2602. The CPU operation rate of the storage node corresponding to the entry is stored in the CPU operation rate 2603. The remaining available capacity of the storage node corresponding to the entry is stored in the available capacity 2604.
[0315] In addition, the operation information is stored in each storage node's own memory 202 and can be acquired from these storage nodes. In addition, the operation information is not limited to the CPU operation rate, and may be, for example, IOPS (Input / Output Operations Per Second).
[0316] According to the second record of the operation information table 2506, it can be known that at 10:00, the CPU operation rate of the storage node 102 (storage node 1) with the node ID 1 in the storage system 101 with the system ID 2 is 45% and the available capacity is 10 TB.
[0317] Next, the service level information table 2505 will be described in detail.
[0318] Fig. 27 It is a diagram showing the structure of the service level information table according to the fifth embodiment.
[0319] The service level information table 2505 is a table for managing the service level of the volume, and stores a record for each volume. The record of the service level information table 2505 includes fields for system ID 2700 , volume ID 2701 , RPO 2702 , and maximum capacity 2703 .
[0320] The ID of the storage system to which the volume corresponding to the entry belongs is stored in the system ID 2700. The ID of the volume corresponding to the entry is stored in the volume ID 2701. The value of an index indicating the extent to which data rewinding is permitted when the volume corresponding to the entry fails is stored in the RPO 2702. The maximum capacity available for the volume corresponding to the entry is stored in the maximum capacity 2703.
[0321] According to the first record of the service level information table 2505, the volume with volume ID 1 in the storage system 100 with system ID 1 can allow data rewinding for 1 minute in case of failure, and the maximum capacity is 500 GB.
[0322] Next, an optimal pair creation process in which the plurality of device management terminals 2400 create replication pairs between the storage system 100 and the storage system 101 will be described.
[0323] Fig.28 This is a flowchart of the optimal pair generation process according to the fifth embodiment.
[0324] Here, when executing the optimal pairing generation processing, the management terminal 105 sends a pairing generation instruction (the ID of the CTG as the guaranteed range of the write order, the system ID of the primary side storage system, the ID of the primary side volume, and the system ID of the secondary side storage system 101) to multiple device management terminals 2400, for example, based on the user's instructions.
[0325] The optimal pairing generation program 2502 of the plurality of device management terminals 2400 (strictly speaking, the CPU 201 that executes the optimal pairing generation program 2502 ) receives a pairing generation instruction from the current management terminal 105 ( S2800 ).
[0326] Next, the optimal pair generation program 2502 refers to the replication pair management table 2503, the volume management table 2504, and the service level information table 2505, and selects a node that satisfies the service level of the volume designated as the primary side and maximizes the resource utilization efficiency of the storage system on the secondary side (S2801). Typically, the optimal pair generation program 2502 first selects from the replication pair management table 2503 whether the action mode 609 is intra-node or inter-node, so as to satisfy the index value of RPO 2702.
[0327] For example, when the index value of RPO2702 is short, configuration within the node is necessary, so the storage node that manages the volume within the node as the storage node for configuring the secondary volume is temporarily determined as a candidate for the configuration destination. Here, when the available capacity 2604 of the candidate storage node is greater than the maximum capacity 2703 of the maximum capacity of the positive side volume, the candidate storage node is determined as the volume configuration destination. When the available capacity is less than the maximum capacity, it means that configuration is not possible, so it is set to end in error.
[0328] On the other hand, when the index value of RPO2702 is long, cross-node selection can be performed, so all nodes in the storage system of the secondary side are used as candidates to determine whether the range of the volume configuration destination can be narrowed and whether it can be configured. For example, it is determined whether the available capacity of available capacity 2604 is greater than the maximum capacity of maximum capacity 2703 and is a storage node that can store volumes. If there are multiple storage nodes that can store volumes, refer to the CPU operating rate of CPU operating rate 2603 and determine the storage node with the most spare capacity as the configuration destination.
[0329] Next, the optimal pair generation program 2502 executes a pair volume generation process for generating a secondary volume based on the determined arrangement and generating a log volume as needed (see Fig.29 )(S2802).
[0330] Next, the optimal pair generation program 2502 determines whether the pair formation is in the same node mode based on the result of the volume configuration in step S2801 (S2803). If the result is that the pair addition is in the same node mode (if it is "true": S2803: Yes), the optimal pair generation program 2502 advances the processing to step S2804. On the other hand, if it is in the cross-node mode (if it is "false": S2803: No), the processing advances to step S2805.
[0331] In step S2804, the optimal pair generation program 2502 instructs the storage system on the leading side to perform pair addition processing in the same node mode. Fig.14 The following figure shows the pairing addition process of the patterns within the same node.
[0332] In step S2805, the optimal pair generation program 2502 instructs the storage system on the leading side to perform pair addition processing in the cross-node mode. Fig.15 Paired addition processing for cross-node mode is shown.
[0333] Next, the paired volume creation process (S2802) will be described in detail.
[0334] Fig.29 This is a flowchart of paired volume creation processing according to the fifth embodiment.
[0335] The optimal pair generation program 2502 refers to the replication pair management table 2503 and obtains information of all records (S2900). Next, the optimal pair generation program 2502 determines whether the designated CTG ID matches the CTG ID of the existing replication pair (S2901).
[0336] As a result, when the designated CTG ID is consistent with the existing CTG ID (“true”: S2901: Yes), the optimal pair generation program 2502 advances the processing to step S2902. On the other hand, when the designated CTG ID is inconsistent with the existing CTG ID (“false”: S2901: No), the processing advances to step S2905.
[0337] In step S2902, the optimal pair generation program 2502 specifies a storage node to the secondary storage system and instructs it to generate a paired volume. The secondary storage system receives the instruction and performs the same process as step S1204 to generate a paired volume.
[0338] Next, the optimal pair generation program 2502 determines whether there is a log volume having a secondary log ID associated with the specified CTG ID in the storage node to be generated (S2903). As a result, if there is a log volume having a secondary log ID associated with the specified CTG ID (if it is "true": S2903: Yes), the optimal pair generation program 2502 ends the processing. On the other hand, if there is no log volume having a secondary log ID associated with the specified CTG ID (if it is "false": S2903: No), the processing proceeds to S2904.
[0339] In step S2904, the optimal pair generation program 2502 specifies a storage node to the secondary storage system and sends a log volume generation instruction. The secondary storage system receiving the instruction then performs the same process as step S1206 to generate a log volume.
[0340] In step S2905, the optimal pair creation program 2502 instructs creation of a paired volume in the same manner as in step S2902. The secondary storage system, having received the instruction, performs the same process as in step S1204 to create a paired volume.
[0341] Next, the optimal pair generation program 2502 instructs the generation of a log volume in the same manner as step S2904 (S2906). The secondary storage system that has received the instruction performs the same process as step S1206 to generate a log volume. The optimal pair generation program 2502 then terminates the process.
[0342] In addition, the present invention is not limited to the above-mentioned embodiment, and can be implemented with appropriate modifications within the scope not departing from the gist of the present invention.
[0343] For example, in the above-mentioned embodiments, part or all of the processing performed by the processor may also be performed by a hardware circuit. In addition, the program in the above-mentioned embodiments may be installed from a program source. The program source may also be a program distribution server or a recording medium (e.g., a removable recording medium).
[0344] Explanation of symbols
[0345] 10.10A Computer System
[0346] 100, 101 storage system
[0347] 102, 111 storage nodes
[0348] 103 Inter-node network
[0349] 104 Host computer
[0350] 105 Management Terminal
[0351] 106 positive and side network
[0352] 107 Secondary Host Computer
[0353] 108 secondary management terminal
[0354] 110 Storage System Network
[0355] 111 storage nodes
[0356] 112, 113CTG.
Claims
1. A computer system comprising a first storage system and a second storage system, characterized in that: The first storage system manages a plurality of first volumes belonging to a consistency group that ensures the writing order of data. The second storage system has multiple storage nodes. The second storage system generates a plurality of second volumes as replication destinations of the plurality of first volumes in a dispersed manner on the plurality of storage nodes, and a second log volume storing log data indicating the contents written in the first volume as the replication source of the second volume exists in each of the plurality of storage nodes for generating the plurality of second volumes. The first storage system generates a plurality of first log volumes corresponding to the respective second log volumes and storing log data about the plurality of first volumes, A write order guarantee process is executed to control the processes for the plurality of first volumes so that the journal data representing the contents written to the plurality of first volumes can be stored in the plurality of first journal volumes in a state in which the write order of the plurality of first journal volumes can be guaranteed.
2. The computer system according to claim 1, characterized in that: The second storage system generates a second volume as a copy destination of the first volume at any storage node, and generates the second log volume when the storage node does not have a second log volume storing log data indicating the written content in the first volume as a copy source of the second volume. The first storage system determines whether the plurality of second volumes that are the replication destinations of the plurality of first volumes are generated in one storage node of the second storage system or in a plurality of storage nodes, and in the case of generating the plurality of second volumes in a plurality of storage nodes, prepares first log volumes corresponding to the plurality of second log volumes generated in the plurality of storage nodes, respectively. The write order assurance process is executed when a plurality of the second volumes serving as the copy destinations of a plurality of the first volumes are generated in a plurality of storage nodes of the second storage system.
3. The computer system according to claim 2, characterized in that: The first storage system receives identification information of a consistency group storing a volume to be replicated and identification information of the first volume to be replicated, The received first volume of the consistency group is used as an object to generate the second volume.
4. The computer system according to claim 2, characterized in that: The second storage system determines a storage node for generating the second volume from among the plurality of storage nodes according to the number of volumes of the storage node or the remaining capacity of the storage node.
5. The computer system according to claim 2, characterized in that: The computer system further comprises a management device, The management device determines a storage node for configuring the second volume based on a target recovery time point for the first volume.
6. The computer system according to claim 1, characterized in that: The write order guarantee process includes a process of suppressing write processes to the plurality of first volumes when updating time information for adding the log data.
7. The computer system according to claim 6, characterized in that: The computer system further includes: a storage unit that stores restriction information indicating whether to suppress write processing when updating the time information for the plurality of volumes; When updating the time information, the first storage system refers to the restriction information and determines whether to suppress a write process to the volume.
8. The computer system according to claim 6, characterized in that: The first storage system has a plurality of storage nodes that act independently. When updating time information assigned to the log data, one storage node of the first storage system instructs all other storage nodes to suppress write processing to the plurality of first volumes.
9. A remote copy control method for a computer system, the computer system comprising a first storage system and a second storage system, characterized in that: The second storage system has multiple storage nodes. The first storage system manages a plurality of first volumes belonging to a consistency group that ensures the writing order of data. The second storage system generates a plurality of second volumes as replication destinations of the plurality of first volumes in a dispersed manner in the plurality of storage nodes, and a second log volume storing log data indicating the contents written in the first volume as the replication source of the second volume exists in each of the plurality of storage nodes for generating the plurality of second volumes. The first storage system generates a plurality of first log volumes corresponding to the respective second log volumes and storing log data about the plurality of first volumes, A write order guarantee process is executed to control the processes for the plurality of first volumes so that the journal data representing the contents written to the plurality of first volumes can be stored in the plurality of first journal volumes in a state in which the write order of the plurality of first journal volumes can be guaranteed.
10. A remote copy control program product, used to enable multiple computers in a computer system including a first storage system and a second storage system to execute, characterized in that: The second storage system has multiple storage nodes. The remote replication control program product enables the computer of the first storage system to manage multiple first volumes belonging to a consistency group that ensures the writing order of data, and enables the computer of the second storage system to generate multiple second volumes that serve as the replication destinations of the multiple first volumes in a dispersed manner on the multiple storage nodes, and enables the multiple storage nodes that generate the multiple second volumes to respectively have a second log volume that stores log data of the write content in the first volume that is the replication source of the second volume, and enables the computer of the first storage system to generate multiple first log volumes corresponding to each of the second log volumes and storing log data about the multiple first volumes, and execute write order guarantee processing for controlling the processing of the multiple first volumes so that the log data representing the write content to the multiple first volumes can be stored in the multiple first log volumes in a state that can ensure the writing order of the multiple first log volumes.
Citation Information
Patent Citations
Storage system and remote copy control method of storage system
JP2007264946A
Storage system and method for controlling the same
JP2019101702A
Cited By
Arbitration method and system and electronic equipment
CN120950428A