Computer system, remote copy control method, and remote copy control program

The computer system addresses the challenges of asynchronous remote copying by distributing volumes and journal volumes across multiple storage nodes, ensuring efficient and appropriate data copy operations while maintaining data integrity and write order.

JP2025089894AActive Publication Date: 2025-06-16HITACHI VANTARA LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023204858
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-16
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

Existing data copy technologies face challenges in easily and appropriately setting up and executing asynchronous remote copying between a source storage system and a destination secondary storage system, especially when the secondary storage system is a distributed system with multiple storage nodes.

Method used

A computer system that includes a first storage system managing volumes within a consistency group ensuring write order, and a second storage system with multiple storage nodes, where each node creates copy destination volumes and corresponding journal volumes for write order guarantee, allowing for efficient distribution of volumes across nodes.

Benefits of technology

Enables easy and appropriate setup and execution of asynchronous remote copy operations, ensuring data integrity and write order across multiple storage nodes, thereby simplifying user settings and improving operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089894000001_ABST
    Figure 2025089894000001_ABST
Patent Text Reader

Abstract

To make it possible to easily and appropriately perform setting of asynchronous remote copy from a copy source storage system to a copy destination sub-storage system including a plurality of storage nodes and execution of the asynchronous remote copy.SOLUTION: In a computer system 10, a storage system 100 manages a plurality of first volumes that belong to a CTG. A storage system 101 creates a plurality of second volumes in a distributed manner in a plurality of storage nodes 102, and causes second journal volumes to exist in the plurality of storage nodes. The storage system 100 creates a first journal volumes corresponding to each of the second journal volumes, and executes write order guarantee processing of controlling processing for the plurality of first volumes so that journal data can be stored in the plurality of first journal volumes in a state in which a write order is guaranteed.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data copy technology between storage systems.

Background Art

[0002] Techniques for copying data between storage systems are known. For example, Patent Document 1 discloses a method for guaranteeing the update order of data across devices by determining an updatable point in each device based on write order information in asynchronous remote copy between a plurality of storage devices.

[0003] Further, Patent Document 2 discloses a method in which, in a distributed storage system composed of a plurality of storage nodes, in order to ensure data redundancy while maintaining response performance, I / O processing is shared and executed by each node, and a physical area owned by a certain node is preferentially allocated as a storage area handled by the node.

[0004] Also, in a computer system, in remote copy from a primary storage system (copy source storage system) to a secondary storage system (copy destination storage system), a CTG (consistency group) may be configured that guarantees the write order of data for a plurality of volumes.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] For example, when asynchronously remotely copying the volume of the primary storage system in a computer system to the secondary storage system, the secondary storage system may be a distributed storage system composed of a plurality of storage nodes. In this case, the volumes that are the copy destinations of the respective volumes belonging to the same CTG in the primary storage system may be created by being distributed among the plurality of storage nodes.

[0007] Thus, when distributing a plurality of volumes belonging to the same CTG in the secondary storage system among a plurality of storage nodes, it is necessary to create a journal volume for storing data indicating the updated contents of the volume for each storage node, and also create corresponding journal volumes in the primary storage system, which is troublesome for user settings. Also, when a plurality of volumes belonging to the same CTG in the secondary storage system are distributed among a plurality of storage nodes, processing for guaranteeing the update order of data for the plurality of volumes belonging to the CTG must be performed.

[0008] On the other hand, it may be to create the volumes that are the copy destinations of the respective volumes belonging to the same CTG in the primary storage system in one storage node. In this case, the user must make settings different from the case of distributing a plurality of volumes belonging to the same CTG in the secondary storage system among a plurality of storage nodes, which is a complicated process for the user.

[0009] The present invention has been made in view of the above circumstances, and an object thereof is to provide a technology capable of easily and appropriately setting and executing asynchronous remote copying from a source storage system to a destination secondary storage system including a plurality of storage nodes. Means for Solving the Problems

[0010] To achieve the above object, a computer system according to one aspect is a computer system including a first storage system and a second storage system, wherein the first storage system manages a plurality of first volumes belonging to a consistency group that guarantees the write order of data, and the second storage system has a plurality of storage nodes, and the second storage system distributes and creates a plurality of second volumes that are copy destinations of the plurality of first volumes on the plurality of storage nodes, and for each of the plurality of storage nodes that create the plurality of second volumes, a second journal volume is provided to store journal data indicating the write content in the first volume that is the copy source of the second volume. The first storage system creates a plurality of first journal volumes corresponding to the respective second journal volumes and storing journal data for the plurality of first volumes, and executes a write order guarantee process for controlling the processing of the plurality of first volumes so that journal data indicating the write content for the plurality of first volumes can be stored in the plurality of first journal volumes in a state where the order of writing to the plurality of first journal volumes can be guaranteed.

Advantages of the Invention

[0011] According to the present invention, it is possible to easily and appropriately set up and execute asynchronous remote copy from a copy source storage system to a copy destination storage system including a plurality of storage nodes.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Embodiments for Carrying Out the Invention

[0013] The embodiments will be described with reference to the drawings. It should be noted that the embodiments described below do not limit the invention according to the claims, and not all of the elements and combinations thereof described in the embodiments are essential for the solution means of the invention.

[0014] In the following description, information may be described in terms of the "AAA table", but the information may be represented in any data structure. That is, in order to indicate that the information is independent of the data structure, the "AAA table" can be referred to as "AAA information".

[0015] Also, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be combined into one table.

[0016] Also, in the following description, the "program" may be used as the main body of the operation to describe the process. However, the program is executed by a processor (e.g., a CPU (Central Processing Unit)) to perform the defined process while appropriately using a memory or a communication I / F. Therefore, the main body of the process may be the processor (or a device such as a controller or a computer having the processor).

[0017] Also, the program may be installed from a program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0018] Also, in the following description, when describing elements of the same type without distinction, reference numerals may be used, and when describing elements of the same type separately, the ID (e.g., identification number) of the element may be used. For example, when describing the storage nodes without particular distinction, it may be described as "storage node 102", and when describing individual nodes separately, it may be described as "storage node 1", "storage node 2", etc. Also, in the following description, the name of an element within a storage node n (where n is a natural number) may be suffixed with n to distinguish which node the element belongs to.

[0019] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the present invention is not limited to the embodiments described below.

[0020] [First Embodiment] FIG. 1 is an overall configuration diagram of a computer system according to the first embodiment.

[0021] The computer system 10 includes a storage system 100, a primary host computer 104, a primary management terminal 105, a storage system 101, a secondary host computer 107, and a secondary management terminal 108.

[0022] The storage system 100 and the storage system 101 are connected via a storage system - to - storage system network 110. The storage system - to - storage system network 110 is, for example, a network such as a wired LAN (Local Area Network), a wireless LAN, or a WAN (Wide Area Network), and may be a network utilizing, for example, Ethernet (registered trademark) or FibreChannel.

[0023] The storage system 100, the primary host computer 104, and the primary management terminal 105 are connected via a primary - side network 106. In this embodiment, the range connected to the primary - side network 106 is referred to as the primary - side system, and each component is treated as a primary - side component. The primary - side network 106 is, for example, a LAN or the like.

[0024] The storage system 100 is an example of a first storage system and has one or more storage nodes 111. The storage node 111 is an example of a computer and provides a plurality of volumes. Each storage node 111 is provided for redundancy, and in the storage system 100, it operates as one storage node as a whole.

[0025] The primary host computer 104 executes applications and the like, and performs various processes involving I / O requests on the storage system 100. The primary management terminal 105 performs a process of managing the storage system 100 by giving instructions such as volume creation to the storage system 100.

[0026] The storage system 101, the secondary host computer 107, and the secondary management terminal 108 are connected via the secondary network 109. In this embodiment, the range connected to the secondary network 109 is referred to as the secondary system, and each component is treated as a secondary component. The secondary network 109 is, for example, a LAN or the like.

[0027] The storage system 101 is an example of a second storage system and has a plurality of storage nodes 102. In this embodiment, the storage system 101 is a distributed storage system such as a scale-out storage system composed of a plurality of storage nodes 102 communicably connected by an inter-node network 103, and operates in cooperation as one cluster to provide a plurality of volumes.

[0028] The secondary host computer 107 executes applications and the like, and performs various processes involving I / O requests on the storage system 101. The secondary management terminal 108 performs a process of managing the storage system 101 by giving instructions such as volume creation to the storage system 101.

[0029] In the computer system 10, one or both of the primary system and the secondary system may be operated on the cloud.

[0030] In this embodiment, for disaster prevention and backup purposes, remote copy is performed on the storage system 101 on the secondary side for the data written by the primary host computer 104 to the storage system 100 on the primary side. In the event of a failure of the primary side system or the like, it is assumed that the secondary host computer 107 in the secondary side system performs a recovery process based on the data stored in the storage system 101 and resumes the process.

[0031] An example of the configuration of a pair of volumes (copy pair) of the copy source and the copy destination related to the remote copy in this embodiment will be described. Note that this copy pair is configured by a pair creation process (see FIG. 11) described later.

[0032] In the example of FIG. 1, the storage system 100 has a system ID set to No. 1, and the storage system 101 has a system ID set to No. 2. The storage system 101 is composed of three storage nodes 102 (storage node 1) having an ID of No. 1, a storage node 102 (storage node 2) having an ID of No. 2, and a storage node 102 (storage node 3) having an ID of No. 3.

[0033] In the computer system 10, a CTG (consistency group) indicating a range for guaranteeing the write order of data among a plurality of volumes in reflecting data to the secondary side system by remote copy is managed. Each copy pair is managed so as to belong to any one of the CTGs. Specifically, in the computer system 10, there are CTG112 (CTG1) and CTG113 (CTG2). The same CTG (112, 113) has the same ID assigned in the storage systems 100 and 101.

[0034] Regarding CTG1, in the positive storage system 100, there are PVOL (positive volume) 114 (PVOL1) and PVOL 115 (PVOL2) as volumes for which the positive host computer 104 performs I / O processing. Also, in the storage system 100, when a write I / O request is issued to these PVOLs 114 and 115, there is JVOL (journal volume) 116 (JVOL5) as a volume for storing differential data (journal data) indicating the write content used to perform remote copy to the secondary storage system 101 asynchronously with respect to the I / O processing to the PVOLs.

[0035] In the secondary storage system 101, there is JVOL 117 (JVOL5) which pairs with JVOL 116 (JVOL5) of the storage system 100 in storage node 1 and temporarily stores the differential data received by remote copy. Also, in storage node 1, there are SVOL (secondary volume) 118 (SVOL1) and SVOL 119 (SVOL2) as volumes that are the copy destinations of PVOL 114 and PVOL 115 and to which the differential data stored in JVOL 117 is reflected.

[0036] In CTG1, in the secondary storage system 101, the SVOL exists within one storage node and does not exist distributed across multiple storage nodes (does not span storage nodes). Therefore, the computer system 10 operates in a mode that guarantees the data write order by using only one JVOL to reflect data to multiple SVOLs.

[0037] Regarding CTG2, in the positive storage system 100, there are PVOL120 (PVOL3) and PVOL121 (PVOL4) as volumes for which the positive host computer 104 performs I / O processing. Also, in the storage system 100, when a write I / O request is issued to PVOL120, there is JVOL122 (JVOL6) as a volume for storing differential data used to perform remote copy to the secondary storage system 101 asynchronously with the I / O processing to PVOL120. Also, in the storage system 100, when a write I / O request is issued to PVOL121, there is JVOL123 (JVOL7) as a volume for storing differential data used to perform remote copy to the secondary storage system 101 asynchronously with the I / O processing to PVOL121.

[0038] In the secondary storage system 101, in storage node 2, there is JVOL124 (JVOL6) as a volume for temporarily storing differential data received and remotely copied in pair with JVOL122 (JVOL6) of the storage system 100. Also, in storage node 2, there is SVOL126 (SVOL3) as a volume that is the copy destination of PVOL120 (PVOL3) and to which the differential data stored in JVOL124 is reflected. Also, in storage node 3, there is JVOL125 (JVOL7) as a volume for temporarily storing differential data received and remotely copied in pair with JVOL123 (JVOL7) of the storage system 100. Also, in storage node 3, there is SVOL127 (SVOL4) as a volume that is the copy destination of PVOL121 (PVOL4) and to which the differential data stored in JVOL125 is reflected.

[0039] Since the JVOL on the secondary side does not perform data reflection processing on the SVOL across the storage node 102, in CTG2 that spans storage node 2 and storage node 3, a plurality of JVOLs are prepared in the primary storage system 100 and the secondary storage system 101, and the computer system 10 executes a process of guaranteeing the data writing order across the plurality of JVOLs.

[0040] In the computer system 10 according to the present embodiment, in an environment where an environment for guaranteeing the writing order limited within a storage node like CTG1 and an environment for guaranteeing the writing order across storage nodes like CTG2 coexist, the operation method by the user can be facilitated (for example, in the same manner as when such environments do not coexist).

[0041] Next, the configurations of the storage nodes 111 and 102 will be described.

[0042] FIG. 2 is a configuration diagram of a storage node according to the first embodiment. Since the storage nodes 111 and 102 have similar configurations, for convenience, the description will be made using FIG. 2.

[0043] The storage node 111 (102) is an example of a computer and includes a CPU 201 as an example of a processor, a memory 202 as an example of a storage unit, a storage device 203, and a communication (interface) I / F 204. These components 201 to 204 are communicably connected to each other via an internal bus or the like. The CPU 201, the memory 202, the storage device 203, and the communication I / F 204 may each be one or more.

[0044] The CPU 201 controls the overall operation of the storage node 111 (102). The CPU 201 executes various processes based on programs and management information stored in the memory 202. The CPU 201 may be a physical CPU of a physical computer or a virtual CPU obtained by virtually allocating a physical CPU of a physical computer using the virtualization function of the cloud.

[0045] The memory 202 is a volatile semiconductor memory such as SRAM (Static RAM (Random Access Memory)) or DRAM (Dynamic RAM), and stores various programs executed by the CPU 201 and management information referred to or updated by the CPU 201. The memory 202 may be a physical memory or a virtual memory obtained by virtually allocating a physical memory using the cloud virtualization function.

[0046] The storage device 203 is a storage device that stores user data used by the primary host computer 104, the secondary host computer 107, etc. Typically, the storage device 203 may be a non-volatile storage device. The storage device 203 may be, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage device 203 may be a physical storage device or a virtual storage device obtained by virtually allocating a physical storage device using the cloud virtualization function.

[0047] The communication I / F 204 is an interface for performing communication via a network (communication between storage nodes via the inter-node network 103, communication between the primary host computer 104 and the primary management terminal 105 via the primary network 106, communication between the secondary host computer 107 and the secondary management terminal 108 via the secondary network 109, communication between storage systems via the storage system network 110), and is, for example, a NIC (Network Interface Card) or an FC (Fibre Channel) card. The communication I / F 204 may be a physical communication I / F or a communication I / F obtained by virtually allocating a physical communication I / F using the cloud virtualization function.

[0048] Next, the configurations of the primary host computer 104, the primary management terminal 105, the secondary host computer 107, and the secondary management terminal 108 will be described.

[0049] FIG. 3 is a configuration diagram of the host computer and the management terminal according to the first embodiment. Since the main host computer 104, the main management terminal 105, the sub-host computer 107, and the sub-management terminal 108 have similar configurations, they will be described using FIG. 3 for convenience. Also, the same reference numerals are given to the configurations similar to those described in FIG. 2.

[0050] The main host computer 104 (main management terminal 105, sub-host computer 107, sub-management terminal 108) includes a CPU 201 as an example of a processor, a memory 202 as an example of a storage unit, and a communication I / F 204. These components 201, 202, 204 are connected to be communicable with each other via an internal bus or the like. The CPU 201, the memory 202, and the communication I / F 204 may each be one or more.

[0051] The CPU 201 performs processing to control the main host computer 104 (main management terminal 105, sub-host computer 107, sub-management terminal 108) based on the programs and management information stored in the memory 202. The memory 202 stores programs executed by the CPU 201 and management information referred to or updated by the CPU 201. The communication I / F 204 is an interface for communicating with the storage system via a network (for communicating with the storage system 100 via the main-side network 106 or with the storage system 101 via the sub-side network 109).

[0052] Next, it is a diagram for explaining the configuration of the memory 202 of the storage node 111 of the main-side storage system 100.

[0053] FIG. 4 is a diagram for explaining the configuration of the memory of the storage node of the main-side storage system according to the first embodiment.

[0054] The memory 202 of the storage node 111 stores a copy pair management program 401, a host I / O processing program 402, a write order management program 403, a journal data transfer program 404, a volume management program 405, a copy pair management table 406, a volume management table 407, and a write order management table 408.

[0055] When the copy pair management program 401 is executed by the CPU 201, it performs processes such as creation, state change, and deletion of a copy pair (a pair of a primary volume and a secondary volume related to remote copy) according to an instruction from the primary management terminal 105.

[0056] When the host I / O processing program 402 is executed by the CPU 201, it performs I / O processing (Read processing, Write processing) according to an I / O request (Read request, Write request) issued from the primary host computer 104.

[0057] When the write order management program 403 is executed by the CPU 201, it performs a process of storing differential data (journal data) with time information added to the Write data issued to the primary volume (PVOL) constituting the copy pair in the journal volume (JVOL).

[0058] Here, when there are multiple journal volumes of the copy destination belonging to the same CTG, the order of adding records to the table and the transfer order of differential data may change due to differences in the processing time and communication time, so the write order of data may not be guaranteed when reflecting the data in multiple SVOLs via multiple journal volumes. Therefore, when the write order management program 403 is executed by the CPU 201, in order to ensure the write order of data via multiple journal volumes and align the time slices of the data, it suppresses the Write processing corresponding to the Write request from the primary host computer 104, and during that time, it performs a process of updating the time information to be added to the differential data.

[0059] When the journal data transfer program 404 is executed by the CPU 201, it transfers the differential data stored in the journal volume to the secondary storage system 101 via the storage system network 110.

[0060] When the volume management program 405 is executed by the CPU 201, it performs volume management processes such as volume creation and deletion according to instructions from the primary management terminal 105.

[0061] The copy pair management table 406 holds information about copy pairs. The volume management table 407 holds information about volumes. The write order management table 408 holds information about write data corresponding to Write requests issued to the primary volume configured as a copy pair. Details of the copy pair management table 406, the volume management table 407, and the write order management table 408 will be described later.

[0062] Next, it is a diagram for explaining the configuration of the memory 202 of the storage node 102 of the secondary storage system 101.

[0063] FIG. 5 is a diagram for explaining the configuration of the memory of the storage node of the secondary storage system according to the first embodiment.

[0064] The memory 202 of the storage node 102 stores a journal data reception program 501, a host I / O processing program 502, a data reflection mediation program 503, a copy pair management program 504, a volume management program 505, a copy pair management table 506, a volume management table 507, and a write order management table 508. Here, each of the programs 401 to 405 in the memory 202 of the storage node 111 and each of the programs 501 to 505 in the memory 202 of the storage node 102 are examples of a remote copy control program.

[0065] The journal data reception program 501, when executed by the CPU 201, receives the data transferred by the journal data transfer program 404 of the positive storage system 100 and performs a process of storing it in the journal volume.

[0066] The host I / O processing program 502, when executed by the CPU 201, performs I / O processing (Read processing, Write processing) according to I / O requests (Read requests, Write requests) from the secondary host computer 107.

[0067] The data reflection arbitration program 503, when executed by the CPU 201, reflects the data stored in the journal volume to the secondary volume (SVOL) while guaranteeing the write order within the CTG based on the time information.

[0068] The copy pair management program 504, when executed by the CPU 201, performs processes such as creation, state change, deletion, etc. of copy pairs according to instructions from the secondary management terminal 108.

[0069] The volume management program 505, when executed by the CPU 201, performs volume management processes such as creation and deletion of volumes according to instructions from the secondary management terminal 108 and the storage system 100.

[0070] The copy pair management table 506 holds information about copy pairs. The volume management table 507 holds information about volumes. The write order management table 508 holds information about Write data corresponding to Write requests issued to the positive volume configured as a copy pair. Details of the copy pair management table 506, the volume management table 507, and the write order management table 508 will be described later.

[0071] In the storage system 101, these programs and data may be stored in all storage nodes 102, or a specific storage node 102 may be used as a representative node (representative node), and the representative node may hold a table related to all storage nodes 102.

[0072] Next, the copy pair management tables 406 and 506 will be described in detail.

[0073] FIG. 6 is a configuration diagram of the copy pair management table according to the first embodiment. The configurations of the copy pair management table 406 and the copy pair management table 506 are the same, and FIG. 6 shows a specific example of the copy pair management table 406.

[0074] The copy pair management table 406 (similarly for 506) stores records for each copy pair. The records in the copy pair management table 406 (506) include fields of CTG ID 601, primary system ID 602, primary journal ID 603, primary volume ID 604, status 605, secondary system ID 606, secondary journal ID 607, secondary volume ID 608, operation mode 609, and copy path ID 610.

[0075] The CTG ID 601 stores the ID (identification information) of the CTG that guarantees the write order. Since the CTGs containing the volumes that constitute the copy pairs of the remote copy are the same, the same ID is assigned to this CTG in the primary storage system 100 and the secondary storage system 101. In the example of FIG. 6, the CTG ID 601 stores 1, which is the ID of CTG112 (CTG1), or 2, which is the ID of CTG113 (CTG2).

[0076] The primary system ID 602 stores the ID of the storage system that is the primary side in the remote copy. In the example of FIG. 6, the primary system ID 602 stores 1, which is the ID of the storage system 100.

[0077] The positive-side journal ID 603 stores the ID of the journal (journal ID) that stores the data to be written before asynchronously transferring the data to be written when performing a Write process on the positive-side volume (PVOL). One CTG requires one or more positive-side journal volumes and one or more negative-side journal volumes. When multiple negative-side volumes (SVOL) of the negative-side storage system span multiple storage nodes (are distributed across multiple storage nodes), a number of journal volumes are required for the positive-side storage system and the negative-side storage system according to the number of storage nodes involved. In the example of FIG. 6, in the record corresponding to CTG1, the ID 5 which is the volume ID of JVO116 (JVOL5) is stored, and in the record corresponding to CTG2, there are a record in which the ID 6 which is the volume ID of JVOL122 (JVOL6) is stored and a record in which the ID 7 which is the volume ID of JVOL123 (JVOL7) is stored. In the example of FIG. 6, the volume ID of the journal volume is used as the journal ID, but the journal ID is not limited to this, and an ID different from the volume ID may be used.

[0078] The positive-side volume ID 604 stores the ID (identification information) of the volume that is the positive side in remote copy. In the first record of FIG. 6, the positive-side volume ID 604 stores PVOL1 which is the ID of volume 114.

[0079] In state 605, the state of the copy pair is stored. Examples of the state of the copy pair include "pair forming", which is the state during the execution of the process of forming a pair that transfers the existing stored data of the positive-side volume to the secondary-side volume that becomes the copy pair of the secondary-side storage system immediately after creating the copy pair; "paired", which is the state where the formation of the pair is completed and the transfer process of the updated Write data is started; "pair split", which is the state where the transfer process of data to the secondary-side storage system is interrupted and only the data that can guarantee the write order in the CTG is reflected; "failure", which is the state where the write order of the data is lost due to a failure, etc. Note that there may be other states as the state of the copy pair.

[0080] In the secondary-side system ID 606, the ID of the storage system that becomes the secondary side in remote copy is stored. In the example of FIG. 6, in the secondary-side system ID 606, 2, which is the ID of the storage system 101, is stored.

[0081] In the secondary-side journal ID 607, the ID (journal ID) of the journal that stores the differential data received from the primary-side storage system is stored. In the example of FIG. 6, in the record corresponding to CTG1, 5, which is the volume ID of JVO117 (JVOL5), is stored, and in the record corresponding to CTG2, there are a record in which 6, which is the volume ID of JVOL124 (JVOL6), is stored and a record in which 7, which is the volume ID of JVOL125 (JVOL7), is stored. In the example of FIG. 6, as the journal ID, the volume ID of the journal volume is used, but the journal ID is not limited to this, and an ID different from the volume ID may be used.

[0082] The operation mode 609 stores information on the operation mode that guarantees the update order of asynchronous remote copy. As the information on the operation mode, there are "within node", which indicates the within-node mode that is the operation mode when the journal volume on the secondary side of the CTG is within the same storage node, and "across nodes", which indicates the across-node mode that is the operation mode when the journal volume on the secondary side of the CTG spans multiple storage nodes. In the present embodiment, the write order management program 403 switches the method for guaranteeing the update order of data based on the information of the operation mode 609.

[0083] The copy path ID 610 stores the ID of a pre-constructed path (copy path) in the inter-storage-system network 110 used for data transfer between copy pairs between the storage system 100 and the storage system 101. Multiple copy paths may be constructed between the primary and secondary storage systems.

[0084] Next, the volume management table 407 of the primary storage system 100 will be described in detail.

[0085] FIG. 7 is a configuration diagram of the volume management table of the primary storage system according to the first embodiment.

[0086] The volume management table 407 manages records for each volume. The records in the volume management table 407 include fields of a volume ID 701, a volume attribute 702, a maximum capacity 703, a data storage destination storage device 704, and a write I / O limit flag 705.

[0087] The volume ID 701 stores the ID of the volume within the storage system 100.

[0088] The volume attribute 702 stores the attributes of the volume corresponding to the record. As volume attributes, there are "normal", which indicates that it is a normal volume that can be connected to a host computer for I / O, and "journal", which indicates that it is a journal volume (JVOL) that temporarily stores write data to a normal volume for asynchronous remote copy.

[0089] The maximum capacity 703 stores the maximum capacity of the volume corresponding to the record.

[0090] The data storage destination storage device 704 stores the name of the storage device that is the data storage destination assigned to the volume corresponding to the record. As a method of managing the name of the storage device, typically, the function of Thin Provisioning is applied, but other methods may also be used.

[0091] The write I / O restriction flag 705 stores the state (ON or OFF) of the write I / O restriction flag indicating whether Write to the volume corresponding to the record is acceptable. The state of this write I / O restriction flag is an example of restriction information. When the write I / O restriction flag is ON, Write to the volume corresponding to the record is suppressed. In this embodiment, by coordinating the suppression of writes to the volume according to the write I / O restriction flag and the update control of time information, the write order of data is guaranteed.

[0092] For example, the record in the first row of the volume management table 407 corresponds to the volume with volume ID 1, that is, PVOL114, indicating that the volume type is normal, the maximum capacity is 100 GB, the data storage destination storage device is storage device 1, and the write I / O restriction flag is OFF.

[0093] Next, the write order management table 408 of the positive storage system 100 will be described in detail.

[0094] FIG. 8 is a configuration diagram of a write order management table of the positive storage system according to the first embodiment.

[0095] The write order management table 408 is a table for managing the state from the Write process to the positive volume in the pair state in the storage system 100 to the reflection to the temporary storage destination of the data and the secondary volume of the storage system 101, and stores records for each Write process to the PVOL. The records of the write order management table 408 include fields of an ID801, a CTG ID802, time information (generation number) 803, a write target volume ID804, a write address 805, a journal volume ID806, an address on the journal volume 807, and a reflection status 808.

[0096] The ID801 stores the write order of the data written to the volume in the pair state of the copy pair in the storage system 100.

[0097] The CTG ID802 stores the ID of the CTG targeted by the write process corresponding to the record.

[0098] The time information (generation number) 803 stores time information (for example, time zone) regarding the time when the write process corresponding to the record was performed. In the present embodiment, the time information is managed so as to be incremented every certain time, and the time zone in which the write process corresponding to the record was performed can be grasped by the time information. Note that in the present embodiment, the time information is a real number, but it is not limited thereto, and may be time information provided by an NTP server or timer information in the storage system.

[0099] The write target volume ID804 stores the ID of the volume that is in the pair state and is the target of the Write request by the positive host computer 104.

[0100] The write address 805 stores the address (LBA: Logical Block Address) that is the target of the write process corresponding to the record.

[0101] The journal volume ID 806 stores the volume ID of the journal volume in which the data that is the target of the write process corresponding to the record is stored.

[0102] The journal volume upper address 807 stores the address (LBA) of the journal volume in which the data that is the target of the write process corresponding to the record is stored.

[0103] The reflection status 808 stores the reflection status to the secondary volume for the data that is the target of the write process corresponding to the entry. As the reflection status, there are "not transferred" indicating a situation where the target data has not been transferred to the secondary storage system, and "transferred" indicating a situation where the target data has been transferred to the secondary storage system but has not been reflected to the secondary volume. Other situations may exist in the reflection status.

[0104] For example, the record in the first row of the write order management table 408 is the oldest write process record with a write order of 1. The write process targets the CTG with an ID of 1, the time information when the write process was performed is 1, the ID of the volume to be written is 1 (that is, the volume to be written is PVOL114 (PVOL1)), the write target address is 0 to 255, the volume ID of the journal volume in which the data to be written is stored is 5 (that is, JVOL116 (JVOL5)), the address of the journal volume to be written is 0 to 255, and it indicates that the data to be written has been transferred to the secondary storage system.

[0105] Next, the volume management table 507 of the secondary storage system 101 will be described in detail.

[0106] FIG. 9 is a configuration diagram of a volume management table of the secondary storage system according to the first embodiment.

[0107] The volume management table 507 stores entries for each volume of the secondary storage system. The records of the volume management table 507 include fields of a volume ID 901, a volume attribute 902, a maximum capacity 903, a volume storage node 904, a data storage destination storage device 905, and a write I / O limit flag 906.

[0108] The volume ID 901 stores the ID of the volume within the storage system 101.

[0109] The volume attribute 902 stores the attribute of the volume corresponding to the record. As the attributes of the volume, there are "normal", which indicates that it is a normal volume that can be connected to a host computer for I / O, and "journal", which indicates that it is a journal volume (JVOL) that temporarily stores write data to a normal volume for asynchronous remote copy.

[0110] The maximum capacity 903 stores the maximum capacity of the volume corresponding to the record.

[0111] The volume storage node 904 stores the ID of the storage node where the volume corresponding to the record is stored.

[0112] The data storage destination storage device 905 stores the name of the storage device that is the data storage destination assigned to the volume corresponding to the record. As a method of managing the name of the storage device, typically, the function of Thin Provisioning is applied, but other methods may also be used.

[0113] The write I / O limit flag 906 stores the state (ON or OFF) of the write I / O limit flag indicating whether a Write to the volume corresponding to the record is acceptable. When the write I / O limit flag is ON, the Write to the volume corresponding to the record is suppressed. In the present embodiment, in the secondary storage system 101, by setting the write I / O limit flag to ON in the record, the write process is suppressed to guarantee the data at a certain point on the primary side.

[0114] For example, the record in the first row of the volume management table 507 corresponds to the volume with volume ID 1, that is, SVOL118, the type of the volume is normal, the maximum capacity is 100 GB, the storage node storing the volume is storage node 1, the data storage destination storage device is storage device 1, and it indicates that the write I / O limit flag is ON.

[0115] Next, the write order management table 508 of the secondary storage system 101 will be described in detail.

[0116] FIG. 10 is a configuration diagram of the write order management table of the secondary storage system according to the first embodiment.

[0117] The write order management table 508 is a table for managing the state until the data is reflected from the journal volume in the storage system 101 to the paired secondary volume, and stores a record for each write process to the SVOL to which the data to be written in the secondary storage system 101 is transmitted. The record of the write order management table 508 includes fields of ID1001, CTG ID1002, time information (generation number) 1003, write target volume ID1004, write address 1005, journal volume ID1006, journal volume upper address 1007, and reflection status 1008.

[0118] In ID1001, the writing order of data transferred to the secondary volume in a paired state of the copy pair within the storage system 101 is stored.

[0119] In CTG ID1002, the ID of the CTG targeted by the writing process corresponding to the record is stored.

[0120] In the time information (generation number) 1003, time information (for example, time zone) regarding the time when the writing process corresponding to the record was performed is stored. In this embodiment, the time information is managed so as to be incremented every certain period of time, and by the time information, it is possible to grasp in which time zone the writing process corresponding to the record was performed. Note that in this embodiment, the time information is a real number, but it is not limited to this, and it may be time information provided by an NTP server or timer information within the storage system.

[0121] In the write target volume ID1004, the ID of the volume that is the target of the writing corresponding to the record is stored.

[0122] In the write address 1005, the address (LBA: Logical Block Address) that is the target of the writing process corresponding to the record is stored.

[0123] In the journal volume ID1006, the volume ID of the journal volume in which the data that is the target of the writing process corresponding to the record is stored is stored.

[0124] In the journal volume upper address 1007, the address (LBA) of the journal volume in which the data that is the target of the writing process corresponding to the record is stored is stored.

[0125] In the reflection status 1008, the reflection status of the data to be written corresponding to the entry to the secondary volume is stored. As the reflection status, there are "not reflected" indicating the situation where the target data is not reflected to the secondary volume, and "reflected" indicating the situation where the target data has been reflected to the secondary volume. There may be other situations in the reflection status.

[0126] For example, the write order management table 508 shows a table of the situation where the data to be written corresponding to the record with ID801 being 8 in the write order management table 408 has not reached the storage system 101. The record in the first row of the write order management table 508 is the record of the oldest write process with the write order being 1. The write process targets the CTG with ID 1, the time information when the write process was performed is 1, the ID of the volume to be written is 1 (that is, the volume to be written is SVOL118 (SVOL1)), the write target address is 0 to 255, the volume ID of the journal volume where the data to be written is stored is 5 (that is, JVOL117 (JVOL5)), the address of the journal volume to be written is 0 to 255, and it indicates that the data to be written is not reflected in the SVOL.

[0127] Next, the pair creation process for creating a copy pair between the storage system 100 and the storage system 101 of the computer system 10 will be described.

[0128] FIG. 11 is a flowchart of the pair creation process according to the first embodiment.

[0129] Here, when executing the pair creation process, the positive management terminal 105 transmits, for example, according to the user's instruction, a pair creation instruction (the ID of the CTG which is the guaranteed range of the write order, the ID of the positive side volume (the first volume), the system ID of the secondary storage system 101) to the storage node 111 of the storage system 100.

[0130] The copy pair management program 401 of the positive storage node 111 (strictly speaking, the CPU 201 that executes the copy pair management program 401) receives a pair creation instruction from the positive management terminal 105 (S1100).

[0131] Next, the copy pair management program 401 refers to the volume management table 407 and checks the maximum capacity of the positive volume specified in the pair creation instruction (S1101).

[0132] Next, the copy pair management program 401 sends a pair volume creation instruction to the storage system 101 having the system ID specified in the pair creation instruction (S1102). Here, the pair volume creation instruction includes the specified CTG ID and the capacity of the secondary volume to be created as a pair.

[0133] The volume management program 505 of any storage node 102 of the storage system 101 receives the pair volume creation instruction (S1103), determines the storage node 102 where the volume (SVOL) serving as the copy pair is to be placed based on the CTG ID specified in the pair volume creation instruction, and executes a pair volume creation process (see FIG. 12) for creating a journal volume as necessary (S1104).

[0134] The volume management program 505 sends a response including the CTG ID, the volume ID of the SVOL serving as the copy pair, and the volume ID of the journal volume associated with the SVOL to the storage node 111 (S1105).

[0135] The copy pair management program 401 receives the response from the storage system 101 (S1106), and based on the CTG ID received as the response and the volume ID of the journal volume, executes a pair creation preparation process (see FIG. 13) for creating a new journal volume in the storage node 111 as necessary (S1107).

[0136] Next, the copy pair management program 401 determines whether the secondary volumes of all copy pairs belonging to the CTG with the specified CTG ID belong to the same storage node (S1108). Specifically, the copy pair management program 401 refers to the copy pair management table 406 and determines whether all of the secondary journal IDs of all copy pairs included in the CTG with the same CTG ID are the same as the volume ID of the journal volume included in the response received in step S1106.

[0137] As a result, if it is true (S1108: YES), the copy pair management program 401 advances the process to step S1109, while if it is false (S1108: NO), the copy pair management program 401 advances the process to step S1112.

[0138] In step S1109, the copy pair management program 401 adds copy pair information to the copy pair management table 406 as the in-node mode, and executes the pair addition process in the in-node mode (see FIG. 14) that instructs the storage system 101 to add pair information to the copy pair management table 506.

[0139] Next, the copy pair management program 401 copies all the data of the primary volume of the newly created copy pair to the secondary volume that is the pair of the storage system 101 (S1110). Here, in the process of copying all the data, for example, the copy pair management program 401 performs processing corresponding to the I / O request from the primary host computer 104, and typically allocates and manages a bitmap for each address to check whether there is uncopied data.

[0140] Next, the copy pair management program 401 updates the state 605 of the record corresponding to the created copy pair in the copy pair management table 406 to paired, and issues an instruction to update the state of the copy pair to the storage system 101 (S1111).

[0141] On the other hand, in step S1112, the copy pair management program 401 adds pair information to the copy pair management table 406 in the cross-node mode, and executes a cross-node mode pair addition process (see FIG. 15) for instructing the storage system 101 to add pair information to the copy pair management table 506.

[0142] Next, the copy pair management program 401 copies all the data of the positive-side volume of the newly created copy pair to the negative-side volume that is the pair of the storage system 101 (S1113). Here, in the process of copying all the data, for example, while the copy pair management program 401 performs processing corresponding to the I / O request from the positive host computer 104, typically, in order to copy all the data, a bitmap for whether there is uncopied data at each address is allocated and managed.

[0143] Next, the copy pair management program 401 updates the state 605 of the record in the copy pair management table 406 corresponding to the created copy pair to the pair, and issues an instruction to update the state of the copy pair to the storage system 101 (S1114). Accordingly, the storage system 101 updates the state 605 of the record in the copy pair management table 506 corresponding to the created copy pair to the pair.

[0144] Next, the pair volume creation process (S1104) in the storage system 101 will be described.

[0145] FIG. 12 is a flowchart of the pair volume creation process according to the first embodiment.

[0146] The volume management program 505 of the negative-side storage system 101 refers to the copy pair management table 506 (S1200), and checks whether the ID of the CTG specified in the pair volume creation instruction is registered (S1201).

[0147] As a result, if the CTG ID is registered in the copy pair management table 506 (when true: S1201: YES), the volume management program 505 proceeds with the process to step S1202, and if not registered (when false: S1201: NO), the process proceeds to step S1207.

[0148] In step S1202, the volume management program 505 refers to the volume management table 507 to check for available IDs and which storage nodes have a large number of volumes placed.

[0149] Next, the volume management program 505 determines which storage node to create a copy pair volume (pair volume) based on the information checked in step S1202 (S1203). As a method for determining the storage node where the volume will be created, it may be determined based on the number of volumes it belongs to (volume count) or the remaining capacity of the storage device. Typically, as the destination for creating the volume, it may be determined to be a storage node with a small number of volumes it belongs to or a storage node with a large remaining capacity of the storage device.

[0150] Next, the volume management program 505 creates a pair volume (second volume) for the determined storage node and adds a record corresponding to the pair volume to the volume management table 507 (S1204).

[0151] Next, the volume management program 505 determines whether there is a journal volume belonging to the CTG ID specified in the pair volume creation instruction within the storage node where the pair volume was created (S1205).

[0152] As a result, if there is a journal volume belonging to the CTG ID specified in the pair volume creation instruction in the storage node where the pair volume was created (when true: S1205: YES), the volume management program 505 ends the process. On the other hand, if there is no journal volume belonging to the CTG ID specified in the pair volume creation instruction in the storage node where the pair volume was created (when false: S1205: NO), the volume management program 505 creates a journal volume (second journal volume) in that storage node, adds a record corresponding to the journal volume to the volume management table 507 (S1206), and ends the process.

[0153] In step S1207, the volume management program 505 refers to the volume management table 507 to check for available IDs and which storage nodes have a large number of volumes arranged.

[0154] Next, based on the information confirmed in step S1207, the volume management program 505 determines which storage node to create the volume (pair volume) to be the copy pair (S1208). As a method for determining the storage node where the volume is to be created, it may be determined based on the number of volumes it belongs to or the remaining capacity of the storage device. Typically, as the destination for creating the volume, it may be determined to be a storage node with fewer volumes or a storage node with a large remaining capacity of the storage device.

[0155] Next, the volume management program 505 creates a pair volume for the determined storage node and adds a record corresponding to the pair volume to the volume management table 507 (S1209).

[0156] Next, the volume management program 505 creates a journal volume for the determined storage node, adds a record corresponding to the journal volume to the volume management table 507 (S1210), and ends the process.

[0157] Next, the pair creation preparation process (S1107) in the storage system 100 will be described.

[0158] FIG. 13 is a flowchart of the pair creation preparation process according to the first embodiment.

[0159] The copy pair management program 401 of the storage system 100 refers to the copy pair management table 406 and checks the existing CTG ID (S1300). Next, the copy pair management program 401 refers to the volume management table 407 and checks the existing journal volume (S1301).

[0160] Next, the copy pair management program 401 determines whether the CTG ID specified in the pair creation instruction exists based on the check result of step S1300 (S1302).

[0161] As a result, if the CTG ID specified in the pair creation instruction exists (true: S1302: YES), the copy pair management program 401 proceeds to step S1303, while if the CTG ID specified in the pair creation instruction does not exist (false: S1302: NO), the copy pair management program 401 proceeds to step S1305.

[0162] In step S1303, the copy pair management program 401 determines whether the ID of the journal volume of the storage system 101 received in S1106 exists in the copy pair management table 406.

[0163] As a result, if the ID of the journal volume of the storage system 101 exists in the copy pair management table 406 (true: S1303: YES), the copy pair management program 401 ends the process.

[0164] On the other hand, if the ID of the journal volume of the storage system 101 does not exist in the copy pair management table 406 (false: S1303: NO), the copy pair management program 401 creates a corresponding journal volume (the first journal volume) in the storage system 100, adds a record of the created journal volume to the volume management table 407 (S1304), and ends the process.

[0165] In step S1305, the copy pair management program 401 creates a corresponding journal volume in the storage system 100 and adds a record of the created journal volume to the volume management table 407. Then, the copy pair management program 401 ends the process.

[0166] Next, the pair addition process (S1109) in the same-node mode in the computer system 10 will be described.

[0167] FIG. 14 is a flowchart of the pair addition process in the same-node mode according to the first embodiment.

[0168] The copy pair management program 401 adds a record corresponding to the copy pair to be created to the copy pair management table 406 with the operation mode 609 being within the node (S1400). Next, the copy pair management program 401 sends a pair addition instruction to the secondary storage system 101 (S1401).

[0169] In response to this, the copy pair management program 504 of the storage system 101 receives the pair addition instruction (S1402), adds a record of the copy pair to be created to the copy pair management table 506 with the operation mode 609 being within the node (S1403), and sends a completion response of the pair addition to the storage system 100 (S1404).

[0170] Next, the copy pair management program 401 of the storage system 100 receives the completion response (S1405) and ends the pair addition process in the in-node mode.

[0171] Next, the pair addition process in the cross-node mode in the computer system 10 (S1112) will be described.

[0172] FIG. 15 is a flowchart of the pair addition process in the cross-node mode according to the first embodiment.

[0173] The copy pair management program 401 of the storage system 101 adds a record of the copy pair to be created to the copy pair management table 406 with the operation mode 609 set to cross-node (S1500).

[0174] Next, the copy pair management program 401 checks the copy pair management table 406 and determines whether the operation mode 609 of all the copy pair records with the same CTG ID is in-node (S1501). As a result, if the operation mode 609 of the records of other copy pairs with the same CTG ID is in-node (true in S1501: YES), the copy pair management program 401 proceeds to step S1502, while if the operation mode 609 of the records of other copy pairs is cross-node (false in S1501: NO), the process proceeds to step S1503.

[0175] In step S1502, the copy pair management program 401 checks the copy pair management table 406, switches the operation mode 609 of all the copy pair records with the same CTG ID to cross-node, and proceeds to step S1503.

[0176] In step S1503, the copy pair management program 401 sends a pair addition instruction to the storage system 101.

[0177] The copy pair management program 504 of the storage system 101 receives a pair addition instruction (S1504), and adds a record of the copy pair to be added to the copy pair management table 506 with the operation mode 609 across nodes (S1505).

[0178] Next, the copy pair management program 504 checks the copy pair management table 506 and determines whether the operation mode 609 of the records of other copy pairs with the same CTG ID is within the node (S1506). As a result, if the operation mode 609 of the records of other copy pairs with the same CTG ID is within the node (true: S1506: YES), the copy pair management program 504 proceeds to step S1507, while if the operation mode 609 of the records of other copy pairs is across nodes (false: S1506: NO), the process proceeds to step S1508.

[0179] In step S1507, the copy pair management program 504 checks the copy pair management table 506, switches the operation mode 609 of all records of copy pairs with the same CTG ID to across nodes, and proceeds to step S1508.

[0180] In step S1508, the copy pair management program 504 sends a completion response for pair addition to the storage system 100.

[0181] Next, the copy pair management program 401 of the storage system 100 receives the completion response (S1509) and ends the pair addition process in the across-node mode.

[0182] Next, the host I / O processing in the computer system 10 will be described.

[0183] FIG. 16 is a flowchart of the host I / O processing according to the first embodiment.

[0184] The host I / O processing program 402 of the storage system 100 receives an I / O request issued by the primary host computer 104 (S1600). Here, the I / O request includes a LUN (Logical Unit Number), the type of I / O (whether it is a Write process or a Read process), the start address, the amount of data, etc. Further, in the case of a Write process, the actual data to be written is included. Note that the correspondence between the LUN and the volume ID is managed in the storage system 100 by an existing technique.

[0185] Next, the host I / O processing program 402 determines whether the I / O request is a Write process (S1601). As a result, if the I / O request is a Write process (true: S1601: YES), the host I / O processing program 402 advances the process to step S1603, while if the I / O process is a Read process (false: S1601: NO), the process advances to step S1602.

[0186] In step S1602, the host I / O processing program 402 performs a Read process. Typically, if past data of the address targeted by the I / O request exists in the cache, the host I / O processing program 402 reads out the cache data, if the past data does not exist, reads out data from the storage device of the storage destination, and transmits the read data to the primary host computer 104.

[0187] In step S1603, the host I / O processing program 402 refers to the copy pair management table 406, determines whether the volume ID corresponding to the LUN of the I / O request is recorded as the positive-side volume ID and whether its state is a pair. As a result, if the volume ID corresponding to the LUN is recorded as the positive-side volume ID and its state is a pair (true: S1603: YES), the host I / O processing program 402 advances the process to step S1606. If the volume ID corresponding to the LUN is recorded as the positive-side volume ID and its state is not a pair (false: S1603: NO), the host I / O processing program 402 advances the process to step S1604.

[0188] In step S1604, the host I / O processing program 402 performs a Write process. Typically, if past data of the address targeted by the I / O request exists in the cache, the host I / O processing program 402 updates the cache data. If no past data exists, the host I / O processing program 402 writes new data to the cache. Next, the host I / O processing program 402 sends a completion response of the Write process to the positive host computer 104 (S1605).

[0189] In step S1606, the host I / O processing program 402 refers to the volume management table 407 and determines whether the write I / O limit flag 705 of the volume targeted by the Write process is ON. As a result, if the write I / O limit flag 705 of the target volume is ON (true: S1606: YES), the host I / O processing program 402 advances the process to step S1607. If the write I / O limit flag 705 of the target volume is not ON (false: S1606: NO), the process advances to step S1608.

[0190] In step S1607, the host I / O processing program 402 waits for a certain period of time to update the journal time information to align the data time slices, and then advances the process to step S1606.

[0191] In step S1608, the host I / O processing program 402 performs a Write process.

[0192] Next, the host I / O processing program 402 issues a journal creation instruction to the write order management program 403 by specifying the parameters specified during the Write process for the data for which the Write process was performed in step S1608 (S1609). Here, the parameters include the volume ID and the start address.

[0193] The write order management program 403 receives the journal creation instruction (S1610). Next, the write order management program 403 refers to the copy pair management table 406, confirms the ID of the positive journal volume corresponding to the volume to be written, and stores the data written in the free area of the journal volume with the confirmed ID (S1611).

[0194] Next, the write order management program 403 adds a record including the address of the data written in step S1611 to the write order management table 408 (S1612). Next, the write order management program 403 sends a completion of addition to the journal to the host I / O processing program 402 (S1613).

[0195] The host I / O processing program 402 receives a completion response (S1614). Next, the host I / O processing program 402 sends a completion response of the Write process to the main host computer 104 (S1615).

[0196] According to the above-described host I / O processing, it is possible to appropriately suppress the execution of the Write process in order to wait for the process of updating the time information of the journal for the volume in which the write I / O restriction flag 705 is ON.

[0197] Next, the write order management process in the computer system 10 will be described.

[0198] FIG. 17 is a flowchart of the write order management process according to the first embodiment.

[0199] The write order management process is started, for example, when a copy pair is created in the computer system 10. Here, the host I / O process and the write order management process are examples of the write order guarantee process.

[0200] The write order management program 403 refers to the copy pair management table 406 and acquires information on all existing copy pairs (S1700).

[0201] Next, the write order management program 403 determines whether there is a positive-side volume in which the state 605 is a pair and the operation mode 609 is across nodes based on the acquired copy pair information (S1701). As a result, if such a positive-side volume exists (true: S1701: YES), the write order management program 403 advances the process to S1702, while if such a positive-side volume does not exist (false: S1701: NO), the process advances to step S1705.

[0202] In step S1702, the write order management program 403 turns on the write I / O restriction flag 705 in the volume management table 407 for all positive-side volumes in which the state 605 is a pair and the operation mode 609 is across nodes. As a result, the Write process by the host computer to these volumes is suppressed.

[0203] Next, the write order management program 403 increases the set value of the time information given at the time of journal addition by 1 (updates by 1 generation) (S1703). As a result, this set value will be given to the Write process hereafter.

[0204] Next, for all positive-side volumes where the state 605 is a pair and the operation mode 609 is across nodes, the write order management program 403 turns off the write I / O limit flag 705 in the volume management table 407 (S1704) and advances the process to step S1705. As a result, the Write process by the host computer for these volumes will resume.

[0205] In step S1705, the write order management program 403 waits for a certain period until the timing of the next time information update, and after the wait, advances the process to step S1700. In the above-described write order management process, steps S1702 to S1704 are executed for all positive-side volumes where the state 605 is a pair and the operation mode 609 is across nodes. However, the present invention is not limited to this. For example, the processing of steps S1702 to S1704 may be performed with the timing shifted for each CTG. By doing so, the influence on the computer system 10 due to the suppression of the Write process can be suppressed.

[0206] Next, the journal data transfer process in the computer system 10 will be described.

[0207] FIG. 18 is a flowchart of the journal data transfer process according to the first embodiment.

[0208] The journal data transfer program 404 of the storage system 100 refers to the write order management table 408 and acquires information about all write processes (S1800). Next, the journal data transfer program 404 determines whether there is untransferred journal data based on the reflection status 808 (S1801).

[0209] As a result, if there is untransferred journal data (when true: S1801: YES), the journal data transfer program 404 proceeds with the process to step S1802. On the other hand, if there is no untransferred journal data (when false: S1801: NO), the process ends.

[0210] In step S1802, the journal data transfer program 404 transfers the untransferred journal data to the storage system 101 which is the secondary storage system. Specifically, the journal data transfer program 404 refers to the journal volume ID of record ID1006 in the write order management table 408 corresponding to the untransferred journal data and the address of the journal volume upper address 1007, reads data from the corresponding address of the journal volume corresponding to this ID, refers to the copy pair management table 406, selects the copy path of copy path ID610 of the corresponding record, specifies the volume ID of the secondary volume ID608 of the record as a parameter, and transfers the journal data.

[0211] When the journal data reception program 501 of the storage system 101 receives the transferred journal data (S1803), it refers to the copy pair management table 506, identifies the secondary journal ID corresponding to the specified secondary volume ID, and stores the journal data in the identified journal volume (S1804).

[0212] Next, the journal data reception program 501 registers, as records, information on the received journal data, information such as the address of the journal volume where the journal data was written, etc. in the write order management table 508 (S1805). Then, the journal data reception program 501 sends a completion response to the storage system 100 (S1806).

[0213] The journal data transfer program 404 of the storage system 100 receives a completion response (S1807). Next, the journal data transfer program 404 updates the reflection status 808 to transferred for the record corresponding to the journal data for which the transfer has been completed in the write order management table 408 (S1808), and advances the process to step S1800.

[0214] In the journal data transfer process of FIG. 18, an example is shown in which the positive storage system transfers journal data to the secondary storage system starting from the positive side. However, the present invention is not limited to this, and the secondary storage system may periodically request the transfer of journal data to the primary storage system, and in response, the primary storage system may transmit the journal data.

[0215] Next, the data reflection arbitration process in the storage system 101 will be described.

[0216] FIG. 19 is a flowchart of the data reflection arbitration process according to the first embodiment.

[0217] The data reflection arbitration process is executed periodically, for example. The data reflection arbitration program 503 of the storage system 101 refers to the write order management table 508 and acquires information on records corresponding to all journal data (S1900). Next, the data reflection arbitration program 503 determines whether there is a record in which the reflection status 1008 is not reflected among the records, thereby determining whether there is unreflected journal data in the volume (S1901). As a result, if there is unreflected journal data (true: S1901: YES), the data reflection arbitration program 503 advances the process to step S1902, and if there is no unreflected journal data (false: S1901: NO), the data reflection arbitration program 503 ends the process.

[0218] In step S1902, the data reflection arbitration program 503 refers to the copy pair management table 506 based on the volume ID of the write target corresponding to each journal data, and determines whether the operation mode 609 of the record of the copy pair corresponding to that volume ID is cross-node.

[0219] As a result, when the operation mode 609 is cross-node (true: S1902: YES), the data reflection arbitration program 503 advances the process to step S1903, while when the operation mode 609 is intra-node (false: S1902: NO), the process advances to step S1907.

[0220] In step S1903, the data reflection arbitration program 503 refers to the write order management table 508 and collects the latest time information 1003 of the journal data corresponding to each secondary journal ID belonging to the same CTG ID (S1903).

[0221] Here, all the journal data up to the previous generation of the latest arrived generation of each secondary journal belonging to the same CTG (that is, the generation with the generation number obtained by subtracting 1 from the latest generation number of the time information) has arrived at the secondary journal volume. Therefore, the data reflection arbitration program 503 reflects the unreflected journal data up to the time information of the previous generation to the address of the write address 1005 of the volume (SVOL) corresponding to the volume ID of the volume ID 1004 of the write target that is the secondary volume (S1904).

[0222] The data reflection arbitration program 503 updates the reflection status 1008 of the record corresponding to the journal data reflected to the secondary volume to reflected (S1905). At this time, the data reflection arbitration program 503 may notify the storage system 100 that this journal data has been reflected, and delete the record corresponding to this journal data from the write order management tables 408 and 508.

[0223] After the reflection of journal data is completed, the data reflection mediation program 503 waits for a certain period of time (S1906), and then advances the process to step S1901.

[0224] In step S1907, the data reflection mediation program 503 collects the journal data of each secondary journal ID belonging to the same CTG ID.

[0225] Here, since the data in the journal volume is in the order of the journal ID that matches the writing order, mediation between multiple journals is not necessary. Therefore, the data reflection mediation program 503 reflects the unreflected journal data in order from the youngest ID number to the address corresponding to the writing address 1005 of the volume corresponding to the volume ID of the writing target volume ID1004 that becomes the secondary volume (S1908).

[0226] Next, the data reflection mediation program 503 updates the reflection status 1008 of the record corresponding to the journal data that has been reflected in the secondary volume to reflected (S1909), and advances the process to step S1901. At this time, the data reflection mediation program 503 may notify the storage system 100 that the journal data has been reflected, and delete the record corresponding to this journal data from the write order management tables 408 and 508.

[0227] In this embodiment, the data reflection mediation process is performed in the secondary storage system. However, the primary storage system 100 refers to the time information of the records in the write order management table 408 whose transfer status 808 has been transferred, and periodically notifies the secondary storage system that up to one generation before the latest generation number that has arrived can be reflected. The secondary storage system that receives this may reflect the journal data in the secondary volume.

[0228] Also, in this embodiment, journal data is created during I / O processing and continuously transferred to the secondary storage system. However, for example, a Snapshot may be created in a state where WriteI / O is suppressed at regular intervals such as several minutes, several hours, several days, etc., and the difference from the previous Snapshot may be transferred to the secondary storage system to reflect (copy) the data to the secondary side.

[0229] [Second Embodiment] Next, a computer system according to the second embodiment will be described. Here, in this embodiment, basically, the parts different from the computer system according to the first embodiment will be described. Note that the same reference numerals will be used to describe the functional parts similar to those of the computer system according to the first embodiment.

[0230] In the second embodiment, the functional sharing of each storage node 102 in the storage system 101, the way of holding tables, and the information on pair creation instructions from the user are different.

[0231] First, the functional sharing of each storage node 102 and the way of holding tables will be described. Each storage node 102 operates independently in the same manner as the storage system 100, but has the same ID as the storage system.

[0232] In the storage node 102 of the storage system 101, there are a representative node and other nodes (general nodes). In the representative node, all the same programs as in the first embodiment are operating, and there are tables having records for all the storage nodes of the storage system 101 in each table. On the other hand, in the general node, it only has records related to the volumes and copy pairs belonging to itself and operates under the instruction from the representative node, so only some programs are operating.

[0233] Next, the information specified by the user as the pairing instruction will be described. When creating a pair, it is assumed that the paths between the primary side, the secondary side, and the storage system have been constructed in advance, and that the journal volume and the volumes of the primary side and the secondary side have already been created by the user.

[0234] In this configuration, the CTG ID, the journal volume ID, the primary volume ID, and the secondary volume ID are specified by the user as the pairing instruction. In the communication device system of this embodiment, it cannot be grasped from the primary storage system that the secondary storage system is a storage system composed of a plurality of nodes, and direct operations cannot be executed on appropriate storage nodes. For this reason, in this embodiment, the operation instructions are transferred between the storage nodes of the secondary storage system so as to enable operations on the storage nodes having the volume to be operated on.

[0235] FIG. 20 is a flowchart of the pairing process according to the second embodiment. The same processing steps as those in the pairing process according to the first embodiment shown in FIG. 11 are denoted by the same reference numerals.

[0236] Here, when executing the pairing process, the primary management terminal 105, for example, in accordance with a user instruction, transmits a pairing instruction (the CTG ID, the primary journal volume ID, the secondary journal volume ID, the primary volume ID, and the secondary volume ID, which are the guaranteed range of the writing order) to the storage node 111 of the storage system 100.

[0237] The copy pair management program 401 of the primary storage node 111 (strictly speaking, the CPU 201 that executes the copy pair management program 401) receives the pairing instruction from the primary management terminal 105 (S2000).

[0238] Next, the copy pair management program 401 refers to the copy pair management table 406 and determines whether the positive journal ID belonging to the CTG ID specified in the pair creation instruction is single and the same as the positive journal ID specified in the pair creation instruction (S2001). As a result, if the positive journal ID belonging to the specified CTG ID is single and the same as the positive journal ID specified in the pair creation instruction (when it is true: S2001: YES), the copy pair management program 401 proceeds to step S2002, while if the positive journal ID belonging to the specified CTG ID is single and not the same as the positive journal ID specified in the pair creation instruction (when it is false: S2001: NO), the process proceeds to step S2005.

[0239] In step S2002, the copy pair management program 401 executes the pair addition process in the same node mode (see FIG. 21), and executes the data copy of all addresses (S1100) and the update of the copy pair management table 406 (S1111).

[0240] In step S2005, the copy pair management program 401 executes the pair addition process in the cross-node mode (see FIG. 22), and executes the data copy of all addresses (S1113) and the update of the copy pair management table 406 (S1114).

[0241] Next, the pair addition process (S2002) in the same node mode in the computer system 10 will be described.

[0242] FIG. 21 is a flowchart of the pair addition process in the same node mode according to the second embodiment. In FIG. 21, the same steps as those in the pair addition process in the same node mode according to the first embodiment shown in FIG. 14 are denoted by the same reference numerals. Also, in the second embodiment, the processes of steps S1402, S1403, and S1404 are executed by the representative storage node (representative node) among the plurality of storage nodes 102.

[0243] The copy pair management program 504 of the representative node refers to the volume management table 507 and checks which storage node 202 the journal volume of the copy pair added in step S1403 belongs to (S2100).

[0244] Next, the copy pair management program 504 of the representative node instructs the copy pair management table 506, which only has records of volumes related to that storage node, to add a record related to the copy pair to the specified storage node 202 (S2101).

[0245] Next, the pair addition process (S2005) across nodes in the computer system 10 will be described.

[0246] FIG. 22 is a flowchart of the pair addition process in the cross-node mode according to the second embodiment. In FIG. 22, the same steps as those in the pair addition process in the cross-node mode according to the first embodiment shown in FIG. 15 are denoted by the same reference numerals. Also, in the second embodiment, the processes of steps S1504, S1505, S1506, S1507, and S1508 are executed by the representative storage node (representative node) among the storage nodes 102.

[0247] The copy pair management program 504 of the representative node refers to the volume management table 507 and identifies all storage nodes belonging to the CTG with the specified CTG ID and the storage node to which the journal volume with the specified secondary journal ID belongs (S2200).

[0248] Next, the copy pair management program 504 of the representative node instructs the copy pair management table 506 of the storage node where the specified journal volume exists to add a record of the copy pair (S2201).

[0249] Also, the copy pair management program 504 of the representative node issues an instruction to rewrite the operation mode 609 of the existing pairs in the copy pair management table 506 of these storage nodes across nodes from within the nodes for all the storage nodes to which the copy pairs within the same CTG belong (S2202).

[0250] According to the computer system according to the second embodiment, in the secondary storage system, since the representative storage node executes the process of notifying various information to other storage nodes, in the primary storage system, the process can be performed without being aware of the configuration of the secondary storage nodes.

[0251] [Third Embodiment] Next, the computer system according to the third embodiment will be described. Here, in this embodiment, basically, the parts different from the computer system according to the second embodiment will be described. Note that the same reference numerals will be used to describe the functional parts similar to those of the computer system according to the first embodiment.

[0252] In the computer system according to the third embodiment, the storage system 101 in the second embodiment operates as the primary storage system (first storage system), and the storage system 100 operates as the secondary storage system (second storage system). In this embodiment, the storage node 102 of the storage system 101 further stores the programs necessary as the primary side (such as the write order management program 403) in the program of the storage node 111, and the storage node 111 stores the programs necessary as the secondary side (such as the data reflection arbitration program 503) in the program of the storage node 102.

[0253] Next, the write order management process in the computer system 10 will be described.

[0254] FIG. 23 is a flowchart of the write order management process according to the third embodiment. In FIG. 23, steps similar to those of the write order management process according to the first embodiment shown in FIG. 17 are denoted by the same reference numerals. Further, in the third embodiment, the processes of steps S1700, S1701, and S1705 are executed by the representative storage node (representative node) among the storage nodes 102.

[0255] The write order management program 403 of the representative node simultaneously issues an instruction to turn on the write I / O restriction flag 906 for all storage nodes having the volumes that are copy pairs, in order to turn on the write I / O restriction flag 906 of the volume management table 507 for all copy pairs operating in the node-crossing mode (S2300). As a result, each storage node having the volume that is a copy pair receives the instruction, sets the write I / O restriction flag 906 of the entry corresponding to the corresponding volume in the volume management table 507 to ON, and transmits to the representative node that the setting has been completed.

[0256] The write order management program 403 of the representative node waits until completion is received from all the storage nodes 102 to which the instruction has been issued (S2301). When completion is received from all the storage nodes 102, the set value of the time information given at the time of journal addition is incremented by 1 (updated by one generation), and each storage node is simultaneously instructed to turn off the write I / O restriction flag 906 of the record corresponding to the volume that is a copy pair in the volume management table 507 (S2302). As a result, hereinafter, this set value will be given to the Write process in each storage node.

[0257] [Fourth Embodiment] Next, a computer system according to the fourth embodiment will be described. Here, in this embodiment, basically, parts different from the computer system according to the third embodiment will be described. Note that, for functional parts similar to those of the computer system according to the first embodiment, the same reference numerals are used for description.

[0258] The computer system according to the fourth embodiment is an example in which, in the computer system according to the third embodiment, the sub-management terminal 108 undertakes a part of the processing when performing a copy pair operation on the storage system 101.

[0259] In this embodiment, at the time of pair creation, the sub-management terminal 108 holds information equivalent to the copy pair management table 506 in the memory 202 and performs processing similar to the write order management processing shown in FIG. 23.

[0260] Here, although the storage system 101 is composed of a plurality of storage nodes 102, since the ID as a storage system is one, the sub-management terminal 108 is treated as one node. For this reason, the sub-management terminal 108 cannot perform processing that is conscious of a plurality of storage nodes as in step S2300, and an instruction is sent to one storage node.

[0261] Therefore, in this embodiment, when a certain storage node receives an instruction from the sub-management terminal 108, it transfers the received instruction to the representative node, and the representative node refers to the information in the copy pair management table 506 and the volume management table 507 to identify the position of the storage node to which the volume belongs and transfer the instruction to the storage note, and collects the completion responses obtained from each storage node into one. Next, the representative node instructs the storage node that has received the instruction from the sub-management terminal 108 to respond to the sub-management terminal 108 with the completion response.

[0262] By such processing, the same control is possible from the sub-management terminal 108 whether the positive storage system is the storage system 100 or the storage system 101 composed of a plurality of nodes.

[0263] [Fifth Embodiment] Next, a computer system according to the fifth embodiment will be described. Here, in this embodiment, basically, parts different from the computer system according to the first embodiment will be described. Note that for functional parts similar to those of the computer system according to the first embodiment, the same reference numerals will be used for description.

[0264] In this embodiment, a plurality of device management terminals having a function of managing a plurality of storage systems and connected to both a primary storage system and a secondary storage system are further provided. By collecting and utilizing operation information and the like from the plurality of storage systems, this plurality of device management terminals execute processes such as determination of volume allocation for efficiently using resources such as CPU and capacity while maintaining the RPO (Recovery Point Objective: target recovery time) specified by the user.

[0265] For example, when performing write order control across a plurality of storage nodes, since journal data at a time that has not arrived in common for all journals is not reflected in the secondary volume, the RPO increases. In this embodiment, volume allocation that can increase resource utilization efficiency while maintaining the RPO specified by the user is performed.

[0266] FIG. 24 is an overall configuration diagram of the computer system according to the fifth embodiment.

[0267] The computer system 10A further includes a plurality of device management terminals 2400 in the computer system 10. The plurality of device management terminals 2400 are connected to the storage system 100 via the network 2401 and connected to the storage system 101 via the network 2402. The plurality of device management terminals 2400 are an example of a management device, and may have the hardware configuration shown in FIG. 3, or may be a virtual machine on the cloud.

[0268] FIG. 25 is a configuration diagram of the memory of the plurality of device management terminals according to the fifth embodiment.

[0269] The memory 202 of the plurality of device management terminals 2400 includes an operation information collection program 2500, a service level maintenance program 2501, an optimal pair creation program 2502, a copy pair management table 2503, a volume management table 2504, a service level information table 2505, and an operation information table 2506.

[0270] When the operation information collection program 2500 is executed by the CPU 201, it periodically collects operation information such as the CPU operation rate in each system from the storage system 100 and the storage system 101, and stores it in the operation information table 2506.

[0271] When the service level maintenance program 2501 is executed by the CPU 201, it stores indicators such as the RPO specified by the user in the service level information table 2505, periodically checks whether these indicators are satisfied, and when they are not satisfied, notifies the user of an alert or the like.

[0272] When the optimal pair creation program 2502 is executed by the CPU 201, it receives an instruction to create a copy pair from the user, determines the volume arrangement of the remote copy pair so as to maximize the utilization of resources such as the CPU while maintaining the specified service level, and instructs the storage system 100 and the storage system 101 to create volumes and copy pairs.

[0273] The copy pair management table 2503 holds information about the copy pairs of the storage systems 100 and 101. The copy pair management table 2503 has records similar to the copy pair management table 406 and the copy pair management table 506.

[0274] The volume management table 2504 holds information about the volumes of the storage systems 100 and 101. The volume management table 2504 has records similar to the volume management table 407 and the volume management table 507.

[0275] The service level information table 2505 holds service level metrics such as the RPO specified by the user. The operation information table 2506 holds operation information such as the CPU operation rate of the storage systems 100 and 101. Details of the service level information table 2505 and the operation information table 2506 will be described later.

[0276] Next, the operation information table 2506 will be described in detail.

[0277] FIG. 26 is a configuration diagram of the operation information table according to the fifth embodiment.

[0278] The operation information table 2506 is a table for managing operation information and stores records for each storage node. The records of the operation information table 2506 include fields of a system ID 2600, a node ID 2601, a time 2602, a CPU operation rate 2603, and an available capacity 2604.

[0279] The system ID 2600 stores the ID of the storage system to which the storage node corresponding to the entry belongs. The node ID 2601 stores the node ID of the storage node corresponding to the entry. Note that although there are multiple storage nodes in the storage system 100, they are nodes for redundancy, and in the storage system 100, multiple nodes operate as one storage node, so the node ID 2601 does not have a node ID set.

[0280] The time 2602 stores the time when the operation information of the storage node corresponding to the entry was acquired. The CPU operation rate 2603 stores the CPU operation rate in the storage node corresponding to the entry. The available capacity 2604 stores the remaining available capacity in the storage node corresponding to the entry.

[0281] Note that the operation information is held in the memory 202 of each storage node and can be obtained from these storage nodes. The operation information is not limited to the CPU operation rate and may be, for example, IOPS (Input / Output Operations Per Second).

[0282] According to the second record of the operation information table 2506, it can be seen that the storage node 102 (storage node 1) with node ID 1 in the storage system 101 with system ID 2 has a CPU operation rate of 45% and an available capacity of 10 TB at 10::00.

[0283] Next, the service level information table 2505 will be described in detail.

[0284] FIG. 27 is a configuration diagram of the service level information table according to the fifth embodiment.

[0285] The service level information table 2505 is a table for managing the service level for volumes and stores records for each volume. The records of the service level information table 2505 include fields of a system ID 2700, a volume ID 2701, an RPO 2702, and a maximum capacity 2703.

[0286] The system ID 2700 stores the ID of the storage system to which the volume corresponding to the entry belongs. The volume ID 2701 stores the ID of the volume corresponding to the entry. The RPO 2702 stores the value of an index indicating how far the data rollback at the time of a failure for the volume corresponding to the entry can be tolerated. The maximum capacity 2703 stores the maximum available capacity of the volume corresponding to the entry.

[0287] According to the first record of the service level information table 2505, it can be seen that for the volume with volume ID 1 in the storage system 100 with system ID 1, the data rollback during a failure can be tolerated for 1 minute and the maximum capacity is 500 GB.

[0288] Next, the optimal pair creation process for creating a copy pair between the storage system 100 and the storage system 101 by the multiple device management terminals 2400 will be described.

[0289] FIG. 28 is a flowchart of the optimal pair creation process according to the fifth embodiment.

[0290] Here, when executing the optimal pair creation process, the main management terminal 105 transmits a pair creation instruction (the ID of the CTG which is the guaranteed range of the write order, the system ID of the primary storage system, the ID of the primary volume, the system ID of the secondary storage system 101) to the multiple device management terminals 2400 according to, for example, a user's instruction.

[0291] The optimal pair creation program 2502 of the multiple device management terminals 2400 (strictly speaking, the CPU 201 that executes the optimal pair creation program 2502) receives the pair creation instruction from the main management terminal 105 (S2800).

[0292] Next, the optimal pair creation program 2502 refers to the copy pair management table 2503, the volume management table 2504, and the service level information table 2505, and selects a node that satisfies the service level of the volume designated as the primary side and maximizes the resource utilization efficiency of the secondary storage system (S2801). Typically, the optimal pair creation program 2502 first selects whether the operation mode 609 is within a node or across nodes from the copy pair management table 2503 so as to satisfy the index value of the RPO 2702.

[0293] For example, when the index value of RPO2702 is short, since the placement within the node is essential, the storage node that manages the volume within the node is tentatively determined as a candidate for the destination where the secondary volume is to be placed. Here, if the available capacity of the available capacity 2604 of the candidate storage node is equal to or greater than the maximum capacity of the positive-side volume 2703, the candidate storage node is determined as the volume placement destination. If the available capacity is less than the maximum capacity, it means that placement is not possible, so an error occurs and the process ends.

[0294] On the other hand, when the index value of RPO2702 is long, since cross-node selection is possible, all nodes within the secondary storage system are considered candidates, and the narrowing down of the volume placement destination and the determination of placement feasibility are judged. For example, it is determined whether the available capacity of the available capacity 2604 is equal to or greater than the maximum capacity of the maximum capacity 2703 and whether it is a storage node capable of storing the volume. When there are multiple storage nodes capable of storing the volume, the CPU utilization rate 2603 of the CPU utilization rate is referred to, and the storage node with the most remaining capacity is determined as the placement destination.

[0295] Next, based on the determined placement, the optimal pair creation program 2502 executes a pair volume creation process (see Figure 29) to create a secondary volume and, if necessary, a journal volume (S2802).

[0296] Next, based on the result of the volume placement in step S2801, the optimal pair creation program 2502 determines whether pair formation is in the same-node mode (S2803). As a result, if it is pair addition in the same-node mode (true: S2803: YES), the optimal pair creation program 2502 proceeds to step S2804, while if it is in the cross-node mode (false: S2803: NO), the process proceeds to step S2805.

[0297] In step S2804, the optimal pair creation program 2502 instructs the positive - side storage system to perform pair addition processing for the in - same - node mode. As a result, the instructed positive - side storage system executes the pair addition processing for the in - same - node mode shown in FIG. 14.

[0298] In step S2805, the optimal pair creation program 2502 instructs the positive - side storage system to perform pair addition processing for the cross - node mode. As a result, the instructed positive - side storage system executes the pair addition processing for the cross - node mode shown in FIG. 15.

[0299] Next, the pair volume creation process (S2802) will be described in detail.

[0300] FIG. 29 is a flowchart of the pair volume creation process according to the fifth embodiment.

[0301] The optimal pair creation program 2502 refers to the copy pair management table 2503 and acquires the information of all records (S2900). Then, the optimal pair creation program 2502 determines whether the specified CTG ID matches the CTG ID of the existing copy pair (S2901).

[0302] As a result, if the specified CTG ID matches the existing CTG ID (true: S2901: YES), the optimal pair creation program 2502 advances the process to step S2902. On the other hand, if the specified CTG ID does not match the existing CTG ID (false: S2901: NO), the process advances to step S2905.

[0303] In step S2902, the optimal pair creation program 2502 instructs the secondary - side storage system to create a pair volume by specifying a storage node. As a result, the instructed secondary - side storage system creates a pair volume by performing the same processing as in step S1204.

[0304] Next, the optimal pair creation program 2502 determines whether there is a journal volume having a secondary journal ID related to the CTG ID specified in the storage node for which the volume is to be created (S2903). As a result, if there is a journal volume having a secondary journal ID related to the specified CTG ID (when it is true: S2903: YES), the optimal pair creation program 2502 ends the process. On the other hand, if there is no journal volume having a secondary journal ID related to the specified CTG ID (when it is false: S2903: NO), the process proceeds to S2904.

[0305] In step S2904, the optimal pair creation program 2502 transmits a journal volume creation instruction by designating a storage node to the secondary storage system. Thereby, the instructed secondary storage system performs the same process as in step S1206 to create a journal volume.

[0306] In step S2905, the optimal pair creation program 2502 instructs pair volume creation in the same manner as in step S2902. Thereby, the instructed secondary storage system performs the same process as in step S1204 to create a pair volume.

[0307] Next, the optimal pair creation program 2502 instructs journal volume creation in the same manner as in step S2904 (S2906). Thereby, the instructed secondary storage system performs the same process as in step S1206 to create a journal volume. After this, the optimal pair creation program 2502 ends the process.

[0308] Note that the present invention is not limited to the above-described embodiments, and can be appropriately modified and implemented without departing from the gist of the present invention.

[0309] For example, in the above embodiment, part or all of the processing performed by the processor may be performed by a hardware circuit. Also, the program in the above embodiment may be installed from a program source. The program source may be a program distribution server or a recording medium (for example, a portable recording medium).

Explanation of Signs

[0310] 10, 10A... computer system, 100, 101... storage system, 102, 111... storage node, 103... inter-node network, 104... primary host computer, 105... primary management terminal, 106... primary-side network, 107... secondary host computer, 108... secondary management terminal, 110... network between storage systems, 111... storage node, 112, 113... CTG

Claims

1. A computer system including a first storage system and a second storage system, The first storage system manages a plurality of first volumes belonging to a consistency group that guarantees the write order of data, The second storage system has a plurality of storage nodes, The second storage system distributes and creates a plurality of second volumes that are copy destinations of the plurality of first volumes on the plurality of storage nodes, and for each of the plurality of storage nodes that created the plurality of second volumes, a second journal volume that stores journal data indicating the write content in the first volume that is the copy source of the second volume is made to exist, The first storage system creates a plurality of first journal volumes corresponding to the respective second journal volumes and storing journal data for the plurality of first volumes, A write order guarantee process is executed to control the processing of the plurality of first volumes so that journal data indicating the write content for the plurality of first volumes can be stored in the plurality of first journal volumes in a state where the order of writing to the plurality of first journal volumes can be guaranteed. Computer system.

2. When the second storage system creates a second volume that is a copy destination of the first volume on any storage node and there is no second journal volume that stores journal data indicating the write content in the first volume that is the copy source of the second volume, the second journal volume is created. The first storage system determines whether the plurality of second volumes, which are the copy destinations of the plurality of first volumes, are generated on one storage node of the second storage system or on a plurality of storage nodes. When they are generated on a plurality of storage nodes, a first journal volume corresponding to each of the plurality of second journal volumes generated on the plurality of storage nodes is prepared. When the plurality of second volumes, which are the copy destinations of the plurality of first volumes, are generated on a plurality of storage nodes of the second storage system, the write order guarantee process is executed. The computer system according to claim 1.

3. The first storage system receives identification information of a consistency group storing volumes to be copied and identification information of the first volume to be copied. A process of creating the second volume is performed on the first volume of the received consistency group. The computer system according to claim 2.

4. The second storage system determines a storage node for creating the second volume from among the plurality of storage nodes based on the number of volumes or the remaining capacity of the storage nodes of the storage node. The computer system according to claim 2.

5. Further comprising a management device, The management device determines a storage node for arranging the second volume based on a target recovery time point for the first volume. The computer system according to claim 2.

6. The write order guarantee process includes a process of suppressing write processes for the plurality of first volumes when updating time information to be given to the journal data. The computer system according to claim 1.

7. The storage unit further stores restriction information indicating whether to suppress the writing process when updating the time information for a plurality of volumes. When updating the time information, the first storage system determines whether to suppress the writing process for the volume by referring to the restriction information. The computer system according to claim 6.

8. The first storage system has a plurality of storage nodes that operate independently. When updating the time information to be attached to the journal data, one storage node of the first storage system instructs all other storage nodes to suppress the writing process for the plurality of first volumes. The computer system according to claim 6.

9. A remote copy control method by a computer system including a first storage system and a second storage system, The second storage system has a plurality of storage nodes. The first storage system manages a plurality of first volumes belonging to a consistency group that guarantees the writing order of data. The second storage system distributes and creates a plurality of second volumes that are copy destinations of the plurality of first volumes on the plurality of storage nodes, and stores a second journal volume indicating the written content in the first volume that is the copy source of the second volume in each of the plurality of storage nodes that created the plurality of second volumes. The first storage system creates a plurality of first journal volumes corresponding to the respective second journal volumes and storing journal data for the plurality of first volumes. Execute a write order guarantee process that controls the processing for the plurality of first volumes so that journal data indicating the write content for the plurality of first volumes can be stored in the plurality of first journal volumes in a state where the write order for the plurality of first journal volumes can be guaranteed. Remote copy control method.

10. A remote copy control program for causing a plurality of computers in a computer system including a first storage system and a second storage system to execute, The second storage system has a plurality of storage nodes, The remote copy control program, Cause the computer of the first storage system to manage a plurality of first volumes belonging to a consistency group that guarantees the data write order, Cause the computer of the second storage system to create a plurality of second volumes that are copy destinations of the plurality of first volumes on the plurality of storage nodes in a distributed manner, and cause a second journal volume that stores journal data indicating the write content in the first volume that is the copy source of the second volume to exist in each of the plurality of storage nodes that created the plurality of second volumes, Cause the computer of the first storage system to create a plurality of first journal volumes that correspond to the respective second journal volumes and store journal data for the plurality of first volumes, Execute a write order guarantee process that controls the processing for the plurality of first volumes so that journal data indicating the write content for the plurality of first volumes can be stored in the plurality of first journal volumes in a state where the write order for the plurality of first journal volumes can be guaranteed. Remote copy control program.

Citation Information

Patent Citations

  • Write set boundary management for heterogeneous storage controllers that support asynchronous updates of secondary storage.

    JP2007537519A

  • Remote copy system and method of deciding target recovery point in remote copy system

    JP2009193208A

  • Storage system and control method therefor

    WO2015173859A1

  • Computer system and management method for computer system

    WO2016194096A1

  • Storage system and remote copy control method of storage system

    JP2007264946A