Storage system and management method for storage system

By predicting load conditions for new volumes and journal volumes in a secondary storage system, the management program ensures efficient node allocation, mitigating performance bottlenecks and maintaining system performance in geographically separated data centers with SDS.

JP2025146388APending Publication Date: 2025-10-03HITACHI VANTARA LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024047132
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The allocation of new volumes in a secondary storage system can lead to increased load on specific nodes, causing performance bottlenecks and reducing the overall storage system performance, especially when using Software Defined Storage (SDS) with geographically separated data centers.

Method used

A management program predicts the data processing load for new secondary volumes and journal volumes at the secondary site, considering conditions with and without redundancy processing, to accurately select the optimal node for placement based on load prediction results.

Benefits of technology

This approach allows for high-accuracy prediction of performance impacts from new volume allocation, preventing bottlenecks and maintaining optimal storage system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146388000001_ABST
    Figure 2025146388000001_ABST
Patent Text Reader

Abstract

To predict the performance impact due to new volume allocation, with high accuracy.SOLUTION: A storage system comprises a primary site that holds a primary volume, and a secondary site that holds a secondary volume and a journal volume. The primary site is configured with one or more nodes. The secondary site is configured with a plurality of nodes. When a management program running on any node or a predetermined management apparatus creates a new secondary volume and a new journal volume on the secondary site, the management program predicts the load of a processor for the new secondary volume under the operating conditions for performing redundancy processing within the secondary site, predicts the load for the new journal volume under the operating conditions for not performing the redundancy processing within the secondary site, and selects a node on which to place a new volume based on the results of load prediction.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a storage system and a method for managing a storage system. [Background technology]

[0002] Conventionally, a remote copy function has been used as a technology for replicating a storage system between multiple geographically separated data centers. Regarding remote copy, for example, a technology described in Japanese Patent Laid-Open Publication No. 2005-18736 (Patent Document 1) includes the following description: "A primary storage system and a secondary storage system are installed, for example, more than 100 miles apart," and "When a first storage system receives a write request from a first host associated with the first storage system, it stores the write data in a first data volume and generates a journal including controller data and journal data. A second storage system includes a journal volume and receives and stores the journal generated by the first storage system in the journal volume. A third storage system includes a second data volume and receives the journal from the second storage system according to information provided by the controller data and stores the journal data of the journal in the second storage system." [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-18736 Summary of the Invention [Problem to be solved by the invention]

[0004] Of the primary storage system (storage system at the primary site) and the secondary storage system (storage system at the secondary site), at least the secondary storage system may employ Software Defined Storage (SDS). SDS is based on one or more (typically multiple) storage nodes. These storage nodes may be located, for example, in an on-premise environment or a cloud environment. A storage node (hereinafter referred to as a node) is, for example, a general-purpose computer, and has a cache and a VOL (logical volume). The cache is typically located in volatile memory, and the VOL is typically based on persistent storage.

[0005] The primary-site storage system has a primary volume (PVOL) as a logical volume to be accessed by a host. The secondary-site storage system has a secondary volume (SVOL) that is a logical volume corresponding to the PVOL, and a journal volume (JVOL) that is a logical volume that stores a journal indicating the content and order of writes made to the PVOL.

[0006] When allocating a new volume (SVOL or JVOL) to a secondary-site storage system, it is important to decide on which node of the secondary-site storage system to allocate it to, because the load on the node where the new volume is allocated may increase, causing the node's performance to become a bottleneck, resulting in a decline in storage system performance. Therefore, an object of the present invention is to predict with high accuracy the performance impact caused by new volume allocation. [Means for solving the problem]

[0007] In order to achieve the above object, one representative storage system of the present invention comprises a primary site that holds a primary volume that stores host data, and a secondary site that holds a secondary volume that stores a copy of the host data stored in the primary volume by remote copy, and that holds a journal volume that temporarily stores data that is remotely copied from the primary volume to the secondary volume as journal data, wherein the primary site is configured with one or more nodes each having a processor, memory, and non-volatile storage medium, and the secondary site is configured with a plurality of nodes each having a processor, memory, and non-volatile storage medium, and wherein when a management program running on one of the nodes or a specified management device creates a new secondary volume and a new journal volume at the secondary site, the management program predicts the data processing load of the processor for the new secondary volume under operating conditions that perform redundancy processing within the secondary site, and predicts the load for the new journal volume under operating conditions that do not perform redundancy processing within the secondary site, and selects a node to which the new secondary volume and new journal volume are to be placed based on the results of the load prediction. Furthermore, one representative storage system management method of the present invention comprises a primary site that maintains a primary volume for storing host data, and a secondary site that maintains a secondary volume that stores a copy of the host data stored in the primary volume by remote copy, and that maintains a journal volume that temporarily stores data remotely copied from the primary volume to the secondary volume as journal data, wherein the primary site is configured with one or more nodes each having a processor, memory, and non-volatile storage medium, and the secondary site is configured with a plurality of nodes each having a processor, memory, and non-volatile storage medium, and is characterized in that, when a management program running on one of the nodes or a predetermined management device creates a new secondary volume and a new journal volume at the secondary site, the management program predicts a data processing load by the processor for the new secondary volume under operating conditions in which redundancy processing is performed within the secondary site, and predicts the load for the new journal volume under operating conditions in which redundancy processing is not performed within the secondary site, and selects a node to which the new secondary volume and new journal volume are to be placed based on the results of the load prediction. [Effects of the Invention]

[0008] According to the present invention, it is possible to predict with high accuracy the performance impact caused by the new allocation of volumes. Problems, configurations, and effects other than those described above will become clear from the following description of the embodiment. [Brief explanation of the drawings]

[0009] [Figure 1] Illustration of volume placement selection [Figure 2] A diagram showing an example of the physical configuration of a storage system. [Figure 3] A diagram showing an example of the software platform configuration for the site [Figure 4] Image showing an overview of the remote copy configuration in a storage system [Figure 5] Image showing an overview of I / O request processing in a storage system [Figure 6] FIG. 1 is a diagram showing an example of data and programs stored in a memory. [Figure 7] FIG. 10 is a diagram illustrating an example of a system configuration management table. [Figure 8] Figure 1 shows an example of a performance information management table [Figure 9] Figure 2 shows an example of a performance information management table [Figure 10] FIG. 10 is a diagram showing an example of a pair configuration management table. [Figure 11] Flowchart showing the procedure for new SVOL allocation [Figure 12] Flowchart explaining the details of performance information collection processing [Figure 13] Flowchart for explaining details of back-end data input / output flow rate calculation processing [Figure 14] Flowchart explaining the details of the network data input / output flow rate calculation process [Figure 15] Flowchart for explaining the details of the performance bottleneck determination process for a target node [Figure 16] Flowchart for explaining details of the SVOL placement candidate determination process (S113) [Figure 17] A diagram showing the flow of pair formation processing [Figure 18] Figure showing the flow of the initial journal creation process on the main site DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment will be described with reference to the drawings. In the following description, an "interface device" may refer to one or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., an NIC and an HBA (Host Bus Adapter)).

[0011] In the following description, "memory" refers to one or more memory devices, which are an example of one or more storage devices, and may typically be a primary storage device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.

[0012] In the following description, a "persistent storage device" may refer to one or more persistent storage devices, which are an example of one or more storage devices. A persistent storage device may typically be a non-volatile storage device (e.g., an auxiliary storage device), and specifically may be, for example, a hard disk drive (HDD), a solid state drive (SSD), or a non-volatile memory express (NVMe) drive.

[0013] Furthermore, in the following description, a "processor" may refer to one or more processor devices. The at least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be another type of processor device such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may also be a processor core. The at least one processor device may also be a processor device in a broader sense, such as a hardware circuit that performs part or all of the processing (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).

[0014] In the following description, information that provides an output for an input may be described using expressions such as "xxx table." However, this information may be data of any structure (for example, structured data or unstructured data), or may be a neural network that generates an output for an input, or a learning model such as a genetic algorithm or random forest. Therefore, the "xxx table" may be referred to as "xxx information." In the following description, the structure of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.

[0015] In the following description, processing may be described using a "program" as the subject. However, because a program is executed by a processor to perform a predetermined process using a storage device and / or an interface device, etc., as appropriate, the subject of the process may also be the processor (or a device such as a controller having the processor). A program may be installed in a device such as a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0016] In the following description, when describing elements of the same type without distinguishing between them, common parts of the reference symbols will be used, and when describing elements of the same type with distinction between them, reference symbols or identifiers of the elements will be used. For example, for PVOL, a reference symbol such as "PVOL102P1" may be used, or an identifier such as "PVOL1" may be used. [Example]

[0017] FIG. 1 is an explanatory diagram of the selection of the location of a volume. FIG. 1 shows a primary site 201a and a secondary site 201b. The primary site 201a is a site that holds a plurality of primary volumes that store host data. The primary site is also called the primary site. The secondary site 201b is a site that holds a secondary volume that stores a copy of the host data stored in the primary volume by remote copying, and also holds a journal volume that temporarily stores data that is remotely copied from the primary volume to the secondary volume as journal data. The secondary site is also called a secondary site.

[0018] The primary site 201a is configured with one or more nodes each having a processor, memory, and non-volatile storage medium, while the secondary site 201b is configured with multiple nodes each having a processor, memory, and non-volatile storage medium. 1, the secondary site 201b has four nodes. The primary site 201a has four PVOLs, and the four nodes of the secondary site 201b have SVOLs that correspond to any of the four PVOLs. Furthermore, the secondary site 201b has JVOLs that correspond to the SVOLs.

[0019] Each node in the secondary site 201b creates its own SVOL and JVOL using a drive, which is a nonvolatile storage medium of the node. Each node in the secondary site 201b can make its own SVOL redundant to another node. Each node in the secondary site 201b does not provide redundancy for its own JVOL.

[0020] The management program for managing the storage system can run on a node at either the primary site or the secondary site, for example. Alternatively, a specified management device may be provided, and the management program may be executed by the management device.

[0021] The management program collects data flow information for the PVOL in the primary site 201a, and load information and node configuration information for each node in the secondary site 201b. When creating a new volume in the secondary site 201b, the management program uses the data flow information of the PVOL, the load information of each node in the secondary site 201b, and the node configuration information to calculate the effect of the new volume on performance, and based on the calculation results, determines which node the new volume should be placed on. When calculating the impact on performance, the management program predicts the load for the new secondary volume under operating conditions in which redundancy processing is performed within the secondary site 201b, and predicts the load for the new journal volume under operating conditions in which redundancy processing is not performed within the secondary site 201b.Then, based on the results of the load prediction, it selects a node on which to place the new volume.

[0022] Here, we explain why redundancy processing is not required for the journal volume at the secondary site, and only the data on the secondary volume can be made redundant. The journal data at the primary site and the journal data at the secondary site are identical, and the contents of the two volumes are synchronized by copying the journal data from the primary volume to the secondary volume. Even without making the journal volume at the secondary site redundant, if the journal data at the primary site can be used, it is possible to restore the data before it was copied to the secondary volume. By retaining the data in the journal volume at the primary site on the secondary volume at the secondary site until the data in the primary volume is destaged, the secondary site notifies the primary site that the secondary volume destage is complete upon data destaging on the secondary volume, and the primary site discards the journal data upon that notification, this process eliminates journal data redundancy. Therefore, even if journal data redundancy is omitted, loss of journal data due to a secondary site failure can be avoided. Furthermore, if journal data at the secondary site is lost due to a secondary site failure, recovery is possible by resynchronizing the data in the primary site's primary volume and journal volume at the primary site with the data in the journal volume.

[0023] Fig. 2 is a diagram showing an example of the physical configuration of the storage system 101. As shown in Fig. 2, the storage system 101 may be provided with one or more sites 201.

[0024] The sites 201 are communicably connected via a network 202. The network 202 is, for example, a wide area network (WAN), but is not limited to a WAN. The site 201 is a data center or the like, and is configured to include one or more nodes 210.

[0025] The node 210 may have the configuration of a general server computer. The node 210 is configured to include, for example, one or more processor packages 213 including a processor 211 and a memory 212, one or more drives 214, and one or more ports 215. These components are connected via an internal bus 216.

[0026] The processor 211 is, for example, a CPU (Central Processing Unit) and performs various types of processing.

[0027] The memory 212 stores control information and data necessary to realize the functions of the node 210. The memory 212 also stores, for example, a program executed by the processor 211. The memory 212 may be a volatile dynamic random access memory (DRAM), a non-volatile storage class memory (SCM), or any other storage device.

[0028] The drive 214 stores various types of data, programs, etc. The drive 214 may be a SAS (Serial Attached SCSI) or SATA (Serial Advanced Technology Attachment) connected hard disk drive (HDD) or SSD (Solid State Drive), an NVMe (Non-Volatile Memory Express) connected SSD, an SCM, or a drive box equipped with multiple HDDs or SSDs, and is an example of a storage device.

[0029] The port 215 is connected to a network 220, and connects the node to other nodes 210 in the site 201 so as to be able to communicate with each other via the network 220. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.

[0030] The physical configuration of the storage system 101 is not limited to the above. For example, the networks 202 and 220 may be made redundant. For example, the network 220 may be separated into a management network and a storage network, the connection standard may be Ethernet (registered trademark), Infiniband, or wireless, and the connection topology is not limited to the configuration shown in Fig. 2. For example, the drive 214 may be configured independent of the node 210.

[0031] FIG. 3 is a diagram showing an example of the configuration of the software platform of the site 201.

[0032] For example, a software platform having the configuration illustrated in Fig. 3 can be employed at site 201, which serves as a secondary site. Site 201 has a network storage service 30 that provides multiple persistent stores 32 to multiple nodes 210 via network 220. The persistent stores 32 are storage areas based on one or more drives 214. The persistent stores 32 are non-volatile storage media that do not lose data even if a failure occurs in the network storage service 30.

[0033] The node 210 includes an instance store 65 , a hypervisor 64 , and a virtual machine 61 .

[0034] The instance store 65 provides temporary block-level storage for instances. This storage may reside on a drive 214 physically attached to the node 210. The instance store 65 is a volatile medium, where data is lost if the node 210 loses power.

[0035] The hypervisor 64 dynamically creates and deletes the virtual machines 61 .

[0036] The virtual machine 61 manages one or more virtual drives 63 and runs storage control software (SCS) 62 .

[0037] The SCS 62 controls I / O (Input / Output) for the virtual drive 63. The storage control software 62 is made redundant between the nodes 210. That is, if a failure occurs in a node 210, the SCS (Standby) 62 of another node 210 changes from Standby to Active, replacing the SCS (Active) 62 of the node 210.

[0038] The virtual drive 63 is a storage area to which the instance store 65 or the persistent store 32 is allocated. The virtual drive 63 may be treated as a VOL 102.

[0039] In this way, the site 201 uses an instance store 65 using DAS (Direct Attached Storage) and storage (network storage service 30) via a network 220 such as iSCSI. The hypervisor 64 may not be required, and for example, the DAS and network storage service 30 may be configured using bare metal.

[0040] Fig. 4 is an image diagram showing an overview of the remote copy configuration in the storage system 101. In detail, Fig. 4 shows an overview when a remote copy pair is established between multiple volumes in the storage system 101 between the primary site 201a and the secondary site 201b.

[0041] In the example of FIG. 4, two consistency groups 401a and 401b are constructed between the primary site 201a and the secondary site 201b. A consistency group is made up of multiple remote copy pair volumes, and multiple volumes at the primary site within a consistency group are copied to the secondary site while maintaining consistency. Specifically, update differential data up to the same time for multiple volumes 102 within the consistency group 401 is copied to the secondary site 201b. Consistency group control is managed by a journal volume (JNL). The journal volume stores update differential data for multiple PVOLs along with metadata such as the write time. When transferring data from a PVOL to the secondary site 201b, the storage cluster at the primary site 201a transfers to the secondary site 201b only the update differential data for the multiple PVOLs up to the same time that has been written to the journal volume. This allows data to be copied to the SVOL at the secondary site 201b while maintaining the consistency of update times between the multiple PVOLs.

[0042] 4, consistency group 401a copies data to volumes 102i and 102j in node 210d of secondary site 201b (i.e., SVOLs 102i and 102j in consistency group 401c) while maintaining the consistency of volumes 102a and 102b in node 210a of primary site 201a. Consistency group 401b copies data to volume 102l in node 210e of secondary site 201b and volume 102n in node 210f (i.e., SVOLs 102l and 102n in consistency group 401d) while maintaining the consistency of volume 102d in node 210b of primary site 201a and volume 102f in node 210c.

[0043] As can be seen from the specific configuration described above, the consistency group 401 may be made up of volumes within a specific node within a site, or may be made up of volumes in multiple nodes within a site.

[0044] Furthermore, although not shown, the storage system 101 may construct a remote copy pair by directly associating the PVOL and SVOL without a journal volume. In this case, update differential data for the PVOL is transferred directly to the node having the SVOL without going through the journal volume, and written to the SVOL. When the PVOL and SVOL are directly paired in this way, the update differential data from the PVOL to the SVOL is reflected without going through the journal volume, so the update differential data from the PVOL can be reflected to the SVOL at high speed. This is useful in cases where remote copying is performed in synchronization with I / O processing from the host. On the other hand, when the PVOL and SVOL are directly paired, there is no journal volume, so consistency control such as reflecting update differential data up to the same time in the SVOL is not possible.

[0045] Fig. 5 is a conceptual diagram showing an overview of I / O request processing in the storage system 101. More specifically, Fig. 5 shows an overview of I / O processing in the storage system 101 when a remote copy pair is established between PVOL 102a of the primary site 201a and SVOL 102d of the secondary site 201b.

[0046] First, an application 502 running on a host 501 issues a write request to a node 210a to write data 503a (data A) and data 503b (data B) to a PVOL 102a. Receiving the write request, the node 210a writes data A and data B to a PVOL 102a, and further writes data A and data B to a journal volume 102b (JNL1) as update differential data.

[0047] Next, node 210a transfers the update differential data written in journal volume 102b to journal volumes 102c and 102e of secondary site 201b. At this time, if multiple communication paths have been established between primary site 201a and secondary site 201b, any of the communication paths may be used for data transfer. Normally, node 210a of primary site 201a transfers the update differential data to node 210d, which has ownership of SVOL 102d paired with PVOL 102a. However, if a failure occurs on the communication path with ownership, the update differential data may be transferred to a node such as 210e that does not have ownership. For example, when node 210a of primary site 201a transfers update differential data to node 210e of secondary site 201b, which does not have ownership, node 210e of secondary site 201b transfers the received update differential data to node 210d, which does have ownership, and node 210d writes the transferred update differential data to journal volume 102c.

[0048] Next, node 210d periodically writes the update differential data written to journal volume 102c of secondary site 201b to SVOL 102d. Data 503a (data A) and 403b (data B) written to SVOL 102d are then written to drive 214a via storage pool 504a. If drive 214a is configured as a Direct Attached Storage (DAS) in which nodes (servers) are connected one-to-one, the data is written to a local drive (drive 214a) installed in node 210d. By writing all data to be written to SVOL 102d to drive 214a of node 210d, which has ownership of SVOL 102d, later when data is read from SVOL 102d, there is no need to read the data from another node. This allows storage system 101 to eliminate inter-node transfer processing and achieve high-speed read processing.

[0049] In addition, the storage pool 504 (for example, storage pool 504a) provides storage functions such as thin provisioning, compression, and deduplication, and executes processing of the storage functions required for the written data. Furthermore, in order to protect data from node failure when writing to drive 214a, the storage system 101 also writes redundant data of the data to be written (data A, B) to drive 214b of another node (for example, node 210e, which is the standby node). Regarding the writing of redundant data, if the data protection policy is replication, the storage system 101 writes a replica of the write data to drive 214b as redundant data. On the other hand, if the data protection policy is erasure coding, the storage system 101 calculates parity from the write data and writes the calculated parity to drive 214b as redundant data.

[0050] Also, in Figure 5, the internal configuration of nodes 210b, 210c, and 210f is omitted, but each of these nodes 210 may also have a PVOL and SVOL, like the above-mentioned nodes 210a, 210d, and 210e, and may process I / O from host 501.

[0051] Furthermore, the I / O processing flow shown in Figure 5 is an example of push-type I / O processing in which data is distributed from the primary site 201a to the secondary site 201b, but the storage system 101 can also execute pull-type I / O processing in which data is read from the secondary site 201b to the primary site 201a.

[0052] FIG. 6 is a diagram showing an example of data and programs stored in the memory 212. As shown in FIG. Information is read from the drive 214 to the memory 212. For example, the various tables included in the control information table 610 and the various programs included in the SCS 62 are deployed in the memory 212 while the processes in which they are used are being executed, but at other times they are stored in a non-volatile storage area such as the drive 214 in case of a power outage or the like.

[0053] The control information table 610 includes a system configuration management table 611 , a pair configuration management table 612 , a cache management table 613 , and a performance information management table 614 .

[0054] The storage program 620 includes a pair creation processing program 621, an initial journal creation processing program 622, a restore processing program 623, a journal read processing program 624, a journal purge processing program 625, a cache storage processing program 626, a destage processing program 627, an update journal creation processing program 628, a pair recovery processing program 629, a performance information collection program 630, a performance calculation program 631, and a volume placement destination determination program 632. The storage program 620 uses the operations of a performance information collection program 630, a performance calculation program 631, and a volume placement destination determination program 632 to realize the function of a management program that manages the storage system.

[0055] Each program constituting the storage program 620 is an example of a program used when the various functions of the node 210 are realized by software (storage control software 62), and specifically, these functions are realized by the processor 211 reading these programs stored in the drive 214 into the memory 212 and executing them. Note that in the storage system 101 according to this embodiment, the various functions of the node 210 may be realized by hardware such as a dedicated circuit having functions corresponding to the above-mentioned programs, or may be realized by a combination of software and hardware. Furthermore, some of the various functions of the node 210 may be realized by another computer capable of communicating with the node 210.

[0056] 7 is a diagram showing an example of the system configuration management table 611. The system configuration management table 611 stores information for managing the configuration of the nodes 210, drives 214, and ports 215 within the site 201.

[0057] The system configuration management table 611 is configured to include a node configuration management table 710, a drive configuration management table 720, a port configuration management table 730, and a volume configuration management table 740. The storage system 101 manages the node configuration management table 710 for each site 201 with respect to the multiple nodes 210 present in each site 201, and the node 210 manages the drive configuration management table 720, the port configuration management table 730, and the volume configuration management table 740 with respect to the multiple drives 214 within its own node 210.

[0058] The node configuration management table 710 is provided for each site 201, and stores information indicating the configuration of the nodes 210 provided in the site 201 (such as the relationship between the nodes 210 and the drives 214). More specifically, the node configuration management table 710 stores information in which a node ID 711, a status 712, a drive ID list 713, a port ID list 714, and data protection are associated with each other.

[0059] The node ID 711 is identification information that can identify the node 210. The status 712 is status information (for example, NORMAL, WARNING, FAILURE, etc.) that indicates the status of the node 210. The drive ID list 713 is identification information that can identify the drive 214 provided in the node 210. The port ID list 714 is identification information that can identify the port 215 provided in the node 210. The data protection 715 indicates a method for protecting redundant data, such as "mirror" or "distributed parity."

[0060] The drive configuration management table 720 is provided for each node 210, and stores information indicating the configuration of the drives 214 provided in the node 210. More specifically, the drive configuration management table 720 stores information in which a drive ID 721, a status 722, and a size 723 are associated with each other.

[0061] The drive ID 721 is identification information capable of identifying the drive 214. The status 722 is status information (for example, NORMAL, WARNING, FAILURE, etc.) indicating the status of the drive 214. The size 723 is information (for example, TB (terabytes) or GB (gigabytes)) indicating the capacity of the drive 214.

[0062] The port configuration management table 730 is provided for each node 210, and stores information indicating the configuration of the ports 215 provided in the node 210. More specifically, the port configuration management table 730 stores information in which a port ID 731, a state 732, and an address 733 are associated with each other.

[0063] The port ID 731 is identification information that can identify the port 215. The status 732 is status information (for example, NORMAL, WARNING, FAILURE, etc.) that indicates the status of the port 215. The address 733 is information that indicates an address (identification information) on the network that is assigned to the port 215. The address may be in the form of an IP (Internet Protocol), a WWN (World Wide Name), a MAC (Media Access Control) address, etc.

[0064] The volume configuration management table 740 is provided for each node 210, and stores information indicating the configuration related to the volume. More specifically, the volume configuration management table 740 stores information in which a volume ID 741, a status 742, a size 743, a data reduction 744, and a journal setting 745 are associated with each other.

[0065] The volume ID 741 is identification information that uniquely identifies a volume created in the node 210. The status 742 is status information (for example, Normal, Failure) that indicates the status of the volume. The size 743 is information that indicates the capacity of the volume (for example, TB (terabytes) or GB (gigabytes)). The data reduction indicates the setting related to data reduction of the volume (for example, disabled, compressed, compressed + deduplication). The journal setting 745 indicates whether or not the setting as a journal volume is applied.

[0066] 8 and 9 are diagrams showing an example of the performance information management table 614. The performance information management table 614 includes a maximum performance information management table 614a and a performance history information management table 614b.

[0067] The maximum performance information management table 614a indicates the maximum performance of the CPU, drive, and port. The management program queries each node to collect maximum performance information held by the node, manages this information as a maximum performance information management table 614a, and uses it to calculate the node's performance limit. The maximum performance information management table 614a is expanded in memory as data in table or list format, but the format is not limited to this.

[0068] The maximum performance information management table 614 a includes a CPU information management table 810 , a drive performance information management table 820 , and a port performance information management table 830 .

[0069] The CPU information management table 810 stores information in which a node ID 811, a CPU core 812, a CPU generation 813, a CPU frequency 814, and an initial copy processing time 815 are associated with each other. The node ID 811 is identification information that can identify the node 210. The CPU core 812 indicates the CPU core installed in the node 210. The CPU generation 813 indicates CPU generation information. The CPU frequency 814 indicates the CPU operating clock. The initial copy processing time 815 indicates the time required for the process of restoring data already stored in the primary volume to the secondary volume.

[0070] The drive performance information management table 820 stores information in which a drive ID 821, a type 822, a maximum throughput performance 823, and a latency 824 are associated with each other. The drive ID 821 is identification information capable of identifying the drive 214. The type 822 is information indicating the type of drive (for example, NVMeSSD, HDD, Persistent Store, Instance Store). The maximum throughput performance 823 indicates the maximum throughput of the drive. The latency 824 indicates the delay time of the drive.

[0071] The port performance information management table 830 stores information in which a port ID 831, a NIC bandwidth 832, and a port sharing with a persistent store 833 are associated with each other. The port ID 831 is identification information that can identify the port 215. The NIC bandwidth 832 is information that indicates the bandwidth of the port 215 (for example, 15 Gb / s, 0.1 Gb / s, etc.). The port sharing with persistent store 833 indicates whether or not a port is shared with the persistent store 32, and if a port is shared, further indicates identification information (drive ID) of the persistent store 32.

[0072] The performance history information management table 614b is created by the management software, which queries each node to collect and store performance history information held by the node. The management software uses the performance history information management table 614b to calculate the current load. There is room for discretion as to which data to use. For example, the software may look at the average load over the past hour, or the maximum load since the device started operating. The performance history information management table 614b is stored in memory as table- or list-format data, but the format is not limited.

[0073] The performance history information management table 614 b includes a CPU performance history information management table 910 , a drive performance history information management table 920 , a port performance history information management table 930 , and a volume performance history information management table 940 .

[0074] The CPU performance history information management table 910 stores information in which a data ID 911, a CPU core 912, a collection time 913, a CPU frequency 914, and a CPU operating rate 915 are associated with each other. The data ID 911 is identification information that can identify the collected data. The CPU core 912 indicates which CPU core the data pertains to. The collection time 913 indicates the time when the data was collected. The CPU frequency 914 indicates the frequency at which the CPU core was operating. The CPU operating rate 915 indicates the operating rate of the CPU core, for example, as a percentage.

[0075] The drive performance history information management table 920 stores information in which a data ID 921, a drive ID 922, a collection time 923, a throughput 924, and a latency 925 are associated with each other. Data ID 921 is identification information that can identify collected data. Drive ID 922 indicates which drive the data is for. Collection time 923 indicates the time when the data was collected. Throughput 924 indicates the throughput of the drive at the collection time. Latency 925 indicates the latency of the drive at the collection time.

[0076] The port performance history information management table 930 stores information in which a data ID 931, a port ID 932, a collection time 933, and a throughput 934 are associated with each other. The data ID 931 is identification information that can identify collected data. The port ID 932 indicates which port the data pertains to. The collection time 933 indicates the time when the data was collected. The throughput 934 indicates the throughput of the port at the collection time.

[0077] The volume performance history information management table 940 stores information in which a data ID 941, a volume ID 942, a collection time 943, and a throughput 944 are associated with each other. The data ID 941 is identification information that can identify collected data. The volume ID 942 indicates which volume the data pertains to. The collection time 943 indicates the time when the data was collected. The throughput 944 indicates the throughput of the volume at the collection time.

[0078] FIG. 10 is a diagram showing an example of the pair configuration management table 612. The pair configuration management table 612 includes a volume management table 1010 , a pair management table 1020 , and a journal management table 1030 .

[0079] The volume management table 1010 stores information indicating the configuration of the volume 102. More specifically, the volume management table 1010 stores information in which a volume ID 1011, an owner node ID 1012, a retreat node ID 1013, a size 1014, and an attribute 1015 are associated with each other.

[0080] The volume ID 1011 is identification information that can identify the volume 102. The owner node ID 1012 is information that indicates the node 210 that has ownership of the volume 102. The fallback node ID 1013 is information that indicates the node 210 that will take over processing in the event of a failure in the node 210 that has ownership of the SVOL. The size 1014 is information that indicates the capacity of the volume 102 (for example, TB (terabytes) or GB (gigabytes)). The attribute 1015 is information that indicates the attribute of the volume 102, and includes NML_VOL (normal volume), PAIR_VOL (pair volume), JNL_VOL (journal volume), etc.

[0081] The pair management table 1020 stores information indicating the configuration of a remote copy pair. More specifically, the pair management table 1020 stores information in which a pair ID 1021, a primary journal volume ID 1022, a primary volume ID 1023, a secondary journal volume 1024, a secondary volume ID 1025, and a status 1026 are associated with each other.

[0082] The pair ID 1021 is identification information that can identify the remote copy pair. The primary journal volume ID 1022 is the ID of the volume 102 that records journal information at the primary site 201a of the remote copy pair. The primary volume ID 1023 is the ID of the volume 102 that serves as the copy source at the primary site 201a of the remote copy pair. The secondary journal volume 1024 is the ID of the volume 102 that records journal information at the secondary site 201b of the remote copy pair. The secondary volume ID 1025 is the ID of the volume 102 that serves as the copy destination at the secondary site 201b of the remote copy pair. The status 1026 is status information (e.g., PAIR, COPY, SUSPEND, etc.) that indicates the status of the remote copy pair. "PAIR" is a status in which writing to the PVOL is periodically reflected in the SVOL. "COPY" is a status in which initial copying is in progress. "SUSPEND" is a status in which the pair is suspended (a status in which synchronization between the PVOL and SVOL is not performed).

[0083] The journal management table 1030 stores information related to journals. More specifically, the journal management table 1030 stores, for each journal, information such as a pair group ID 1031, a journal ID 1032, a P / S volume ID 1033, a P / S volume address 1034, a size 1035, and a cache segment ID 1036.

[0084] The pair group ID 1031 is the ID of the consistency group to which the journal belongs. The journal ID 1032 is the ID of the journal. The journal ID corresponds to SEQ# and is, for example, a consecutive number in the consistency group. In other words, the journal ID indicates the order of writing, and the data in the journal is stored in the SVOL in the consistency group in the order of the journal IDs.

[0085] The P / S volume ID 1033 includes the ID of the PVOL to which the data in the journal is written, and the ID of the SVOL to which the data in the journal is written. The P / S volume address 1034 includes the storage address of the data in the PVOL to which the data in the journal is written, and the storage address of the data in the SVOL to which the data in the journal is written.

[0086] The size 1035 indicates the size of the journal. For example, one journal contains one or more pieces of data. The cache segment ID 1036 is the ID of the cache segment into which the data in the journal is written.

[0087] 11 is a flowchart showing the processing steps for new allocation of an SVOL. When new allocation of an SVOL occurs, the storage program 620 executes the following steps S101 to S114 in order. In step S101, the storage program 620 receives from the user the designation of the secondary site where the SVOL is to be located, and then proceeds to step S102. In S102, the storage program 620 determines whether or not an existing PVOL exists in the primary site. If an existing PVOL does not exist in the primary site (S102; No), that is, if a new PVOL and a new SVOL are to be allocated, proceed to S103. If an existing PVOL exists in the primary site (S102; Yes), that is, if a new SVOL corresponding to an existing PVOL is to be allocated, proceed to step S105. Note that even if an existing PVOL exists, if a new PVOL and a new SVOL are to be allocated rather than an SVOL corresponding to the existing PVOL, proceed to step S103.

[0088] In step S103, the storage program 620 receives input from the user of the number and capacity of PVOLs to be newly created and the expected PVOL throughput performance, and then proceeds to step S104. In step S104, the storage program 620 creates a PVOL in the primary site, and then proceeds to step S107.

[0089] In step S105, the storage program 620 receives a designation of a PVOL to be remotely copied from the user, and acquires information on the number of PVOLs and their capacity, and then proceeds to step S106. In step S106, the storage program 620 acquires the throughput performance history information of the target PVOL, and then proceeds to step S107.

[0090] In step S107, the storage program 620 determines the number and capacity of SVOLs and JVOLs, and then proceeds to step S108. In step S108, the performance information collection program 630 selects a redundant node corresponding to the destination node of the secondary site and collects performance information, and then proceeds to step S109. In step S109, the performance calculation program 631 performs a calculation process for the input / output flow rate of the back-end data of the target node, and then the process proceeds to step S110. In step S110, the performance calculation program 631 performs a calculation process for the input / output flow rate of network data of the target node, and then the process proceeds to step S111.

[0091] In step S111, the performance calculation program 631 performs a performance bottleneck determination process for the target node, and then proceeds to step S112. In step S112, the performance calculation program 631 determines whether or not calculation has been completed for all nodes on the secondary site. If there are any nodes that have not yet been calculated (S112; No), the program returns to step S108. If calculation has been completed for all nodes (step S112; Yes), the program proceeds to step S113.

[0092] In step S113, the volume placement decision program 632 performs a process to determine SVOL placement candidate nodes, and then proceeds to step S114. Details will be described later, but in this process, the SVOL placement candidate nodes are notified to the user as placement nodes. In step S114, the storage program 620 receives confirmation and correction input from the user regarding the number, capacity, and placement destination of SVOLs and JBOLs, and creates volumes.

[0093] In this way, the storage program 620 collects the information necessary for creating and allocating an SVOL, predicts performance, and determines the location of the SVOL. The storage program 620 changes the parameters used to create the SVOL and JVOL depending on whether the PVOL to be remote copied is a new PVOL or an existing PVOL is reused. It also changes the PVOL flow rate information used in flow rate calculations. The PVOL throughput performance (data flow rate) can be input from the track record of other volumes, or it can be input by the user. Once the SVOL placement is determined, the JVOL placement location is also determined. This is because the SVOL and JVOL must be installed on the same node. Therefore, the load caused by the JVOL is also taken into consideration.

[0094] FIG. 12 is a flowchart illustrating the details of the performance information collection process (S108) shown in FIG. The performance information collection program 630 selects a collection destination node and requests the redundant node corresponding to the selected node to acquire maximum performance information from the maximum performance information management table 614a (step S201). Thereafter, a request is made to the redundant node corresponding to the selected node to acquire current performance history information (step S202).

[0095] Fig. 13 is a flowchart illustrating in detail the backend data input / output flow rate calculation process (S109) shown in Fig. 11. In the backend data input / output flow rate calculation process, the performance information collection program 630 sequentially executes the following steps S201 to S309.

[0096] In step S301, the performance information collection program 630 identifies the volume ID of the PVOL to be remote copied and acquires flow rate information, and then proceeds to step S302. In step S302, the performance information collection program 630 determines whether the attribute of the volume to be estimated is a JVOL. If the attribute of the volume to be estimated is a JVOL (S302; Yes), the program proceeds to step S303. If the attribute of the volume to be estimated is not a JVOL (S302; No), the program proceeds to step S306.

[0097] In step S303, the performance information collection program 630 determines whether the cache mode is on. If the cache mode is not on (S303; No), the program proceeds to step S304. If the cache mode is on (S303; Yes), the program proceeds to step S305. In step S304, the back-end flow rate is equal to the JVOL flow rate, and the JOVL flow rate is equal to the PVOL flow rate. Therefore, the performance information collection program 630 sets the back-end flow rate to the PVOL flow rate. Then, the process proceeds to step S309. S305: JVOL destaging does not occur. Therefore, the performance information collection program 630 sets the back-end flow rate to 0. Then, the process proceeds to step S309.

[0098] In step S306, the performance information collection program 630 determines whether the data protection setting is mirror. If the data protection setting is mirror (S306; Yes), the program proceeds to step S307. If the data protection setting is not mirror (S306; No), the program proceeds to step S308.

[0099] In step S307, since it is a mirror, a copy of the host data is transferred to the redundant node. The performance information collection program 630 calculates the back-end flow rate using the following formula, and then proceeds to step S309. Back-end flow rate = SVOL flow rate + log data writing + SVOL redundancy flow rate = 2*PVOL flow rate + log data writing

[0100] In step S308, since it is distributed parity, the parity transfer data generated from the host data is transferred to the redundant node. The performance information collection program 630 calculates the back-end flow rate using the following formula, and proceeds to step S309. Back-end traffic = SVOL traffic + log data write + mDnP parity transfer volume =(1+n / m)*PVOL flow rate + log data writing

[0101] In S309, the performance information collection program 630 determines whether the process has been completed for all VOLs. If there are any volumes that have not yet been completed (S309; ​​No), the process returns to step S301. If the process has been completed for all VOLs (S309; ​​Yes), the process in the figure ends.

[0102] Fig. 14 is a flowchart illustrating in detail the network data input / output flow rate calculation process (S110) shown in Fig. 11. In the network data input / output flow rate calculation process, the performance information collection program 630 sequentially executes the following steps S401 to S407.

[0103] In step S401, the performance information collection program 630 identifies the volume ID of the PVOL to be remote copied and acquires flow rate information, and then proceeds to step S402. In step S402, the performance information collection program 630 determines whether the attribute of the volume to be estimated is a JVOL. If the attribute of the volume to be estimated is a JVOL (S402; Yes), the program proceeds to step S403. If the attribute of the volume to be estimated is not a JVOL (S402; No), the program proceeds to step S404.

[0104] In step S403, the JVOL is not subjected to redundancy processing. Therefore, the performance information collection program 630 sets the network flow rate to 0. Then, the process proceeds to step S407.

[0105] In step S404, the performance information collection program 630 determines whether the data protection setting is mirror. If the data protection setting is mirror (S404; Yes), the program proceeds to step S405. If the data protection setting is not mirror (S404; No), the program proceeds to step S406.

[0106] In step S405, since it is a mirror, a copy of the host data is transferred to the redundant node. The performance information collection program 630 calculates the network flow rate using the following formula, and then proceeds to step S407. Network flow rate = SVOL flow rate =PVOL flow rate

[0107] In step S406, since it is distributed parity, the parity transfer data generated from the host data is transferred to the redundant node. The performance information collection program 630 calculates the network flow rate using the following formula, and then the process proceeds to step S407. Network flow rate = SVOL flow rate + parity transfer rate =(1+n / m)*PVOL flow rate

[0108] In S407, the performance information collection program 630 determines whether or not the process has been completed for all VOLs. If there are any volumes that have not yet been completed (S407; No), the program returns to step S401. If the process has been completed for all VOLs (S407; Yes), the processing in the figure ends.

[0109] Fig. 15 is a flowchart illustrating in detail the target node performance bottleneck determination process (S111) shown in Fig. 11. In the target node performance bottleneck determination process, the performance calculation program 631 sequentially executes the following steps S501 to S511.

[0110] In step S501, the performance calculation program 631 calculates the expected data flow rate by adding the estimated value of the port calculated in the previous process (S110) to the data flow rate currently occurring at the port of the node (redundant node corresponding to the destination node) for which performance information was collected in step S108. Then, the program proceeds to step S502. In step S502, the performance calculation program 631 determines whether the performance limit (the performance value recorded in the performance management information) is greater than the predicted flow rate in the previous step (S501). If the performance limit is equal to or less than the predicted flow rate (S502; No), the program proceeds to step S503. If the performance limit is greater than the predicted flow rate (S502; Yes), the program proceeds to step S504.

[0111] In step S503, the performance calculation program 631 sets a flag indicating that the network is a bottleneck, and then proceeds to step S505. In step S504, the performance calculation program 631 sets a flag indicating that the network is not a bottleneck, and then the process proceeds to step S505.

[0112] In step S505, the performance calculation program 631 calculates the expected data flow rate by adding the estimated data flow rate calculated in the previous step (S109) to the data flow rate currently occurring in the disk of the node (redundant node corresponding to the destination node) for which performance information was collected in step S108. Then, the program proceeds to step S506. In step S506, the performance calculation program 631 determines whether the performance limit (the performance value recorded in the performance management information) is greater than the predicted flow rate in the previous step (S505). If the performance limit is equal to or less than the predicted flow rate (S506; No), the program proceeds to step S507. If the performance limit is greater than the predicted flow rate (S506; Yes), the program proceeds to step S508.

[0113] In step S507, the performance calculation program 631 sets a flag indicating that the disk is a bottleneck, and then the process proceeds to step S509. In step S508, the performance calculation program 631 sets a flag indicating that the disk is not a bottleneck, and then the process proceeds to step S509.

[0114] In step S509, the performance calculation program 631 determines whether or not both the disk and the port are bottlenecks. If a bottleneck has occurred in either one (S509; No), the program proceeds to step S510. If neither is a bottleneck (S509; Yes), the program proceeds to step S511.

[0115] In step S510, the performance calculation program 631 sets a flag indicating that a bottleneck has occurred in the node due to the placement of the SVOL and JVOL, and then ends the performance bottleneck determination process for the target node. In step S511, the performance calculation program 631 sets a flag indicating that the placement of the SVOL and JVOL will not cause a bottleneck in the node, and then ends the performance bottleneck determination process for the target node.

[0116] Fig. 16 is a flowchart illustrating in detail the SVOL placement candidate determination process (S113) shown in Fig. 11. In the SVOL placement candidate determination process, the volume placement decision program 632 sequentially executes the following steps S601 to S607. In step S601, the volume placement destination determination program 632 acquires the bottleneck determination results for all nodes, and then proceeds to step S602. In step S602, the volume placement destination determination program 632 determines whether or not there is a node where a bottleneck will not occur. If there is no node where a bottleneck will not occur (S602; No), the program proceeds to step S603. If there is a node where a bottleneck will not occur (S602; Yes), the program proceeds to step S606.

[0117] In step S603, the volume placement destination determination program 632 determines whether or not there is a node where only the port is a bottleneck. If there is no node where only the port is a bottleneck (S603; No), the program proceeds to step S604. If there is a node where only the port is a bottleneck (S603; Yes), the program proceeds to step S605.

[0118] In step S604, the volume placement destination decision program 632 selects the node where the disk is bottlenecked as the placement destination node for the SVOL and JVOL, and then proceeds to step S607. In step S605, the volume placement destination decision program 632 selects the node whose port is the bottleneck as the placement destination node for the SVOL and JVOL, and then proceeds to step S607. In step S606, the volume placement destination decision program 632 selects a node where no bottleneck will occur as the placement destination node for the SVOL and JVOL, and then proceeds to step S607.

[0119] In step S607, the volume placement destination determination program 632 notifies the user of the result of the selection of the placement node, and ends the SVOL placement candidate determination process. In this way, the volume placement destination decision program 632 checks the bottleneck decision results for all nodes, and gives priority to selecting a node that will not become a bottleneck, and determines it as the placement destination. When comparing network bottlenecks and disk bottlenecks, the disk has a larger impact due to its narrower bandwidth, so if either becomes a bottleneck, the node where the network is the bottleneck will be selected as the priority.

[0120] FIG. 17 is a diagram showing the flow of the pair forming process. According to the pair creation process, a remote copy pair is created through communication between a primary base system and a secondary base system. In the explanation of Fig. 17, the pair creation process program 621P is a program in the primary node that has a PVOL candidate. The pair creation process program 621S is a program in the secondary node that has an SVOL candidate.

[0121] The pair forming processing program 621P sends a pre-check request to the pair forming processing program 621S (S701). The pair forming processing program 621S receives the request (S801) and performs predetermined pre-checks, such as checking whether a VOL to be paired exists and whether the information on the partner device is correct. The pair forming processing program 621S returns a response to the pre-check request (S802). The response indicates the results of the pre-check. The pair forming processing program 621P receives the response (S702).

[0122] If the response is a predetermined response, the pair creation processing program 621P sends a pair creation request to the pair creation processing program 621S, specifying PVOL102P as the PVOL candidate and SVOL102S as the SVOL candidate (S703). The pair creation processing program 621S receives the request (S803), creates a VOL pair (registers the information in the pair management table 1020), and sets the pair status 1026 to "COPY" (S804). The pair creation processing program 621S starts restore processing (S805) and returns a response to the pair creation request (S806). The pair creation processing program 621S waits for completion of the initial copy (S807). When the initial copy is complete (S808: Yes), specifically, when the restore processing program 623S has completed synchronization of the data in the PVOL and the data in the SVOL, the pair creation processing program 621S sets the status 1026 of the pair to "PAIR" (S809).

[0123] The pair creation processing program 621P receives the response sent in S806 (S704). If the response is a predetermined response, the pair creation processing program 621P creates a VOL pair (registers the information in the pair management table 1020) and sets the status 1026 of the pair to "COPY" (S705). If the pair has a resynchronization option (S706: Yes), the pair creation processing program 621P sets the resynchronization option (S707).

[0124] The pair creation processing program 621P starts the initial journal creation processing (S708). The pair creation processing program 621P waits for the completion of the initial copy (S709). When all initial journals have been purged (S710: Yes), specifically when the restore processing program 623S has completed synchronization of the PVOL data and the SVOL data, the pair creation processing program 621P sets the status 1026 of the pair to "PAIR" (S711).

[0125] FIG. 18 is a diagram showing the flow of the initial journal creation process at the main site. When the initial journal creation process is started, the process shown in Figure 18 is performed. In this process, data from the PVOL is generated as the initial journal. The journal is basically stored in the cache, but if the cache is full, it is destaged to the PJVOL and the destaged data is purged from the cache. There are two types of initial copies: full copy and differential copy. If the pair status 1026 is "SUSPEND", a journal is created by differential copy only for the update difference while the pair is suspended. A redundancy flag and a non-volatile flag are set for the cache segment, and writing is controlled according to these flags.

[0126] If the resynchronization option is not set (S901: No), the initial journal creation processing program 622P sets the data in all areas of the PVOL as the target for creating the initial journal (S902).If the resynchronization option is set (S901: Yes), the initial journal creation processing program 622P references a difference management table (not shown) that indicates the differences between the PVOL and SVOL, and sets the data in the difference area as the target for creating the initial journal (S903).

[0127] The initial journal creation processing program 622P reserves a cache segment for the initial journal from the free queue (S904), and obtains the VOL address corresponding to that segment (the PVOL address of the PVOL area where the data to be included in the initial journal is located) (S905). The initial journal creation processing program 622P reads data from that address (S906), creates metadata for the initial journal (S907), creates an initial journal including the data read in S906 and the metadata created in S907, and sets this initial journal as the storage target (S908). The metadata includes, for example, the journal ID, LBA (VOL address), transfer length, pair ID, PVOL_ID, and SVOL_ID.

[0128] The initial journal creation processing program 622P sets both the redundancy flag and the non-volatile flag in the redundancy-non-volatile flag for the segment secured in S904 to "1" (S909). The initial journal creation processing program 622P starts the cache storage processing (S910).

[0129] If the cache usage rate (for example, the ratio of the total capacity of dirty segments and clean segments to the total capacity of the cache 55P) exceeds a predetermined value (S911: Yes), the initial journal creation processing program 622P targets the initial journal in the cache segment for destaging (S912) and starts destaging processing (S913). The initial journal creation processing program 622P releases the cache segment that has the initial journal destaged to the PJVOL, that is, makes it a free segment (S914).

[0130] If there are any initial journals that have not been created (S915: No), the process returns to S904. If all initial journals have been created (S915: Yes), the initial journal creation process ends.

[0131] As described above, the storage system disclosed in the embodiments comprises a primary site (201a) that holds a primary volume that stores host data, and a secondary site (201b) that holds a secondary volume that stores a copy of the host data stored in the primary volume by remote copy, and that holds a journal volume that temporarily stores data that is remotely copied from the primary volume to the secondary volume as journal data, wherein the primary site is configured with one or more nodes 210 each having a processor, memory, and non-volatile storage medium, and the secondary site is configured with multiple nodes 210 each having a processor, memory, and non-volatile storage medium, and when a management program running on any of the nodes or a predetermined management device creates a new secondary volume and a new journal volume at the secondary site, it predicts the data processing load of the processor for the new secondary volume under operating conditions where redundancy processing is performed within the secondary site, and predicts the load for the new journal volume under operating conditions where redundancy processing is not performed within the secondary site, and selects a node to which the new secondary volume and new journal volume will be placed based on the results of the load prediction. This configuration and operation allows the storage system to predict with high accuracy the performance impact caused by the new allocation of volumes.

[0132] The new journal volume stores data related to remote copying to the new secondary volume, and the new secondary volume and the new journal volume are located in the same volume. This allows the storage system to predict with high accuracy the performance impact when a new secondary volume and a new journal volume are placed in the same volume.

[0133] In addition, the management program acquires load information and performance information for each node at the secondary site, and when it becomes the destination node, predicts the network flow rate, which is the data flow rate generated by communication with the outside of the node, and the disk flow rate, which is the data flow rate generated by writing to the non-volatile storage medium within the node, and predicts the data processing load of the processor of the destination node from the predicted network flow rate and disk flow rate, and for a redundant destination node that is the redundant destination of the destination node, predicts the data processing load of the processor of the redundant node after the new secondary volume and new journal volume are placed on the destination node by adding the network flow rate and disk flow rate of the destination node to the load information of the redundant node, and selects as the destination node and redundant node, respectively, the nodes whose loads on the destination node and the redundant node are within the maximum performance range indicated in the performance information. This makes it possible to predict the network traffic and the disk traffic.

[0134] Furthermore, when predicting the network flow rate, the management program does not include the network flow rate for the new journal volume, and predicts that the network flow rate for the new secondary volume will be generated according to the load on the corresponding primary volume. This allows for highly accurate prediction of network traffic load based on volume attributes.

[0135] Furthermore, when predicting the disk flow rate, the management program does not count the load of the new journal volume if the secondary volume is set to directly reflect the data received by the secondary site from the primary site, but predicts that a data processing load corresponding to the load of the corresponding primary volume will be incurred on the processor of the node to which the new journal volume is to be placed if the data received from the primary site is set to be temporarily stored in the journal volume, and predicts that a disk flow rate corresponding to the load of the primary volume will be generated for the new secondary volume, and predicts that a data processing load corresponding to the predicted disk flow rate will be incurred on the processor of the node to which the new secondary volume is to be placed. This allows for highly accurate prediction of disk traffic load based on volume attributes and settings.

[0136] Furthermore, the management program predicts whether a bottleneck will occur at any node due to the load exceeding performance if that node is selected as the placement destination for all nodes at the secondary site, and determines the node where no bottleneck will occur as the placement destination. This effectively avoids the occurrence of bottlenecks.

[0137] The management program also presents the nodes selected as candidates for the placement destination of the new volume to the user, and accepts the user's designation of the placement destination. This allows volumes to be allocated to nodes that meet the user's needs while avoiding bottlenecks.

[0138] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, not only can the configurations be deleted, but also replacements and additions of configurations are possible. For example, in the above embodiment, when a bottleneck occurs, a node where only the port becomes a bottleneck is preferentially selected, but a node where only the disk becomes a bottleneck may be preferentially selected, or the predicted results may be output and user selection may be accepted. [Explanation of symbols]

[0139] 101: storage system, 102: volume, 201: site, 202: network, 210: node, 214: drive, 215: port, 216: internal bus, 220: network, 610: control information table, 611: system configuration management table, 612: pair configuration management table, 613: cache management table, 614: performance information management table, 620: storage program, 630: performance information collection program, 631: performance calculation program, 632: volume placement destination determination program

Claims

1. a primary site that holds a primary volume that stores host data; a secondary site that holds a secondary volume that stores a copy of host data stored in the primary volume by remote copy, and that holds a journal volume that temporarily stores data that is remote copied from the primary volume to the secondary volume as journal data; Equipped with the primary site is configured to include one or more nodes each having a processor, a memory, and a non-volatile storage medium; the secondary site is configured to include a plurality of nodes each having a processor, a memory, and a non-volatile storage medium; A management program running on any node or a predetermined management device When creating a new secondary volume and a new journal volume at the secondary site, the load of data processing by the processor is predicted for the new secondary volume under operating conditions in which redundancy processing is performed within the secondary site, and the load is predicted for the new journal volume under operating conditions in which redundancy processing is not performed within the secondary site, and a node to which the new secondary volume and new journal volume are to be placed is selected based on the results of the load prediction. A storage system comprising:

2. 2. The storage system according to claim 1, The new journal volume stores the data to be remotely copied to the new secondary volume. The new secondary volume and the new journal volume are placed in the same volume. A storage system comprising:

3. 2. The storage system according to claim 1, The management program Acquire load information and performance information for each node at the secondary site; When the node becomes the destination node, a network flow rate, which is a flow rate of data generated by communication with the outside of the node, and a disk flow rate, which is a flow rate of data generated by writing to the non-volatile storage medium inside the node, are predicted, and a data processing load of the processor of the destination node is predicted from the predicted results of the network flow rate and the disk flow rate; predicting a data processing load by a processor of the redundant node after the new secondary volume and new journal volume are placed in the destination node by adding the network flow rate and the disk flow rate of the destination node to the load information of the redundant node, for the destination node that is a redundant destination of the destination node; A storage system characterized in that nodes are selected as the destination node and the redundant node such that the loads of the destination node and the redundant node are within the range of maximum performance indicated in the performance information.

4. 4. The storage system according to claim 3, A storage system characterized in that, when predicting the network flow, the management program does not include the network flow for the new journal volume, and predicts that the network flow for the new secondary volume will be generated in accordance with the load of the corresponding primary volume.

5. 4. The storage system according to claim 3, When predicting the disk flow rate, the management program Regarding the new journal volume, if the secondary volume is set to directly reflect the data received by the secondary site from the primary site, the load of the new journal volume is not counted, but if the data received from the primary site is set to be temporarily stored in the journal volume, it is predicted that a data processing load corresponding to the load of the corresponding primary volume will be generated on the processor of the node where the new journal volume is placed, A storage system characterized by predicting that for the new secondary volume, disk traffic will occur in accordance with the load of the primary volume, and predicting that a data processing load in accordance with the predicted disk traffic will occur on the processor of the node where the volume is placed.

6. 2. The storage system of claim 1, The storage system is characterized in that the management program presents to a user nodes selected as candidates for the placement destination of the new volume, and receives a designation of the placement destination from the user.

7. a primary site that holds a primary volume that stores host data; a secondary site that holds a secondary volume that stores a copy of host data stored in the primary volume by remote copying, and that holds a journal volume that temporarily stores data that is remotely copied from the primary volume to the secondary volume as journal data, wherein the primary site is configured with one or more nodes that have a processor, memory, and non-volatile storage media, and the secondary site is configured with a plurality of nodes that have a processor, memory, and non-volatile storage media, comprising: A management program running on any node or a predetermined management device When creating a new secondary volume and a new journal volume at the secondary site, predicting a data processing load by the processor under operating conditions for performing redundancy processing within the secondary site for the new secondary volume; predicting the load on the new journal volume under operating conditions in which redundancy processing is not performed in the secondary site; Selecting a node on which to place the new secondary volume and the new journal volume based on the results of the load prediction A storage system management method comprising:

Citation Information

Patent Citations

  • Remote copy system

    JP2005018736A