Management device and management method for storage system

The management device addresses the issue of access load management in storage systems by monitoring and calculating access rates, thereby preventing resource shortages and ensuring processing performance during secondary data use in cloud storage systems.

JP2025080994APending Publication Date: 2025-05-27HITACHI VANTARA LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023194449
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Conventional disaster recovery technologies do not adequately consider the access load on storage systems when secondary use of copied data is performed on a cloud site, leading to potential resource shortages and decreased processing performance.

Method used

A management device for a storage system that monitors access loads and throughput on a second storage system, calculates access numbers for each server, and identifies servers with excessive increase rates, displaying relevant information to manage access loads effectively.

Benefits of technology

This solution enables effective management of access loads during secondary use of copied data, preventing resource shortages and maintaining processing performance in storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080994000001_ABST
    Figure 2025080994000001_ABST
Patent Text Reader

Abstract

To take an access load into consideration when copy data to another storage is used secondarily on the other storage side.SOLUTION: A management node manages an access load on a second storage that provides second servers with a second volume to which a first volume provided to first servers by a first storage is copied, and a snapshot created from the second volume. When the number of accesses to a storage device storing the second volume or throughput thereof exceeds a threshold, the management node calculates an increase rate of the number of accesses on the basis of a first number of writes and a first number of reads to the second volume for each of the first servers. In addition, the management node calculates an increase rate of the number of accesses on the basis of a second number of writes and a second number of reads to the second volume for each of the second servers. Then, the management node specifies the first and second servers in which the increase rate exceeds the threshold, and displays them on a display unit.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a management device and a management method for a storage system.

Background Art

[0002] Conventionally, in order to prevent data loss at the primary site in the event of a large-scale disaster such as an earthquake or a fire, a technology called disaster recovery (DR) that copies and stores data at a secondary site in a remote location is known.

[0003] In recent years, DR that has been performed between on-premises sites is now being performed between an on-premises site and the cloud. Here, in the cloud, since resources can be used in the required amount at the required time, secondary use such as analyzing the data copied to the cloud on the cloud side becomes possible.

[0004] Regarding this point, for example, Patent Document 1 discloses a computer system that can perform migration control according to user requirements by controlling a related volume pair or group, combining the source volume and the destination volume.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in the above-described conventional technology, the access load on the storage when the copied data to the storage of the cloud site is secondarily used on the cloud site side is not considered.

[0007] The present invention has been made in consideration of the above circumstances, and an object thereof is to consider the access load when secondary use of copied data to another storage is performed on the other storage side.

Means for Solving the Problems

[0008] According to one aspect of the present invention, there is provided a management device for a storage system that manages an access load in a second storage that provides a second volume obtained by remotely copying a first volume provided by a first storage to a second server, a snapshot created from the second volume, and the management device has a processor and a memory, and the processor monitors the number of accesses or throughput to a storage device that stores the second volume, and when the number of accesses or the throughput exceeds a threshold, based on the first write number and the first read number for each first server with respect to the second volume, calculates the first access number for each first server to the storage device, calculates the second access number for each second server to the storage device based on the second write number and the second read number for each second server with respect to the second volume, calculates an increase rate of the first access number for each first server and an increase rate of the second access number for each second server, identifies the first server and the second server for which the increase rate exceeds the threshold, and is characterized by displaying information related to the identified first server and second server on a display unit.

Effects of the Invention

[0009] According to the present invention, it is possible to consider the access load when secondary use of copied data to another storage is performed on the other storage side. Therefore, it is possible to suppress a decrease in processing performance due to resource shortage in another storage due to the access load when secondary use of copied data is performed on the other storage side.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Modes for Carrying Out the Invention

[0011] In the following description, the "interface device" may be one or more interface devices. The one or more interface devices may be at least one of the following. · One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface device is an interface device for at least one of an I / O device and a remote display computer. The I / O interface device for the display computer may be a communication interface device. At least one I / O device may be either an input device such as a user interface device, for example, a keyboard and a pointing device, or an output device such as a display device. · One or more communication interface devices. The one or more communication interface devices may be one or more of the same type of communication interface devices (for example, one or more Network Interface Cards (NICs)) or two or more different types of communication interface devices (for example, a NIC and a Host Bus Adapter (HBA)).

[0012] Also, in the following description, "memory" is one or more memory devices that are an example of one or more storage devices, and may typically be a main memory device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.

[0013] Also, in the following description, "storage device" may be one or more persistent storage devices that are an example of one or more storage devices. Persistent storage devices may typically be non-volatile storage devices (such as auxiliary storage devices), and specifically, for example, HDD (Hard Disk Drive), SSD (Solid State Drive), NVME (Non-Volatile Memory Express) drive, or SCM (Storage Class Memory).

[0014] Also, in the following description, "CPU (Central Processing Unit)" is an example of one or more processor devices. At least one processor device may typically not be limited to a CPU, but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core.

[0015] At least one processor device may also be a circuit that is an aggregate of gate arrays described in a hardware description language for performing part or all of the processing. A circuit is a processor device in a broad sense such as, for example, an FPGA (Field-Programmable Gate Array), CPLD (Complex Programmable Logic Device), or ASIC (Application Specific Integrated Circuit).

[0016] In the following description, the function may be described in terms of a "yyy functional unit". The function may be realized by one or more computer programs being executed by a processor, or by one or more hardware circuits (such as an FPGA or ASIC), or by a combination thereof.

[0017] When the function is realized by a program being executed by a processor, since the defined processing is performed while appropriately using a storage device and / or an interface device, etc., the function may be regarded as at least part of the processor. The processing described with the functional unit as the subject may also be the processing performed by the processor or a device having the processor.

[0018] The program may be installed from a program source. The program source may be, for example, a program distribution computer or a computer-readable recording medium (such as a non-transitory recording medium). The description of each function is an example, and a plurality of functions may be combined into one function, or one function may be divided into a plurality of functions. The "yyy functional unit" may also be called the "yyy unit".

[0019] In the following description, "volume" (VOL) indicates a storage area of a storage, and these may be implemented by a physical storage device or a logical storage device. Also, the VOL may be a physical VOL or a virtual VOL (VVOL). A snapshot VOL may be a VOL as a snapshot of a VOL.

[0020] In the following description, when describing elements of the same kind without distinction, common reference signs among the reference signs are used, and when describing elements of the same kind while distinguishing them, reference signs may be used.

[0021] [Embodiment 1] FIG. 1 is a diagram showing an example of the configuration of the overall system S according to Embodiment 1. The overall system S is configured by connecting an on-premises data center 1 and a cloud site 2 via a WAN (Wide Area Network).

[0022] The on-premises data center 1 is a data center where an on-premises environment is installed. In the on-premises data center 1, an on-premises DB (Data Base) node 11 and a storage 13 connected via a SAN (Storage Area Network) 12 are arranged.

[0023] The cloud site 2 is a site where a cloud environment is constructed. The cloud site 2 includes a compute node 21 connected via a SAN 22, a storage cluster 23, copy nodes 25a, 25b, and a management node 27.

[0024] The compute node 21 is a server that executes various processes, and performs data input / output with respect to a storage node 24 included in the storage cluster 23 during processing.

[0025] The storage cluster 23 is configured to include a plurality of storage nodes 24a, 24b, ….

[0026] The copy nodes 25a, 25b are servers in which secondary utilization systems 26a, 26b that secondarily utilize data remotely copied from the storage 13 of the on-premises data center 1 onto the cloud site 2 operate. The secondary utilization systems 26a, 26b are, for example, an analysis system that analyzes data, a test environment of a system that conducts tests using the data, and the like. In the present embodiment, the compute node 21 and the copy nodes 25a, 25b are separately shown for convenience, but the copy nodes 25a, 25b may be included in the compute node 21.

[0027] Details of the management node 27 will be described later with reference to FIG. 12.

[0028] (Configuration of DB Node 11) FIG. 2 is a diagram showing an example of the configuration of DB node 11 according to Embodiment 1. DB node 11 is an example of a computer node (server) arranged in on-premises data center 1. DB node 11 has a network I / F 31, a CPU 32, volumes 33a and 33b, and a memory 34. Memory 34 stores a database function unit 35 and storage connection information 36.

[0029] Network I / F 31 is an interface for DB node 11 to communicate with storage 13 (FIG. 1) via SAN 12.

[0030] CPU 32 controls the overall operation of DB node 11 and executes a predetermined program to realize database function unit 35. Volumes 33a and 33b are storage areas provided to DB node 11 by storage 13.

[0031] Storage connection information 36 has columns of a storage IP address 37 and a storage iSCSI name 38. Database function unit 35 connects to storage 13 based on storage connection information 36 to realize data access.

[0032] (Configuration of Storage 13) FIG. 3 is a diagram showing an example of the configuration of storage 13 according to Embodiment 1. Storage 13 has a network I / F 41, a CPU 42, storage devices 43a and 43b, and a memory 44. Memory 44 stores a storage control function unit 45, storage configuration information 46, storage performance information 47, volume performance information 48, and a remote copy function unit 49.

[0033] Network I / F 41 is an interface for storage 13 to communicate with DB node 11 via SAN 12.

[0034] The CPU 42 controls the overall control of the storage 13 and executes a predetermined program to realize the storage control function unit 45 and the remote copy function unit 49.

[0035] The storage configuration information 46 is the configuration information of the storage 13. The storage performance information 47 is performance information such as the number of accesses per unit time and throughput of the storage 13 observed in time series. The volume performance information 48 is performance information such as the number of accesses per unit time and throughput of the volumes 33a and 33b observed in time series.

[0036] The storage control function unit 45 processes I / O to the storage devices 43a and 43b in response to an I / O request from the DB node 11. The remote copy function unit 49 copies the volumes and journals stored in the storage devices 43a and 43b to the cloud site 2 in response to a remote copy request.

[0037] (Configuration of the compute node 21) FIG. 4 is a diagram showing an example of the configuration of the compute node 21 according to Embodiment 1. The compute node 21 includes a network I / F 51, a CPU 52, volumes 53a and 53b, and a memory 54. The memory 54 stores a database function unit 55 and storage connection information 56.

[0038] The network I / F 51 is an interface for the compute node 21 to communicate with the storage cluster 23 via the SAN 22.

[0039] The CPU 52 controls the overall control of the compute node 21 and executes a predetermined program to realize the database function unit 55. The volumes 53a and 53b are storage areas provided to the compute node 21 by the storage cluster 23.

[0040] The storage connection information 56 has columns of a storage IP address 57 and a storage iSCSI name 58. The database function unit 55 connects to the storage node 24 based on the storage connection information 56 to realize data access.

[0041] (Configuration of Storage Node 24) FIG. 5 is a diagram showing an example of the configuration of the storage node 24 according to Embodiment 1. One storage cluster 23 is configured to include a plurality of storage nodes 24. The storage node 24 has a network I / F 61, a CPU 62, storage devices 63a, 63b, and a memory 64. The memory 64 stores a storage control function unit 65. The memory 64 also stores storage configuration information 66, volume configuration information 67, storage device configuration information 68, volume performance information 69, storage device performance information 70, a remote copy function unit 71, and remote copy configuration information 72.

[0042] The network I / F 61 is an interface for the storage node 24 to communicate with the compute node 21 and the copy nodes 25a, 25b via the SAN 22.

[0043] The CPU 62 controls the overall operation of the storage node 24 and executes a predetermined program to realize the storage control function unit 65 and the remote copy function unit 71.

[0044] (Configuration of Storage Configuration Information 66) FIG. 6 is a diagram showing an example of the configuration of the storage configuration information 66 according to Embodiment 1. The storage configuration information 66 manages information on the data protection type and node size of the storage node 24. The storage configuration information 66 has columns of a storage node ID 661, a data protection type 662, and a node size 663.

[0045] The storage node ID 661 is the identification information of storage node 24. The data protection type 662 is the protection type of the corresponding storage node 24. The data protection types include "Mirror", "mDnP", etc. "Mirror" is a protection method in which the same data is stored in two storage devices 63. "mDnP" is a protection method in which data is stored in m storage devices 63 and parity data is stored in n storage devices 63. Data protection may be realized, for example, by EC (Erasure Coding) or by RAID (Redundant Array of Independent Disks).

[0046] The node size 663 indicates the storage capacity of storage node 24.

[0047] (Configuration of volume configuration information 67) FIG. 7 is a diagram showing an example of the configuration of volume configuration information 67 according to Embodiment 1. The volume configuration information 67 has columns of a volume ID 671, a storage node ID 672, a data protection storage node ID 673, a journal group ID 674, a volume type 675, and a snapshot source volume 676.

[0048] The volume ID 671 is the identification information of the volume. The storage node ID 672 is the identification information of the storage node 24 where the corresponding volume is arranged. The data protection storage node ID 673 is the identification information of the storage node 24 where the protection data of the corresponding volume is arranged.

[0049] The journal group ID 674 is the journal group to which the corresponding volume belongs. The volume type 675 is information indicating whether the corresponding volume is a regular volume or a snapshot volume. The snapshot source volume 676 indicates the identification information of the volume from which the snapshot is created when the corresponding volume is a snapshot volume.

[0050] (Configuration of Storage Device Configuration Information 68) FIG. 8 is a diagram showing an example of the configuration of storage device configuration information 68 according to Embodiment 1. The storage device configuration information 68 has columns of storage node IDs 681 and storage device IDs 682.

[0051] The storage node ID 681 is identification information of the storage node 24. The storage device ID 682 is identification information of the storage devices 63a and 63b. That is, the storage device configuration information 68 indicates the storage devices 63 connected to each storage node 24.

[0052] (Configuration of Volume Performance Information 69) FIG. 9 is a diagram showing an example of the configuration of volume performance information 69 according to Embodiment 1. The volume performance information 69 has columns of storage node ID 691, volume ID 692, date and time 693, number of reads 694, number of writes 695, read throughput 696, and write throughput 697.

[0053] The storage node ID 691 is identification information of the storage node 24. The volume ID 692 is identification information of the volume. The date and time 693 is the date and time when the corresponding record was recorded.

[0054] The number of reads 694 indicates the number of simultaneous reads from the volume identified by the corresponding storage node ID 691 and volume ID 692 at the corresponding date and time 693. The number of writes 695 indicates the number of simultaneous writes to the volume identified by the corresponding storage node ID 691 and volume ID 692 at the corresponding date and time 693.

[0055] The read throughput 696 and the write throughput 697 indicate the respective throughputs of read / write such as IOPS for the volume identified by the corresponding storage node ID 691 and volume ID 692 at the corresponding date and time 693.

[0056] (Configuration of Storage Device Performance Information 70) FIG. 10 is a diagram showing an example of the configuration of storage device performance information 70 according to Embodiment 1. The storage device performance information 70 has columns of a storage node ID 701, a storage device ID 702, a date and time 703, a read count 704, a write count 705, a read throughput 706, and a write throughput 707.

[0057] The storage node ID 701 is identification information of storage node 24. The storage device ID 702 is identification information of storage device 63. The date and time 703 is the date and time when the corresponding record was recorded.

[0058] The read count 704 indicates the number of simultaneous reads from the storage device 63 identified by the corresponding storage node ID 701 and storage device ID 702 at the corresponding date and time 703. The write count 705 indicates the number of simultaneous writes to the storage device 63 identified by the corresponding storage node ID 701 and storage device ID 702 at the corresponding date and time 703.

[0059] The read throughput 706 and the write throughput 707 indicate the respective throughputs of read / write such as IOPS for the storage device 63 identified by the corresponding storage node ID 701 and storage device ID 702 at the corresponding date and time 703.

[0060] (Configuration of Remote Copy Configuration Information 72) FIG. 11 is a diagram showing an example of the configuration of remote copy configuration information 72 according to Embodiment 1. The remote copy configuration information 72 has columns of a storage ID 721, a volume ID 722, a storage node ID 723, and a volume ID 724.

[0061] The storage ID 721 is the identification information of storage 13. The volume ID 722 is the identification information of the volume on storage 13. The storage node ID 723 is the identification information of storage node 24. The volume ID 724 is the identification information of the volume on storage node 24. That is, the remote copy configuration information 72 indicates the correspondence between the volume on storage 13 and the volume on storage node 24 that remotely copies this volume.

[0062] Return to the description of FIG. 5. The storage control function unit 65 processes the I / O to the volume on storage node 24 in response to the I / O request from compute node 21. Also, the storage control function unit 65 cooperates with other storage nodes 24 according to the specified protection type, and redundantizes the volumes and journals stored in its own storage node with other storage nodes 24.

[0063] The remote copy function unit 71 copies the volume and journal on storage 13 to its own storage node 24 in cooperation with the remote copy function unit 49 of storage 13 in response to the remote copy request.

[0064] (Configuration of management node 27) FIG. 12 is a diagram showing an example of the configuration of management node 27 according to Embodiment 1. The management node 27 is a management device that manages the compute node 21, the storage cluster 23 (storage system), and the copy node 25.

[0065] The management node 27 has a network I / F 81, a CPU 82, volumes 83a, 83b, a memory 84, and a display unit 93. The display unit 93 is a display device or the like. The memory 84 stores a performance degradation cause identification function unit 85, an improvement proposal function unit 86, a performance degradation cause display control unit 87, an improvement proposal display control unit 88, DB configuration information 89, copy server configuration information 90, volume group configuration information 91, and volume group performance information 92.

[0066] The network I / F 81 is an interface for the management node 27 to communicate with the storage node 24 and the copy nodes 25a and 25b via the SAN 22.

[0067] The CPU 82 controls the overall control of the management node 27 and executes a predetermined program to implement the performance degradation cause identification functional unit 85, the improvement proposal functional unit 86, the performance degradation cause display control unit 87, and the improvement proposal display control unit 88.

[0068] (Configuration of the DB configuration information 89) FIG. 13 is a diagram showing an example of the configuration of the DB configuration information 89 according to Embodiment 1. The DB configuration information 89 has columns of a DB node ID 891, a storage node ID 892, and a volume ID 893.

[0069] The DB node ID 891 is the identification information of the DB node 11. The storage node ID 892 is the identification information of the storage node 24. The volume ID 893 is the identification information of the volume on the storage node 24 identified by the storage node ID 892. That is, the DB configuration information 89 shows the correspondence between the DB node 11 and the volume on the storage node 24 accessed by this DB node 11.

[0070] (Configuration of the copy server configuration information 90) FIG. 14 is a diagram showing an example of the configuration of the copy server configuration information 90 according to Embodiment 1. The copy server configuration information 90 has columns of a copy node ID 901, a storage node ID 902, and a volume ID 903.

[0071] The copy node ID 901 is the identification information of the copy node 25. The storage node ID 902 is the identification information of the storage node 24. The volume ID 903 is the identification information of the volume on the storage node 24 identified by the storage node ID 902. That is, the copy server configuration information 90 shows the correspondence between the copy node 25 and the volume on the storage node 24 accessed by this copy node 25.

[0072] (Volume group configuration information 91) FIG. 15 is a diagram showing an example of the configuration of the volume group configuration information 91 according to Embodiment 1. The volume group configuration information 91 has columns of a volume group ID 911, a volume ID 912, a compute type 913, and a compute ID 914.

[0073] The volume group ID 911 is identification information of a group of volumes created on the storage node 24. The volume ID 912 is identification information of a volume belonging to the volume group ID 911. The compute type 913 indicates the type of computing resources for accessing the volume identified by the combination of the volume group ID 911 and the volume ID 912. For example, there are "DB" and "copy". "DB" indicates that it is accessed by the DB node 11. "Copy" indicates that it is accessed by the copy node 25. The compute ID 914 indicates the computing resources for accessing the volume identified by the combination of the volume group ID 911 and the volume ID 912. For example, "DB11a" indicates the DB node 11a. Also, for example, "CN25a" indicates the copy node 25a.

[0074] (Volume group performance information 92) FIG. 16 is a diagram showing an example of the configuration of the volume group performance information 92 according to Embodiment 1. The volume group performance information 92 has columns of a volume group ID 921, a date and time 922, the number of accesses 923, and a throughput 924.

[0075] The volume group ID 921 is identification information for a group of volumes created on storage node 24. The date and time 922 is the date and time when the corresponding record was recorded. The access count 923 is the total number of accesses to all volumes belonging to the corresponding volume group ID 921 at the corresponding date and time 922. The throughput 924 indicates the read / write throughput such as IOPS for all volumes belonging to the corresponding volume group ID 921 at the corresponding date and time 922.

[0076] (Regarding the volume group) FIG. 17 is a diagram showing an example of a volume group according to Embodiment 1. The arrangement of each volume of the main volume, sub-volume, journal volume, and snapshot volume in FIG. 17 is based on the exemplary contents of the tables in FIGS. 6 to 11 and FIGS. 13 to 16. In FIG. 17, "P" is the main volume, "S" is the sub-volume paired with the main volume, "J" is the journal volume, and "SS" is the snapshot volume of the sub-volume.

[0077] As shown in FIG. 17, in the on-premises data center 1, the DB nodes 11a and 11b are operating. The DB node 11a accesses the main volumes 212 and 213. The DB node 11b accesses the main volume 211. The main volumes 211, 212, and 213 belong to the journal group 201. Journal data is written to the journal volume 214 for the main volumes 211, 212, and 213. A journal group is a logical group of volumes whose history is managed by one journal volume.

[0078] At one cloud site 2, in storage node 24a, a journal volume 221 which is a copy of journal volume 214 is created. Secondary volumes 222, 223, 224 are created based on the journal data stored in journal volume 221 as remote copies of primary volumes 211, 212, 213. Secondary volumes 222, 223, 224 belong to journal group 203.

[0079] Depending on the data protection type of storage node 24a (storage configuration information 66 (Figure 6)), the number N (N is a positive integer) of storage devices 63 to which volumes are to be written is different. As described above, when the data protection type is "Mirror", N = 2, and when it is "mDnP", N = (m + n). The "DB write count" described later is multiplied by N when the access load is calculated.

[0080] Snapshot volume 225 is a snapshot of secondary volume 224 belonging to journal group 203. Snapshot volume 225 is accessed by copy node 25b where secondary use system 26b operates.

[0081] Similarly, the primary journal group 202 in on-premises data center 1 and the secondary journal group 204 in cloud site 2 form a pair. Snapshot volume 226 is a snapshot of a certain secondary volume xxx (see Figure 7) belonging to journal group 204. Snapshot volume 226 is accessed by copy node 25a where secondary use system 26a operates.

[0082] Also, the journal group 205 on the storage node 24b configures mirroring for data protection using the storage device 63 of the storage node 24a. This indicates that the site for creating a remote copy on the cloud site 2 is not limited to the on-premises data center 1, but may be other cloud sites. That is, in addition to a hybrid cloud like the present embodiment, the embodiment can also be applied to a multi-cloud.

[0083] In such a volume arrangement, the following access loads (1) to (2) become problems. (1) The access load on the storage device 63 due to remote copying from the journal groups 201 and 202 in the on-premises data center 1 and the journal group 205 on the storage node 24b in the cloud site 2 to the storage node 24a. (2) The access load on the storage device 63 via the snapshot volumes 225 and 226 by the secondary use system 26 on the copy node 25.

[0084] (Processing for Identifying the Cause of Performance Degradation According to Embodiment 1) FIG. 18 is a flowchart showing an example of the processing for identifying the cause of performance degradation according to Embodiment 1.

[0085] Prior to the processing for identifying the cause of performance degradation, the performance degradation cause identification functional unit 85 (FIG. 12) of the management node 27 refers to the storage device performance information 70 (FIG. 10). Then, for each storage node ID 701 and date / time 703, the performance degradation cause identification functional unit 85 refers to the read count 704, write count 705, read throughput 706, and write throughput 707.

[0086] Then, the performance degradation cause identification function unit 85 determines, for each storage node ID 701 and date / time 703, whether or not the total access count value X1 obtained by summing the read count 704 and write count 705 exceeds a predetermined threshold value. Further, the performance degradation cause identification function unit 85 determines whether or not the total throughput value X2 obtained by summing the read throughput 706 and write throughput 707 exceeds a predetermined threshold value. Then, for the storage node 24 at the date / time 703 when the total access count value X1 or the total throughput value X2 exceeds the predetermined threshold value, a performance degradation cause identification process is executed. In the present embodiment, it is assumed that for the storage node 24a, the total access count value X1 or the total throughput value X2 exceeds the predetermined threshold value.

[0087] The "access amplification factor" is a factor that amplifies the values of each index defined for each read and write and for each index, such as the IOPS during read, the IOPS during write, the throughput during read, and the throughput during write, for each data protection type.

[0088] First, in step S11, the performance degradation cause identification function unit 85 acquires the read count R1 and write count W1 for each DB node 11. That is, the performance degradation cause identification function unit 85 refers to the DB configuration information 89 (FIG. 13) and acquires the storage node ID 892 and volume ID 893 for each DB node (DB node ID 891). Then, the performance degradation cause identification function unit 85 refers to the volume performance information 69 (FIG. 9) of the corresponding storage node 24a based on the acquired storage node ID 892 and volume ID 893, and acquires the read count 694 and / or write count 695 at the latest date / time 693. In this way, the read count R1 and write count W1 when each DB node 11 accesses each volume are acquired.

[0089] Next, in step S12, the performance degradation cause identification function unit 85 acquires the read count R2 and write count W2 for each copy server. That is, the performance degradation cause identification function unit 85 refers to the copy server configuration information 90 (FIG. 14) and acquires the storage node ID 902 and volume ID 903 for each copy server (copy node ID 901). Then, the performance degradation cause identification function unit 85 refers to the volume performance information 69 (FIG. 9) of the corresponding storage node 24a based on the acquired storage node ID 902 and volume ID 903, and acquires the read count 694 at the latest date and time 693. In this way, the read count R2 and write count W2 when each copy server accesses each volume are acquired.

[0090] Next, in step S13, the performance degradation cause identification function unit 85 calculates the storage device access count A1 based on the read count R1 and write count W1 for each DB node 11 according to Equation (1). Also, the performance degradation cause identification function unit 85 calculates the storage device access count A2 based on the read count R2 and write count W2 for each copy server according to Equation (2). Here, "N1" and "N2" are the number of storage devices 63 to be written according to the data protection type, and "α1", "β1", "α2", and "β2" are predetermined "access amplification factors". A1 = α1 × R1 + N1 × β1 × W1 ··· (1) A2 = α2 × R2 + N2 × β2 × W2 ··· (2)

[0091] Next, in step S14, the performance degradation cause identification function unit 85 identifies, as a performance degradation cause, a DB node 11 in which the storage device access count A1 calculated in step S13 has increased beyond a predetermined value (for example, 40%) within a certain period. The "certain period" may be the acquisition interval of the date and time 693 in the volume performance information 69 (FIG. 9), or may be a time obtained by grouping a certain number of these acquisition intervals. In step S14, the increase rate of the latest storage device access count A1 with respect to the past storage device access count A1 is calculated for a certain period, and it is determined whether this increase rate has exceeded a predetermined value.

[0092] Next, in step S15, the performance degradation cause identification function unit 85 identifies, as the cause of performance degradation, a copy server (copy node 25) in which the storage device access count A2 calculated in step S13 has increased by a predetermined amount (for example, 40%) or more over a certain period. This "certain period" is the same as the "certain period" in step S14. In step S15, the increase rate of the latest storage device access count A2 with respect to the past storage device access count A2 is calculated over a certain period, and it is determined whether this increase rate exceeds a predetermined value.

[0093] Next, in step S16, the performance degradation cause display control unit 87 outputs the DB node 11 identified in step S14 and the copy server (copy node 25) identified in step S15 to the performance degradation cause display screen 87D (FIG. 19).

[0094] (Configuration of the performance degradation cause display screen 87D) FIG. 19 is a diagram showing an example of the configuration of the performance degradation cause display screen 87D according to Embodiment 1. The performance degradation cause display screen 87D has display items for each of the threshold-exceeded storage node 301, compute node ID 311, access count increase rate 312, and performance degradation factor 313.

[0095] The threshold-exceeded storage node 301 indicates the storage node 24 (storage node 24a in this embodiment) in which the total access count value X1 or throughput total value X2 that triggered the execution of the performance degradation cause identification process (FIG. 18) has exceeded a predetermined threshold.

[0096] The compute node ID 311 indicates the DB node 11 or copy node 25 that accesses the storage node 24a in which the total access count value X1 or throughput total value X2 that triggered the execution of the performance degradation cause identification process (FIG. 18) has exceeded a predetermined threshold. The DB node 11 indicated in the compute node ID 311 can be obtained from the DB configuration information 89 (FIG. 13). The copy node 25 indicated in the compute node ID 311 can be obtained from the copy server configuration information 90 (FIG. 14).

[0097] The access number increase rate 312 is calculated in steps S13, S14, and S15 of the performance degradation cause identification process (Figure 18). The performance degradation factor 313 has an "○" input for the DB node 11 or the copy node 25 identified as the performance degradation cause in steps S14 and S15 of the performance degradation cause identification process (Figure 18).

[0098] (Improvement proposal process according to Embodiment 1) Figure 20 is a flowchart showing an example of the improvement proposal process (during overload) according to Embodiment 1. In this embodiment, it is monitored whether the above-mentioned total access number value X1 or the total throughput value X2 exceeds a predetermined threshold. When the storage node 24 is overloaded, an expansion of resources such as scale-out (addition of storage nodes 24) or scale-up is proposed.

[0099] The improvement proposal process according to this embodiment is executed when the above-mentioned total access number value X1 or the total throughput value X2 exceeds a predetermined threshold, similar to the performance degradation cause identification process (Figure 18) according to Embodiment 1. The improvement proposal process is sequenced after the execution of the performance degradation cause identification process.

[0100] First, in step S21, the improvement proposal function unit 86 (Figure 12) of the management node 27 groups the volume groups within the journal groups 203 and 204 and the snapshot volumes related to this volume group for the target node. In this embodiment, the "target node" is the storage node 24a for which the total access number value X1 or the total throughput value X2 exceeds a predetermined threshold.

[0101] Specifically, in the example shown in FIG. 17, the storage node 24a is the target node. In this case, the volume group (journal volume 221, secondary volumes 222, 223, 224) belonging to the journal group 203 and the snapshot volume 225 of the secondary volume 224 are grouped into the volume group ID 911 "VG1". Similarly, the journal group 204 and the snapshot volume 226 are grouped into the group with the volume group ID 911 being "VG2". The result of this grouping is as shown in the volume group configuration information 91 (FIG. 15).

[0102] Next, in step S22, the improvement proposal function unit 86 groups the volume groups of other nodes that are mirroring the target node. Specifically, in the example shown in FIG. 17, the target node is the storage node 24a, and the storage node that is mirroring the storage node 24a is the storage node 24b. Therefore, the volume group within the journal group in the storage node 24b and the snapshot volume related to this volume group are grouped. The result of this grouping is stored in the volume group configuration information 91 (FIG. 15) in the same manner as in step S21.

[0103] Next, in step S23, the improvement proposal function unit 86 calculates the storage device access counts of the volume groups grouped in steps S21 and S22. Specifically, the storage device access count for each volume of the storage node 24a accessed by the DB node 11 is calculated from the above formula (1). Also, the storage device access count for each volume of the storage node 24a accessed by the copy node 25 is calculated from the above formula (2). From these, the storage device access count for each volume group is calculated. The storage device access count for each volume group is stored in the volume group performance information 92 (FIG. 16) together with the date and time 922 which is the time stamp at the time of calculation.

[0104] Next, in step S24, the improvement proposal function unit 86 determines whether the latest and maximum number of storage device accesses per volume group calculated in step S23 is less than or equal to the threshold value. If the latest and maximum number of storage device accesses is less than or equal to the threshold value (step S24 Yes), the improvement proposal function unit 86 transfers the process to step S26. If it is greater than the threshold value (step S24 No), the improvement proposal function unit 86 transfers the process to step S25.

[0105] In step S25, the improvement proposal function unit 86 creates a volume migration plan for scale-out. The volume migration plan is a proposal to perform scale-out to other storage nodes 24 other than the storage node 24a in volume group units so that the total access count value X1 or the total throughput value X2 becomes less than or equal to a predetermined threshold value, and to distribute the load of the storage node 24.

[0106] Following step S24 or S25, in step S26, the improvement proposal function unit 86 creates a "scale-up plan". The scale-up plan is a proposal to scale up the storage node 24 (storage node 24a in this embodiment) with an overloaded load so that the total access count value X1 or the total throughput value X2 becomes less than or equal to a predetermined threshold value. The scale-up plan may also include a proposal to scale up other storage nodes 24 belonging to the storage cluster 23 including the storage node 24 with an overloaded load.

[0107] Next, in step S27, the improvement proposal function unit 86 creates a copy server load reduction plan. The copy server load reduction plan is a proposal to stop one or more copy servers (copy nodes 25) so that the total access count value X1 or the total throughput value X2 becomes less than or equal to a predetermined threshold value.

[0108] Next, in step S28, the improvement proposal function unit 86 determines whether the cause of the threshold exceedance of the total access count value X1 or the total throughput value X2 is the DB (DB node 11). The improvement proposal function unit 86 can determine whether the cause of the threshold exceedance (performance degradation) of the total access count value X1 or the total throughput value X2 is the DB (DB node 11) by referring to the execution result of the performance degradation cause identification process (Fig. 18) (for example, the performance degradation cause display screen 87D (Fig. 19)). If the cause of the threshold exceedance is the DB (step S28 Yes), the improvement proposal function unit 86 transfers the process to step S29; if the cause of the threshold exceedance is not the DB (step S28 No), the improvement proposal function unit 86 transfers the process to step S30.

[0109] In step S29, the improvement proposal function unit 86 creates a DB load reduction plan. The improvement proposal function unit 86 obtains, with reference to the volume group configuration information 91 (Fig. 15), the number of volumes in the storage node 24 accessed by the DB node 11 determined as the cause of the threshold exceedance in step S28, and divides "100%" by this number of volumes. The DB load reduction plan is to display each of the division results as the load reduction rate when each DB node 11 is deleted.

[0110] Next, in step S30, the improvement proposal display control unit 88 displays each countermeasure on the improvement proposal screen 88D (Fig. 21) on the display unit 93. The countermeasures are the volume movement plan (scale-out plan) at the time of scale-out created in step S25, the scale-up plan created in step S26, and the load reduction plan created in step S27.

[0111] The user selects one of the countermeasures, i.e., the scale-out plan, the scale-up plan, and the load reduction plan, displayed on the improvement proposal screen 88D, and presses the selection button 88D1 (Fig. 21). Then, under the instruction of the management node 27, the countermeasure is executed by the DB node 11, the storage cluster 23, the storage node 24, and / or the copy node 25.

[0112] When the scale-out plan is executed, the copy setting between the storage 13 and the storage node 24 included in the storage cluster 23 is changed, and the connection is set to the volume on the storage node 24 where the scale-out destination is located.

[0113] (Configuration of the improvement proposal screen 88D (when load exceeds limit)) FIG. 21 is a diagram showing an example of the configuration of the improvement proposal screen 88D (when load exceeds limit) according to Embodiment 1. The improvement proposal screen 88D shown in FIG. 21 shows a "scale-out plan", a "scale-up plan", and a "load reduction plan" so that the user can select them.

[0114] In the improvement proposal screen 88D shown in FIG. 21, the "scale-out plan" has a display of the current number of storage nodes 321 and the number of storage nodes 322 after scale-out.

[0115] The current number of storage nodes 321 indicates the number of storage nodes 24 before scale-up by the process of step S26 (FIG. 20). The number of storage nodes 322 after scale-out indicates the number of storage nodes 24 after scale-up by the process of step S26.

[0116] Also, in the improvement proposal screen 88D shown in FIG. 21, the "scale-out plan" has a display of the volume ID 331, the compute ID 332, the transferability 333, and the destination storage node ID 334.

[0117] The volume ID 331 indicates a list of identification information of the volumes on the storage node 24. The compute ID 332 indicates the identification information of the computing resources (DB node 11 or copy node 25) that access the volume identified by the volume ID 331. The transferability 333 indicates whether the corresponding volume can be moved to another storage node 24 in the volume migration plan created in step S25 (FIG. 20). When the transferability 333 is "○", it indicates that it can be moved, and when it is "‐", it indicates that it cannot be moved. The destination storage node ID 334 indicates the identification information of the destination storage node 24 when the corresponding volume can be moved.

[0118] Also, on the improvement proposal screen 88D shown in FIG. 21, the "scale-up plan" has displays of storage node ID 341, current node size 342, and node size after scale-up 343.

[0119] The storage node ID 341 is the identification information of storage node 24. The current node size 342 indicates the node size before scale-up of the corresponding storage node 24 in the scale-up plan created in step S26 (FIG. 20). Also, the node size after scale-up 343 indicates the node size after scale-up of the corresponding storage node 24 in the scale-up plan created in step S26.

[0120] Also, on the improvement proposal screen 88D shown in FIG. 21, the "load reduction plan" has displays of compute ID 351, volume ID 352, and load reduction rate 353.

[0121] The compute ID 351 is the identification information of the DB node 11 or copy node 25 proposed to be stopped in step S27 (FIG. 20). The volume ID 352 is the identification information of the volume accessed by the corresponding DB node 11 or copy node 25. The load reduction rate 353 indicates the ratio of the load reduced by stopping the corresponding DB node 11 or copy node 25 to the total load. When the user selects a load reduction plan, the compute ID 351 can also be specified to indicate which of the DB nodes 11 or copy nodes 25 is to be stopped.

[0122] (Modification Example of Embodiment 1) In step S21 of this embodiment, it is assumed that a snapshot is included in the volume group, and the volume group including the snapshot is moved to the storage node 24 at the scale-out destination. However, the present invention is not limited to this. A countermeasure may be created in which the snapshot is not included in the volume group, deleted at the time of scale-out, and recreated from the volume moved to the storage node 24 at the scale-out destination. By causing the storage cluster 23 or the storage node 24 to execute this countermeasure, volume movement associated with scale-out or scale-in can be performed promptly.

[0123] (Effect of Embodiment 1) In Embodiment 1, the DB nodes 11 and copy nodes 25 in which the increase rate of the number of accesses to the storage device 63 per DB node 11 and the increase rate of the number of accesses to the storage device 63 per copy node 25 exceed the threshold are specified. Then, the information related to the specified DB nodes 11 and copy nodes 25 is displayed on the display unit 93.

[0124] Therefore, according to Embodiment 1, the user can know whether the cause of the performance degradation of the storage node 24 is due to remote copy or secondary use of data, and can appropriately improve it.

[0125] Also, in Embodiment 1, the number of accesses to the storage device 63 per DB node 11 is calculated based on the number of reads and writes to the storage device 63 per DB node 11, the number of storage devices 63 per data protection type, and the access amplification factor. Also, the number of accesses to the storage device 63 per copy node 25 is calculated based on the number of reads and writes per copy node 25 and the number of storage devices 63 per data protection type and the access amplification factor.

[0126] Therefore, according to Embodiment 1, the number of accesses to the storage device 63 can be appropriately estimated based on the data protection type and the access amplification factor.

[0127] In Embodiment 1, a volume group including the volume at the remote copy destination is created, and when the number of accesses to any volume group exceeds a threshold value, the storage at the remote copy destination is scaled out to add a storage node 24. Then, a countermeasure including a scale-out plan for moving the volume group whose number of accesses has exceeded the threshold value to the newly added storage node 24 is created and displayed.

[0128] Therefore, according to Embodiment 1, since scaling out is performed in units of volume groups, scaling out can be performed while maintaining data consistency of related volumes.

[0129] In Embodiment 1, a snapshot is included in the volume group, and a scale-out plan for moving the volume group including the snapshot to the newly added storage node 24 is created.

[0130] Therefore, in Embodiment 1, scaling out can be performed while maintaining data consistency of related volumes in units of volume groups including snapshots.

[0131] In Embodiment 1, the volume group excluding the snapshot is moved to the newly added storage node. Then, a scale-out plan is created to newly create a snapshot based on the volumes included in the volume group at the storage node 24, which is the destination where the volume group has been moved.

[0132] Therefore, according to Embodiment 1, scaling out can be quickly completed by a volume group that does not include a snapshot.

[0133] Also, in Embodiment 1, a volume group including a storage node 24 having a storage device 63 whose access count or throughput exceeds a threshold and a volume on another storage node 24 having a redundant configuration is created. Then, the access count for the storage device 63 for each volume group is calculated and compared with the threshold. When the access count for any second volume group exceeds the threshold, the corresponding storage is scaled out to add storage nodes 24. Then, a scale-out plan is created to move the volume group whose access count exceeds the threshold to the added storage nodes 24.

[0134] Therefore, according to Embodiment 1, the overload of access by remote copy in a multi-cloud environment can also be eliminated by scale-out.

[0135] Also, in Embodiment 1, a countermeasure plan is created to scale up a storage node 24 having a storage device 63 whose access count or throughput exceeds a threshold.

[0136] Therefore, according to Embodiment 1, the overload of access to the storage device 63 can be eliminated by scale-up.

[0137] Also, in Embodiment 1, a server reduction plan is created to reduce a copy node 25 whose increase rate of access count or throughput to the storage device 63 exceeds a threshold.

[0138] Therefore, according to Embodiment 1, the overload of access to the storage device 63 can be eliminated by reducing the copy node 25.

[0139] [Embodiment 2] In Embodiment 1, it is monitored whether the total access count value X1 or the total throughput value X2 exceeds a predetermined threshold. When the load on the storage node 24 exceeds the limit, resource expansion such as scale-out (adding more storage nodes 24) or scale-up is proposed. On the other hand, there may be a case where the load on the storage node 24 is low and the resources allocated to the storage node 24 are excessive. In Embodiment 2, when there is a resource surplus, the resources are reduced to suppress unnecessary resource charging.

[0140] (Improvement Proposal Processing (When Reducing Storage Nodes) According to Embodiment 2) FIG. 22 is a flowchart showing an example of the improvement proposal processing (when reducing storage nodes) according to Embodiment 2. The improvement proposal processing (when reducing storage nodes) is executed, for example, at a fixed period.

[0141] First, in step S31, the improvement proposal function unit 86 (FIG. 12) initializes the reduction number n of the storage nodes 24 to be reduced during scale-in to n = 0. Next, in step S32, the improvement proposal function unit 86 increments the reduction number n by 1.

[0142] Next, in step S33, the improvement proposal function unit 86 creates a "combination of reducing n storage nodes 24" during scale-in. Next, in step S34, the improvement proposal function unit 86 selects one unselected combination from the "combination of reducing n storage nodes 24" created in step S33. Next, in step S35, the improvement proposal function unit 86 creates a volume transfer plan to transfer the volume of the storage nodes 24 to be reduced to other storage nodes 24 that are not to be reduced in the "combination of reducing n storage nodes 24".

[0143] Next, in step S36, the improvement proposal function unit 86 determines whether volume migration plans have been created for all "combinations of reducing the number of n storage nodes 24". If the improvement proposal function unit 86 has created volume migration plans for all "combinations of reducing the number of n storage nodes 24" (step S36 Yes), the process proceeds to step S37. On the other hand, if there is a "combination of reducing the number of n storage nodes 24" for which a volume migration plan has not been created (step S36 No), the process returns to step S34.

[0144] In step S37, the improvement proposal function unit 86 determines whether a volume migration plan can be created. In step S37, it is determined whether one or more volume migration plans can be created to move the volume of the storage node 24 to be reduced to other storage nodes 24 that are not to be reduced. "A volume migration plan can be created" means that when the volume of the storage node 24 to be reduced is moved to a storage node 24 that is not to be reduced, the number of accesses to the storage device 63 at this storage node 24 that is not to be reduced is equal to or less than the threshold value. The number of accesses refers to the above-mentioned storage device access number A1 and storage device access number A2.

[0145] If the improvement proposal function unit 86 can create one or more volume migration plans (step S37 Yes), the process proceeds to step S38, and if it cannot be created (step S37 No), the process returns to step S32.

[0146] In step S38, the improvement proposal function unit 86 determines the candidates for the storage nodes 24 to be reduced. Among the "combinations of reducing the number of n storage nodes 24", the combination with the maximum reduction number n, where the number of accesses of each storage node 24 to the storage device 63 is equal to or less than the predetermined threshold value and the maximum value of the number of accesses is the minimum compared to other volume migration plans, is used as the candidate. That is, in step S38, the combination of "reducing the number of n storage nodes 24" with the most balanced number of accesses to the storage device 63 and the maximum reduction number n is determined as the candidate.

[0147] Next, in step S39, the improvement proposal function unit 86 selects a node size that is slightly smaller during downscaling. In step S39, when reducing the number of storage nodes 24, the node size of the storage nodes 24 that are not reduced is decreased by one level for downscaling.

[0148] Next, in step S40, the improvement proposal function unit 86 acquires the number of accesses to the storage device 63 in the storage node 24 whose node size was downscaled in step S39. This number of accesses is the above-described storage device access number A1 and storage device access number A2.

[0149] Next, in step S41, the improvement proposal function unit 86 determines whether the number of accesses to the storage device acquired in step S40 is equal to or less than the threshold value. When the number of accesses to the storage device is equal to or less than the threshold value (step S41 Yes), the improvement proposal function unit 86 returns the process to step S39, and when it exceeds the threshold value (step S41 No), the process proceeds to step S42.

[0150] In step S42, the improvement proposal function unit 86 determines the node size before the reduction in the last executed step S39 as the node size during downscaling.

[0151] Next, in step S43, the improvement proposal display control unit 88 displays the storage node 24 to be reduced, the reduction number n, and the volume movement plan (scale-in plan) determined in step S38, and the node size (scale-down plan) during downscaling determined in step S42. The improvement proposal display control unit 88 displays these countermeasures on the improvement proposal screen 88D2 (FIG. 23) on the display unit 93. The user selects one of the improvement plans of the scale-in plan and the scale-down plan displayed on the improvement proposal screen 88D2, and the selection button 88D1 (FIG. 23) is pressed. Then, under the instruction of the management node 27, the corresponding measures are executed by the DB node 11, the storage cluster 23, the storage node 24, and / or the copy node 25.

[0152] (Configuration of Improvement Proposal Screen 88D2 (When Reducing Storage Nodes)) FIG. 23 is a diagram showing an example of the configuration of the improvement proposal screen 88D (when reducing storage nodes) according to Embodiment 2. The improvement proposal screen 88D2 shown in FIG. 23 shows a "scale-in plan" and a "scale-down plan" so that the user can select them.

[0153] In the improvement proposal screen 88D2 shown in FIG. 23, the "scale-in plan" has a display of the current number of storage nodes 371 and the number of storage nodes 372 after scale-in.

[0154] The current number of storage nodes 371 indicates the number of storage nodes 24 before scale-in determined in the process of step S38 (FIG. 22). The number of storage nodes 372 after scale-in indicates the number of storage nodes 24 after scale-in determined in the process of step S38.

[0155] Also, in the improvement proposal screen 88D shown in FIG. 23, the "scale-in plan" has a display of volume ID 381, compute ID 382, transferability 383, and destination storage node ID 384.

[0156] The volume ID 381 indicates a list of identification information of volumes on the storage node 24. The compute ID 382 indicates the identification information of the computing resources (DB node 11 or copy node 25) that access the volume identified by the volume ID 381. The transferability 383 indicates whether the corresponding volume can be moved to another storage node 24 in the volume movement plan created in step S35 (FIG. 22). "○" for transferability 383 indicates that it can be moved, and "‐" indicates that it cannot be moved. The destination storage node ID 384 indicates the identification information of the destination storage node 24 when the corresponding volume can be moved.

[0157] Also, in the improvement proposal screen 88D2 shown in FIG. 23, the "scale-down plan" has displays of the storage node ID 391, the current node size 392, and the node size after scale-down 393.

[0158] The storage node ID 391 is the identification information of the storage node 24. The current node size 392 indicates the node size of the storage node 24 before the change to the node size at the time of scale-down determined in step S42 (FIG. 22). Also, the node size 393 after scale-down indicates the node size at the time of scale-down determined in step S42.

[0159] Note that the scale-down plan may include proposals for scale-down of other storage nodes 24 belonging to the storage cluster 23 including the storage node 24 with surplus resources.

[0160] (Effect of Embodiment 2) In Embodiment 2, a scale-in plan is created in which the volume on the storage node 24 to be reduced among the plurality of storage nodes 24 is moved to a storage node 24 that is not to be reduced among the plurality of storage nodes 24, and the corresponding storage node 24 is reduced.

[0161] Therefore, according to Embodiment 2, by appropriately reducing the surplus resources for the access load on the storage device 63 by scale-in, wasteful resource consumption can be eliminated and the cloud usage fee can be reduced.

[0162] Also, in Embodiment 2, a scale-down plan is created in which the node size of the storage node 24 that is not to be reduced among the plurality of storage nodes 24 is scaled down.

[0163] Therefore, according to Embodiment 2, by appropriately reducing the surplus resources for the access load on the storage device 63 by scale-down, wasteful resource consumption can be eliminated and the cloud usage fee can be reduced.

[0164] Also, in Embodiment 2, in response to a user instruction, the DB node 11, the copy node 25, the storage node 24, or the storage cluster 23 is instructed to execute any of the countermeasures displayed on the display unit 93. The countermeasures are a scale-out plan, a scale-up plan, a load reduction plan, a scale-in plan, and a scale-down plan.

[0165] Therefore, according to Embodiment 2, from the creation to the implementation of the countermeasures for the performance degradation cause identification process (FIG. 18), the improvement proposal process (when the load exceeds) (FIG. 18), and the improvement proposal process (when the number of storage nodes is reduced) (FIG. 22), it can be seamlessly executed by the GUI operation displayed on the display unit 93.

[0166] As described above, the embodiments according to the present disclosure have been described in detail. However, the present disclosure is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, it is possible to add, delete, or replace a part of the configuration of the above-described embodiments with other configurations.

[0167] Also, each of the above-described configurations, functional units, processing units, etc. may be realized in hardware by designing a part or all of them, for example, by an integrated circuit. Also, each of the above-described configurations, functions, etc. may be realized in software by a processor interpreting and executing a program for realizing each function. Information such as programs, tables, files, etc. for realizing each function can be placed in a storage device such as a memory, HDD, SSD, or in a recording medium such as an IC card, SD card, DVD.

[0168] Also, in each of the above-described figures, control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines in actual implementation. For example, it may be considered that almost all the configurations are actually connected to each other.

[0169] Also, the above-described processing functions and data arrangement forms are merely examples. The processing functions and data arrangement forms can be changed to an optimal arrangement form from the viewpoints of hardware and software performance, processing efficiency, communication efficiency, etc.

Explanation of Signs

[0170] S: Entire system, 1: On-premises data center, 2: Cloud site, 11, 11a, 11b: DB nodes, 13: Storage, 21: Compute node, 23: Storage cluster, 24, 24a, 24b: Storage nodes, 26, 26a, 26b: Reuse system, 27: Management node, 82: CPU, 84: Memory, 87D: Performance degradation cause display screen, 88D, D2: Improvement proposal screen.

Claims

1. A management device for a storage system that manages an access load in a second storage that provides a second volume remotely copied from a first volume provided by a first storage and a snapshot created from the second volume to a second server, wherein: The management device has a processor and a memory; The processor: Monitors the number of accesses or throughput to the storage device storing the second volume; When the number of accesses or the throughput exceeds a threshold value; Calculates a first number of accesses for each first server to the storage device based on a first number of writes and a first number of reads for each first server to the second volume; Calculates a second number of accesses for each second server to the storage device based on a second number of writes and a second number of reads for each second server to the second volume; Calculates an increase rate of the first number of accesses for each first server and an increase rate of the second number of accesses for each second server; Identifies the first server and the second server for which the increase rate exceeds the threshold value; Displays information related to the identified first server and second server on a display unit. The management device is characterized by the above.

2. The management device according to Claim 1, wherein: The processor: Calculates the first number of accesses for each first server based on the first number of reads and the first number of writes for each first server and the number of storage devices and an access amplification factor for each data protection type; Calculates the second number of accesses for each second server based on the second number of reads and the second number of writes for each second server and the number of storage devices and an access amplification factor for each data protection type. The management device is characterized by the above.

3. The management device according to Claim 1, wherein: The second storage is a storage cluster configured to include a plurality of storage nodes; The processor: When the number of accesses or the throughput exceeds a threshold value; Creates a volume group including the second volume; Calculates a third number of accesses to the storage device for each volume group; Compares the third number of accesses with a threshold value. When the third access count for any of the volume groups exceeds a threshold value, scale out the second storage to add storage nodes, and create a countermeasure including a scale-out plan to move the volume group for which the third access count has exceeded the threshold value to the added storage node, and display it on the display unit A management device characterized by the above

4. The management device according to claim 3, wherein The processor Includes the snapshot in the volume group Creates a countermeasure including a scale-out plan to move the volume group including the snapshot to the added storage node A management device characterized by the above

5. The management device according to claim 3, wherein The processor Excludes the snapshot from the volume group Moves the volume group from which the snapshot has been excluded to the added storage node Deletes the snapshot Creates a countermeasure including a scale-out plan to newly create the snapshot based on the second volume included in the volume group at the storage node to which the volume group has been moved A management device characterized by the above

6. The management device according to claim 3, wherein The processor When the access count or the throughput exceeds a threshold value Creates a second volume group including volumes on other storage nodes that form a redundant configuration with the storage node having the storage device for which the access count or the throughput has exceeded the threshold value Calculates a fourth access count for the storage device for each of the second volume groups Compares the fourth access count with the threshold value When the fourth access count for any of the second volume groups exceeds the threshold value, scale out the second storage to add storage nodes, and create a countermeasure including a scale-out plan to move the second volume group for which the fourth access count has exceeded the threshold value to the added storage node, and display it on the display unit A management device characterized by the above

7. The management device according to claim 3, wherein The processor When the access count or the throughput exceeds a threshold value Create a countermeasure to scale up the storage node having the storage device in which the access number or the throughput exceeds a threshold value, and display it on the display unit. A management device characterized by the above.

8. The management device according to claim 3, The processor is, When the access number or the throughput exceeds a threshold value, Create a countermeasure including a server reduction plan to reduce the second server in which the increase rate exceeds a threshold value, and display it on the display unit. A management device characterized by the above.

9. The management device according to claim 1, The processor is, Move the second volume on the storage node to be reduced among the plurality of storage nodes to a storage node not to be reduced among the plurality of storage nodes, and create a countermeasure including a scale-in plan to reduce the storage node to be reduced, and display it on the display unit. A management device characterized by the above.

10. The management device according to claim 9, The processor is, Create a countermeasure including a scale-down plan to scale down the node size of the storage node not to be reduced among the plurality of storage nodes, and display it on the display unit. A management device characterized by the above.

11. The management device according to any one of claims 3 to 10, The processor is, In response to a user instruction, instruct the first server, the second server, the storage node, or the storage cluster to execute any of the countermeasures displayed on the display unit. A management device characterized by the above.

12. A management method executed by a management device of a storage system that manages an access load in a second storage that provides a second volume obtained by remotely copying a first volume provided by a first storage to a first server, and a snapshot created from the second volume to a second server, The management device has a processor and a memory, The processor is, Monitor the access number or throughput to the storage device storing the second volume, When the access number or the throughput exceeds a threshold value, Calculate the first access number for each first server to the storage device based on the first write number and the first read number for each first server with respect to the second volume. Calculate the second access number for each second server to the storage device based on the second write number and the second read number for each second server with respect to the second volume. Calculate the increase rate of the first access number for each first server and the increase rate of the second access number for each second server. Identify the first server and the second server for which the increase rate exceeds the threshold value. Display the information related to the identified first server and second server on a display unit. A management method characterized by having each process.

Citation Information

Patent Citations

  • CDN attack detection methods, devices, storage media, and electronic equipment

    CN112367324B

  • An architecture for dynamically autoscaling network security microservices based on load

    JP2019523949A

  • Storage system and analytical method of storage system

    JP2021149550A

  • Storage system and method for analyzing storage system

    US20210294497A1

  • Data migration system and data migration control method

    WO2018131133A1