Storage system deployment method and computer system
The system improves storage system reliability by configuring redundant instances in subzones with risk boundaries, addressing high availability challenges in cloud SDS services.
Patent Information
- Application Number
- JP2024109481
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2042-06-20
AI Technical Summary
There is a demand for improved reliability, such as high availability, in storage systems, particularly in cloud SDS services where data is distributed across multiple data centers and subzones.
A computer system with a storage controller that configures redundant instances in multiple subzones separated by risk boundaries, utilizing redundancy groups, storage controllers, and capacity pools to enhance reliability.
This approach enhances the reliability of storage systems by ensuring continued operation even in the event of failures within subzones, maintaining data availability and integrity.
Smart Images

Figure 0007733778000001 
Figure 0007733778000002 
Figure 0007733778000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a configuration technology for components related to data management in a computer system including a storage system capable of arranging multiple instances in multiple subzones separated by risk boundaries. [Background technology]
[0002] For example, a cloud SDS service has been proposed that runs on a public cloud by applying SDS (Software Defined Storage) technology, which configures storage by running storage software on a general-purpose server.
[0003] In order to achieve high availability in a public cloud, for example, a configuration is being carried out in which a plurality of data centers called availability zones or zones are prepared and data and services are distributed across the plurality of data centers.
[0004] In recent years, it has also become possible to improve availability within an availability zone by utilizing subzones separated by risk boundaries such as power supply boundaries and rack boundaries.
[0005] As related technologies, Patent Document 1 discloses a spread placement group (SPG) technology consisting of multiple instances each placed in a different subzone. Patent Document 2 discloses a technology for placing replicas of volume partitions across power supply boundaries. Patent Document 3 discloses a technology for building a storage system with a high degree of flexibility while ensuring a certain level of fault tolerance. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] U.S. Patent No. 10,536,340 [Patent Document 2] U.S. Patent No. 9,826,041 [Patent Document 3] Japanese Patent Application Publication No. 2020-64473 Summary of the Invention [Problem to be solved by the invention]
[0007] In storage systems, there is a demand for improved reliability, such as high availability of managed data.
[0008] The present invention has been made in view of the above circumstances, and its object is to provide a technique that can easily and appropriately improve the reliability of a storage system. [Means for solving the problem]
[0009] In order to achieve the above-mentioned object, a computer system according to one aspect is a computer system including a storage system capable of placing multiple instances in any of multiple subzones separated by risk boundaries, and a processor of the computer system configures a storage controller that controls I / O processing for volumes based on capacity pools provided by multiple storages to be redundant with multiple instances placed in the multiple subzones. [Effects of the Invention]
[0010] According to the present invention, it is possible to easily and appropriately improve the reliability of a storage system. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing the overall configuration of a computer system according to one embodiment. [Figure 2] FIG. 2 is a diagram illustrating the configuration of an availability zone according to one embodiment. [Figure 3]FIG. 3 is a configuration diagram of a storage client and an SDS cluster in an availability zone according to one embodiment. [Figure 4] FIG. 4 is a diagram illustrating the arrangement of spread placement groups according to one embodiment. [Figure 5] FIG. 5 is a configuration diagram of a bare metal server according to an embodiment. [Figure 6] FIG. 6 is a configuration diagram of an SDS cluster according to one embodiment. [Figure 7] FIG. 7 is a diagram illustrating scaling out an SDS cluster according to an embodiment. [Figure 8] FIG. 8 is a diagram showing another example of the configuration of a redundant group according to an embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the configuration of a redundant capacity pool function according to an embodiment. [Figure 10] FIG. 10 is a diagram illustrating another example of the configuration of the redundant capacity pool function according to an embodiment. [Figure 11] FIG. 11 is a diagram showing the configuration of a global capacity pool according to one embodiment. [Figure 12] FIG. 12 is a configuration diagram of a redundant storage controller function according to one embodiment. [Figure 13] FIG. 13 is a diagram illustrating a failover of a redundant storage controller function according to one embodiment. [Figure 14] FIG. 14 is a configuration diagram of a redundant SDS cluster management function according to an embodiment. [Figure 15] FIG. 15 is a configuration diagram of a memory of an instance having an SDS cluster deployment management function according to an embodiment. [Figure 16] FIG. 16 is a configuration diagram of a memory of an instance having redundant SDS cluster management functionality according to one embodiment. [Figure 17] FIG. 17 is a diagram illustrating a configuration of a redundancy group-spread placement group mapping table according to one embodiment. [Figure 18]FIG. 18 is a diagram illustrating a spread placement group-instance mapping table according to one embodiment. [Figure 19] FIG. 19 is a diagram showing the structure of a spread placement group configuration table according to one embodiment. [Figure 20] FIG. 20 is a flowchart of an SDS cluster deployment process according to one embodiment. [Figure 21] FIG. 21 is a flowchart of an option calculation process according to one embodiment. [Figure 22] FIG. 22 is a flowchart of an instance deployment process according to an embodiment. [Figure 23] FIG. 23 is a flowchart of an SDS cluster expansion process according to one embodiment. [Figure 24] FIG. 24 is a flowchart of a service volume generation process according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] The following description of the embodiments will be given with reference to the drawings. Note that the embodiments described below do not limit the scope of the invention as claimed, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention.
[0013] In the following explanation, information may be described using the expression "AAA table", but the information may be expressed in any data structure. In other words, to show that the information does not depend on the data structure, the "AAA table" can be called "AAA information".
[0014] In the following description, processing may be described with a "program" as the subject of operation. However, since a program is executed by a processor (e.g., a CPU (Central Processing Unit)) to perform a predetermined process using a storage unit (e.g., a memory) and / or an interface (e.g., a port) as appropriate, the program may also be the subject of the processing operation. Processing described with a program as the subject of operation may also be processing performed by a processor or a computer having the processor (e.g., a server). It may also include a hardware circuit that performs some or all of the processing performed by the processor. A program may also be installed from a program source. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0015] FIG. 1 is a diagram showing the overall configuration of a computer system according to one embodiment.
[0016] The computer system 1 is an example of an environment in which a storage system operates, and includes one or more user terminals 3 and one or more regions 10. The user terminals 3 and the regions 10 are connected via the Internet 2, which is an example of a network.
[0017] The user terminal 3 is a terminal of a user who uses various services provided by the region 10. When there are multiple regions 10, the regions 10 are located in, for example, different countries.
[0018] The region 10 includes a region gateway 11 and one or more availability zones 20. The region gateway 11 connects each availability zone 20 to the Internet 2 so that they can communicate with each other.
[0019] The multiple availability zones 20 in the region 10 are located in different buildings, for example, and each has independent power sources, air conditioning, and networks so as not to affect the other availability zones 20.
[0020] An availability zone 20 includes an availability zone gateway 21 and resources 29 such as compute, network, and storage.
[0021] FIG. 2 is a diagram illustrating the configuration of an availability zone according to one embodiment.
[0022] The availability zone 20 includes an availability zone gateway 21, a spine network switch 22, a cloud storage service 23, and multiple sub-zones 30.
[0023] The availability zone gateway 21 connects the region gateway 11 and the spine network switch 22 so that they can communicate with each other.
[0024] The cloud storage service 23 provides the subzone 30 with the functionality to store various types of data, including multiple storage devices.
[0025] The spine network switch 22 mediates communication between the availability zone gateway 21, the cloud storage service 23, and the multiple subzones 30. The spine network switch 22 and the multiple subzones 30 are connected in a spine-leaf configuration.
[0026] Each subzone 30 includes a leaf network switch 31, one or more bare metal servers 32, and a power distribution unit (PDU) 33. The subzone 30 may also be equipped with an uninterruptible power supply (UPS). Each subzone 30 is divided by a risk boundary for power supplies and networks. That is, each subzone 30 is provided with an independent PDU 33 and leaf network switch 31 to prevent a failure from affecting other subzones 30. The subzone 30 may be configured as, for example, one server rack or multiple server racks.
[0027] The bare metal server 32 includes one or more instances 40. The instance 40 is a functional unit that executes processing, and may be, for example, a virtual machine (VM) that runs on the bare metal server 32.
[0028] The leaf network switch 31 connects each bare metal server 32 to the spine network switch 22 so that they can communicate with each other. The PDU 33 is a unit that distributes and supplies power from a power source (for example, a commercial power source) to each bare metal server 32.
[0029] FIG. 3 is a configuration diagram of a storage client and an SDS cluster in an availability zone according to one embodiment.
[0030] The availability zone 20 includes an instance 40 configured with a storage client 42, an instance 40 configured with an SDS (Software Defined Storage) cluster deployment management function 41, multiple instances 40 that constitute an SDS cluster 45, and a virtual network switch 44 that communicatively connects each instance 40.
[0031] The storage client 42 executes various processes using the storage area provided by the SDS cluster 45. The SDS cluster deployment management function 41 executes a process for managing the deployment of the SDS cluster 45.
[0032] The SDS cluster 45 is composed of multiple instances 40. The SDS cluster 45 performs various processes as a storage system that manages data. Each instance of the SDS cluster 45 operates a redundant SDS cluster management function 71, a redundant storage controller function 72, a redundant capacity pool function 73, and the like, which will be described later. The data managed by the SDS cluster 45 may be stored in a storage device inside the instance 40, or in a storage device external to the instance 40. The external storage device is, for example, a volume provided by a cloud storage service 23.
[0033] Note that Figure 3 shows an example in which the instance 40 on which the storage client 42 runs and the instance 40 that constitutes the SDS cluster 45 are placed in the same availability zone 20, but the present invention is not limited to this, and they may also be placed in different availability zones 20.
[0034] FIG. 4 is a diagram illustrating the arrangement of spread placement groups according to one embodiment.
[0035] Here, a spread placement group (SPG) 50 is a group consisting of a plurality of instances 40 of a plurality of subzones 30, with one instance 40 per subzone 30.
[0036] 4 shows a state in which four SPGs 50, each made up of six instances 40, are configured in the availability zone 20. In such a configuration, even if a power or network failure occurs in one of the subzones 30 in each SPG 50, the power or network of the other subzones 30 is not affected, and the instances 40 in the other subzones 30 can continue processing.
[0037] Next, the hardware configuration of the bare metal server 32 will be described.
[0038] FIG. 5 is a configuration diagram of a bare metal server according to an embodiment.
[0039] The bare metal server 32 includes a NIC (Network Interface Card) 61, a CPU 62 as an example of a processor, a memory 63, and a storage device 64.
[0040] The NIC 61 is an interface such as a wired LAN card or a wireless LAN card.
[0041] The CPU 62 executes various processes according to programs stored in the memory 63 .
[0042] The memory 63 is, for example, a RAM (Random Access Memory), and stores the programs executed by the CPU 62 and necessary information.
[0043] The storage device 64 may be an NVMe drive, a SAS drive, or a SATA drive. The storage device 64 stores programs executed by the CPU 62 and data used by the CPU 62. Note that the bare metal server 32 does not necessarily have to be equipped with a storage device 64.
[0044] Each configuration of the bare metal server 32 is virtually allocated to an instance 40 configured on the bare metal server 32 .
[0045] FIG. 6 is a diagram illustrating an example of a configuration of an SDS cluster according to an embodiment.
[0046] The SDS cluster 45 includes one or more (three in the example of FIG. 6) redundancy groups 70. A redundancy group 70 is a group of instances 40 that completes the redundancy of the capacity pool function, the storage controller function, and the management function of the SDS cluster.
[0047] The redundancy group 70 includes one or more (one in the example of FIG. 6) SPGs 50. The SPGs 50 include multiple (six in the example of FIG. 6) instances 40.
[0048] In one redundant group 70 (RG#1 in FIG. 6), a redundant SDS cluster management function 71, which is a function that makes the SDS cluster management function redundant, is configured by multiple instances 40. In each RG 70, a redundant storage controller function 72, which is a function that makes the storage controller redundant, and a redundant capacity pool function 73, which is a function that makes the capacity pool redundant, are configured by multiple instances 40.
[0049] In the example of Figure 6, a redundant SDS cluster management function 71, a redundant storage controller function 72, and a redundant capacity pool function 73 are configured by instances 40 within the same SPG 50, and even if a failure occurs in the subzone 30 of any of the instances 40, the function can be failed over to an instance 40 in another subzone 30.
[0050] Next, scaling out the SDS cluster 45 will be described.
[0051] Fig. 7 is a diagram illustrating scaling out an SDS cluster according to one embodiment. Fig. 7 shows an example of scaling out an SDS cluster 45 having three RGs 70 (RG#1, RG#2, RG#3) shown in Fig. 7(1).
[0052] When scaling out the SDS cluster 45, the redundant SDS cluster management function 71 adds an RG 70 as a unit, as shown in Fig. 7(2). Specifically, the redundant SDS cluster management function 71 determines the SPG 50 that constitutes the RG 70 (RG#4), and constructs a redundant storage controller function 72 and a redundant capacity pool function 73 using multiple instances 40 of the SPG 50.
[0053] Next, another example of the configuration of the redundancy group 70 will be described.
[0054] FIG. 8 is a diagram showing another example of the configuration of a redundant group according to an embodiment.
[0055] While Figure 6 shows an example in which the RG70 is configured with one SPG50, if the redundancy of the SDS cluster management function, storage controller, and capacity pool is 2 or greater, the RG70 can be configured with multiple SPG50s to cope with some failures, so it may be configured with multiple SPG50s. For example, if the redundancy is n, the RG70 may be configured with n SPG50s. Here, redundancy refers to the number of SPGs that can withstand simultaneous failures.
[0056] Furthermore, if the number of instances 40 required for the RG 70 is greater than the number of instances that can be supported by the SPG, or if the number of instances is greater than the number of instances that can actually be allocated by the SPG, the RG 70 may be configured with multiple SPGs 50. In this case, failures in deploying or expanding the SDS cluster can be suppressed.
[0057] 8, the RG 70 is configured with two SPGs 50 (SPG#a1, SPG#a2). Each SPG 50 includes three instances 40. Note that the number of instances 40 in each SPG 50 is the same, but this is not limiting, and the number of instances 40 may be different.
[0058] Next, an example of the configuration of a redundant capacity pool mechanism will be described.
[0059] 9 is a diagram showing an example of the configuration of a redundant capacity pool function according to an embodiment of the present invention, which shows an example in which redundancy is achieved by mirroring in a redundant capacity pool 80.
[0060] In the example of FIG. 9, one or more storage devices 24 provided by the cloud storage service 23 are connected to each of the multiple instances 40 that make up the RG 70. In these instances 40, pairs of primary chunks 81 and secondary chunks 82 are configured using chunks based on the storage area of a storage device 24 connected to another instance 40. A redundant capacity pool 80 is configured based on the storage area of the primary chunks 81. A redundant capacity pool function 73 provides a service volume 83, which is a volume that can be used by clients, to a redundant storage controller function 72 based on the storage area of the redundant capacity pool 80. Note that instead of the storage device 24, a storage device 64 inside each instance 40 may be used.
[0061] Next, another example of the configuration of the redundant capacity pool mechanism will be described.
[0062] Fig. 10 is a diagram showing another example of the configuration of the redundant capacity pool function according to an embodiment of the present invention, which shows an example in which redundancy is achieved in a redundant capacity pool 80 by erasure cording (4D2P).
[0063] In the example of FIG. 10, one or more storage devices 24 provided by the cloud storage service 23 are connected to each of the multiple instances 40 that make up the RG 70. In these instances 40, a 4D2P configuration is constructed using the storage areas of the storage devices 24 connected to multiple (six in the figure) instances 40. The four Ds represent data chunks, and the two Ps represent parity chunks that store parity data. A redundant capacity pool 80 is configured based on the data areas of the data chunks. A redundant capacity pool function 73 provides a service volume 83, which is a volume that clients can use, to a redundant storage controller function 72 based on the storage area of the redundant capacity pool 80. Note that instead of the storage devices 24, a storage device 64 inside each instance 40 may be used.
[0064] Next, the global capacity pool 84 in the case where the SDS cluster 45 is made up of multiple RGs 70 will be described.
[0065] FIG. 11 is a diagram showing the configuration of a global capacity pool according to one embodiment.
[0066] Each redundant capacity pool function 73 manages a plurality of redundant capacity pools 80 configured from each SPG 50 as one global capacity pool 84 provided by the SDS cluster 45 .
[0067] Next, the redundant storage controller function 72 will be described.
[0068] FIG. 12 is a configuration diagram of a redundant storage controller function according to one embodiment.
[0069] In the example of FIG. 12, a redundant storage controller function 72 is configured by a plurality of instances 40 that constitute an RG 70. In the redundant storage controller function 72, an active controller 90 (an example of a storage controller: an active controller) that provides a service volume 83 to a client is arranged in each instance 40, and a standby controller 90 (an example of a storage controller: a standby controller) that serves as a failover destination for the active controller 90 is arranged in one or more other instances 40 (two in the example of FIG. 12). The controller 90 executes I / O processing and the like for the service volume 83. The active controller and the standby controller synchronize management information for the service volumes they manage (such as the ID of the active controller, connection information for the service volumes, and the data layout of the service volumes in the redundant capacity pool).
[0070] 12, an active controller #1 that manages service volume #1-x is located in instance #1, and standby controllers #1-1 and #1-2 of the standby system corresponding to the active controller #1 are located in instances #2 and #3, respectively. Also, an active controller #2 that manages service volume #2-x is located in instance #2, and standby controllers #2-1 and #2-2 of the standby system corresponding to the active controller #2 are located in instances #3 and #1, respectively. Also, an active controller #3 that manages service volume #3-x is located in instance #3, and standby controllers #3-1 and #3-2 of the standby system corresponding to the active controller #3 are located in instances #1 and #2, respectively.
[0071] Next, a failover of redundant storage controller functionality according to one embodiment will be described.
[0072] 13 is a diagram illustrating a failover of a redundant storage controller function according to an embodiment of the present invention, showing an example in which an instance failure occurs in instance #1 shown in FIG.
[0073] If a failure occurs in instance #1, the standby controller #1-1, which is the standby system for active controller #1, is promoted to active controller #1 and begins operating. Then, service volume #1-x, which was managed by instance #1, is managed by active controller #1 of instance #2.
[0074] When a failover occurs, the access destination from the storage client 42 to the service volume 83 can be switched using a mechanism such as ALUA (Asymmetric Logical Unit Access), which is a well-known technology. Therefore, the storage client 42 issues an IO request to the service volume #1-x that was issued to the instance #1 to the post-switchover instance #2.
[0075] If a further failure occurs in instance #2 after this, standby controller #1-2 will be promoted to active controller #1, and standby controller #2-1 will be promoted to active controller #2, and these active controllers for instance #3 will manage service volumes #1-x and #2-x.
[0076] Next, the redundant SDS cluster management function 71 will be described.
[0077] FIG. 14 is a configuration diagram of a redundant SDS cluster management function according to an embodiment.
[0078] In the example of Fig. 14, a redundant SDS cluster management function 71 is configured by multiple instances 40 that make up the RG 70. In the redundant SDS cluster management function 71, an active cluster manager 92 (primary master) that manages the SDS cluster is placed in one instance 40, and standby cluster managers 92 (secondary master 1, secondary master 2) that serve as failover destinations for the active cluster manager 92 are placed in one or more other instances 40 (two in the example of Fig. 14). The cluster manager 92 is an example of a storage cluster manager.
[0079] When all of instances #1, #2, and #3 are running, the cluster management unit 92 (primary master) of instance #1 manages the cluster. In this case, cluster management information (information on the instances that make up the cluster, service volumes, etc.) is synchronized with instances #2 and #3.
[0080] Here, if instance #1 stops, the cluster management unit 92 (secondary master 1) of instance #2 is promoted to primary master and takes over cluster management, and if instance #2 stops, secondary master 2 of instance #3 is promoted to primary master and takes over cluster management.
[0081] Next, the configuration of the memory 63 of the instance 40 having the SDS cluster deployment management function 41 will be described.
[0082] FIG. 15 is a configuration diagram of a memory of an instance having an SDS cluster deployment management function according to an embodiment.
[0083] The memory 63 of the instance 40 having the SDS cluster deployment management function 41 stores a redundancy group (RG)-spread placement group (SPG) mapping table 101, a spread placement group (SPG)-instance mapping table 102, a spread placement group configuration table 103, an SDS cluster deployment program 104, and an SDS cluster expansion program 105.
[0084] The RG-SPG mapping table 101 stores the correspondence between RGs and the SPGs included in the RGs. The SPG-instance mapping table 102 stores the correspondence between SPGs and the instances that make up the SPGs. The SPG configuration table 103 stores information on candidate SPG configurations. The SDS cluster deploy program 104 is executed by the CPU 62 of the instance 40 to execute an SDS cluster deploy process (see FIG. 20) that deploys an SDS cluster. The SDS cluster expansion program 105 is executed by the CPU 62 of the instance 40 to execute an SDS cluster expansion process (see FIG. 23) that expands the SDS cluster 45.
[0085] Next, the configuration of the memory 63 of the instance 40 having the redundant SDS cluster management function 71 will be described.
[0086] FIG. 16 is a configuration diagram of a memory of an instance having redundant SDS cluster management functionality according to one embodiment.
[0087] The memory 63 of the instance 40 having the redundant SDS cluster management function 71 stores a service volume generation program 111. The service volume generation program 111 is executed by the CPU 62 of the instance 40 to execute a service volume generation process (see FIG. 24) that generates a service volume.
[0088] Next, the RG-SPG mapping table 101 will be described.
[0089] FIG. 17 is a diagram illustrating a configuration of a redundancy group-spread placement group mapping table according to one embodiment.
[0090] An entry in the RG-SPG mapping table 101 includes fields for a redundancy group ID 101a and a spread placement group (SPG) ID 101b.
[0091] The redundancy group ID 101a stores an identifier of a redundancy group (redundancy group ID: RG ID). The SPG ID 101b stores an identifier of an SPG (SPG ID) included in the redundancy group of the same entry.
[0092] For example, according to the example of the RG-SPG mapping table 101 in FIG. 17, it can be seen that the RG with an RG ID of 2 includes SPGs with SPG IDs of 22 and 31.
[0093] Next, the SPG-instance mapping table 102 will be described.
[0094] FIG. 18 is a diagram illustrating a spread placement group-instance mapping table according to one embodiment.
[0095] An entry in the SPG-instance mapping table 102 includes fields for spread placement group ID 102a, instance ID 102b, and redundant SDS cluster management function 102c.
[0096] The spread placement group ID 102a stores an SPG ID. The instance ID 102b stores the ID of an instance (instance ID) included in the SPG of the SPG ID in the entry. The redundant SDS cluster management function 102c stores the type of cluster manager 92 of the redundant SDS cluster management function 71 operating in the instance corresponding to the entry. The types include a primary master indicating an operating cluster manager, and a secondary master indicating a standby cluster manager (if there are multiple secondary masters, secondary master 1, secondary master 2, ...). Note that the information on the redundant SDS cluster management function 102c may be managed separately from the SPG-instance mapping table 102.
[0097] Next, the SPG configuration table 103 will be described.
[0098] Fig. 19 is a diagram showing the configuration of a spread placement group configuration table according to one embodiment. Fig. 19(1) is an example where the number of RG instances is divisible by the number of SPGs, and Fig. 19(2) is an example where the number of RG instances is not divisible by the number of SPGs.
[0099] The SPG configuration table 103 is provided in association with each RG that configures the SDS cluster 45. The SPG configuration table 103 stores an entry for each option of the SPG configuration in the RG. The entry of the SPG configuration table 103 includes an option # 103a, a group number 103b, and an element number 103c.
[0100] The option number corresponding to the entry is stored in the option #103a. The number of groups 103b stores the number of SPGs in the configuration of the option corresponding to the entry. The number of elements 103c stores the number of elements (components, in this example, instances) included in the SPG in the option corresponding to the entry.
[0101] Here, if the number of RG instances in the option (6) is divisible by the number of SPGs (2) as shown in Figure 19(1), the number of elements 103c stores the number of elements for one, as shown in the second entry. If the number of RG instances in the option (7) is not divisible by the number of SPGs (2) as shown in Figure 19(2), the number of elements 103c stores the number of elements for multiple groups, as shown in the second entry.
[0102] Next, the SDS cluster deployment process will be described. The SDS cluster deployment process is executed when the SDS cluster deployment program 104 receives an instruction to deploy an SDS cluster from the user terminal 3, for example.
[0103] FIG. 20 is a flowchart of an SDS cluster deployment process according to one embodiment.
[0104] The SDS cluster deployment program 104 (more precisely, the CPU 62 that executes the SDS cluster deployment program 104) executes an option calculation process (see FIG. 21) that calculates the option for the number of groups in the SPG and the number of elements (instances) in the SPG (S11).
[0105] Next, the SDS cluster deployment program 104 executes an instance deployment process (see FIG. 22) for deploying an instance using an SPG (S12).
[0106] Next, the SDS cluster deployment program 104 creates a mapping (RG-SPG mapping table 101) of the correspondence between the SPG in which the instance is deployed and the RG that includes that SPG (S13).
[0107] Next, the SDS cluster deployment program 104 constructs a redundant SDS cluster management function 71 using multiple instances 40 in any one of the redundant groups (S14). Here, each SDS cluster management unit of the redundant SDS cluster management function 71 does not have to be constructed in all instances 40 in the redundant group; for example, it may be constructed in an instance 40 with the required redundancy + 1.
[0108] Next, the SDS cluster deploy program 104 executes the processing of loop 1 (S15 to S17) for each redundant group. Here, the redundant group to be processed in loop 1 is referred to as the target redundant group.
[0109] In the processing of loop 1, the SDS cluster deploy program 104 constructs a redundant storage controller function 72 with multiple instances 40 of the target redundancy group (S15). Next, the SDS cluster deploy program 104 constructs a redundant capacity pool function 73 with multiple instances 40 of the target redundancy group (S16). Next, the SDS cluster deploy program 104 registers the redundant capacity pool 80 of the target redundancy group in the global capacity pool 84 (S17).
[0110] The SDS cluster deployment program 104 performs loop 1 processing for each redundant group, and when processing has been completed for all redundant groups, ends the SDS cluster deployment processing.
[0111] This SDS cluster deployment process allows you to create an SDS cluster with redundancy that can withstand failures.
[0112] Next, the option calculation process (S11) will be described.
[0113] FIG. 21 is a flowchart of an option calculation process according to one embodiment.
[0114] The SDS cluster deployment program 104 determines the number of redundant groups and the number of instances in each redundant group. R Here, the number of redundant groups and the number of instances I are determined (S21). R For example, this may be determined based on the total number of instances in the SDS cluster specified by the user and the minimum number of instances (minimum required number of instances) required to achieve protection (redundancy) of the specified data.
[0115] For example, if the total number of instances is 12 and a 4D2P configuration is specified for data protection, the minimum number of instances required is 6, so the number of redundant groups may be 12 / 6=2. Also, if the total number of instances is 20, the number of redundant groups may be 20 / 6=3.33·, rounded down to 3. In this case, the number of instances in each redundant group may be 7,7,6 or 8,6,6. Also, if the total number of instances is less than the minimum number of instances required, the number of redundant groups may be 1, and the number of instances may be the minimum number of instances required. Also, if the number of redundant groups and the number of instances in each redundant group are 1, R The specification may be accepted from the user.
[0116] Next, the SDS cluster deployment program 104 determines the maximum number of groups G max (S22) where G max =max(R d ,R c ,R m ) and R d is the data redundancy for the specified data protection, and R c is the redundancy of the redundant storage controller function, and R m is the redundancy of the redundant SDS cluster management function.
[0117] Here, redundancy refers to the number of failures that can be tolerated. For example, if data is configured as a mirror, the data redundancy is 1, and if it is configured as a 4D2P, the data redundancy is 2. Also, if the storage controller is configured as three controllers, active, standby, and standby, the redundancy of the redundant storage controller function is 2.
[0118] Next, the SDS cluster deployment program 104 determines the number of groups i (i ranges from 1 to G max For each of the first to fifth values (in ascending order), the processing of loop 2 (S23 to S26) is executed.
[0119] In the processing of loop 2, the SDS cluster deployment program 104 increases the number of instances in the redundancy group by I R Divide by i and use the value to get the number of basic elements of SPG, e i (S23). R If it is not divisible by i, the decimal part of the divided value is rounded up and used as the number of basic elements of SPG, e i Let's say.
[0120] Next, the SDS cluster deployment program 104 determines whether or not a predetermined option exclusion condition is met (S24). Here, the option exclusion condition is a condition that determines that an option cannot be used to build an appropriate SDS cluster. Specifically, for example, i is greater than the maximum number of elements supported by the SPG (maximum number of supported elements), and the redundancy group that builds the SDS cluster management function, i is R m i is greater than R C ,R D There are things that are bigger, etc.
[0121] As a result, if it is determined that the predetermined option exclusion condition is not met (S24: No), the SDS cluster deploy program 104 registers an option entry in the SPG configuration table 103 in which the number of groups is the number of groups i and the number of elements is the number of basic elements ei (S25). R If is not divisible by i, then for some SPGs, the number of basic elements e i For one SPG, the remaining number of elements is used. For example, if the number of instances is I R If is 7 and i is 2, the number of basic elements e i is 4, so for one SPG, the number of elements is the basic element number e i For the remaining SPGs, the number of elements is set to the number of instances I R to the number of basic elements e i This is subtracted to make it 3.
[0122] On the other hand, if it is determined that the specified option exclusion condition is met (S24: Yes), the SDS cluster deployment program 104 excludes the case of this group number i from the option (S26) and ends the processing of loop 2 for this group number i.
[0123] The SDS cluster deployment program 104 performs the processing of loop 2 for each group number i, and when the processing has been completed for all group numbers, ends the option calculation processing.
[0124] Next, the instance deployment process (S12) will be described.
[0125] FIG. 22 is a flowchart of an instance deployment process according to an embodiment.
[0126] The SDS cluster deployment program 104 executes the processing of loop 3 (loop 4 (S31, S32)) for each redundant group of all redundant groups.
[0127] In the processing of loop 3, the SDS cluster deployment program 104 executes the processing of loop 4 (S31, S32) for each option of option number i (in ascending order of option number).
[0128] In the processing of loop 4, the SDS cluster deployment program 104 specifies an SPG with the number of elements of option number i for the number of groups of option number i, and attempts to deploy an instance (S31).
[0129] Next, the SDS cluster deployment program 104 determines whether the deployment of the instance in step S31 was successful (S32).
[0130] As a result, if the instance deployment is successful (S32: Yes), the SDS cluster deployment program 104 exits loop 4 and executes loop 3 processing on the next redundant group, and if processing has been executed for all redundant groups, it exits loop 3 processing and terminates the instance deployment processing.
[0131] On the other hand, if the instance deployment is not successful even after trying all option numbers (S32: No), the SDS cluster deployment program 104 exits loop 4 and terminates the instance deployment process as an abnormal end.
[0132] Next, the SDS cluster expansion process will be described. The SDS cluster expansion process is executed when the SDS cluster expansion program 105 receives an instruction to expand the SDS cluster from the user terminal 3, for example.
[0133] 23 is a flowchart of an SDS cluster expansion process according to an embodiment. Note that steps similar to those in the SDS cluster deployment process shown in FIG. 20 are denoted by the same reference numerals.
[0134] The SDS cluster expansion program 105 (more precisely, the CPU 62 executing the SDS cluster expansion program 105) executes an option calculation process (see Figure 21) that calculates options for the number of groups in the SPG and the number of elements (instances) in the SPG for the expansion portion of the cluster (S41).
[0135] Next, the SDS cluster expansion program 105 executes an instance deployment process (see FIG. 22) for deploying an instance using an SPG for the expansion portion of the cluster (S42).
[0136] Next, the SDS cluster expansion program 105 adds the correspondence between the SPG in which the instance that is the expansion part in the cluster is deployed and the RG that includes that SPG to the mapping (RG-SPG mapping table 101) (S43).
[0137] Next, the SDS cluster expansion program 105 executes the processing of loop 5 (S44, S15 to S17) for the added redundant group. Here, the redundant group to be processed in loop 5 is referred to as the target redundant group.
[0138] In the processing of loop 5, the SDS cluster expansion program 105 adds the instance of the target redundant group to the management targets of the redundant SDS cluster management function 71 (S44). Next, the SDS cluster expansion program 105 executes the processing of steps S15, S16, and S17.
[0139] The SDS cluster expansion program 105 performs the processing of loop 5 for each added redundant group, and when the processing has been completed for all redundant groups, ends the SDS cluster expansion processing.
[0140] This SDS cluster expansion process allows the SDS cluster to be expanded to tolerate failures.
[0141] Next, the service volume generation process will be described.
[0142] FIG. 24 is a flowchart of a service volume generation process according to an embodiment.
[0143] The service volume creation program 111 (strictly speaking, the CPU 62 that executes the service volume creation program 111) selects from the global capacity pool 84 a redundant capacity pool 80 that has free capacity designated as the service volume to be created (S51).
[0144] Next, the service volume creation program 111 allocates capacity to the service volume 83 from the selected redundant capacity pool 80 (S52).
[0145] Next, the service volume creation program 111 selects one of the active controllers of the instances 40 belonging to the redundant group 70 to which the selected redundant capacity pool 80 is provided, allocates the created service volume 83 (S53), and terminates the processing. After this processing, a process is performed using known technology to register a path or the like in the storage client 42 so that the service volume 83 can be accessed.
[0146] The present invention is not limited to the above-described embodiment, and can be modified appropriately without departing from the spirit of the present invention.
[0147] For example, in the above-described embodiments, some or all of the processing performed by the CPU may be performed by a hardware circuit. Also, the programs in the above-described embodiments may be installed from a program source. The program source may be a program distribution server or a storage medium (e.g., a portable storage medium). [Explanation of symbols]
[0148] 1...Computer system, 2...Internet, 10...Region, 20...Availability zone, 30...Subzone, 32...Bare metal server, 40...Instance, 45...SDS cluster, 50...SPG, 62...CPU, 63...Memory, 70...Redundant group, 71...Redundant SDS cluster management mechanism, 72...Redundant storage controller function, 73...Redundant capacity pool function
Claims
1. A method for deploying a storage system in a computer system, the method deploying a storage system in a plurality of instances arranged in a plurality of subzones separated by a risk boundary, comprising: The storage system includes a redundant group (RG) having a redundant capacity pool function that provides a capacity pool based on a plurality of storage devices, and a controller function that controls I / O processing for the capacity pool, The processor of the computer system configuring a spread placement group (SPG) having a plurality of instances, each of the plurality of instances being placed in a different subzone; The controller function and the redundant capacity pool function of the redundancy group (RG) are configured to be redundant among the plurality of instances arranged in different subzones within a spread placement group (SPG); When deploying a storage system to the multiple instances, a processor of the computer system Deploying the spread placement group (SPG) with multiple instances; creating a mapping between a spread placement group (SPG) having the plurality of instances and the redundancy group (RG); establishing a management function for managing the redundancy group (RG) with multiple instances of the spread placement group (SPG); constructing a controller function for the redundancy group (RG) with multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG); Using multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG), a redundant capacity pool function of the redundancy group (RG) is constructed. Storage system deployment method.
2. 2. The storage system deployment method according to claim 1, Deploying a plurality of the redundant groups (RGs) each having the controller function and the redundant capacity pool function, A single deployed management function manages multiple redundancy groups (RGs). Storage system deployment method.
3. 3. The storage system deployment method according to claim 2, The plurality of redundancy groups (RG) are configured using a plurality of the spread placement groups (SPG), The management function is configured with one of the spread placement groups (SPG) mapped to one of the redundancy groups (RG). Storage system deployment method.
4. 2. The storage system deployment method according to claim 1, Deploying a plurality of the redundant groups (RGs) each having the controller function and the redundant capacity pool function, The plurality of capacity pools related to the plurality of redundancy groups (RGs) that have been constructed are registered in a global capacity pool. Storage system deployment method.
5. 2. The storage system deployment method according to claim 1, Determine the number of redundancy groups (RGs) to be deployed and the number of instances each of the redundancy groups (RGs) will use; determining a data redundancy, a controller function redundancy, and a management function redundancy of the redundancy group (RG) to be deployed; determining the number of redundancy groups (RGs) and the number of instances used by each of the redundancy groups (RGs); Determine the number of spread placement groups (SPGs) to be deployed based on the determined number of redundancy groups (RGs), the number of instances used by each of the redundancy groups (RGs), the redundancy of data in the redundancy groups (RGs), the redundancy of the controller function, and the redundancy of the management function. Storage system deployment method.
6. 2. The storage system deployment method according to claim 1, A controller function of the redundancy group (RG) and a redundant capacity pool function that provides the capacity pool used by the controller function are constructed in the same spread placement group (SPG). Storage system deployment method.
7. 2. The storage system deployment method according to claim 1, A plurality of spread placement groups (SPGs) are used to form one redundancy group (RG). Storage system deployment method.
8. 8. The storage system deployment method according to claim 7, The redundancy of the controller function is set to redundancy n, which is a redundancy that can withstand failures in n subzones, and n spread placement groups (SPGs) are used to configure one redundancy group (RG). Storage system deployment method.
9. 2. The storage system deployment method according to claim 1, When deploying a storage system to a cluster in which the storage system has already been deployed, the processor of the computer system Deploying the spread placement group (SPG) with multiple instances; Adding a mapping for the redundancy group (RG) to be additionally deployed to the mapping between the spread placement group (SPG) having the plurality of instances and the redundancy group (RG); Add the newly deployed redundancy group (RG) to the management targets of the established management function, constructing a controller function for the redundancy group (RG) with multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG); Using multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG), a redundant capacity pool function of the redundancy group (RG) is constructed. Storage system deployment method.
10. 5. The storage system deployment method according to claim 4, The processor of the computer system Selecting the capacity pool from the global capacity pool and allocating it to the service volume to be accessed by the client Storage system deployment method.
11. 5. The storage system deployment method according to claim 4, When a failure occurs in any of the subzones in which the multiple controller functions and multiple redundant capacity pool functions of the redundancy group (RG) are arranged, the controller function and the redundant capacity pool function arranged in another subzone will fail over the function. Storage system deployment method.
12. 1. A computer system comprising: A storage system that provides multiple instances arranged in multiple subzones separated by a risk boundary, The storage system includes a redundant group (RG) having a redundant capacity pool function that provides a capacity pool based on a plurality of storage devices, and a controller function that controls I / O processing for the capacity pool, The processor of the computer system configuring a spread placement group (SPG) having a plurality of instances, each of the plurality of instances being placed in a different subzone; The controller function and the redundant capacity pool function of the redundancy group (RG) are configured to be redundant among the plurality of instances arranged in different subzones within a spread placement group (SPG); When deploying a storage system to the multiple instances, a processor of the computer system Deploying the spread placement group (SPG) with multiple instances; creating a mapping between a spread placement group (SPG) having the plurality of instances and the redundancy group (RG); establishing a management function for managing the redundancy group (RG) with multiple instances of the spread placement group (SPG); constructing a controller function for the redundancy group (RG) with multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG); Using multiple instances of the spread placement group (SPG) mapped to the redundancy group (RG), a redundant capacity pool function of the redundancy group (RG) is constructed. Computer system.
Citation Information
Patent Citations
Storage system and data placement method in storage system
JP2020064473A
Storage system and method for controlling the same
JP2021036450A
Information processing system and method
JP2021135703A
Spread placement groups
US10536340B1
Relative placement of volume partitions
US9826041B1