Management device, storage system, and information processing method
The management device in cluster systems addresses workload concentration by strategically placing workloads and replicas based on network traffic margins, enhancing throughput and resource utilization.
Patent Information
- Application Number
- JP2021095239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-06-07
AI Technical Summary
In cluster systems, workload concentration on specific storage can lead to decreased throughput due to uncontrolled load distribution.
A management device that includes a scheduler to select server and storage nodes based on network traffic margins, ensuring balanced workload placement and replica positioning to disperse load across the system.
Improves throughput by effectively utilizing resources and preventing performance degradation through fair resource allocation and load balancing.
Smart Images

Figure 0007700520000004 
Figure 0007700520000005 
Figure 0007700520000006
Abstract
Description
Technical Field
[0001] The present invention relates to a management device, a storage system, and an information processing method.
Background Art
[0002] There exists a cluster system having storage via a network.
[0003] In a cluster system, a plurality of servers are prepared as a cluster connected by a network, and applications share and utilize the hardware. As one form of the cluster configuration, there is a system in which computing resources (CPUs) and storage resources (storage) are separated. As an execution form of an application, the adoption of low-overhead containers is progressing.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] FIG. 1 is a diagram for explaining an example of data reading in a cluster system.
[0006] In the cluster system shown in FIG. 1, for data A and B stored in storage #1 and data C stored in storage #2, replica data A' is stored in storage #2 and replica data B' and C' are stored in storage #3, respectively.
[0007] As shown by symbol A1, server #1 is reading data A from storage #1. Also, as shown by symbol A2, server #2 is reading data B from storage #1. Further, as shown by symbol A3, server #3 is reading data C from storage #2.
[0008] In this way, in a cluster system, if the workload is not controlled, the workload may concentrate on a specific storage (storage #1 in the example shown in FIG. 1), and there is a risk that the throughput will decrease.
[0009] On one aspect, it aims to improve throughput by dispersing the load in the storage system.
Means for Solving the Problem
[0010] On one aspect, the management device Comprising a plurality of server devices and a plurality of storage devices connected to each other via a network is a management device for a storage system, Each of the plurality of server devices includes a processor and an accelerator, and each of the plurality of storage devices includes a volume. when executing a container, Select a first server device from among the plurality of server devices, where the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the accelerator and the volume is equal to or less than the margin in the network. Select a second server device from among the plurality of server devices, where the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the processor and the accelerator is equal to or less than the margin in the network. Determine the first server device or the second server device as the workload placement destination. .
Effect of the Invention
[0011] On one aspect, it is possible to improve throughput by dispersing the load in the storage system.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0013] 〔A〕Embodiment Hereinafter, an embodiment will be described with reference to the drawings. However, the embodiments shown below are merely examples, and there is no intention to exclude various modifications and applications of technologies not explicitly shown in the embodiments. That is, the present embodiment can be variously modified and implemented without departing from its gist. Also, each figure is not intended to include only the components shown in the figure, and can include other functions and the like.
[0014] Hereinafter, in the drawings, since the same reference numerals indicate the same parts, the description thereof will be omitted.
[0015] 〔A-1〕Configuration Example FIG. 2 is a block diagram for briefly explaining the acquisition and accumulation of workload information in the embodiment.
[0016] When a workload is executed, workload information 131 is acquired and saved in association with the workload ID. In other words, when a container is executed, the access status to resources is acquired. And resource requirement information is accumulated in association with the container image.
[0017] As the load during the execution of a workload, CPU load, memory load, Graphics Processing Unit (GPU) usage volume / accelerator load, I / O load, etc. are observed. The I / O load may be the average, maximum value, and variance of the total data volume and speed.
[0018] The I / O load may be acquired by classifying it into between CPU 101 - storage 103, between CPU 101 - network 102, between CPU 101 - accelerator 104, and between accelerator 104 - storage 103.
[0019] CPU 101 and accelerator 104 may incorporate memory.
[0020] FIG. 3 is a block diagram for explaining the operation of determining the replica position in the storage system 100 as an embodiment.
[0021] The storage system 100 includes a management node 1, a plurality (three in the illustrated example) of compute nodes 2, and a plurality (three in the illustrated example) of storage nodes 3.
[0022] The management node 1 is an example of a management device. When it receives a workload execution request, it collects workload load information 131, node resource information 132, compute node load information 133, accelerator load information 134, storage node load information 135, and volume placement information 136. Then, the management node 1 schedules the workload (WL) 210 and the accelerator (ACC) 220 to determine the placement of the workload 210 and the replica location of the used volume.
[0023] The workload load information 131 indicates the load caused by the workload 210 executed on the compute node 2.
[0024] The node resource information 132 is static information indicating the extent to which each node has memory, CPU, accelerator 220, etc.
[0025] The compute node load information 133 indicates the load of the CPU, memory (MEM), and network (NET) in the compute node 2.
[0026] The accelerator load information 134 indicates the load of the accelerator 220 in the compute node 2.
[0027] The storage node load information 135 indicates the load of the disk 13 in the storage node 3.
[0028] The volume placement information 136 indicates which volumes exist in each storage node 3.
[0029] The management node 1 allows the scheduler 110, which will be described later with reference to FIG. 4, to grasp the resource utilization status. When the workload 210 (or equivalently, the container) is started, the compute node 2 (or equivalently, the CPU node and the accelerator node) and the storage node 3 are determined based on the load requirements and resource utilization status of the workload 210. In the determination, the selection of the storage node 3 is dynamically controlled according to the load status, taking into account that the increase in I / O load does not exceed the network slack (or equivalently, the margin) of the placement node.
[0030] FIG. 4 is a block diagram schematically showing a configuration example of the storage system 100 shown in FIG. 3.
[0031] The storage system 100 is, for example, a cluster system, and includes a management node 1, a plurality (two in the example shown in FIG. 4) of compute nodes 2, and a plurality (two in the example shown in FIG. 4) of storage nodes 3. The management node 1, the plurality of compute nodes 2, and the plurality of storage nodes 3 are connected via a network 170.
[0032] The management node 1 is an example of a management device, and includes a scheduler 110, information 130, and a Network Interface Card (NIC) 17. The scheduler 110 determines the placement of the workload 210 in the compute node 2 and the disk 13 in the storage node 3. The information 130 includes the workload load information 131, node resource information 132, compute node load information 133, accelerator load information 134, storage node load information 135, and volume placement information 136 shown in FIG. 3. The NIC 17 connects the management node 1 to the network 170.
[0033] Each compute node 2 is an example of a server device and includes a plurality (three in the example shown in FIG. 4) of workloads 210 and a NIC 17. The workload 210 is arranged by the management node 1 and is executed to access the data of the storage node 3. The NIC 17 connects the compute node 2 to the network 170.
[0034] Each storage node 3 is an example of a storage device and includes a plurality (three in the example shown in FIG. 4) of disks 13 and a NIC 17. The disk 13 is a storage device that stores data to be accessed from the compute node 2. The NIC 17 connects the storage node 3 to the network 170.
[0035] FIG. 5 is a block diagram schematically showing a hardware configuration example of the information processing apparatus 10 as an embodiment.
[0036] The hardware configuration example of the information processing apparatus 10 shown in FIG. 5 shows the hardware configuration examples of the management node 1, the compute node 2, and the storage node 3 shown in FIG. 4, respectively.
[0037] The information processing apparatus 10 includes a processor 11, a Random Access Memory (RAM) 12, a disk 13, a graphic Interface (I / F) 14, an input I / F 15, a storage I / F 16, and a network 1 / F 17.
[0038] The processor 11 is, by way of example, a processing device that performs various controls and calculations, and realizes various functions by executing an Operating System (OS) and programs stored in the RAM 12.
[0039] Note that a program for realizing the functions of the processor 11 may be provided in a form recorded on a computer-readable recording medium such as a flexible disk, CD (CD-ROM, CD-R, CD-RW, etc.), DVD (DVD-ROM, DVD-RAM, DVD-R, DVD+R, DVD-RW, DVD+RW, HD DVD, etc.), Blu-ray disk, magnetic disk, optical disk, magneto-optical disk, etc. Then, a computer (processor 11 in this embodiment) may read the program from the above-described recording medium via a reading device (not shown), transfer it to an internal recording device or an external recording device, store it, and use it. Further, the program may be recorded in a storage device (recording medium) such as a magnetic disk, optical disk, magneto-optical disk, etc., and provided to the computer via a communication path from the storage device.
[0040] When realizing the functions of the processor 11, a program stored in the internal storage device (RAM 12 in this embodiment) may be executed by a computer (processor 11 in this embodiment). Also, the computer may read and execute a program recorded on a recording medium.
[0041] The processor 11 controls the operation of the entire information processing apparatus 10. The processor 11 may be a multi-processor. The processor 11 may be, for example, any one of a Central Processing Unit (CPU), Micro Processing Unit (MPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Programmable Logic Device (PLD), and Field Programmable Gate Array (FPGA). Also, the processor 11 may be a combination of two or more types of elements among the CPU, MPU, DSP, ASIC, PLD, and FPGA.
[0042] The RAM 12 may be, for example, a Dynamic RAM (DRAM). The software program of the RAM 12 may be appropriately loaded into and executed by the processor 11. Also, the RAM 12 may be used as a primary recording memory or a working memory.
[0043] The disk 13 is, by way of example, a device that can read and write data, and for example, a Hard Disk Drive (HDD), a Solid State Drive (SSD), or a Storage Class Memory (SCM) may be used.
[0044] The graphic I / F 14 outputs video to the display device 140. The display device 140 is, for example, a liquid crystal display, an Organic Light-Emitting Diode (OLED) display, a Cathode Ray Tube (CRT), an electronic paper display, etc., and displays various types of information for an operator or the like.
[0045] The input I / F 15 receives the input of data from the input device 150. The input device 150 is, for example, a mouse, a trackball, or a keyboard, and through this input device 150, an operator performs various input operations. The input device 150 and the display device 140 may be combined, for example, a touch panel.
[0046] The storage I / F 16 performs input / output of data to and from the medium reader 160. The medium reader 160 is configured such that a recording medium can be attached. The medium reader 160 is configured to be able to read the information recorded on the recording medium when the recording medium is attached. In this example, the recording medium has portability. For example, the recording medium is a flexible disk, an optical disk, a magnetic disk, a magneto-optical disk, or a semiconductor memory, etc.
[0047] Network I / F17 connects the information processing apparatus 10 to the network 170 and is an interface device for communicating with other information processing apparatuses 10 (in other words, the management node 1, the compute node 2, or the storage node 3) or an external device (not shown) via this network 170. As the network I / F17, for example, various interface cards corresponding to the standards of the network 170 such as a wired Local Area Network (LAN), a wireless LAN, or a Wireless Wide Area Network (WWAN) can be used.
[0048] FIG. 6 is a block diagram schematically showing an example of the software configuration of the management node 1 shown in FIG. 3.
[0049] The management node 1 functions as a scheduler 110 and an information exchange unit 111.
[0050] The information exchange unit 111 acquires workload load information 131, compute node load information 133, storage node load information 135, and accelerator load information 134 from other nodes (in other words, the compute node 2 and the storage node 3) via the network 170.
[0051] In other words, the information exchange unit 111 acquires the workload load information 131 and the system load information (in other words, the compute node load information 133, the accelerator load information 134, and the storage node load information 135) when the container is executed.
[0052] Based on the node resource information 132, the volume placement information 136, the workload load information 131, the compute node load information 133, the storage node load information 135, and the accelerator load information 134 acquired by the information exchange unit 111, the scheduler 110 determines the placement of the workload 210.
[0053] In other words, when starting the workload 210, the scheduler 110 determines the destination of the workload 210 and the replica position of the volume based on the workload load information 131 and the system load information.
[0054] The scheduler 110 may select a first compute node 2 from among a plurality of compute nodes 2, where the sum of the traffic between the processor 11 and the network 170, the traffic between the processor 11 and the volume, and the traffic between the accelerator 220 and the volume is equal to or less than the margin in the network 170. Further, the scheduler 110 may select a second compute node 2 from among a plurality of compute nodes 2, where the sum of the traffic between the processor 11 and the network 170, the traffic between the processor 11 and the volume, and the traffic between the processor 11 and the accelerator 220 is equal to or less than the margin in the network 170. Then, the scheduler 110 may determine the first compute node 2 or the second compute node 2 as the destination of the workload 210.
[0055] The scheduler 110 may select one or more first storage nodes 3 from among a plurality of storage nodes 3 provided in the storage system 100, where the sum of the traffic between the processor 11 and the volume and the traffic between the accelerator 220 and the volume is equal to or less than the margin in the network 170. Then, the scheduler 110 may determine the one or more first storage nodes 3 as the replica position.
[0056] The scheduler 110 may determine the replica position when the load bias among the plurality of storage nodes 3 provided in the storage system 100 exceeds a threshold.
[0057] FIG. 7 is a block diagram schematically showing an example of the software configuration of the compute node 2 shown in FIG. 3.
[0058] Compute node 2 includes workload deployment unit 211, information exchange unit 212, and load information acquisition unit 213 as agents.
[0059] Load information acquisition unit 213 acquires load information 230 including workload load information 131, compute node load information 133, and accelerator load information 134 shown in FIG. 6 from OS 20.
[0060] Information exchange unit 212 transmits load information 230 acquired by load information acquisition unit 213 to management node 1 via Virtual Switch (VSW) 214 and network 170.
[0061] Workload deployment unit 211 deploys workload (WL) 210 based on the decision in management node 1.
[0062] FIG. 8 is a block diagram schematically showing a software configuration example of storage node 3 shown in FIG. 3.
[0063] Storage node 3 includes information exchange unit 311 and load information acquisition unit 312 as agents.
[0064] Load information acquisition unit 312 acquires load information 330 including storage node load information 135 shown in FIG. 6 from OS 30.
[0065] Information exchange unit 311 transmits load information 330 acquired by load information acquisition unit 312 to management node 1 via VSW 313 and network 170.
[0066] 〔A-2〕Operation example The replica position determination process in the embodiment will be described according to the flowchart (steps S1 to S5) shown in FIG. 9.
[0067] The management node 1 arranges the CPU and the accelerator 220 (step S1). The details of the arrangement process of the CPU and the accelerator 220 will be described later with reference to FIG. 10.
[0068] The management node 1 determines whether the CPU and the accelerator 220 can be arranged (step S2).
[0069] If the arrangement cannot be made (refer to the NO route in step S2), the process proceeds to step S5.
[0070] On the other hand, if the arrangement can be made (refer to the YES route in step S2), the management node 1 arranges the storage (step S3). The storage arrangement process will be described later with reference to FIG. 11.
[0071] The management node 1 determines whether the storage can be arranged (step S4).
[0072] If the arrangement cannot be made (refer to the NO route in step S4), the management node 1 sets the workload 210 to the standby state (step S5). Then, the replica position determination process ends.
[0073] On the other hand, if the arrangement can be made (refer to the YES route in step S4), the replica position determination process ends.
[0074] Next, the details of the arrangement process of the CPU and the accelerator 220 shown in FIG. 9 will be described according to the flowchart (steps S11 to S17) shown in FIG. 10.
[0075] The management node 1 sets the set of nodes that satisfy the requirements of the CPU and the memory (MEM) as X, the set of nodes that satisfy the requirements of the accelerator (ACC) 220 as Y, and the intersection X ∩ Y of X and Y as Z (step S11).
[0076] The management node 1 selects one node in the set Z that satisfies the network requirement "CPU_NET + CPU_VOL + ACC_VOL <= network slack" (step S12). Here, CPU_NET indicates the communication volume between the CPU and the network, CPU_VOL indicates the communication volume between the CPU and the volume, and ACC_VOL indicates the communication volume between the accelerator 220 and the volume. Also, the network slack indicates the margin of the network volume of a certain node.
[0077] The management node 1 determines whether a node that satisfies the network requirement is found (step S13).
[0078] If a node that satisfies the network requirement is found (refer to the YES route in step S13), it is considered deployable, and the placement process of the CPU and the accelerator 220 ends.
[0079] On the other hand, if no node that satisfies the network requirement is found (refer to the NO route in step S13), the management node 1 selects one node in the set X that satisfies the network requirement "CPU_NET + CPU_VOL + CPU_ACC <= network slack" (step S14). Here, CPU_NET indicates the communication volume between the CPU and the network, CPU_VOL indicates the communication volume between the CPU and the volume, and CPU_ACC indicates the communication volume between the CPU and the accelerator 220. Also, the network slack indicates the margin of the network volume of a certain node.
[0080] The management node 1 determines whether a node that satisfies the network requirement is found (step S15).
[0081] If no node that satisfies the network requirement is found (refer to the NO route in step S15), it is considered non-deployable, and the placement process of the CPU and the accelerator 220 ends.
[0082] On the one hand, when a node that meets the network requirements is found (refer to the YES route in step S15), management node 1 selects one node from set Y that meets the network requirement "ACC_VOL + CPU_ACC <= network slack" (step S16). Note that ACC_VOL indicates the communication volume between the accelerator 220 and the volume, and CPU_ACC indicates the communication volume between the CPU and the accelerator 220. Also, network slack indicates the margin of the network volume of a certain node.
[0083] Management node 1 determines whether a node that meets the network requirements is found (step S17).
[0084] If a node that meets the network requirements is found (refer to the YES route in step S17), it is considered that the placement is possible, and the placement process of the CPU and the accelerator 220 ends.
[0085] On the other hand, if a node that meets the network requirements is not found (refer to the NO route in step S17), it is considered that the placement is impossible, and the placement process of the CPU and the accelerator 220 ends.
[0086] Next, the details of the storage placement process shown in FIG. 9 will be described according to the flowchart shown in FIG. 11 (steps S21 to S26).
[0087] Management node 1 sets the set of storage nodes 3 having replicas of the volume as V (step S21).
[0088] Management node 1 selects one node from set V that meets the network requirement "CPU_VOL + ACC_VOL <= network slack" (step S22). Note that CPU_VOL indicates the communication volume between the CPU and the volume, and ACC_VOL indicates the communication volume between the accelerator 220 and the volume. Also, network slack indicates the margin of the network volume of a certain node.
[0089] The management node 1 determines whether a node that satisfies the network requirements has been found (step S23).
[0090] If a node that satisfies the network requirements is found (refer to the YES route in step S23), it is considered that placement is possible, and the storage placement process ends.
[0091] On the other hand, if no node that satisfies the network requirements is found (refer to the NO route in step S23), the management node 1 selects a plurality of nodes from the set V that satisfy the network requirement "CPU_VOL + ACC_VOL <= network slack" by combining them (step S24). Here, CPU_VOL indicates the communication volume between the CPU and the volume, ACC_VOL indicates the communication volume between the accelerator 220 and the volume, and network slack indicates the margin of the network volume of a certain node.
[0092] The management node 1 places the volume among the selected nodes (step S25).
[0093] The management node 1 determines whether a node that satisfies the network requirements has been found (step S26).
[0094] If a node that satisfies the network requirements is found (refer to the YES route in step S26), it is considered that placement is possible, and the storage placement process ends.
[0095] On the other hand, if no node that satisfies the network requirements is found (refer to the NO route in step S26), it is considered that placement is impossible, and the storage placement process ends.
[0096] FIG. 12 is a table for explaining the bandwidth target value in the storage placement process shown in FIG. 11.
[0097] There are three usage volumes V1, V2, V3, and it is assumed that the replica is in storage node 3. In the example shown in FIG. 12, the replicas of volume V1 are arranged in storage nodes #1 and #2, the replicas of volume V2 are arranged in storage nodes #2 and #3, and the replicas of volume V3 are arranged in storage nodes #1 and #3. Also, let the loads for each of volumes V1, V2, V3 be R1, R2, R3 respectively.
[0098] Let the network Slack for each storage node 3 be S1, S2, S3.
[0099] The storage allocation is performed according to the following steps (1) to (4).
[0100] (1) Allocate R 11 , R 31 to the maximum extent without exceeding S1. At this time, prioritize R 11 . Note that R 11 = min(S1, R1), and R 31 = min(S1 - R 11 , R3). R 11 + R 31 ≤ S1, R 11 ≤ R1, R 31 ≤ R3
[0101] (2) Allocate R 12 , R 22 to the maximum extent without exceeding S2. At this time, prioritize R 12 . R 12 + R 21 ≤ R2, R 12 = R1 - R 11 , R 22 ≤ R2
[0102] (3) Allocate R 23 , R 33 within the range not exceeding S3. R 23 + R 33 ≤ S3, R 23 = R2 - R22 , R 33 = R3 - R 31
[0103] If none of the above (1) to (3) can be satisfied, it becomes impossible to allocate.
[0104] Then, access to the volume is controlled so that the bandwidth target value is satisfied. When the volume of storage node 3 is composed of N blocks, the access bandwidth R to each block is distributed to R1 and R2 (R = R1 + R2).
[0105] When deploying the workload 210, the corresponding number of blocks is divided into N1 and N2 as follows for two nodes having replicas of the volume.
Equation
[0106] Also, when executing the workload 210, the access bandwidth to the volume is restricted to R at the execution node of the workload 210 that accesses the volume. In a situation where access to each block is performed uniformly, the access bandwidths to each replica are R1 and R2.
[0107] Next, the rebalancing process of storage node 3 in the embodiment will be described according to the flowchart shown in FIG. 13 (steps S31, S32).
[0108] The management node 1 determines at regular intervals whether the load imbalance of each storage node 3 exceeds a threshold value (step S31).
[0109] If the load imbalance of each storage node 3 does not exceed the threshold value (see the NO route in step S31). The process in step S31 is repeatedly executed.
[0110] On the other hand, when the load bias of each storage node 3 exceeds the threshold (see the YES route in step S31), the management node 1 performs load balance selection for the storage node 3 (step S32). As a result, it is possible to prevent the load on a specific storage node 3 from increasing and the performance from degrading, and by reducing the load bias, it is possible to achieve fair resource allocation to the workload. Then, the balance processing of the storage node 3 ends.
[0111] The average and difference d and variance D of the network Slack are obtained by the following formula, and balance is implemented when the variance D exceeds the threshold t.
[0112]
Equation
[0113] Then, the balance is performed according to the following steps (1) to (4).
[0114] (1) Define the following sets G and L.
Equation
[0115] (2) Extract the set V of volumes to which replicas belong in both sets G and L.
[0116] (3) Select one volume from the set V and move the load assignment from the set G to the set L.
[0117] (4) Repeat until the difference in bandwidth of the volumes belonging to the sets G and L becomes below a certain value or there are no more candidate volumes for movement.
[0118] 〔B〕Effect According to the management node 1, the storage system 100, and the information processing method in an example of the above-described embodiment, for example, the following operational effects can be achieved.
[0119] When the management node 1 executes a container, it acquires the workload information 131 and the system load information (in other words, the compute node load information 133, the accelerator load information 134, and the storage node load information 135). When the workload 210 is started, the management node 1 determines the placement destination of the workload 210 and the replica position of the volume based on the workload information 131 and the system load information.
[0120] Thereby, the load in the storage system 100 can be dispersed to improve the throughput. Specifically, the resources of the cluster, including communication and storage, can be effectively utilized. Therefore, more applications can be executed on the same system.
[0121] The management node 1 selects a first compute node 2 from among a plurality of compute nodes 2, where the sum of the communication volume between the processor 11 and the network 170, the communication volume between the processor 11 and the volume, and the communication volume between the accelerator 220 and the volume is equal to or less than the margin in the network 170. The management node 1 selects a second compute node 2 from among a plurality of compute nodes 2, where the sum of the communication volume between the processor 11 and the network 170, the communication volume between the processor 11 and the volume, and the communication volume between the processor 11 and the accelerator 220 is equal to or less than the margin in the network 170. The management node 1 determines the first compute node 2 or the second compute node 2 as the placement destination of the workload 210.
[0122] Thereby, an appropriate compute node 2 can be selected as the placement destination of the workload 210.
[0123] The management node 1 selects, from among a plurality of storage nodes 3, one or more first storage nodes 3 such that the sum of the traffic volume between the processor 11 and the volume and the traffic volume between the accelerator 220 and the volume is equal to or less than the margin in the network 170. The management node 1 determines one or more first storage nodes 3 as replica positions.
[0124] Thereby, an appropriate storage node 3 can be selected as the replica position of the volume.
[0125] The management node 1 determines the replica position when the load imbalance among the plurality of storage nodes 3 provided in the storage system 100 exceeds a threshold value.
[0126] Thereby, it is possible to prevent the load on a specific storage node 3 from increasing and the performance from deteriorating, and by reducing the load imbalance, it is possible to achieve a fair resource allocation to the workload.
[0127] 〔C〕Others The disclosed technology is not limited to the above-described embodiments, and can be implemented with various modifications without departing from the spirit of the present embodiment. Each configuration and each process of the present embodiment can be selectively adopted as necessary, or can be appropriately combined.
[0128] 〔D〕Supplementary Note Regarding the above embodiments, the following supplementary notes are further disclosed.
[0129] (Supplementary Note 1) A management device for a storage system, acquires workload load information and system load information when a container is executed, and determines a workload placement destination and a replica position of a volume based on the workload load information and the system load information when the workload is started. Management device.
[0130] (Supplementary Note 2) The storage system includes a plurality of server devices and a plurality of storage devices that are connected to each other via a network. Each of the plurality of server devices includes a processor and an accelerator. Each of the plurality of storage devices includes a volume. Select a first server device from the plurality of server devices, where the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the accelerator and the volume is less than or equal to the margin in the network. Select a second server device from the plurality of server devices, where the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the processor and the accelerator is less than or equal to the margin in the network. Determine the first server device or the second server device as the workload placement destination. The management device according to Supplementary Note 1.
[0131] (Supplementary Note 3) The storage system includes a plurality of server devices and a plurality of storage devices that are connected to each other via a network. Each of the plurality of server devices includes a processor and an accelerator. Each of the plurality of storage devices includes a volume. Select one or more first storage devices from the plurality of storage devices, where the sum of the traffic between the processor and the volume and the traffic between the accelerator and the volume is less than or equal to the margin in the network. Determine the one or more first storage devices as the replica location. The management device according to Supplementary Note 1 or 2.
[0132] (Supplementary Note 4) When the load bias among a plurality of storage devices provided in the storage system exceeds a threshold value, the replica position is determined. The management device according to any one of Appendices 1 to 3.
[0133] (Appendix 5) A storage system including a management device, a server device, and a storage device, The server device transmits system load information in the server device to the management device, The storage device transmits system load information in the storage device to the management device, The management device When a container is executed, workload load information and system load information transmitted from the server device and the storage device are acquired, When a workload is started, based on the workload load information and the system load information, the determination of the workload placement destination and the replica position of the volume is performed. Storage system.
[0134] (Appendix 6) The server device and the storage device are connected to each other via a network, Each of the plurality of server devices includes a processor and an accelerator, Each of the plurality of storage devices includes a volume, A first server device in which the sum of the communication volume between the processor and the network, the communication volume between the processor and the volume, and the communication volume between the accelerator and the volume is equal to or less than the margin in the network is selected from among the plurality of server devices, A second server device in which the sum of the communication volume between the processor and the network, the communication volume between the processor and the volume, and the communication volume between the processor and the accelerator is equal to or less than the margin in the network is selected from among the plurality of server devices, Determine the first server device or the second server device as the workload placement destination. The storage system according to appended note 5.
[0135] (Appended note 7) The server device and the storage device are connected to each other via a network. Each of the plurality of server devices includes a processor and an accelerator. Each of the plurality of storage devices includes a volume. Select, from among the plurality of storage devices, one or more first storage devices for which the sum of the communication volume between the processor and the volume and the communication volume between the accelerator and the volume is equal to or less than the margin in the network. Determine the one or more first storage devices as the replica positions. The storage system according to appended note 5 or 6.
[0136] (Appended note 8) When the load bias among the plurality of storage devices exceeds a threshold, determine the replica positions. The storage system according to any one of appended notes 5 to 7.
[0137] (Appended note 9) An information processing method in a storage system including a management device, a server device, and a storage device, The server device transmits system load information in the server device to the management device. The storage device transmits system load information in the storage device to the management device. The management device When a container is executed, obtain workload load information and system load information transmitted from the server device and the storage device. When a workload is started, determine a workload placement destination and a replica position of a volume based on the workload load information and the system load information. Information processing method
[0138] (Appendix 10) The server device and the storage device are connected to each other via a network, Each of the plurality of server devices includes a processor and an accelerator, Each of the plurality of storage devices includes a volume, Select a first server device from among the plurality of server devices, wherein the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the accelerator and the volume is equal to or less than the margin in the network, Select a second server device from among the plurality of server devices, wherein the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the processor and the accelerator is equal to or less than the margin in the network, Determine the first server device or the second server device as the workload placement destination, The information processing method according to Appendix 9
[0139] (Appendix 11) The server device and the storage device are connected to each other via a network, Each of the plurality of server devices includes a processor and an accelerator, Each of the plurality of storage devices includes a volume, Select one or more first storage devices from among the plurality of storage devices, wherein the sum of the traffic between the processor and the volume and the traffic between the accelerator and the volume is equal to or less than the margin in the network, Determine the one or more first storage devices as the replica location, The information processing method according to Appendix 9 or 10
[0140] (Appendix 12) When the load bias among the plurality of storage devices exceeds a threshold value, determine the replica position. The information processing method according to any one of Appendices 9 to 11.
Explanation of Signs
[0141] 1: Management node 2: Compute node 3: Storage node 10: Information processing device 11: Processor 12: RAM 13: Disk 14: Graphics I / F 15: Input I / F 16: Storage I / F 17: Network I / F 100: Storage system 101: CPU 102, 170: Network 103: Storage 104, 220: Accelerator 110: Scheduler 111: Information exchange unit 130: Information 131: Workload information 132: Node resource information 133: Compute node load information 134: Accelerator load information 135: Storage node load information 136: Volume placement information 140: Display device 150: Input device 160: Media reader 210: Workload 211: Workload deployment unit 212: Information exchange unit 213: Load information acquisition unit 214, 313: VSW 230: Load information 311: Information Exchange Unit 312: Load Information Acquisition Unit 330: Load Information
Claims
A management device for a storage system including a plurality of server devices and a plurality of storage devices connected to each other via a network, wherein: each of the plurality of server devices includes a processor and an accelerator; each of the plurality of storage devices includes a volume; when a container is executed, select a first server device from among the plurality of server devices, wherein the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the accelerator and the volume is equal to or less than the margin in the network; select a second server device from among the plurality of server devices, wherein the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the processor and the accelerator is equal to or less than the margin in the network; determine the first server device or the second server device as a workload placement destination; A management device. According to claim 1, select one or more first storage devices from among the plurality of storage devices, wherein the sum of the traffic between the processor and the volume and the traffic between the accelerator and the volume is equal to or less than the margin in the network; determine the one or more first storage devices as replica positions; The management device according to claim 1.
3. When the load bias among the plurality of storage devices provided in the storage system exceeds a threshold value, determine the replica position; The management device according to claim 1 or 2. A storage system including a management device, a plurality of server devices, and a plurality of storage devices connected to each other via a network, wherein: each of the plurality of server devices includes a processor and an accelerator; each of the plurality of storage devices includes a volume; the management device, when a container is executed, select a first server device from among the plurality of server devices, wherein the sum of the traffic between the processor and the network, the traffic between the processor and the volume, and the traffic between the accelerator and the volume is equal to or less than the margin in the network; Select, from among the plurality of server devices, a second server device in which the sum of the traffic volume between the processor and the network, the traffic volume between the processor and the volume, and the traffic volume between the processor and the accelerator is equal to or less than the margin in the network. Determine the first server device or the second server device as a workload placement destination. Storage system.
5. An information processing method in a storage system including a management device, a plurality of server devices, and a plurality of storage devices connected to each other via a network, Each of the plurality of server devices includes a processor and an accelerator. Each of the plurality of storage devices includes a volume. When a container is executed, the management device selects, from among the plurality of server devices, a first server device in which the sum of the traffic volume between the processor and the network, the traffic volume between the processor and the volume, and the traffic volume between the accelerator and the volume is equal to or less than the margin in the network. Select, from among the plurality of server devices, a second server device in which the sum of the traffic volume between the processor and the network, the traffic volume between the processor and the volume, and the traffic volume between the processor and the accelerator is equal to or less than the margin in the network. Determine the first server device or the second server device as a workload placement destination. Information processing method.
Citation Information
Patent Citations
Resource load balancing
JP2016528617A
Volume arrangement management apparatus, volume arrangement management method and volume arrangement management program
JP2020038421A
Method to determine disposition for VM / container and volume in HCI environment, and storage system
JP2020052730A
Computer system and data management method
JP2020154587A
Management device, information processing system and management program
JP2021002125A