A Dynamic Resource Scheduling Method for GPU Clusters

By building a resource-time and resource-performance model and optimizing GPU cluster resource scheduling with the migration mechanism, the problem of low resource utilization rate and task completion time does not meet user requirements in heterogeneous bandwidth environments is solved, and efficient resource utilization and deadline guarantee is achieved.

CN114647515BActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210382828.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-07-04
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

Existing GPU cluster schedulers cannot effectively combine resource configuration, resource layout and deadline requirements in heterogeneous bandwidth environments, resulting in low resource utilization and task completion time not meeting user requirements.

Method used

Build a resource-time model and resource-performance model, combine the optimal resource scheme and migration mechanism of the task, and dynamically schedule tasks to the GPU cluster to optimize resource utilization and meet deadline requirements.

Benefits of technology

It effectively reduces the completion time of deep learning training tasks, improves the resource utilization rate and deadline guarantee rate of the GPU cluster, and improves the working efficiency of the GPU cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114647515B_ABST
    Figure CN114647515B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic resource scheduling method for a GPU cluster. A resource-time model and a resource-performance model are constructed; dynamic resource scheme decision-making for distributed deep learning tasks is carried out; physical resource node allocation is performed according to the optimal scheme of the tasks; before each execution of the task scheduling process by the dynamic resource scheduling algorithm, the situation of the running tasks will be analyzed to decide whether to perform resource migration: the scheduler executes the scheduling algorithm to select new tasks to run on the GPU cluster. The present invention comprehensively considers the completion time of the task itself and the user's deadline for completion, and can dynamically schedule the GPU in real time according to the load situation of the GPU cluster and the running situation of the tasks, effectively reducing the completion time of the deep learning training tasks, maximizing the deadline guarantee rate and effectively improving the working efficiency of the GPU cluster and the resource utilization rate of the GPU cluster nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dynamic resource scheduling method for a GPU cluster. On the GPU cluster, multiple GPU devices are used to perform parallel training on a deep neural network (DNN) model by means of distributed deep learning technology, thereby accelerating the training process. Background Art

[0002] Deep learning technology has been applied to numerous business scenarios in the past few years. R & D personnel construct DNN models according to the target characteristics of business scenarios and repeatedly train them on specific data sets until the model accuracy is maintained at an expected level, so as to improve the complexity of business scenarios. At this time, DNN models with more complex structures and increasing numbers of layers are required to obtain higher accuracy. At the same time, the scale of data sets is also constantly growing, resulting in a large amount of time required to train a DNN model.

[0003] Therefore, the academic and industrial communities need to construct distributed deep learning tasks through distributed parallel computing, and use multiple GPU devices on the GPU cluster to train the DNN model simultaneously to accelerate the training process. Current mainstream machine learning frameworks, such as PyTorch, TensorFlow, etc., all provide complete technical support for distributed deep learning.

[0004] Most enterprises and universities usually purchase multiple GPU devices to form a small and medium-sized GPU cluster to run distributed deep learning tasks of multiple users. They currently use existing GPU cluster schedulers, such as Yarn, Mesos, and Kubernetes, which do not provide good scheduling support for distributed deep learning tasks, resulting in improper resource allocation, low operating efficiency, and inability to meet user requirements. In a GPU cluster using Yarn for resource management in a certain laboratory, wireless broadband technology and Ethernet are respectively used to interconnect GPU devices within the same rack and across racks. Due to the bandwidth differences between GPU devices, different resource layout methods will result in different training efficiencies of the DNN model. The historical scheduling logs on this GPU cluster show that the average resource utilization rate of this GPU cluster is only 50%. In addition, for GPU cluster users, the task deadline is a key indicator to measure user satisfaction. In most cases, users can accept tasks completed before the deadline, while when the task end time exceeds the deadline, users' satisfaction with the performance of the GPU cluster will drop significantly.

[0005] Many experts and scholars have conducted research on resource scheduling in GPU clusters for different optimization metrics. The existing related work mainly focuses on the resource scheduling process from two aspects: reducing task completion time and improving the performance metrics of GPU clusters. Although the existing related work can effectively solve the resource scheduling problem of GPU clusters, there is less research work that attempts to combine resource allocation, resource layout, and deadline requirements. However, there are still certain limitations in maximizing the deadline guarantee rate in heterogeneous bandwidth environments. Summary of the Invention

[0006] An object of the present invention is to propose a dynamic scheduling method for GPU clusters aiming at the multi-task scheduling problem with deadline requirements in heterogeneous bandwidth environments, combining resource configuration, resource layout, and deadline requirements to obtain scheduling decisions, maximizing the deadline guarantee rate and improving the resource utilization rate of GPU cluster nodes.

[0007] The method of the present invention includes the following steps:

[0008] A dynamic resource scheduling method for GPU clusters includes the following steps:

[0009] Step (1), based on the iterative characteristics of the DNN model under the Ring-Allreduce communication architecture of distributed machine learning and the bandwidth differences between GPU devices, construct a resource-time model;

[0010] Step (2), construct a resource-performance model based on the resource quantity used in the resource scheme, task running time, and task deadline;

[0011] Step (3), based on steps (1) and (2), make dynamic resource scheme decisions for distributed deep learning tasks;

[0012] Step (4), based on step (3), perform physical resource node allocation according to the optimal scheme of the task ;

[0013] Step (5), before each execution of the task scheduling process by the dynamic resource scheduling algorithm, analyze the situation of the tasks that have been run and decide whether to perform resource migration:

[0014] Step (6), the scheduler executes the scheduling algorithm to select new tasks to run on the GPU cluster.

[0015] Another object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the above method.

[0016] Another object of the present invention is to provide a computing device, including a memory and a processor, wherein an executable code is stored in the memory, and when the processor executes the executable code, the above method is implemented.

[0017] Advantages of the present invention:

[0018] 1) The running time of a task under different resource scenarios is obtained by using a time model. Secondly, a performance model is used to guide the generation of an optimal resource scenario for a distributed deep learning task. Then, tasks are scheduled based on the principle of the nearest deadline. Finally, resource allocation is performed to determine the physical resource locations of the resource scenario, and a running scenario containing GPU cluster node numbers and the number of GPUs is generated. The resource scheduling process is completed by starting a task running script on a GPU cluster server with the help of a machine learning framework application interface.

[0019] 2) A migration mechanism is introduced to reduce the impact of resource fragmentation scenarios that occur during the scheduling process.

[0020] 3) The Ring-Allreduce communication architecture can effectively reduce the communication time required for the parameter synchronization phase.

[0021] 4) A resource-performance model is used to screen available resource scenarios, and the resource scenarios that can be effectively used for distributed training are retained, reducing the phenomenon of resource waste.

[0022] Compared with the prior art, the present invention comprehensively considers the completion time of the task itself and the user's deadline for completion, and can dynamically schedule the work of GPUs in real time according to the GPU cluster load situation and task running situation, effectively reducing the completion time of deep learning training tasks, maximizing the deadline guarantee rate, and effectively improving the working efficiency of the GPU cluster and the resource utilization rate of GPU cluster nodes.

[0023] Definition of related concepts and symbol explanations

[0024] c free : represents the number of idle GPUs;

[0025] T run : represents the actual running time under a certain resource scenario;

[0026] T cal : represents the computing time on a single GPU device;

[0027] l s : represents a list of single-node resource scenarios;

[0028] l m : represents a list of cross-node resource scenarios. Description of the drawings

[0029] Figure 1 : Schematic diagram of Ring-Allreduce communication architecture.

[0030] Figure 2 : Schematic diagram of resource configuration decision-making.

[0031] Figure 3 : Schematic diagram of physical resource allocation.

[0032] Figure 4 : Schematic diagram of resource migration mechanism.

[0033] Figure 5 : Schematic diagram of dynamic resource scheduling method. Specific implementation manners

[0034] The following further analyzes the present invention with reference to the accompanying drawings.

[0035] The dynamic resource scheduling method for a GPU cluster according to the present invention includes the following steps:

[0036] A dynamic resource scheduling method for a GPU cluster, characterized by including the following steps:

[0037] Step (1), based on the iterative characteristics of the DNN model under the Ring-Allreduce communication architecture for distributed machine learning and the bandwidth differences between GPU devices, construct a resource-time model:

[0038] The Ring-Allreduce communication architecture includes several nodes of a GPU cluster that communicate with each other. Each node includes multiple CPUs and GPUs. The GPU devices on the same node communicate via the Peripheral Component Interconnect Express (PCIe) and QuickPath Interconnect (QPI) (where PCIe communication is used between GPUs, PCIe communication is used between GPUs and CPUs, and QPI communication is used between CPUs), and the nodes in the GPU cluster communicate with each other via the InfiniBand wireless broadband technology;

[0039] The resource-time model includes the following:

[0040] (1.1) The actual running time T of the distributed deep learning task under a certain resource plan run is expressed as follows:

[0041] T run = T step × N step × N epoch Equation (1)

[0042] where T step is the time spent on training a dataset with a batch size of the DNN model, and Nstep is the number of datasets of a batch size that can be input in one iteration of the DNN model, N epoch represents the iteration round;

[0043] (1.2)T step consists of the computing time T on a single CPU device cal , the communication time T between CPU and CPU devices comm and is calculated as follows:

[0044] T step = T cal + T comm Equation (2)

[0045] (1.3)N step will change with the different total number of GPUs included in the resource plan. The more the number, the smaller N step will be correspondingly reduced; N step , the size S of the DNN model training dataset dataset , the batch size S batch and the total number of GPUs N GPU have the following relationship during the distributed data parallel training process:

[0046]

[0047] where N GPU is obtained by accumulating c on each node of the resource plan used , and c used represents the number of GPUs used by the training task on a single node;

[0048] (1.4) By placing the DNN model on a single GPU device for several batches of iterations and recording the corresponding running time, since there is no multi-device communication, this running time only includes the computing time on a single GPU device, which is expressed as follows:

[0049]

[0050] where T' step is the running time of several iterations, and N' step is the corresponding number of iterations;

[0051] (1.5) If there is no communication time, then the running time of the task and the total number of GPUs included in the resource plan will be in an inverse proportion relationship, that is, as the total number of GPUs increases, the running time of the task will decrease proportionally. When there is communication time, it will lead to a decrease in running efficiency; the communication time T under the Ring-Allreduce communication architecture comm is expressed as follows:

[0052]

[0053] Among them, BW is the bandwidth speed between two GPU devices. If the two GPU devices are on the same node, then BW is the bandwidth between the GPU devices within the node. If the two GPU devices are on different nodes, then BW is the network bandwidth between the nodes.

[0054] Step (2): A resource - performance model is constructed based on the resource quantity used in the resource plan, the task running time, and the task deadline:

[0055] (2.1) Deadline modeling:

[0056] (2.1.1) Assume that the user's deadline requirement for the task consists of the task arrival time, the task priority, and the maximum running time of the task. The maximum running time is the running time of the task on a single GPU device only. Define several task priorities and convert the priorities into the expected running time T of the task exp , and its calculation formula is as follows:

[0057]

[0058] Among them, α corresponds to the task priority, represents the running time of the task on a single GPU device;

[0059] (2.1.2) Assume that the arrival time and the running start time of the task are T arr and T start respectively. Then the deadline T dl and the running end time T end of the task can be respectively expressed as:

[0060] T dl = T arr + T exp Equation (7)

[0061] T end = T start + T run Equation (8)

[0062] (2.1.3) When the deadline T dl and the running end time T end of the task satisfy the following Equation (9), it means that the task meets the user's deadline requirement when it ends:

[0063] T end < T dl Equation (9)

[0064] (2.2) When all the GPU devices held by the resource plan are located on the same node, the bandwidth speed is the direct connection bandwidth between the GPU devices. When the GPU devices held by the resource plan are located on different nodes, the bandwidth speed is the network bandwidth between the nodes. As can be seen from Equation (5), when N GPU and N param remain unchanged, T comm increases as BW decreases. Substitute Equation (2) and Equation (3) into Equation (1), and require that the time of multi-machine distributed training is shorter than the running time of single-machine training, then the following inequality can be obtained:

[0065]

[0066] Among them, the first half and the second half of the inequality are the time for the DNN model to train one iteration on multiple nodes and a single node respectively. Simplifying Equation (10) gives:

[0067] T comm <(N GPU -1)×T cal Equation (11)

[0068] When the DNN model is performing multi-machine distributed training, T comm 、N GPU and T cal can only achieve the purpose of accelerating model training by meeting Equation (10);

[0069] (2.3) To measure the performance of tasks under different resource plans and select the resource plan with the highest operating efficiency among multiple resource plans that meet the deadline requirements, and give full play to the resource performance, the performance formula of the resource-performance model is defined as:

[0070]

[0071] Among them, T dl represents the deadline of the task;

[0072] Step (3). On the basis of steps (1) and (2), the process of making dynamic resource plan decisions for distributed deep learning tasks is as follows Figure 2 :

[0073] Generate a list of available resource plans for each task in the waiting queue based on the cluster idle resources and resource layout, and determine the optimal resource plan for each task according to the resource-performance model and combined with the cluster node load conditions; specifically as follows:

[0074] (3.1) Obtain the resource list R, and set the number of resource nodes with c free >0 as n, c freeIndicates the number of idle GPUs in a single node. The maximum value of c in the resource node free is max(c free ), and its cumulative sum is sum(c free ). Initialize the single-node resource plan list l s and the cross-node resource plan list l m ;

[0075] (3.2) If n = 1, obtain the resource plans R used from 1 to max(c free ) in the resource list R and add them to l t . If n > 1, obtain the resource plans R s from 1 to sum(c used ) in the resource list R and add them to l free ; t m s Calculate T

[0076] and T s and l m in l t for R run and T end in l t according to equations (1) and (8), and filter out some inefficient resource plans R

[0077] (3.3) Obtain the resource plan R s in l with the maximum performance and t as the single-node expected plan and obtain the resource plan R s in l end where T dl > T end and T t is the smallest as the single-node unexpected plan

[0078] Obtain the resource plan R m in l with the maximum performance and t as the cross-node expected plan and obtain the resource plan R m in l end where T dl > T end and T t is the smallest as the cross-node unexpected plan Note that among them and may not exist;

[0079] (3.4) Determine whether the cross-node expected scheme is satisfied Exists and there is 0 < c in the GPU cluster free <N GPU resource nodes. If satisfied, it means that the current task has a cross-node resource scheme to utilize local resources and end the operation within T dl At this time, the optimal resource scheme If not satisfied but exists, then the optimal resource scheme If not satisfied and still does not exist, it means that the current idle resources in the GPU cluster cannot enable the current task to end the operation within T dl and it is considered that there is no expected scheme at this time and available for selection, determine whether the cross-node unexpected scheme is satisfied Exists and there is 0 < c in the GPU cluster free <N GPU resource nodes. If satisfied, it means that the current task has a cross-node resource scheme to utilize local resources and end the operation within T dl At this time, the optimal resource scheme If not satisfied, then the optimal resource scheme

[0080] Step (4), based on the optimal scheme of the task in step (3) Execute the physical resource node allocation process as follows Figure 3 :

[0081] (4.1) Obtain the resource list R and sort it in ascending order according to the c of the nodes free ;

[0082] (4.2) If the optimal resource scheme in step (3) is a single-node expected scheme then traverse the resource list R to find the resource node Node(s, c free ≥ N GPU ) and remove N free GPU devices from this resource node Node(s, c free ), and add Node(s, N GPU ) to used where Node(s, c free ) represents the resource node with serial number s in the GPU cluster and having c free ​A node object with idle GPUs, where s represents the serial number of the node object; Node(s, N used ) represents a node object in the GPU cluster with serial number s and having N used GPUs in use;

[0083] (4.3) If the optimal resource plan in step (3) is the cross-node expected plan then set N used = N GPU , where N GPU is the total number of GPUs obtained by accumulating the c used of each node in the resource plan. Traverse the resource list R to find the resource node object Node(s, c free ) with c free > 0. Remove min(c free , N used ) GPU devices from the resource node object Node(s, c free ) and N used respectively, and add the resource node object Node(s, min(c free , N GPU )) to . And so on until N used = 0, and the traversal ends; Node(s, min(c free , N GPU )) represents Node(s, N used ) which represents a node object in the GPU cluster with serial number s and having min(c free , N used ) GPUs;

[0084] Step (5): Before each execution of the task scheduling process in the dynamic resource scheduling algorithm, analyze the running task situation to decide whether to perform the resource migration process as follows Figure 4 ; specifically as follows:

[0085] (5.1) Initialize the task lists l s and l m ; Traverse the running task queue Q run , add the task t originally running on a single node to l s , and add the task t originally running across nodes to l m ;

[0086] (5.2) Sort the tasks t in l s and l m in descending order according to the total number of GPUs N of the optimal resource plan; First, traverse l GPU s ​, perform the physical resource allocation process of step (4) for task t therein, and then traverse l m , also perform the physical resource allocation process of step (4) for task t therein;

[0087] Step (6), the scheduler executes the scheduling algorithm to select a new task to run on the GPU cluster; specifically as follows:

[0088] (6.1) Receive the waiting task queue Q wait , the resource list R and the current time T curr , where T curr increases by unit time. When the arrival time T of the distributed deep learning task arr = T curr , add the task to the queue Q wait , and at this time, pre-calculate the deadline T of the task according to formula (7) dl ;

[0089] (6.2) The process of dynamic resource scheduling is as follows Figure 5 ; the specific steps are as follows:

[0090] (6.2.1) According to the load situation of the GPU cluster resources, attempt to perform the resource migration process of step (5);

[0091] (6.2.2) Traverse the waiting queue Q wait , perform the resource scheme decision of step (3) on task t to obtain the optimal resource scheme of t

[0092] (6.2.3) Initialize the expected task queue Q exp and the unexpected task queue If the running end time T of the task end and the deadline T dl satisfy formula (9), then add task t to the queue Q exp , otherwise add it to the queue ; sort the task t in the queue Q exp in ascending or descending order according to the value of T dl - T end . At this time, the task t at the head of the queue is under the resource scheme with T end closer to T dl ; sort the task t in the queue in ascending order according to the value of T end . The task t at the head of the queue is under the resource scheme with T end closer to T dl ; note that the queue Q exp may be empty;

[0093] (6.2.4) If the queue Q exp is not empty, select the head task t of the queue as the scheduled task t * ; if the queue Q exp is empty, select the head task t in the queue as the scheduled task t * ; perform the physical resource allocation process in step (4) on t * .

[0094] Experimental data

[0095] (1) The time for the DNN model participating in the measurement to train multiple iterative rounds on a single GPU device in the GPU cluster

[0096]

[0097] (2) Comparison of experimental data:

[0098] (2.1) Earliest Deadline First (EDF): Select the task with the smallest deadline from the waiting queue and perform resource allocation using the overall GPU resources;

[0099] (2.2) First In First Out (FIFO): Select the task with the smallest arrival time from the waiting task queue and perform resource allocation using the overall GPU resources;

[0100] (2.3) Themis: Allocate GPU resources to multiple waiting tasks according to the fairness of the completion time and schedule them to the GPU cluster for operation at one time, as much as possible to ensure that the tasks have similar completion times;

[0101] (2.4) No Resource Migration (NoRM): In order to verify the effectiveness of the migration mechanism introduced by DRS, remove the migration mechanism part in DRS and compare its various performance metrics with those of DRS.

[0102] (2.5) The dynamic resource scheduling method for the GPU cluster of the present invention is Dynamic Resource Scheduling (DRS).

[0103] (3) Comparison of the performance of each scheduling algorithm for different task arrival rates:

[0104] (3.1) Deadline guarantee rate / %

[0105]

[0106] (3.2) Average waiting time / h

[0107]

[0108] (3.3) Average completion time / h

[0109]

[0110] (4) Performance comparison of each scheduling algorithm under different numbers of resource nodes:

[0111] (4.1) Deadline guarantee rate / %

[0112]

[0113] (4.2) Average waiting time / h

[0114]

[0115] (4.3) Average completion time / h

[0116]

[0117]

[0118] (5) Performance comparison of each scheduling algorithm under the same number of urgent tasks:

[0119] (5.1) Deadline guarantee rate / %

[0120]

[0121] (5.2) Average waiting time / h

[0122]

[0123] (5.3) Average completion time / h

[0124]

[0125]

[0126] (6) Performance comparison of each scheduling algorithm at different reception times:

[0127] (6.1) Deadline guarantee rate / %

[0128]

[0129] (6.2) Average waiting time / h

[0130]

[0131] (6.3) Average completion time / h

[0132]

Claims

1. A dynamic resource scheduling method for a GPU cluster, characterized in that Including the following steps: Step (1), based on the DNN model iteration characteristics and the bandwidth difference between GPU devices under the Ring-Allreduce communication architecture of distributed machine learning, construct a resource-time model: The resource-time model includes the following: (1.1) The actual running time T of the distributed deep learning task under a certain resource plan run is expressed as follows: T run = T step × N step × N epoch Equation (1) Among them, T step is the time taken for the DNN model to train a dataset of a batch size, N step is the number of datasets of a batch size that can be input by the DNN model in one iteration round, N epoch represents the iteration round; (1.2)T step consisting of the computing time T on a single CPU device cal and the communication time T between CPUs and CPU devices, and its calculation formula is as follows: comm as follows: T step = T cal + T comm Equation (2) (1.3)N step will vary with the total number of GPUs included in the resource plan. The larger the number, the smaller N step will be correspondingly reduced; N step , the size S of the DNN model training dataset dataset , the batch size S batch and the total number of GPUs N GPU have the following relationship during the distributed data parallel training process: Among them, N GPU is obtained by accumulating c of each node in the resource plan used , where c used represents the number of GPUs used by the training task on a single node; (1.4) By placing the DNN model on a single GPU device for several batches of iterations and recording the corresponding running time, since there is no multi-device communication involved, this running time only includes the computing time on a single GPU device, which is expressed as follows: where T' step is the running time of several iterations, and N' step is the corresponding number of iterations; (1.5) If there is no communication time, then the running time of the task and the total number of GPUs included in the resource plan will be inversely proportional, that is, as the total number of GPUs increases, the running time of the task will decrease proportionally. When there is communication time, it will lead to a decrease in running efficiency; the communication time T under the Ring-Allreduce communication architecture comm is expressed as follows: Where BW is the bandwidth speed between two GPU devices. If the two GPU devices are on the same node, then BW is the bandwidth between the GPU devices within the node. If the two GPU devices are on different nodes, then BW is the network bandwidth between the nodes; Step (2), based on the resource quantity used in the resource scheme, the task running time, and the task deadline, construct a resource-performance model: (2.1) Deadline modeling: (2.1.1) Let the user's deadline requirement for a task consist of the task arrival time, task priority, and maximum task running time, where the maximum running time is the running time of the task on a single GPU device only. Define several task priorities and convert the priority to the expected running time T of the task exp , and its calculation formula is as follows: where α corresponds to the task priority, represents the time for the task to run on a single GPU device; (2.1.2) Let the arrival time and start time of the task be T arr and T start respectively. Then the deadline T dl and end time T end of the task can be expressed as follows: T dl = T arr + T exp Equation (7) T end = T start + T run Equation (8) (2.1.3) When the deadline T of the task dl and the running end time T end satisfy the following formula (9), it indicates that the deadline requirement of the user is met at the end of the task: T end <T dl Formula (9) (2.2) When all the GPU devices held by the resource plan are on the same node, the bandwidth speed is the direct connection bandwidth between the GPU devices. When the GPU devices held by the resource plan are on different nodes, the bandwidth speed is the network bandwidth between the nodes. As can be seen from Equation (5), when N GPU and N param remain unchanged, T comm increases as BW decreases. Substitute Equations (2) and (3) into Equation (1), and require that the time for multi-machine distributed training is shorter than the running time of single-machine training, then the following inequality is obtained: Where the first half and the second half of the inequality are the times for the DNN model to train one iteration round on multiple nodes and a single node respectively. Simplifying formula (10) gives: T comm <(N GPU -1)×T cal Formula (11) When the DNN model is performing multi-machine distributed training, T comm 、N GPU and T cal Only by conforming to formula (10) can the purpose of accelerating model training be achieved; (2.3) To measure the performance of the task under different resource schemes and select the resource scheme with the highest running efficiency among multiple resource schemes that meet the deadline requirements, and give full play to the resource performance, define the performance formula of the resource-performance model as: Among which T dl represents the deadline of the task; Step (3), based on steps (1) and (2), make a dynamic resource scheme decision for the distributed deep learning task: Generate a list of available resource schemes for each task in the waiting queue based on the cluster idle resources and the resource layout. According to the resource-performance model and combined with the cluster node load situation, determine the optimal resource scheme for each task; Step (4), based on the optimal solution of the task Perform physical resource node allocation; Step (5), before each execution of the task scheduling process by the dynamic resource scheduling algorithm, analyze the situation of the tasks that have already run and decide whether to perform resource migration; Step (6), the scheduler executes the scheduling algorithm to select a new task to run on the GPU cluster.

2. The dynamic resource scheduling method for a GPU cluster according to claim 1, characterized in that The Ring-Allreduce communication architecture described in step (1) includes several nodes of a GPU cluster that communicate with each other. Each node includes multiple CPUs and GPUs. The GPU devices on the same node communicate via the Peripheral Component Interconnect Express (PCIe) and the QuickPath Interconnect (QPI). Among them, the communication between GPUs uses PCIe, the communication between GPU and CPU uses PCIe, and the communication between CPUs uses QPI; the nodes in the GPU cluster communicate with each other via wireless broadband technology.

3. The dynamic resource scheduling method for a GPU cluster according to claim 1, characterized in that Specifically in step (3) as follows: (3.1) Obtain the resource list R, and set the number of resource nodes with c free > 0 to be n, and c free represents the number of idle GPUs in a single node. Among the resource nodes, c free the maximum value is max(c free ) and its cumulative sum is sum(c free ), and initialize the single-node resource scheme list l s and the cross-node resource scheme list l m ; (3.2) If n = 1, then obtain c from the resource list R used From 1 to the resource scheme R of max(c free ) t Add to l s If n > 1, then obtain c from the resource list R used From 1 to the resource scheme R of sum(c free ) t Add to l m ; Calculate l according to Equation (1) and Equation (8) s and l m where R t of T run and T end and filter out some inefficient resource solutions R according to Equation (11) t ; (3.3) Obtain l according to Equation (12) s Medium performance And The resource plan R when the value is the largest t As the single-node expected plan And obtain l according to Equation (12) s T in end >T dl And T end The resource plan R when the value is the smallest t As the single-node unexpected plan Obtain the resource plan R when the value from l m with the best performance and is the largest as the cross-node expected plan t and obtain the resource plan R when l in m where T end >T dl and the value of T end is the smallest as the cross-node unexpected plan t Note that among them and and may not exist; (3.4) Determine whether the cross-node expected scenario is satisfied Exists and there is 0 < c in the GPU cluster free <N GPU of resource nodes. If satisfied, it means that the current task has a cross-node resource plan to utilize local resources and end the operation within T dl and end the operation within. At this time, the optimal resource plan If not satisfied but exists, then the optimal resource plan If not satisfied and still does not exist, it means that the current idle resources in the GPU cluster cannot enable the current task to end the operation within T dl and end the operation, then it is considered that there is no expected scenario at this time and are available for selection, then determine whether the cross-node unexpected scenario is satisfied Exists and there is 0 < c in the GPU cluster free <N GPU of resource nodes. If satisfied, it means that the current task has a cross-node resource plan to utilize local resources and end the operation within T dl and end the operation within. At this time, the optimal resource plan If not satisfied, then the optimal resource plan 4. The dynamic resource scheduling method for a GPU cluster according to claim 3, wherein Specifically, step (4) is: (4.1) Obtain the resource list R and sort it in ascending order according to the c of the nodes free Ascending order; (4.2) If the optimal resource plan in step (3) is the single-node expected plan then traverse the resource list R to find c free ≥ N GPU of the resource node Node(s, c free ). Remove N free GPU devices from this resource node Node(s, c GPU ), and add Node(s, N used ) to . End the traversal, where Node(s, c free ) represents the node object in the GPU cluster with the serial number s and having c free idle GPUs, and s represents the serial number of the node object; Node(s, N used ) represents the node object in the GPU cluster with the serial number s and having N used in-use GPUs; (4.3) If the optimal resource plan in step (3) is the cross-node expected plan then set N used = N GPU , where N GPU is the total number of GPUs accumulated from each node in the resource plan. Traverse the resource list R to find the resource node object Node(s, c used ), where c free > 0. Remove min(c free , N free ) GPU devices from the resource node object Node(s, c used ) and N free respectively, and add the resource node object Node(s, min(c used , N free )) to GPU ), and so on until N = 0, ending the traversal; Node(s, min(c used , N free )) represents the node object with serial number s in the GPU cluster and having min(c GPU , N used ) GPUs, and Node(s, N free ) represents the node object with serial number s in the GPU cluster and having N used GPUs.

5. The dynamic resource scheduling method for a GPU cluster according to claim 4, wherein Specifically, step (5) is: (5.1) Initialize the task list l s and l m ; Traverse and run the task queue Q run , add the task t that was originally running on a single node to l s , add the task t that was originally running across nodes to l m ; (5.2) Sort the tasks t in l s and l m in descending order according to the total number N of GPUs in the optimal resource plan ; first traverse l GPU and perform the physical resource allocation process of step (4) on the tasks t therein, and then traverse l s and perform the physical resource allocation process of step (4) on the tasks t therein as well. m ​ 6. The dynamic resource scheduling method for a GPU cluster according to claim 5, characterized in that Specifically, step (6) is: (6.1) Receive the waiting task queue Q wait , the resource list R, and the current time T curr , where T curr increases at unit time. When the arrival time T of the distributed deep learning task arr = T curr , add the task to the queue Q wait . At this time, pre-compute the deadline T of the task according to Equation (7) dl ; (6.2) Dynamic resource scheduling.

7. The dynamic resource scheduling method for a GPU cluster according to claim 6, characterized in that Specifically, step (6.2) is: (6.2.1) According to the load situation of the GPU cluster resources, try to execute the resource migration process of step (5); (6.2.2) Traverse the waiting queue Q wait and perform the resource plan decision in step (3) on task t to obtain the optimal resource plan for t (6.2.3) Initialize the expected task queue Q exp and the unexpected task queue If the running end time T of the task end and the deadline T dl satisfy Equation (9), then add the task t to the queue Q exp otherwise add it to the queue ; Sort the task t in the queue Q exp in ascending / descending order according to the value of T dl -T end At this time, the task t at the head of the queue has a T under the resource plan end that is closer to T dl ; Sort the task t in the queue in ascending order according to the value of T end The task t at the head of the queue has a T under the resource plan end that is closer to T dl ; Note that the queue Q exp may be empty; (6.2.4) If the queue Q exp is not empty, select the head task t of the queue as the scheduled task t * ; if the queue Q exp is empty, select the head task t in the queue as the scheduled task t * ; perform the physical resource allocation process in step (4) for t * .

8. An electronic device, characterized in that, Including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1-7.

9. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions that, when called and executed by a processor, cause the processor to implement the method according to any one of claims 1-7.