GPU Heterogeneous Cluster Scheduling Method, System and Medium for Interference Perception
Through the interference-aware GPU heterogeneous cluster scheduling method, the hyperparameters of deep learning applications are dynamically adjusted, solving the problems of resource isolation and performance interference of deep learning applications on heterogeneous clusters, and achieving more efficient GPU resource utilization and application performance.
Patent Information
- Application Number
- CN202111615877.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-27
AI Technical Summary
When running deep learning applications on heterogeneous clusters, it is difficult for the existing technology to effectively predict and manage the runtime characteristics of different deep learning frameworks under different parameters, resulting in resource isolation setting static, unable to dynamically adjust, and low-priority tasks waiting delays are long.
Through the interference-aware GPU heterogeneous cluster scheduling method, the hyperparameters of deep learning applications are dynamically adjusted, Bayesian optimization is used to find the optimal hyperparameter combination, reduce performance interference, and conduct online detection and prediction based on performance interference, optimizing GPU shared combination and hyperparameter configuration.
It realizes more full utilization of GPU resources, reduces the performance interference of deep learning applications when sharing GPUs, and improves the runtime performance of each application and the resource utilization efficiency of heterogeneous clusters.
Smart Images

Figure CN114237913B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a method, system and medium for GPU heterogeneous cluster scheduling with interference awareness. Background Art
[0002] The rapid development of artificial intelligence has led to the emergence of more and more deep learning applications. These applications require a large amount of training data and high-performance computing capabilities during operation, which has promoted the emergence and application of high-performance chips such as GPUs, FPGAs, TPUs, and AISCs. To meet the requirements of upper-layer deep learning applications for computing and storage capabilities, Kubernetes has become an orchestration tool for deep learning container applications widely used in the industrial and academic fields due to its portable, scalable, and self-healing features.
[0003] However, there are many problems and challenges in scheduling deep learning applications on heterogeneous clusters. On the one hand, due to the increasingly complex and diverse network structures of upper-layer applications, it is difficult to predict resource requirements and runtime performance. On the other hand, although the method of containers monopolizing GPUs achieves resource isolation to a certain extent, the price of graphics card resources is expensive, and there is an imbalance between application requests and the supply and demand of computing resources, resulting in the problem of timeout waiting for applications. Summary of the Invention
[0004] An object of this application is to solve at least the above problems and / or deficiencies and provide at least the advantages described later.
[0005] The method for GPU heterogeneous cluster scheduling with interference awareness provided by this application includes: obtaining a target shared GPU combination with the least interference from at least one shared GPU combination; using Bayesian optimization to find a target hyperparameter combination for the target shared GPU combination until convergence; and running a target deep learning workload with the target hyperparameter combination to evaluate the performance of the scheduling algorithm.
[0006] Optionally, in this application embodiment, before obtaining a target shared GPU combination with the least interference from at least one shared GPU combination, the method further includes: obtaining at least one first deep learning workload; using Bayesian optimization to find hyperparameter combinations for the at least one first deep learning workload; respectively running each first deep learning workload application on the shared GPU offline with the hyperparameter combination of each first deep learning workload to obtain at least one performance interference score; and determining the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score.
[0007] Optionally, in the embodiments of the present application, determining the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score includes: in an online scenario, applying the at least one shared GPU combination and the at least one performance interference score to a new deep learning workload to determine the interference corresponding to the at least one shared GPU combination.
[0008] Optionally, in the embodiments of the present application, using the hyperparameter combination of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain at least one performance interference score includes: based on a target algorithm, using the hyperparameter combination of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain the at least one performance interference score.
[0009] Optionally, in the embodiments of the present application, the target deep learning workload includes at least one of the following: natural language processing system, face recognition system, target detection system, image classification system, recommendation system.
[0010] Optionally, in the embodiments of the present application, using Bayesian optimization to find a target hyperparameter combination for the target shared GPU combination until convergence includes: combining the hyperparameters of at least one second deep learning workload application, and selecting the best target hyperparameter configuration for running separately through Bayesian optimization.
[0011] The GPU heterogeneous cluster scheduling system provided by the present application includes an acquisition module and a processing module; the acquisition module is used to obtain a target shared GPU combination with the least interference from at least one shared GPU combination; the processing module is used to use Bayesian optimization to find a target hyperparameter combination for the target shared GPU combination obtained by the acquisition module until convergence; and use the target hyperparameter combination to run a target deep learning workload to implement the performance evaluation of the scheduling algorithm.
[0012] Optionally, in the embodiments of the present application, the acquisition module is further used to obtain at least one first deep learning workload; the processing module is further used to use Bayesian optimization to find the hyperparameter combination of the at least one first deep learning workload obtained by the acquisition module; and use the hyperparameter combination of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain at least one performance interference score; and determine the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score.
[0013] Optionally, in the embodiments of the present application, the processing module is specifically configured to, in an online scenario, apply the at least one shared GPU combination and the at least one performance interference score to a new deep learning workload, and determine the interference corresponding to the at least one shared GPU combination.
[0014] The readable storage medium provided by the embodiments of the present application stores a computer program thereon, and when the computer program is executed by a processor, the steps of the interference-aware GPU heterogeneous cluster scheduling method as described in the above claims are implemented.
[0015] Compared with the prior art, the present application has the following beneficial effects:
[0016] In the interference-aware GPU heterogeneous cluster scheduling method provided by the embodiments of the present application, at the upper application level, most current GPU sharing is designed for a single deep learning framework, without considering the runtime characteristics of deep learning applications implemented by different frameworks under different parameters; in terms of GPU sharing, the implemented computing resource isolation and video memory isolation settings are static and do not have a dynamic adjustment scheme. The present invention dynamically adjusts hyperparameters without restricting resource occupancy, thus making more efficient use of GPU resources; in terms of application scheduling, most current scheduling schemes consider the priority of applications, and tasks with low priority often have long waiting delays. The present invention performs online detection based on the performance interference when different applications share GPUs and predicts unknown application combinations, greatly reducing the performance interference generated by applications when sharing GPUs and ensuring the runtime performance of each application.
[0017] Other advantages, objectives, and features of the present application will be partially reflected by the following description, and partially will be understood by those skilled in the art through the research and practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0019] Figure 1 It is a flowchart of the interference-aware GPU heterogeneous cluster scheduling method provided by the embodiments of the present application;
[0020] Figure 2 It is a structural diagram of the interference-aware GPU heterogeneous cluster scheduling system provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following further elaborates on the present application in conjunction with the accompanying drawings, so that those skilled in the art can implement it with reference to the text of the specification.
[0022] It should be understood that terms such as "having", "including", and "comprising" used herein do not preclude the presence or addition of one or more other elements or combinations thereof.
[0023] In the related art, with the rise of artificial intelligence, deep learning applications have gradually penetrated into fields such as computer vision, natural language processing, and e-commerce recommendation. Since deep learning applications have a very strong demand for computing power, general computing platforms cannot meet their requirements, while high-performance computing provides continuous and stable computing power guarantee for the operation of such applications. However, due to the high purchase price of high-performance chips such as GPUs, the continuous expansion of the cluster scale, the increasing complexity of deep learning applications, and the unpredictable performance interference brought by GPU sharing, these factors have posed great challenges to the scheduling of such applications. Therefore, how to adjust the hyperparameters of deep learning applications, how to select the best application combination to share GPUs, and how to adjust the hyperparameters of applications after sharing are the keys to improving the runtime performance of applications.
[0024] Designing an interference-aware GPU heterogeneous cluster scheduling method and system mainly starts from two aspects. One is to adjust the hyperparameters of deep learning, which have a strong correlation with GPU utilization, video memory utilization, and application training time. The other is to screen the combinations of deep learning applications sharing GPUs, filter out combination schemes with serious interference, and perform fast and accurate interference prediction and scheduling selection for online scenarios, so as to improve the user experience of public cloud services and save the operating costs of operators.
[0025] At present, some GPU sharing solutions have been implemented in the industry. For example, Tencent's GaiaGPU realizes the isolation of GPU computing resources and video memory resources by hijacking the CUDA API. And for computing resources, it provides two methods: soft isolation and hard isolation. The common point of these two methods is that when the utilization rate of the GPU SM used by a task exceeds the resource limit, the API call is postponed. The difference is that if there is free resource, soft isolation allows the task to exceed the setting and dynamically calculate the resource limit. While hard isolation does not allow exceeding the set amount. Alibaba's cGPU builds the sharing module on top of the Nvidia driver. It also completes resource isolation by hijacking the calls to the driver, and allocates the computing power occupied by tasks by setting the length of the time slice occupied by the task, but does not explain to the outside world what method is used to precisely control the time of context switching. Nvidia has also designed a GPU sharing solution, vGPU. Its sharing module is inside the Nvidia driver. vGPU provides a very high-isolation hardware environment through vfio-mdev, which is mainly for virtual machine products, cannot dynamically adjust the resource ratio, and this solution is not open source. In addition, Nvidia also provides the MPS (Multi-Process Service) solution. This method is a way of external service for designers to choose to use, and only supports some NVIDIA architectures.
[0026] However, the existing scheduling schemes for GPU sharing mainly fall into the following two categories: 1) Implement time-division multiplexing and support video memory isolation. 2) Implement time-division and space-division multiplexing and support video memory and computing isolation. The first category of schemes does not support computing isolation, and video memory isolation requires accurate prediction of the video memory for deep learning applications. In practical applications, the video memory size of shared applications cannot be dynamically controlled, and the performance interference generated by shared applications is not considered. The second category of schemes realizes space-time multiplexing and supports the isolation of video memory and computing. However, it is found in specific applications that the computing and video memory resources must be set in advance and cannot be dynamically modified, and only supports specific value sizes. In addition, there is no interference-aware design based on the scheduling layer during sharing. Therefore, in view of these shortcomings, the present invention proposes an interference-aware GPU heterogeneous cluster scheduling method and system. Based on the existing GPU sharing solutions, comprehensively considering the interference characteristics when different shared application combinations, as well as the adjustment and optimization of the hyperparameters of deep learning applications, an iterative scheduling strategy is designed to continuously optimize the shared combination and application hyperparameters, in order to achieve the optimal running state of deep learning applications and improve the application performance and the resource utilization efficiency of the heterogeneous cluster.
[0027] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Such as Figure 1As shown in the figure, the interference-aware GPU heterogeneous cluster scheduling method provided by the embodiments of the present application includes the following steps:
[0029] Step 1: Obtain at least one first deep learning workload.
[0030] Step 2: Use the Bayesian optimization to find the hyperparameter combinations of the at least one first deep learning workload.
[0031] Step 3: Use the hyperparameter combinations of each first deep learning workload to offline run each first deep learning workload application sharing the GPU respectively, so as to obtain at least one performance interference score.
[0032] Among them, Step 3 can be specifically implemented by the following Step 31.
[0033] Step 31: Based on the target algorithm, use the hyperparameter combinations of each first deep learning workload to offline run each first deep learning workload application sharing the GPU respectively, so as to obtain the at least one performance interference score.
[0034] Step 4: Determine the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score.
[0035] Among them, Step 4 can be specifically implemented by the following Step 41.
[0036] Step 41: In the online scenario, apply the at least one shared GPU combination and the at least one performance interference score to the new deep learning workload, and determine the interference corresponding to the at least one shared GPU combination.
[0037] Step 5: Obtain the target shared GPU combination with the minimum interference from the at least one shared GPU combination.
[0038] Step 6: Use the Bayesian optimization to find the target hyperparameter combination for the target shared GPU combination until convergence.
[0039] Step 7: Run the target deep learning workload using the target hyperparameter combination.
[0040] Step 8: Implement the performance evaluation of the scheduling algorithm.
[0041] In the embodiments of this application, the deep learning application network structure is complex, and the required computing power resources have strong performance. However, GPU resources are scarce and expensive. Although existing GPU sharing solutions have, to a certain extent, solved the isolation of computing resources and video memory, there has been little research on the interference problems existing in GPU sharing at the scheduling layer. Therefore, we deeply analyze the hyperparameters of typical deep learning applications for application performance and runtime efficiency, search for the best parameter combinations and application combinations for GPU sharing, and make full use of cluster resources, so as to provide a more optimized solution for cluster scheduling design.
[0042] The different deep learning applications in step 1) include natural language processing, face recognition, object detection, image classification, recommendation systems, etc., and include application implementations under different deep learning frameworks such as (Pytorch, Tensorflow, MXNet, etc.).
[0043] In step 2), the hyperparameters (batchsize and learning rate) of deep learning applications implemented with different types and different frameworks are combined, and the best hyperparameter configuration when the application runs alone is selected through Bayesian optimization.
[0044] In step 3), the performance characterization of different types of deep learning applications when sharing GPUs is run offline, and the scores quantifying performance interference are recorded.
[0045] Among them, the target algorithm is specifically as follows:
[0046]
[0047] JCT = throughput * epoch (2)
[0048] In step 4), for the online scenario, the newly arrived deep learning training application is combined with the applications running in the cluster, and the data prediction and filling of performance interference are carried out through the prior data in step 3), and the stochastic gradient descent method is used for data prediction. Note: At this time, if there are idle GPU resources available in the cluster, the training request can be directly randomly scheduled to an idle GPU node.
[0049] In step 5), the combination with the smallest performance interference is selected from the results of the previous step.
[0050] In step 6), the Bayesian optimization is used again for the selected application combination to find the best hyperparameter combination (that is, the minimum interference when sharing GPUs) until convergence.
[0051] In step 7), the new deep learning application is scheduled to the GPU in a specific cluster node and run with the parameter configuration in 6).
[0052] When implementing the scheduling algorithm in step 8), the most widely used container orchestration framework kubernetes is used as the basis. For the GPU sharing solution, Alibaba's cGPU is selected, and experiments are carried out in a large-scale cluster to evaluate the performance of the scheduling system.
[0053] In the embodiments of this application, benchmarks for deep learning applications implemented with different types and frameworks can be achieved, and the performance characteristics when running alone can be analyzed. The hyperparameters of deep learning applications, especially the batchsize, have a great impact on GPU utilization and also affect the running efficiency of the application at the same time. The present invention adopts a Bayesian optimization strategy to select the best hyperparameters. Stochastic gradient descent is used to predict the performance interference of combinations of different applications sharing GPUs.
[0054] In the embodiments of this application, the adjustment of hyperparameters of deep learning applications during GPU sharing is realized, and the latency during the training of deep learning applications is reduced. A GPU sharing scheduling strategy is designed based on performance interference awareness and is applicable to the offline scenario, providing a relatively complete GPU sharing scheduling solution.
[0055] For the interference-aware GPU heterogeneous cluster scheduling method provided in the embodiments of this application, at the upper application level, most current GPU sharing is designed for a single deep learning framework and does not consider the runtime characteristics of deep learning applications implemented with different frameworks under different parameters; in terms of GPU sharing, the implemented computing resource isolation and video memory isolation settings are static and do not have a dynamic adjustment scheme. The present invention dynamically adjusts hyperparameters without restricting resource occupancy, thus making more full use of GPU resources; in terms of application scheduling, most current scheduling schemes consider the priority of applications, and tasks with low priority often have a long waiting latency. The present invention performs online detection of performance interference when different applications share GPUs and predicts unknown application combinations, greatly reducing the performance interference generated by applications when sharing GPUs and ensuring the runtime performance of each application.
[0056] Figure 2 Fig. shows the interference-aware GPU heterogeneous cluster scheduling system provided in the embodiments of this application. The system 20 includes: an acquisition module 21 and a processing module 22.
[0057] Among them, the acquisition module 21 is used to obtain the target shared GPU combination with the least interference from at least one shared GPU combination. The processing module 22 is used to use Bayesian optimization to find the target hyperparameter combination for the target shared GPU combination obtained by the acquisition module 21 until convergence; and use the target hyperparameter combination to run the target deep learning workload to achieve the performance evaluation of the scheduling algorithm.
[0058] In a possible implementation, the obtaining module 21 is further configured to obtain at least one first deep learning payload. The processing module 22 is further configured to use the Bayesian optimization to find the hyperparameter combinations of the at least one first deep learning payload obtained by the obtaining module; and use the hyperparameter combinations of each first deep learning payload to separately run each first deep learning payload application sharing the GPU offline to obtain at least one performance interference score; and determine the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score.
[0059] In a possible implementation, the processing module 22 is specifically configured to, in an online scenario, apply the at least one shared GPU combination and the at least one performance interference score to a new deep learning payload to determine the interference corresponding to the at least one shared GPU combination.
[0060] In a possible implementation, the processing module 22 is specifically configured to, based on a target algorithm, use the hyperparameter combinations of each first deep learning payload to separately run each first deep learning payload application sharing the GPU offline to obtain the at least one performance interference score.
[0061] In a possible implementation, the target deep learning payload includes at least one of the following: natural language processing system, face recognition system, target detection system, image classification system, recommendation system.
[0062] In a possible implementation, the processing module 22 is specifically configured to combine the hyperparameters of at least one second deep learning payload application, and select the best target hyperparameter configuration for running separately through Bayesian optimization.
[0063] Although the embodiments of the present application have been disclosed as above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present application. For those familiar with the field, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present application is not limited to the specific details and the illustrated examples here.
Claims
1. A GPU heterogeneous cluster scheduling method for interference perception, characterized in that, the method includes: Obtain a target shared GPU combination with the least interference from at least one shared GPU combination; Use Bayesian optimization on the target shared GPU combination to find a target hyperparameter combination until convergence; Run a target deep learning workload using the target hyperparameter combination to achieve performance evaluation of the scheduling algorithm; Before obtaining a target shared GPU combination with the least interference from at least one shared GPU combination, the method further includes: Obtain at least one first deep learning workload; Use Bayesian optimization to find hyperparameter combinations for the at least one first deep learning workload; Use the hyperparameter combinations of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain at least one performance interference score; Determine the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score.
2. The method according to claim 1, characterized in that, the determining the interference corresponding to the at least one shared GPU combination according to the at least one performance interference score includes: In an online scenario, apply the at least one shared GPU combination and the at least one performance interference score to a new deep learning workload to determine the interference corresponding to the at least one shared GPU combination.
3. The method according to claim 1, characterized in that, the using the hyperparameter combinations of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain at least one performance interference score includes: Based on a target algorithm, use the hyperparameter combinations of each first deep learning workload to separately run each first deep learning workload on a shared GPU offline to obtain the at least one performance interference score.
4. The method according to claim 1, characterized in that, the target deep learning workload includes at least one of the following: natural language processing system, face recognition system, object detection system, image classification system, recommendation system.
5. The method according to claim 1, characterized in that, the using Bayesian optimization on the target shared GPU combination to find a target hyperparameter combination until convergence includes: Combine the hyperparameters applied to at least one second deep learning workload, and select the best target hyperparameter configuration for separate operation through Bayesian optimization.
6. A GPU heterogeneous cluster scheduling system for interference perception, characterized in that, the system includes: an acquisition module and a processing module; The acquisition module is used to obtain a target shared GPU combination with the least interference from at least one shared GPU combination; The processing module is used to use Bayesian optimization on the target shared GPU combination obtained by the acquisition module to find a target hyperparameter combination until convergence; and run a target deep learning workload using the target hyperparameter combination to achieve performance evaluation of the scheduling algorithm; The acquisition module is further used to obtain at least one first deep learning workload; The processing module is further configured to use the Bayesian optimization to find the hyperparameter combinations of the at least one first deep learning workload obtained by the acquisition module; and use the hyperparameter combinations of each first deep learning workload to separately run each first deep learning workload application offline sharing the GPU, so as to obtain at least one performance interference score; and, according to the at least one performance interference score, determine the interference corresponding to the at least one shared GPU combination.
7. The system according to claim 6, wherein, when in an online scenario, the processing module is specifically configured to apply the at least one shared GPU combination and the at least one performance interference score to a new deep learning workload, and determine the interference corresponding to the at least one shared GPU combination.
8. A readable storage medium, wherein, a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the interference-aware GPU heterogeneous cluster scheduling method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Dynamic resource allocation method and system under space-based cloud computing architecture, and storage medium
CN110730138A
Automatic machine learning method and system for remote sensing semantic segmentation
CN111797833A