Resource overcommitting scheduling method based on GPU virtualization
By adopting the GPU virtualization-based resource over-distribution scheduling method in the GPU full virtualization environment, the problem of failure to fully utilize GPU resources is solved, and the maximum utilization of GPU resources and the overall resource efficiency are achieved.
Patent Information
- Application Number
- PCT/CN2024/136154
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-19
AI Technical Summary
In a fully virtualized GPU environment, GPU resources are not fully utilized, resulting in the generation of idle GPU resources and low resource utilization.
The resource over-allocation scheduling method based on GPU virtualization is adopted. By creating a GPU virtual machine, setting GPU parameter information, scheduling GPU resources, calculating the available scores of the GPU host and model, the optimal host and GPU model are selected in a comprehensive score to achieve the over-allocation of GPU resources.
Through the GPU resource over-allocation scheduling method, the GPU resources can be maximized, the overall utilization rate of GPU resources can be improved, and the GPU resources can be fully utilized.
Smart Images

Figure CN2024136154_19062025_PF_FP_ABST
Abstract
Description
A resource over-allocation scheduling method based on GPU virtualization
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 14, 2023, with application number 202311722347.2 and invention name “A resource over-allocation scheduling method based on GPU virtualization”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of cloud computing, and in particular to a resource over-allocation scheduling method based on GPU virtualization. Background Art
[0004] There are generally three typical approaches to using GPU resources in virtual machines, including GPU API forwarding, GPU pass-through, and GPU full virtualization. A GPU can only be allocated to one virtual machine, resulting in the lowest GPU resource utilization. In the GPU API forwarding method, theoretically, a GPU can be allocated to many virtual machines, and the GPU resource utilization seems to be the highest, but this method has the lowest performance. GPU full virtualization can slice a GPU into multiple small GPUs and provide them to multiple virtual machines. Compared with GPU pass-through, the utilization rate is much higher. However, in actual use, the GPU resources after GPU full virtualization are not over-allocated, resulting in idle GPU resources and insufficient utilization of GPU resources. Summary of the Invention
[0005] The purpose of this application is to provide a resource over-allocation scheduling method based on GPU virtualization to solve the problems raised in the above background technology.
[0006] To achieve the above objectives, the present application provides the following technical solution: a resource over-allocation scheduling method based on GPU virtualization, the scheduling method comprising the following steps:
[0007] The first step is to create a GPU virtual machine;
[0008] The second step is to set the GPU parameter information of the virtual machine;
[0009] The third step is to schedule GPU resources and select available GPU host sequences based on the GPU parameter information of the virtual machine;
[0010] The fourth step is to perform scheduling calculations based on the stored GPU information. The scheduling calculations include the GPU host scale calculating the available GPU host scores using a calculation formula and the GPU model scale calculating the available scores of each GPU model.
[0011] Step 5: Select the GPU host and GPU model based on the comprehensive scores of the GPU host scale and GPU model scale.
[0012] Optionally, the GPU parameter information of the virtual machine in the second step includes specifying the model of the GPU used in the virtual machine, the size of the video memory used to store and process graphics data, the type of video memory, the number of GPU cores used to process graphics data, and the calculation parameters in GPU scheduling, wherein the calculation parameters include the GPU model to be calculated, the over-provisioning ratio, the online usage, the offline usage, the real available amount, and the virtual available amount.
[0013] Optionally, the real available quantity is calculated as follows: RU=Rt-Ou (1);
[0014] The virtual available amount is calculated as follows: Vu = Rt2*c-(Ou+Fu) (2);
[0015] Where RU represents the real available capacity, Rt represents the real total capacity, Ou represents the online usage, Fu represents the offline usage, Vu represents the virtual available capacity, and c represents the over-allocation coefficient.
[0016] Optionally, the scheduling GPU resource operation includes allocating GPUs according to the allocation type, the GPU availability is scheduled using the virtual availability, and when the virtual machine pre-allocates the GPU, the GPU availability is the minimum of the real availability and the virtual availability, and when the virtual machine is directly allocated, the GPU availability is the real availability, and when the virtual machine is powered on, the GPU resources meet the power-on conditions.
[0017] Optionally, the storage operation of the GPU information includes the following steps:
[0018] A1, obtaining GPU resource information through a GPU resource information acquisition module, wherein the GPU resource information acquisition module and the computing service are both deployed on the computing node;
[0019] A2: The host configures GPU parameters, including CPU over-provisioning coefficients and slice parameters.
[0020] A3, collects GPU model, GPU slot, GPU memory, total amount and usage information of the GPU through the GPU collection module;
[0021] A4, performs GPU resource storage and recycling;
[0022] Optionally, the setting of the slice parameters in A2 includes scanning the GPU information on the computing node to obtain GPU information that meets the requirements, and configuring the slice type of each GPU according to pre-planned parameters. The slice type includes GPU model, GPU slot, GPU video memory, total amount and usage.
[0023] Optionally, the GPU resource storage operation in A4 includes the following steps:
[0024] B1, grouped and stored by GPU model. The storage format of the GPU is as follows:
[0025] B2. After the GPU is allocated to the virtual machine, the GPU usage information is recorded and fields are saved in the GPU. The recorded GPU usage information includes the total amount of GPUs collected and the increased GPU usage. The fields are as follows:
[0026] Among them, gpu_total represents the actual total amount, which is obtained by multiplying the number of pci in the shard and gpu_id; gpu_online_used represents the GPU online usage; gpu_offline_used represents the GPU offline usage.
[0027] Optionally, in the fourth step, the GPU host weigher calculates the score of each available GPU host as follows: W = (Smax-min(U, Smax-1)*t) (3);
[0028] Where W represents the weight; Smax represents the preset maximum value; U represents the available GPU; and t represents the weighing coefficient.
[0029] Optionally, the GPU model weigher calculates the available score of each GPU model, including adding a GPU model weigher and using a scattering algorithm to evenly distribute the GPUs to different GPU models when the GPU model is not specified.
[0030] Optionally, the comprehensive scoring operation of the GPU host scale and the GPU model scale includes the following steps:
[0031] C1, determining scoring indicators, which include performance, scalability, heat dissipation capability, and power consumption;
[0032] C2, set the scoring index weight, the scoring index weight is represented by an integer between 0 and 100;
[0033] C3, normalize the scoring indicators;
[0034] C4, calculate the score of each scoring indicator;
[0035] C5, prioritize by score.
[0036] The technical effects and advantages of this application are:
[0037] This application further over-provisions the GPU resources after full GPU virtualization, stores GPU information and over-provisioning parameters in a specified format, allocates idle GPU resources to virtual machines through a scheduling algorithm to maximize the utilization of GPU resources, and calculates the total amount of virtual GPU resources through the over-provisioning coefficient. Combined with the actual available GPU resources, a list of available hosts is selected, and then the optimal host and GPU model are comprehensively selected through GPU host weighing and GPU model weighing. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] FIG1 is a schematic diagram of the overall method flow of this application. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0040] This application provides a resource over-allocation scheduling method based on GPU virtualization as shown in FIG1 . The scheduling method includes the following steps:
[0041] The first step is to create a GPU virtual machine;
[0042] The second step is to set the GPU parameter information of the virtual machine;
[0043] The third step is to schedule GPU resources and select available GPU host sequences based on the GPU parameter information of the virtual machine;
[0044] The fourth step is to perform scheduling calculations based on the stored GPU information. The scheduling calculations include the GPU host scale using a calculation formula to calculate the available GPU host scores and the GPU model scale using a calculation formula to calculate the available scores of each GPU model.
[0045] Step 5: Select the GPU host and GPU model based on the comprehensive scores of the GPU host scale and GPU model scale.
[0046] Optionally, when creating a virtual machine, you can choose to pass the GPU model parameters or not. If the GPU model is not passed, the scheduler will automatically allocate available GPU resources. During scheduling, the requested GPU model is used to calculate whether the actual usage and virtual available amount of the requested GPU model meet the requirements, thereby filtering out GPU resources that meet the requirements in the GPU scheduler.
[0047] In the second step, the GPU parameter information of the virtual machine includes the model of the GPU used in the virtual machine, the size of the memory used to store and process graphics data, the type of memory, the number of GPU cores used to process graphics data, and the calculation parameters in GPU scheduling. The calculation parameters include the GPU model to be calculated, the over-provisioning ratio, the online usage, the offline usage, the real available capacity, and the virtual available capacity.
[0048] Optionally, a GPU cannot run on multiple virtual machines at the same time. Therefore, after over-provisioning, the total GPU usage of the actually running virtual machines cannot exceed the actual total GPU usage. When the actual total GPU usage is exhausted, only after other GPU virtual machines are shut down and the over-provisioning range is met, can GPU virtual machines be created again.
[0049] The formula for calculating the real available quantity is as follows: RU=Rt-Ou (1);
[0050] The formula for calculating virtual available capacity is as follows: Vu=Rt2*c-(Ou+Fu) (2);
[0051] Where RU represents the real available capacity, Rt represents the real total capacity, Ou represents the online usage, Fu represents the offline usage, Vu represents the virtual available capacity, and c represents the over-allocation coefficient.
[0052] Optional. The real available amount indicates the number of GPUs that can be run online. The real total amount indicates the total number of GPUs that can be run online. The online usage indicates that the GPUs are actually occupied by virtual machines that are powered on. The offline usage indicates that the GPUs are not actually used by powered-off virtual machines, although they are allocated to them. The virtual available amount indicates the number of GPUs that can be run online and offline.
[0053] Scheduling GPU resources includes allocating GPUs based on allocation types. The GPU availability is scheduled based on the virtual availability. When the virtual machine pre-allocates the GPU, the GPU availability is the minimum of the real availability and the virtual availability. When the virtual machine is directly allocated, the GPU availability is the real availability. When the virtual machine is powered on, the GPU resources meet the power-on conditions.
[0054] Optionally, during the filtering phase of the scheduling process, virtual machines without GPUs do not determine whether the host has a GPU. The filtered list of available hosts will include hosts with GPUs, so virtual machines without GPUs may be scheduled to hosts with GPUs. Since GPU resources are special resources, it is necessary to ensure that only virtual machines that really need GPUs run on hosts with GPUs. Therefore, it is necessary to use a GPU host scale to more efficiently utilize GPU resources. Furthermore, the filtering phase refers to the phase of screening and filtering the GPU resource requirements of virtual machines in the virtualization platform. The filtering phase uses a series of conditions and filtering rules to screen out virtual machines that match specific requirements so that GPU resources can be allocated to them. In the filtering phase, GPU type and compatibility, virtual machine specifications and configurations, virtual machine load and performance requirements, and virtual machine priority and scheduling strategies are usually considered.
[0055] The storage operation of GPU information includes the following steps:
[0056] A1, obtain GPU resource information through the GPU resource information acquisition module. The GPU resource information acquisition module and the computing service are both deployed on the computing node;
[0057] A2: The host configures GPU parameters, including CPU over-provisioning coefficients and slice parameters.
[0058] A3, collects GPU model, GPU slot, GPU memory, total amount and usage information of the GPU through the GPU collection module;
[0059] A4, performs GPU resource storage and recycling;
[0060] Setting slice parameters in A2 involves scanning the GPU information on the computing nodes to obtain GPU information that meets the requirements, and configuring the slice type of each GPU according to pre-planned parameters. The slice type includes GPU model, GPU slot, GPU memory, total amount, and usage.
[0061] Optionally, a computing node refers to a node that can perform computing tasks, generally has high computing power and storage resources, and is usually a server or a node in a computing cluster. Deploying the GPU resource information acquisition module and computing services on the computing node can improve resource utilization efficiency and computing performance, because at this time the GPU utilization rate is high, the delay in the data transmission process is small, and the execution efficiency of the entire task is high.
[0062] The GPU resource storage operation in A4 includes the following steps:
[0063] B1 is grouped and stored by GPU model. The GPU storage format is as follows:
[0064] B2, after the GPU is allocated to the virtual machine, the GPU usage information is recorded and saved in the GPU field. The recorded GPU usage information includes the total amount of GPU collected and the increased GPU usage. The fields are as follows:
[0065] Among them, gpu_total represents the actual total amount, which is obtained by multiplying the number of pci in the shard and gpu_id; gpu_online_used represents the GPU online usage; gpu_offline_used represents the GPU offline usage.
[0066] Optional, because GPUs of the same type are the same except for the PCI number parameters, group storage is used for storage so that GPU scheduling can be performed according to the GPU model. Furthermore, the PCI number parameter of the GPU is a standard bus interface for internal expansion components of a computer, and in a GPU device, the PCI number parameter refers to the PCI Express Bus ID, which is a unique identifier of the GPU device on the PCI bus.
[0067] In the fourth step, the GPU host weigher calculates the score of each available GPU host as follows: W = (Smax-min(U,Smax-1)*t) (3);
[0068] Where W represents the weight; Smax represents the preset maximum value; U represents the available GPU; and t represents the weighing coefficient.
[0069] The GPU model weigher calculates the available score of each GPU model. When the GPU model is not specified, the GPU model weigher is added and the GPUs are evenly distributed to different GPU models using a scattering algorithm.
[0070] The comprehensive scoring operation of GPU host scale and GPU model scale includes the following steps:
[0071] C1, determine the scoring indicators, including performance, scalability, heat dissipation and power consumption;
[0072] C2, set the scoring indicator weight, which is expressed as an integer between 0 and 100;
[0073] C3, normalize the scoring indicators;
[0074] C4, calculate the score of each scoring indicator;
[0075] C5, prioritize by score.
[0076] Finally, it should be noted that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent replacements for some of the technical features therein. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A resource over-allocation scheduling method based on GPU virtualization, characterized in that: The scheduling method comprises the following steps: The first step is to create a GPU virtual machine; The second step is to set the GPU parameter information of the virtual machine; The third step is to schedule GPU resources and select available GPU host sequences based on the GPU parameter information of the virtual machine; The fourth step is to perform scheduling calculation according to the stored GPU information during scheduling, wherein the scheduling calculation includes the GPU host weigher calculating the scores of each available GPU host through a calculation formula and the GPU model weigher calculating the available scores of each model of GPU; Step 5. Select the GPU host and GPU model based on the comprehensive scores of the GPU host scale and the GPU model scale.
2. The resource over-allocation scheduling method based on GPU virtualization according to claim 1, characterized in that: The GPU parameter information of the virtual machine in the second step includes specifying the model of the GPU used in the virtual machine, the size of the memory used to store and process graphics data, the type of graphics memory, the number of GPU cores used to process graphics data, and the calculation parameters in GPU scheduling, wherein the calculation parameters include the GPU model to be calculated, the over-provisioning ratio, the online usage, the offline usage, the real available amount, and the virtual available amount.
3. The resource over-allocation scheduling method based on GPU virtualization according to claim 2, characterized in that: The actual available amount is calculated as follows: RU = Rt-Ou (1); The virtual available quantity calculation formula is as follows: Vu=Rt2*c-(Ou+Fu) (2); Among them, RU represents the real available quantity, Rt represents the real total quantity, Ou represents the online usage, Fu represents the offline usage, Vu represents the virtual available quantity, and c represents the over-allocation coefficient.
4. The resource over-allocation scheduling method based on GPU virtualization according to claim 3, characterized in that: The operation of scheduling GPU resources includes allocating GPUs according to the allocation type, the GPU availability is scheduled using the virtual availability, and when the virtual machine pre-allocates the GPU, the GPU availability takes the minimum value of the real availability and the virtual availability, and when the virtual machine is directly allocated, the GPU availability takes the real availability, and when the virtual machine is powered on, the GPU resources meet the power-on conditions.
5. The resource over-allocation scheduling method based on GPU virtualization according to claim 1, characterized in that: The storage operation of the GPU information includes the following steps: A1, obtaining GPU resource information through a GPU resource information acquisition module, wherein the GPU resource information acquisition module and the computing service are both deployed on a computing node; A2, the host configures GPU parameters, wherein the GPU parameters configured by the host include CPU over-allocation coefficient and slice parameters; A3, collects GPU model, GPU slot, GPU memory, total amount and usage information of the GPU through the GPU collection module; A4, performs GPU resource storage and recycling.
6. The resource over-allocation scheduling method based on GPU virtualization according to claim 5, characterized in that: The setting of the slice parameters in A2 includes scanning the GPU information on the computing node, obtaining the GPU information that meets the requirements, and configuring the slice type of each GPU according to the pre-planned parameters. The slice type includes GPU model, GPU slot, GPU video memory, total amount and usage.
7. The resource over-allocation scheduling method based on GPU virtualization according to claim 5, characterized in that: The GPU resource storage operation in A4 includes the following steps: B1, grouped and stored by GPU model, the storage format of the GPU is as follows: B2, after the GPU is allocated to the virtual machine, the GPU usage information is recorded and fields are saved in the GPU. The recorded GPU usage information includes the total amount of GPUs collected and the increased GPU usage. The fields are as follows: Among them, gpu_total represents the actual total amount, which is obtained by multiplying the number of pci in the shard and gpu_id; gpu_online_used represents the online usage of the GPU; gpu_offline_used represents the offline usage of the GPU.
8. The resource over-allocation scheduling method based on GPU virtualization according to claim 1, characterized in that: In the fourth step, the GPU host weigher calculates the score of each available GPU host as follows: W = (S max-min (U, S max-1) * t) (3); Where W represents weight; Smax represents the preset maximum value; U represents available GPU; and t represents the weighing coefficient.
9. The resource over-allocation scheduling method based on GPU virtualization according to claim 1, characterized in that: The GPU model weigher calculates the available score of each model of GPU, including adding a GPU model weigher and using a scattering algorithm to evenly distribute the GPUs to different GPU models when the GPU model is not specified.
10. The resource over-allocation scheduling method based on GPU virtualization according to claim 1, characterized in that: The comprehensive scoring operation of the GPU host weighing device and the GPU model weighing device includes the following steps: C1, determining scoring indicators, wherein the scoring indicators include performance, scalability, heat dissipation capability and power consumption; C2, set the scoring index weight, the scoring index weight is represented by an integer between 0 and 100; C3, normalize the scoring indicators; C4, calculate the score of each scoring indicator; C5, prioritize by score.
Citation Information
Patent Citations
Graphics processing resource allocation method and device, equipment and storage medium
CN113342534A
Method for realizing mixed type virtualization by GPU (Graphics Processing Unit) with multiple virtualization types
CN114816746A
Resource overconfiguration scheduling method based on GPU virtualization
CN117851043A
GPU resource usage display and dynamic GPU resource allocation in a networked virtualization system
US10176550B1
Dynamic allocation of physical graphics processing units to virtual machines
WO2014100558A1