Cloud platform GPU virtual machine scheduling method and device
By dynamically migrating GPU resources in the cloud platform to form an effective channel set, the problems of low GPU resource utilization and high scheduling complexity in the virtual machine environment are solved, and efficient resource scheduling and performance improvement are achieved.
Patent Information
- Application Number
- CN202510724912.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In a virtual machine environment, GPU resource utilization is low and scheduling complexity is high, resulting in high performance computing tasks execution efficiency and resource waste.
By obtaining the number of target GPU resources required for the new task, we judge the remaining resources in the channel, migrate resources dynamically, and prioritize migration from channels with many remaining resources to form an effective channel set, and use preset algorithms to perform resource scheduling to avoid fragmentation of resources in a single channel.
It improves GPU resource utilization, reduces scheduling complexity, meets the needs of high-performance computing tasks, and avoids resource waste and performance bottlenecks.
Smart Images

Figure CN120234099A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a method and device for scheduling GPU virtual machines in a cloud platform. Background Art
[0002] Driven by the continuous advancement of the current digital wave, high-performance computing tasks such as deep learning model training in the field of artificial intelligence and complex data processing in big data analysis are booming. These tasks have extremely stringent requirements for computing resources, especially the increasing dependence on GPUs (Graphics Processing Units). With its powerful parallel computing capabilities, GPUs can significantly improve the processing efficiency of complex computing tasks and have become the core hardware resources supporting high-performance computing.
[0003] However, in the common computing architecture of the virtual machine environment, the scheduling of multi-GPU faces a series of problems that need to be solved urgently. First of all, the problem of low GPU resource utilization is particularly prominent. Due to the inflexible resource allocation mechanism between virtual machines and GPUs, there are often situations where GPU resources are idle or over-allocated. For example, in some business scenarios, some virtual machines are allocated more GPU resources than actually needed, resulting in resource waste; while other virtual machines cannot fully utilize their computing capabilities due to insufficient resources. Secondly, although cross-channel passthrough provides a direct access path for virtual machines to GPUs, this method will bring relatively large performance overhead. During data transmission, frequent cross-channel operations will consume a large amount of system resources, reduce data transmission efficiency, and thus affect the execution speed of the entire computing task. In addition, the high complexity of multi-GPU scheduling cannot be ignored. As the number of GPUs increases, the scheduling algorithm needs to comprehensively consider various factors such as task priority, resource load, and GPU performance differences, which makes the design and implementation of the scheduling algorithm extremely complex and difficult to effectively meet the requirements of high-performance computing tasks for efficient and stable allocation of GPU resources.
[0004] Current traditional scheduling methods cannot balance system resources well in the face of these complex situations. In the scenario of using multi-GPU, this scheduling method significantly affects the performance of virtual machines, and the powerful computing capabilities of GPUs cannot be fully utilized. This not only restricts the execution efficiency of high-performance computing tasks but also causes a great waste of hardware resources. Summary of the Invention
[0005] The object of the present invention is to solve the above problems and propose a method and device for scheduling GPU virtual machines in a cloud platform.
[0006] In the first aspect of the implementation of the present invention, a method for scheduling GPU virtual machines in a cloud platform is first proposed. The method includes:
[0007] Obtain the number of target GPU resources required for the new task. For the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; there is a fixed number of GPU resources in each channel.
[0008] If the total remaining GPU resources are greater than the number of target GPU resources, determine whether there is a single channel that meets the number of target GPU resources.
[0009] If there is no single channel that meets the number of target GPU resources, sort all channels in descending order of the remaining GPU resources in each channel to obtain an initial resource channel sorting set.
[0010] Obtain the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set.
[0011] Dynamically migrate the resources in the effective channel set through a preset algorithm to obtain a migration result; the migration result includes generating a target migration channel or not generating a target migration channel; the target migration channel is a channel that meets the number of target GPU resources.
[0012] If the migration result is to generate a target migration channel, schedule all the number of target GPU resources of the new task to the target migration channel.
[0013] Optionally, after determining whether the total remaining GPU resources in all channels are greater than the number of target GPU resources, it further includes: if the total remaining GPU resources are less than or equal to the number of target GPU resources, determine that the new task execution fails and send an error message to the user.
[0014] Optionally, after determining whether there is a single channel that meets the number of target GPU resources, it further includes:
[0015] If there is a single channel that meets the number of target GPU resources, obtain all the channels that meet the requirements. For all the channels that meet the requirements, determine the usage score of the channel through the GPU video memory occupancy rate, CPU usage rate, and memory usage rate.
[0016] Obtain the channel with the highest usage score as the execution channel of the new task.
[0017] Optionally, dynamically migrating the resources in the effective channel set through a preset algorithm to obtain a migration result includes:
[0018] For each effective channel in the effective channel set, calculate the number of resources missing for the effective channel to meet the new task execution to obtain a resource missing set.
[0019] For each valid channel corresponding to the resource shortage in the valid channel set, determine the virtual machines that meet the conditions in the valid channel according to the resource shortage to obtain the set of virtual machines to be migrated;
[0020] For each virtual machine to be migrated in the set of virtual machines to be migrated, determine the migration cost of the virtual machine to be migrated according to the video memory occupancy rate and usage rate of the virtual machine to be migrated;
[0021] Obtain the virtual machine to be migrated with the lowest migration cost as the initial virtual machine to be migrated for the valid channel;
[0022] For each initial virtual machine to be migrated corresponding to each channel, calculate the migration severity of the initial virtual machine to be migrated, and obtain the initial virtual machine to be migrated with the lowest migration severity as the initial target virtual machine to be migrated;
[0023] If the initial target virtual machine to be migrated can be migrated to a valid channel other than the current channel, the initial target virtual machine to be migrated is used as the target migration channel, and the migration result is to generate the target migration channel; the migration result is that the target migration channel is not generated.
[0024] Optionally, calculating the migration severity of the initial virtual machine to be migrated includes:
[0025] Determine the average video memory occupancy rate, average usage rate of the GPU card being used by the initial virtual machine to be migrated, and the number of resources missing in the channel where the initial virtual machine to be migrated is located to determine the migration severity;
[0026] Through the formula Obtain the migration severity of the initial virtual machine to be migrated;
[0027] Among them, R is the migration severity of the initial virtual machine to be migrated, r1, r2, and r3 are weight coefficients respectively, M is the number of resources missing in the channel where the initial virtual machine to be migrated is located, Q is the average video memory occupancy rate of the GPU card being used, and P is the average usage rate of the GPU card being used.
[0028] In the second aspect of the implementation of the present invention, a cloud platform GPU virtual machine scheduling device is proposed, including:
[0029] The GPU resource total determination module is used to obtain the number of target GPU resources required for the new task, and for the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; there is a fixed number of GPU resources in each channel;
[0030] The channel judgment module is used to, if the total remaining GPU resources are greater than the number of target GPU resources, determine whether there is a single channel that meets the number of target GPU resources;
[0031] A channel sorting module, which is used to sort all channels from more to less remaining GPU resource numbers of each channel to obtain an initial resource channel sorting set if there is no single channel that meets the number of target GPU resources;
[0032] An effective channel determination module, which is used to obtain the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set;
[0033] A resource migration module, which is used to dynamically migrate the resources in the effective channel set through a preset algorithm to obtain a migration result; the migration result includes generating a target migration channel or not generating a target migration channel; the target migration channel is a channel that meets the number of target GPU resources;
[0034] A new task scheduling determination module, which is used to schedule all the target GPU resource numbers of the new task to the target migration channel if the migration result is to generate a target migration channel.
[0035] Optionally, the device further includes:
[0036] An error message sending module, which is used to determine that the new task execution fails and send an error message to the user if the total remaining GPU resources are less than or equal to the number of target GPU resources.
[0037] Optionally, the device further includes:
[0038] A channel scoring determination module, which is used to obtain all channels that meet the number of target GPU resources if there is a single channel that meets the number of target GPU resources, and determine the usage score of the channel through the GPU video memory occupancy rate, CPU usage rate, and memory usage rate for all channels that meet the requirements;
[0039] An execution channel determination module, which is used to obtain the channel with the highest usage score as the execution channel of the new task.
[0040] Optionally, the resource migration module includes:
[0041] A resource shortage set determination module, which is used to calculate the resource quantity lacking for each effective channel in the effective channel set to meet the execution of the new task to obtain a resource shortage set;
[0042] A virtual machine set to be migrated generation determination module, which is used to determine the virtual machines that meet the conditions in the effective channel for the resource shortage corresponding to each effective channel in the effective channel set to obtain a virtual machine set to be migrated generation;
[0043] A migration cost determination module, which is used to determine the migration cost of each virtual machine to be migrated generation in the virtual machine set to be migrated generation according to the video memory occupancy rate and usage rate of the virtual machine to be migrated generation;
[0044] An initial migration virtual machine determination module, configured to obtain the virtual machine to be migrated with the lowest migration cost as the initial migration virtual machine of the effective channel;
[0045] A migration severity determination module, configured to calculate the migration severity of the initial migration virtual machine corresponding to each channel, and obtain the initial migration virtual machine with the lowest migration severity as the initial target migration virtual machine;
[0046] A migration result determination module, configured to: if the initial target migration virtual machine can be migrated to an effective channel other than the current channel, use the initial target migration virtual machine as the target migration channel, and the migration result is to generate a target migration channel; the migration result is that no target migration channel is generated.
[0047] Optionally, the migration severity determination module includes:
[0048] A parameter determination module, configured to determine the average video memory occupancy rate, average usage rate of the GPU card being used by the initial migration virtual machine, and the number of resources missing in the channel where the initial migration virtual machine is located to determine the migration severity;
[0049] A migration severity calculation module, configured to obtain the migration severity of the initial migration virtual machine through the formula ;
[0050] wherein, R is the migration severity of the initial migration virtual machine, r1, r2, and r3 are weight coefficients respectively, M is the number of resources missing in the channel where the initial migration virtual machine is located, Q is the average video memory occupancy rate of the GPU card being used, and P is the average usage rate of the GPU card being used.
[0051] Advantages of the present invention:
[0052] The present invention provides a method for scheduling GPU virtual machines on a cloud platform, which obtains the number of target GPU resources required for a new task. For the number of target GPU resources, it determines whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; if the total remaining GPU resources are greater than the number of target GPU resources, it determines whether there is a single channel that meets the number of target GPU resources; if there is no single channel that meets the number of target GPU resources, it sorts all channels from more to less according to the remaining GPU resources in each channel to obtain an initial resource channel sorting set; it obtains the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set; it performs dynamic migration on the resources in the effective channel set through a preset algorithm to obtain a migration result; if the migration result is to generate a target migration channel, it schedules all the target GPU resources of the new task to the target migration channel. By dynamically migrating idle resources across channels, it avoids resource fragmentation in a single channel, does not rely on the sufficiency of resources in a single channel, preferentially migrates from channels with more remaining resources, can balance the load of each channel, improve the utilization rate of GPU resources, reduce the scheduling complexity, and meet the requirements of high-performance computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The present invention will be further described below with reference to the accompanying drawings.
[0054] Figure 1 It is a flowchart of a method for scheduling GPU virtual machines on a cloud platform provided by an embodiment of the present invention;
[0055] Figure 2 It is a flowchart of a resource migration provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic structural diagram of a device for scheduling GPU virtual machines on a cloud platform provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0058] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] An embodiment of the present invention provides a method for scheduling GPU virtual machines on a cloud platform. Refer to Figure 1 , Figure 1 It is a flowchart of a method for scheduling GPU virtual machines on a cloud platform provided by an embodiment of the present invention. The method includes the following steps:
[0060] S101. Obtain the number of target GPU resources required for the new task. For the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources.
[0061] S102. If the total remaining GPU resources are greater than the number of target GPU resources, then determine whether there is a single channel that meets the number of target GPU resources.
[0062] S103. If there is no single channel that meets the number of target GPU resources, then sort all channels in descending order according to the remaining GPU resources in each channel to obtain an initial resource channel sorting set.
[0063] S104. Obtain the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set.
[0064] S105. Dynamically migrate the resources in the effective channel set through a preset algorithm to obtain a migration result.
[0065] S106. If the migration result is to generate a target migration channel, then schedule all the target GPU resources of the new task to the target migration channel.
[0066] Among them, there is a fixed number of GPU resources in each channel; the migration result includes generating a target migration channel or not generating a target migration channel; the target migration channel is a channel that meets the number of target GPU resources.
[0067] Based on a GPU virtual machine scheduling method for a cloud platform provided by an embodiment of the present invention, by dynamically migrating idle resources across channels, it avoids resource fragmentation in a single channel, does not need to rely on the sufficiency of resources in a single channel, preferentially migrates from channels with more remaining resources, can balance the load of each channel, improve the utilization rate of GPU resources, reduce the scheduling complexity, and meet the requirements of high-performance computing tasks.
[0068] In one implementation, by dynamically migrating idle resources across channels, it avoids resource fragmentation in a single channel (such as when the total remaining resources are sufficient but the resources in a single channel are insufficient), improving the overall resource utilization rate; it does not need to rely on the sufficiency of resources in a single channel, and meets task requirements by combining resources from multiple channels, adapting to diverse resource allocation scenarios.
[0069] In one implementation, this solution includes channel division, virtual machine-channel binding, in-channel resource scheduling, inter-channel load balancing, and intelligent scheduling; Channel division: Divide the physical GPU cards on each server into multiple logical channels, and each channel corresponds to an independent subset of GPU resources; Virtual machine-channel binding: According to the GPU resource requirements of the virtual machine, schedule the virtual machine to a specific channel to ensure that the GPU resources used by the virtual machine come from the same channel; In-channel resource scheduling: Inside the channel, allocate GPU resources to the virtual machine in a direct-pass manner to avoid resource switching overhead and improve GPU utilization and virtual machine performance; Inter-channel load balancing: According to the load conditions of system CPU, memory, GPU, etc., dynamically calculate the remaining resource weights of each channel on each server, and then adjust the virtual machine allocation between channels to achieve load balancing of system resources; Intelligent scheduling (preset algorithm): During the scheduling process, give priority to scheduling virtual machines using multi-GPU to the same channel to avoid cross-channel direct-pass problems. If it is found during the scheduling process that the remaining resources of each channel cannot meet the multi-GPU requirements of a single virtual machine, then turn on the intelligent scheduling (preset algorithm).
[0070] In one implementation, the intelligent scheduling (preset algorithm) includes calculating the number N of remaining GPU resources on the computing platform t of the number N0 of GPU resources that meet the scheduling requirements. If N t < N0, the scheduling fails; if N t > N0, query the number of remaining GPU resources of each channel (channel 1, channel 2,..., channel n), denoted as N1, N2,..., N n , then N t = N1 + N2 +... + N n , and N m < N t (m belongs to (1, 2,..., n)), that is, no single channel meets the scheduling requirements. To avoid cross-channel scheduling, start the virtual machine dynamic migration process, select the 3 channels (the first preset number, determined by technicians, and take 3 as the first preset number as an example) with the most remaining GPU resources on the platform, denoted as N t1 , N t2 , N t3 satisfy (N t1 , N t2 , N t3 ) = sorted(N1, N2,..., N n ), calculate the resource quantities lacking for the virtual machine scheduling to be satisfied by the GPU resources of the corresponding three channels, denoted as M1, M2, M3 respectively. M1 = N t - N t1 , M 2= N t - N t2 , M3 = Nt -N t3 Search for virtual machines that meet the conditions in these three channels respectively. Assume that the optimal channel virtual machines 1 (VM1), 2 (VM2), and 3 (VM3) are found in the three channels respectively, and record the number of GPU resources of each virtual machine as N vm1 , N vm2 , N vm3 , meet the conditions: N vm1 ≥M1, N vm2 ≥M2, N vm3 ≥M3. According to the records of the monitoring system, calculate the average video memory occupancy rate of the GPU cards being used by virtual machines 1 (VM1), 2 (VM2), and 3 (VM3), denoted as Q1, Q2, Q3. Calculate the average utilization rate of the GPU cards being used by virtual machines 1 (VM1), 2 (VM2), and 3 (VM3), denoted as P1, P2, P3; Calculate the migration severity R1, R2, R3 of intelligent scheduling for the above three virtual machines respectively; R1 = r1M1 + r2Q1 + r3P1; R2 = r1M2 + r2Q2 + r3P2; R3 = r1M3 + r2Q3 + r3P3; where r1, r2, r3 are the weight values of the number of GPU resources, the weight value of GPU video memory occupancy, and the weight value of GPU utilization rate respectively. r1, r2, r3 are constants, usually 1, 1, 1, and the specific values are determined by technical personnel. Finally, select the one with the least migration severity among R1, R2, R3, and then execute the migration command on the virtual machine so that the channel where the virtual machine is located can meet the requirements of this scheduling.
[0071] In one implementation, migrate from the channel with more remaining resources first, which can balance the loads of each channel, avoid some channels being overly busy and some channels being idle, and extend the hardware life; The preset algorithm and sorting mechanism can be dynamically adjusted according to the real-time resource status, quickly respond to the resource requirements of sudden tasks, and reduce scheduling delay.
[0072] In one embodiment, after determining whether the total remaining GPU resources in all channels are greater than the target number of GPU resources, it further includes: If the total remaining GPU resources are less than or equal to the target number of GPU resources, determine that the execution of the new task fails and send an error message to the user.
[0073] In one implementation, when the total global resources are insufficient, immediately determine that the task execution fails and give feedback on the error, avoid the system attempting ineffective resource allocation or migration operations, reduce the waste of computing resources and scheduling delay; Expose the resource bottleneck to the user in a timely manner, help the user quickly locate the problem (such as insufficient resources), facilitate the user to adjust the task parameters (such as splitting the task, reducing resource requirements) or apply for expansion, and improve the interaction efficiency.
[0074] In one embodiment, after determining whether there is a single channel that meets the target number of GPU resources, the following steps are further included:
[0075] If there is a single channel that meets the target number of GPU resources, obtain all the channels that meet the requirement. For all the channels that meet the requirement, determine the usage score of the channel based on the GPU video memory occupancy rate, CPU usage rate, and memory usage rate;
[0076] Obtain the channel with the highest usage score as the execution channel for the new task.
[0077] In one implementation, instead of relying solely on the single condition of "sufficient remaining GPU resources", multiple metrics such as the GPU video memory occupancy rate, CPU usage rate, and memory usage rate are combined to comprehensively evaluate the channel load, avoiding scheduling tasks to channels where "GPU resources are sufficient but other resources are fully loaded", and preventing task performance degradation or channel crashes caused by resource imbalance.
[0078] In one implementation, the channel with the lowest overall system load in the current system (such as low video memory occupancy, low CPU, and low memory pressure) is selected through the "usage score" to provide a better operating environment for the new task, reduce resource competition, and improve the task execution efficiency and stability. For example, if a certain channel has sufficient remaining GPU resources but the CPU usage rate has reached 90%, scheduling a new task at this time may cause the GPU resources to be idle due to the CPU bottleneck, while the scoring mechanism will preferentially select a channel with lower CPU and memory loads to achieve coordinated resource utilization.
[0079] In one implementation, the usage score of the channel is determined based on the GPU video memory occupancy rate, CPU usage rate, and memory usage rate. The usage score is obtained through the formula Z = 1 / (aA + bB + cC), where Z is the usage score, and the weights of a, b, and c are 0.4, 0.3, and 0.3 respectively, A is the GPU video memory occupancy rate, B is the CPU usage rate, and C is the memory usage rate.
[0080] In one implementation, assume that the cloud platform now needs to schedule a new GPU virtual machine that requires 4 GPU cards. Then, check whether there is a single channel in the cloud platform whose remaining resources meet the requirement of 4 GPU cards. At this time, it is found that multiple channels can meet the requirement. Then, within the channels that meet the requirement, determine the usage score of the channel based on the three dimensions of the GPU video memory occupancy rate, CPU usage rate, and memory usage rate, and select the optimal channel to schedule the virtual machine.
[0081] In one implementation, the score is calculated dynamically based on real-time monitoring data, which can respond in real time to channel load fluctuations (such as the release of resources when other tasks end), ensuring that each scheduling is based on the latest system state and avoiding resource mismatches caused by static allocation; the automated scoring mechanism replaces the complexity of manual judgment of channel load and reduces the cost of manual adjustment of task allocation by operation and maintenance personnel.
[0082] In one embodiment, refer to Figure 2 , Figure 2 A flowchart of resource migration is provided. Through a preset algorithm, dynamic migration of resources in the set of valid channels is performed to obtain a migration result, including:
[0083] S1051. For each valid channel in the set of valid channels, calculate the amount of resources missing for the execution of the new task by this valid channel to obtain a resource missing set;
[0084] S1052. For the resource missing corresponding to each valid channel in the set of valid channels, determine the virtual machines that meet the conditions in this valid channel according to this resource missing to obtain a set of virtual machines to be migrated;
[0085] S1053. For each virtual machine to be migrated in the set of virtual machines to be migrated, determine the migration cost of this virtual machine to be migrated according to the video memory occupancy rate and usage rate of this virtual machine to be migrated;
[0086] S1054. Obtain the virtual machine to be migrated with the lowest migration cost as the initial virtual machine to be migrated for this valid channel;
[0087] S1055. For the initial virtual machine to be migrated corresponding to each channel, calculate the migration severity of this initial virtual machine to be migrated, and obtain the initial virtual machine to be migrated with the lowest migration severity as the initial target virtual machine to be migrated;
[0088] S1056. If the initial target virtual machine to be migrated can be migrated to a valid channel other than the current channel, then the initial target virtual machine to be migrated is used as the target migration channel, and the migration result is to generate a target migration channel; the migration result is that no target migration channel is generated.
[0089] In one implementation, by identifying and migrating virtual machines (VMs) with low utilization rate, fragmented resources are integrated into a centralized and available resource block, avoiding the purchase of new hardware due to local resource shortage, and significantly reducing the cost of IT infrastructure; determining the migration cost of this virtual machine to be migrated according to the video memory occupancy rate and usage rate of this virtual machine to be migrated is to sum the video memory occupancy rate and usage rate as the migration cost of the virtual machine to be migrated.
[0090] In one implementation, the migration cost assessment (video memory occupancy rate, usage rate) ensures that VMs with the least impact on the business are migrated first. For example, VMs with low video memory occupancy and low CPU usage are migrated first to avoid interrupting critical services. The migration severity assessment further filters out the VMs that have the least impact on the source channel performance, ensuring that the source channel can still operate stably after migration.
[0091] In one implementation, by breaking the resource isolation between channels, cross-channel migration realizes global resource pooling. Even if the resources of a single channel are insufficient, the system can still meet the task requirements through dynamic adjustment, enhancing the ability to handle sudden resource demands; the dual screening mechanism (lowest cost and lowest severity) ensures that each migration decision weighs the resource benefits and business risks, avoiding service interruption or performance fluctuations caused by blind migration.
[0092] In one embodiment, calculating the migration severity of the initial migration virtual machine includes:
[0093] Determine the average video memory occupancy rate, average usage rate of the GPU card used by the initial migration virtual machine, and the amount of resources lacking in the channel where the initial migration virtual machine is located to determine the migration severity;
[0094] Through the formula Obtain the migration severity of the initial migration virtual machine;
[0095] Wherein, R is the migration severity of the initial migration virtual machine, r1, r2, and r3 are weight coefficients respectively, M is the amount of resources lacking in the channel where the initial migration virtual machine is located, Q is the average video memory occupancy rate of the GPU card in use, and P is the average usage rate of the GPU card in use.
[0096] In one implementation, by combining the average video memory occupancy rate, average usage rate of the GPU card, and the initial migration virtual machine into a single metric, the rationality of screening migration virtual machines is improved, and subjective errors are reduced.
[0097] In one implementation, through the mathematical model of this formula, the complex resource scheduling problem is transformed into a quantifiable optimization problem, maximizing resource reuse while ensuring business stability.
[0098] Based on the same inventive concept, the embodiments of the present invention also provide a cloud platform GPU virtual machine scheduling device. Refer to Figure 3 , Figure 3 which is a schematic structural diagram of a cloud platform GPU virtual machine scheduling device provided by the embodiments of the present invention, including:
[0099] The total GPU resource determination module is used to obtain the number of target GPU resources required for a new task, and for the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; there is a fixed number of GPU resources in each channel.
[0100] The channel judgment module is used to, if the total remaining GPU resources are greater than the number of target GPU resources, determine whether there is a single channel that meets the number of target GPU resources.
[0101] The channel sorting module is used to, if there is no single channel that meets the number of target GPU resources, sort all channels from more to less according to the remaining GPU resources in each channel to obtain an initial resource channel sorting set.
[0102] The valid channel determination module is used to obtain the first preset number of channels in the initial resource channel sorting set to obtain a valid channel set.
[0103] The resource migration module is used to dynamically migrate the resources in the valid channel set through a preset algorithm to obtain a migration result; the migration result includes generating a target migration channel or not generating a target migration channel; the target migration channel is a channel that meets the number of target GPU resources.
[0104] The new task scheduling determination module is used to, if the migration result is to generate a target migration channel, schedule all the number of target GPU resources of the new task to the target migration channel.
[0105] Based on the GPU virtual machine scheduling device provided by the embodiment of the present invention, by dynamically migrating idle resources across channels, it avoids resource fragmentation in a single channel, does not rely on the resource sufficiency of a single channel, preferentially migrates from channels with more remaining resources, can balance the load of each channel, improve the utilization rate of GPU resources, reduce the scheduling complexity, and meet the requirements of high-performance computing tasks.
[0106] In one embodiment, the device further includes:
[0107] The error message sending module is used to, if the total remaining GPU resources are less than or equal to the number of target GPU resources, determine that the new task execution fails and send an error message to the user.
[0108] In one embodiment, the device further includes:
[0109] The channel score determination module is used to, if there is a single channel that meets the number of target GPU resources, obtain all the channels that meet, and for all the channels that meet, determine the usage score of the channel through the GPU video memory occupancy rate, CPU usage rate, and memory usage rate.
[0110] An execution channel determination module, configured to obtain the channel with the highest usage score as the execution channel for the new task.
[0111] In one embodiment, the resource migration module includes:
[0112] A resource shortage set determination module, configured to calculate, for each valid channel in the set of valid channels, the number of resources missing for the new task execution to be satisfied by the valid channel, to obtain a resource shortage set;
[0113] A generation migration virtual machine set determination module, configured to determine, for the resource shortage corresponding to each valid channel in the set of valid channels, the virtual machines meeting the conditions in the valid channel to obtain a generation migration virtual machine set;
[0114] A migration cost determination module, configured to determine, for each generation migration virtual machine in the generation migration virtual machine set, the migration cost of the generation migration virtual machine according to the video memory occupancy rate and usage rate of the generation migration virtual machine;
[0115] An initial migration virtual machine determination module, configured to obtain the generation migration virtual machine with the lowest migration cost as the initial migration virtual machine for the valid channel;
[0116] A migration severity determination module, configured to calculate, for the initial migration virtual machine corresponding to each channel, the migration severity of the initial migration virtual machine, and obtain the initial migration virtual machine with the lowest migration severity as the initial target migration virtual machine;
[0117] A migration result determination module, configured to, if the initial target migration virtual machine can be migrated to a valid channel other than the current channel, use the initial target migration virtual machine as the target migration channel, and the migration result is to generate a target migration channel; the migration result is that no target migration channel is generated.
[0118] In one embodiment, the migration severity determination module includes:
[0119] A parameter determination module, configured to determine the average video memory occupancy rate, average usage rate of the GPU card being used by the initial migration virtual machine, and the number of resources missing in the channel where the initial migration virtual machine is located to determine the migration severity;
[0120] A migration severity calculation module, configured to obtain the migration severity of the initial migration virtual machine through the formula ;
[0121] where R is the migration severity of the initial migration virtual machine, r1, r2, and r3 are weight coefficients respectively, M is the number of resources missing in the channel where the initial migration virtual machine is located, Q is the average video memory occupancy rate of the GPU card being used, and P is the average usage rate of the GPU card being used.
[0122] The above has described in detail an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. A GPU virtual machine scheduling method for a cloud platform, characterized in that, The method includes: Obtain the number of target GPU resources required for the new task. For the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; there is a fixed number of GPU resources in each channel. If the total remaining GPU resources are greater than the number of target GPU resources, determine whether there is a single channel that meets the number of target GPU resources. If there is no single channel that meets the number of target GPU resources, sort all channels in descending order according to the remaining GPU resources in each channel to obtain an initial resource channel sorting set. Obtain the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set. Dynamically migrate the resources in the effective channel set through a preset algorithm to obtain a migration result; the migration result includes generating a target migration channel or not generating a target migration channel. The target migration channel is a channel that meets the number of target GPU resources. If the migration result is to generate a target migration channel, schedule all the target GPU resources of the new task to the target migration channel.
2. The cloud platform GPU virtual machine scheduling method according to claim 1, wherein After determining whether the total remaining GPU resources in all channels are greater than the number of target GPU resources, it further includes: If the total remaining GPU resources are less than or equal to the number of target GPU resources, determine that the new task execution fails and send an error message to the user.
3. A cloud platform GPU virtual machine scheduling method according to claim 1, characterized in that, After determining whether there is a single channel that meets the number of target GPU resources, it further includes: If there is a single channel that meets the number of target GPU resources, obtain all the channels that meet the requirements. For all the channels that meet the requirements, determine the usage score of the channel through the GPU video memory occupancy rate, CPU usage rate, and memory usage rate. Obtain the channel with the highest usage score as the execution channel of the new task.
4. A cloud platform GPU virtual machine scheduling method according to claim 1, characterized in that Dynamically migrating the resources in the effective channel set through a preset algorithm to obtain a migration result includes: For each effective channel in the effective channel set, calculate the number of resources lacking for the effective channel to meet the execution of the new task to obtain a resource shortage set. For the resource shortage corresponding to each effective channel in the effective channel set, determine the virtual machines that meet the conditions in the effective channel according to the resource shortage to obtain a virtual machine set to be migrated. For each virtual machine to be migrated in the virtual machine set to be migrated, determine the migration cost of the virtual machine to be migrated according to the video memory occupancy rate and usage rate of the virtual machine to be migrated. Obtain the virtual machine to be migrated with the lowest migration cost as the initial virtual machine to be migrated for the effective channel. For the initial virtual machine to be migrated corresponding to each channel, calculate the migration severity of the initial virtual machine to be migrated, and obtain the initial virtual machine to be migrated with the lowest migration severity as the initial target virtual machine to be migrated. If the initial target virtual machine to be migrated can be migrated to an effective channel other than the current channel, the initial target virtual machine to be migrated is used as the target migration channel, and the migration result is to generate a target migration channel; the migration result is not to generate a target migration channel.
5. A cloud platform GPU virtual machine scheduling method according to claim 4, characterized in that Calculating the migration severity of the initial virtual machine to be migrated includes: Determine the average video memory occupancy rate, average utilization rate of the GPU card used by the initial migration virtual machine, and the number of resources lacking in the channel where the initial migration virtual machine is located to determine the migration severity; Through the formula obtain the migration severity of the initial migrated virtual machine; Among them, R is the migration severity of the initial migration virtual machine, r1, r2, and r3 are weight coefficients respectively, M is the number of resources lacking in the channel where the initial migration virtual machine is located, Q is the average video memory occupancy rate of the GPU card in use, and P is the average utilization rate of the GPU card in use.
6. A GPU virtual machine scheduling device for a cloud platform, characterized in that, The device includes: A total GPU resource determination module, configured to obtain the number of target GPU resources required for a new task, and for the number of target GPU resources, determine whether the total remaining GPU resources in all channels are greater than the number of target GPU resources; there is a fixed number of GPU resources in each channel; A channel determination module, configured to determine whether there is a single channel that meets the number of target GPU resources if the total remaining GPU resources are greater than the number of target GPU resources; A channel sorting module, configured to sort all channels from more to less according to the remaining GPU resources of each channel to obtain an initial resource channel sorting set if there is no single channel that meets the number of target GPU resources; An effective channel determination module, configured to obtain the first preset number of channels in the initial resource channel sorting set to obtain an effective channel set; A resource migration module, configured to perform dynamic migration on the resources in the effective channel set through a preset algorithm to obtain a migration result; the migration result includes generating a target migration channel or not generating a target migration channel; the target migration channel is a channel that meets the number of target GPU resources; A new task scheduling determination module, configured to schedule all the number of target GPU resources of the new task to the target migration channel if the migration result is to generate a target migration channel.
7. The cloud platform GPU virtual machine scheduling device according to claim 6, wherein The device further includes: An error message sending module, configured to determine that the new task execution fails and send an error message to the user if the total remaining GPU resources are less than or equal to the number of target GPU resources.
8. The cloud platform GPU virtual machine scheduling device according to claim 6, characterized in that, The device further includes: A channel score determination module, configured to obtain all the channels that meet the requirements if there is a single channel that meets the number of target GPU resources, and determine the usage score of the channel through the GPU video memory occupancy rate, CPU utilization rate, and memory utilization rate for all the channels that meet the requirements; An execution channel determination module, configured to obtain the channel with the highest usage score as the execution channel of the new task.
9. The cloud platform GPU virtual machine scheduling device according to claim 6, wherein, The resource migration module includes: A resource shortage set determination module, configured to calculate the number of resources lacking for each effective channel in the effective channel set to meet the execution of the new task to obtain a resource shortage set; A generation migration virtual machine set determination module, configured to determine the virtual machines that meet the conditions in the effective channel according to the resource shortage for the resource shortage corresponding to each effective channel in the effective channel set to obtain a generation migration virtual machine set; A migration cost determination module, configured to determine the migration cost of each generation migration virtual machine in the generation migration virtual machine set according to the video memory occupancy rate and utilization rate of the generation migration virtual machine; Initial migration virtual machine determination module, which is used to obtain the virtual machine to be migrated with the lowest migration cost as the initial migration virtual machine of the effective channel; Migration severity determination module, which is used to calculate the migration severity of the initial migration virtual machine corresponding to each channel, and obtain the initial migration virtual machine with the lowest migration severity as the initial target migration virtual machine; Migration result determination module, which is used to, if the initial target migration virtual machine can be migrated to an effective channel other than the current channel, use the initial target migration virtual machine as the target migration channel, and the migration result is to generate a target migration channel; the migration result is that no target migration channel is generated.
10. The cloud platform GPU virtual machine scheduling device according to claim 9, wherein The migration severity determination module includes: Parameter determination module, which is used to determine the average video memory occupancy rate, average usage rate of the GPU card being used by the initial migration virtual machine, and the number of resources lacking in the channel where the initial migration virtual machine is located to determine the migration severity; The migration severity calculation module is used to obtain the migration severity of the initial migration virtual machine through the formula ; Among them, R is the migration severity of the initial migration virtual machine, r1, r2, and r3 are weight coefficients respectively, M is the number of resources lacking in the channel where the initial migration virtual machine is located, Q is the average video memory occupancy rate of the GPU card being used, and P is the average usage rate of the GPU card being used.
Citation Information
Patent Citations
Virtual machine resource scheduling method and system for server clusters
CN103605574A
Method and device for resource scheduling
CN107491352A
Virtual machine allocation method and system, computer equipment and storage medium
CN119645559A
Physical machine resource utilization rate balancing method and device, equipment and medium
CN119987950A