Scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption
By performing priority evaluation of computing tasks and tree resource pool management, dynamic allocation of heterogeneous GPU resources is solved, and the problem of inefficiency of traditional resource pool technology in heterogeneous GPU management and scheduling is achieved, efficient utilization and flexible management of resources are achieved.
Patent Information
- Application Number
- CN202510255226.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional resource pooling technology is difficult to effectively cope with the complexity and diversity of heterogeneous GPUs, resulting in low resource management and scheduling efficiency, especially in mimicry computing scenarios, which is difficult to achieve unified management and efficient utilization of resources.
By evaluating the priority coefficient of the computing task, generating a tree resource pool, configuring drivers for heterogeneous resources, monitoring the state of GPU load in real time, and dynamic allocation and scheduling of resources according to task requirements and trends, including re-division of priority and extended resource pools.
It improves the utilization rate and computing performance of heterogeneous GPU resources, realizes flexible management and efficient scheduling of resources, and adapts to the complexity and diversity of mimicry computing scenarios.
Smart Images

Figure CN120256093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption. Background Art
[0002] In practical applications, different GPUs may come from different manufacturers or have different hardware architectures, computing capabilities, and memory sizes. Heterogeneous GPU resource pooling refers to concentrating multiple different types, brands, and specifications of GPU resources in a pool for unified management and scheduling to achieve resource sharing and dynamic allocation. Through resource pooling, the resource utilization rate of computing tasks can be effectively improved, resource scheduling can be optimized, and computing performance can be enhanced. It has important application values especially in the fields of cloud computing, big data analysis, deep learning, etc.;
[0003] Due to the different performances, power consumptions, and load conditions of different GPUs, heterogeneous resource pooling requires intelligent scheduling algorithms to dynamically allocate tasks to appropriate GPUs to maximize computing efficiency; in traditional computing environments, in order to manage resources in the system, the concept of a resource pool is introduced, organizing computing resources together and abstracting them into a virtual resource pool for easy unified management and dynamic allocation, which not only improves resource utilization but also enhances the flexibility and scalability of the system. However, traditional resource pool technologies are limited by their assumptions about resource types and management strategies and are difficult to effectively handle the complexity and diversity of resources in the mimic computing scenario. Therefore, how to achieve unified management, scheduling, and configuration of heterogeneous resources has become an urgent problem to be solved in the field of mimic computing;
[0004] In view of the above technical deficiencies, a solution is now proposed. Summary of the Invention
[0005] The purpose of the present invention is to: starting from the root resource pool of the tree-shaped resource pool according to the computing power resource value and computing priority required by the computing task, hierarchically match the corresponding resource types, and at the same time calculate the pre-estimated computing power resource value required in the next task requirement through the computing task expansion amount, evaluate the resource expansion trend in combination with the computing power resource value required by the computing task, and re-partition the tree-shaped resource pool according to the resource expansion trend.
[0006] To achieve the above purpose, the present invention adopts the following technical solution: a scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption, including the following steps:
[0007] Step 1: Obtain and store the task requirements submitted by different users, count the computing tasks according to the task requirements, evaluate the priority coefficient of the computing tasks, and obtain the computing power resource value and computing priority required by the computing tasks;
[0008] Step 2: Analyze task-related information according to the computing task, and then evaluate and predict the computing task trend based on the source of the computing task to obtain the estimated computing task expansion volume;
[0009] Step 3: According to the computing power resource value required by the computing task, configure the driver programs of different heterogeneous resources, obtain the configured driver programs, perform corresponding virtualization, generate virtual heterogeneous resources, abstract the virtual heterogeneous resources and then perform hierarchical management to generate a tree-shaped resource pool, where the tree-shaped resource pool includes a root resource pool, a priority resource pool and an extended resource pool;
[0010] Step 4: Based on the tree-shaped resource pool, construct a unified heterogeneous resource management interface, and at the same time, monitor the load status and performance status of each GPU device in real time, and obtain the remaining computing power resource value of each GPU device;
[0011] Step 5: Starting from the root resource pool of the tree-shaped resource pool according to the computing power resource value required by the computing task and the computing priority, hierarchically match the corresponding resource types until the remaining computing power resource value of the GPU device in the resource pool meets the computing power resource value required by the computing task and the computing priority, then allocate the corresponding quantity and specifications of resources to the user, wait for the user to complete the computing task, end the occupation of resources, and release the corresponding resources in a timely manner;
[0012] Step 6: Calculate the pre-estimated computing power resource value required in the next task requirement through the computing task expansion volume, evaluate the resource expansion trend in combination with the computing power resource value required by the computing task, and re-divide the priority resource pool and the extended resource pool according to the resource expansion trend.
[0013] Further, the specific process of obtaining the computing power resource value required by the computing task and the computing priority is as follows:
[0014] S101: Obtain and store the task requirements submitted by different users. The task requirements include task type, number of task items and task completion time limit, and integrate the number of task items and task type into a computing task;
[0015] S102: Input the computing task into a pre-trained neural network model to obtain the computing power resource value required by the computing task;
[0016] S103: Calculate the priority coefficient Yi of the computing task according to the following formula: where e and g are preset proportionality coefficients, α is a GPU demand coefficient preset according to the task type, d is the computing power resource value required by the computing task, and T is the task completion time limit;
[0017] S104. Obtain the preset calculation priority. The calculation priority includes a first-level calculation level, a second-level calculation level, and a third-level calculation level. Each calculation priority has a corresponding priority range preset. According to the priority coefficient of the calculation task, find the corresponding priority range to obtain the calculation priority.
[0018] Further, the specific process of obtaining the estimated expansion amount of the calculation task is as follows:
[0019] S201. Obtain multiple groups of task-related information. The task-related information includes the address of the target virtual machine, the identifier of the target GPU device, the task resource requirements, and the task source.
[0020] S202. Obtain multiple image samples according to the task source. The image samples include at least one labeled result of the task source. Merge the image samples with the same task source label result to obtain a training set for each task source. The name of each training set is the same as the name of the corresponding task source.
[0021] S203. Input each training set into the shared network to obtain the feature map of the image sample; obtain the feature map of the target to be detected according to the feature map of the image sample.
[0022] S204. For each training set, input the feature map of the target to be detected into the task prediction model with the same name as the training set, and iteratively train the task prediction model.
[0023] S205. Obtain the target image according to the task source, input the target image into the multi-task prediction model to obtain the target detection result and the calculation task trend corresponding to the target detection result, and input the calculation task trend into the pre-trained neural network model to obtain the estimated expansion amount of the calculation task.
[0024] Further, all heterogeneous GPU resources in the virtual heterogeneous resources form a tree structure, and the root node and several child nodes are correspondingly marked in the tree structure. The child nodes include priority nodes and standby nodes. The root resource pool is the resource pool corresponding to the root node, the priority resource pool is the heterogeneous resource pool corresponding to the priority node, and the extended resource pool is the heterogeneous resource pool corresponding to the standby node. The heterogeneous GPU resources in the extended resource pool and the priority resource pool can be scheduled with each other.
[0025] Further, the specific process of hierarchically matching the corresponding resource types is as follows:
[0026] S301. Obtain the computing power resource value and computing priority required for the computing task, and obtain the remaining computing power resource value and performance coefficient of each GPU device, and calculate the priority matching coefficient of each GPU device as follows:
[0027] The performance coefficient of the GPU device includes the real-time operation score C of the GPU device and the average operation score of the GPU device obtained based on historical data and the average computing power resource value of the GPU device for completing historical computing tasks d is the computing power resource value required for the computing task;
[0028] Calculate the priority matching coefficient Xi of the GPU device according to the following formula: where β and γ are preset proportionality coefficients;
[0029] S302. Match the priority matching coefficient of each GPU device with the preset priority matching threshold one by one, and mark the GPU devices that meet the matching standard as the GPU devices to which the computing task is to be allocated;
[0030] S303. Obtain the resource pool type where the GPU device to which the computing task is to be allocated is located, and then obtain the current state of the resource pool, and calculate the weight coefficient of the GPU device to which the computing task is to be allocated occupying the resource pool, sort the allocated resource pools according to the weight coefficient, and further select the priority resource pool according to the sorting result and the resource allocation situation.
[0031] Furthermore, the specific process of evaluating the resource expansion trend in combination with the computing power resource value required for the computing task is as follows:
[0032] S401. Obtain the estimated computing task expansion amount. If the task expansion amount shows a downward trend compared to the original computing task amount, maintain the original tree-shaped resource pool structure;
[0033] S402. If the task expansion amount shows an upward trend compared to the original computing task amount, calculate the trend coefficient according to the task expansion amount and the original computing task amount, and calculate the estimated computing power resource value according to the original computing task amount and the trend coefficient;
[0034] S403. Match the appropriate GPU devices from the self-expanding resource pool according to the estimated computing power resource value and mark them as schedulable GPU devices, and divide the schedulable GPU devices into the priority resource pool.
[0035] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0036] The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption evaluates the priority coefficient of computing tasks to obtain the computing power resource value and computing priority required by the computing tasks. According to the computing power resource value required by the computing tasks, the driver programs of different heterogeneous resources are configured, and the virtual heterogeneous resources are abstracted to generate a tree-shaped resource pool. Starting from the root resource pool of the tree-shaped resource pool according to the computing power resource value and computing priority required by the computing tasks, the corresponding resource types are hierarchically matched. At the same time, the task-related information can be parsed according to the computing tasks and the computing task trend can be predicted to obtain the estimated computing task expansion amount. The pre-estimated computing power resource value required in the next task requirement is calculated through the computing task expansion amount, the resource expansion trend is evaluated in combination with the computing power resource value required by the computing tasks, and the tree-shaped resource pool is re-divided according to the resource expansion trend. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 FIG. shows the overall system structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] Embodiment:
[0040] As Figure 1 shown, the scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption includes the following steps:
[0041] Step 1: Obtain and store the task requirements submitted by different users, count the computing tasks according to the task requirements, evaluate the priority coefficient of the computing tasks, and obtain the computing power resource value and computing priority required by the computing tasks;
[0042] The specific process of obtaining the computing power resource value and computing priority required by the computing tasks is as follows:
[0043] S101: Obtain and store the task requirements submitted by different users. The task requirements include task type, number of task items, and task completion time limit, and integrate the number of task items and task type into computing tasks;
[0044] S102: Input the computing tasks into a pre-trained neural network model to obtain the computing power resource value required by the computing tasks;
[0045] S103: Calculate the priority coefficient Yi of the computing tasks according to the following formula: where e and g are preset proportionality coefficients, α is a GPU demand coefficient preset according to the task type, d is the computing power resource value required for the computing task, and T is the task completion time limit;
[0046] S104. Obtain a preset computing priority, where the computing priority includes a first-level computing level, a second-level computing level, and a third-level computing level. Each computing priority is preset with a corresponding priority range. Search for the corresponding priority range according to the priority coefficient of the computing task to obtain the computing priority.
[0047] Step 2. Analyze task-related information according to the computing task, and then evaluate and predict the computing task trend based on the source of the computing task to obtain the estimated computing task expansion amount;
[0048] The specific process of obtaining the estimated computing task expansion amount is as follows:
[0049] S201. Obtain multiple groups of task-related information, where the task-related information includes the address of the destination virtual machine, the identifier of the target GPU device, the task resource requirements, and the task source.
[0050] S202. Obtain multiple image samples according to the task source. Among them, the image samples include at least one labeled result of the task source. Merge the image samples with the same labeled result of the task source to obtain a training set for each task source. Among them, the name of each training set is the same as the name of the corresponding task source;
[0051] S203. Input each training set into the shared network to obtain the feature map of the image sample; obtain the feature map of the target to be detected according to the feature map of the image sample;
[0052] S204. For each training set, input the feature map of the target to be detected into a task prediction model with the same name as the training set, and iteratively train the task prediction model;
[0053] S205. Obtain a target image according to the task source, input the target image into a multi-task prediction model to obtain a target detection result and the computing task trend corresponding to the target detection result, and input the computing task trend into a pre-trained neural network model to obtain the estimated computing task expansion amount.
[0054] Step 3. Configure the driver programs of different heterogeneous resources according to the computing power resource value required for the computing task, obtain the configured driver programs, perform corresponding virtualization to generate virtual heterogeneous resources, hierarchically manage the virtual heterogeneous resources after resource abstraction to generate a tree-shaped resource pool, where the tree-shaped resource pool includes a root resource pool, a priority resource pool, and an extended resource pool;
[0055] All heterogeneous GPU resources in the virtual heterogeneous resources form a tree structure, and the root node and several child nodes are correspondingly marked in the tree structure. The child nodes include priority nodes and standby nodes. Among them, the root resource pool is the resource pool corresponding to the root node, the priority resource pool is the heterogeneous resource pool corresponding to the priority node, and the extended resource pool is the heterogeneous resource pool corresponding to the standby node. The heterogeneous GPU resources in the extended resource pool and the priority resource pool can be scheduled with each other.
[0056] Step 4: According to the tree-shaped resource pool, construct a unified heterogeneous resource management interface, and at the same time, monitor the load status and performance status of each GPU device in real time, and obtain the remaining computing power resource value of each GPU device;
[0057] Step 5: Starting from the root resource pool of the tree-shaped resource pool, hierarchically match the corresponding resource types according to the computing power resource value required by the computing task and the computing priority until the remaining computing power resource value of the GPU device in the resource pool meets the computing power resource value required by the computing task and the computing priority. Then, allocate the corresponding quantity and specifications of resources to the user, wait for the user to complete the computing task, end the occupation of the resources, and release the corresponding resources in time;
[0058] The specific process of hierarchically matching the corresponding resource types is as follows:
[0059] S301: Obtain the computing power resource value and computing priority required by the computing task, and obtain the remaining computing power resource value and performance coefficient of each GPU device, and calculate the priority matching coefficient of each GPU device as follows:
[0060] The performance coefficient of the GPU device includes the real-time operation score C of the GPU device and the average operation score of the GPU device obtained based on historical data and the average computing power resource value of the GPU device to complete the historical computing task d is the computing power resource value required by the computing task;
[0061] Calculate the priority matching coefficient Xi of the GPU device according to the following formula: where β and γ are preset proportionality coefficients;
[0062] S302: Match the priority matching coefficient of each GPU device with the preset priority matching threshold one by one, and mark the GPU device that reaches the matching standard as the GPU device to which the computing task is to be allocated;
[0063] S303. Obtain the resource pool type where the GPU device for the to-be-allocated computing task is located, then obtain the current state of the resource pool, calculate the weight coefficient of the GPU device of the to-be-allocated computing task occupying the resource pool, sort the allocated resource pools according to the weight coefficient, and further select the preferred resource pool according to the sorting result and the resource allocation situation.
[0064] Step 6. Calculate the pre-estimated computing power resource value required in the next task demand through the computing task expansion volume, evaluate the resource expansion trend in combination with the computing power resource value required by the computing task, and re-partition the priority resource pool and the extended resource pool according to the resource expansion trend.
[0065] The specific process of evaluating the resource expansion trend in combination with the computing power resource value required by the computing task is as follows:
[0066] S401. Obtain the estimated computing task expansion volume. If the task expansion volume shows a downward trend compared with the original computing task volume, maintain the original tree-shaped resource pool structure;
[0067] S402. If the task expansion volume shows an upward trend compared with the original computing task volume, calculate the trend coefficient according to the task expansion volume and the original computing task volume, and calculate the pre-estimated computing power resource value according to the original computing task volume and the trend coefficient;
[0068] S403. Match the applicable GPU devices from the extended resource pool according to the pre-estimated computing power resource value, mark them as schedulable GPU devices, and divide the schedulable GPU devices into the priority resource pool.
[0069] The present invention evaluates the priority coefficient of the computing task to obtain the computing power resource value and the computing priority required by the computing task. According to the computing power resource value required by the computing task, configure the driver programs of different heterogeneous resources, and generate a tree-shaped resource pool after abstracting the virtual heterogeneous resources. Starting from the root resource pool of the tree-shaped resource pool according to the computing power resource value and the computing priority required by the computing task, hierarchically match the corresponding resource types. At the same time, the task-related information can be parsed according to the computing task and the computing task trend can be predicted to obtain the estimated computing task expansion volume. Calculate the pre-estimated computing power resource value required in the next task demand through the computing task expansion volume, evaluate the resource expansion trend in combination with the computing power resource value required by the computing task, and re-partition the tree-shaped resource pool according to the resource expansion trend.
[0070] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation;
[0071] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention should cover within the protection scope of the present invention any equivalent substitution or change made according to the technical solution and inventive concept of the present invention.
Claims
1. Scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption, characterized in that It includes the following steps: Step 1: Obtain and store the task requirements submitted by different users, calculate the computing tasks according to the task requirements, evaluate the priority coefficient of the computing tasks, and obtain the computing power resource value and computing priority required for the computing tasks; Step 2: Analyze the task-related information according to the computing tasks, and then evaluate and predict the computing task trend according to the source of the computing tasks to obtain the estimated computing task expansion volume; Step 3: According to the computing power resource value required for the computing tasks, configure the driver programs of different heterogeneous resources, obtain the configured driver programs, perform corresponding virtualization, generate virtual heterogeneous resources, abstract the virtual heterogeneous resources and perform hierarchical management to generate a tree-shaped resource pool. The tree-shaped resource pool includes a root resource pool, a priority resource pool and an expansion resource pool; Step 4: Build a unified heterogeneous resource management interface according to the tree-shaped resource pool, and at the same time monitor the load status and performance status of each GPU device in real time, and obtain the remaining computing power resource value of each GPU device; Step 5: According to the computing power resource value and computing priority required for the computing tasks, start from the root resource pool of the tree-shaped resource pool, hierarchically match the corresponding resource types until the remaining computing power resource value of the GPU device in the resource pool meets the computing power resource value and computing priority required for the computing tasks. Then allocate the corresponding quantity and specification of resources to the user, wait for the user to complete the computing task, end the occupation of the resources, and release the corresponding resources in time; Step 6: Calculate the pre-estimated computing power resource value required in the next task requirement through the computing task expansion volume, evaluate the resource expansion trend in combination with the computing power resource value required for the computing tasks, and re-divide the priority resource pool and the expansion resource pool according to the resource expansion trend.
2. The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption according to claim 1, characterized in that The specific process of obtaining the computing power resource value and computing priority required for the computing tasks is as follows: S101: Obtain and store the task requirements submitted by different users. The task requirements include task type, number of task items and task completion time limit. Integrate the number of task items and task type into a computing task; S102: Input the computing task into a pre-trained neural network model to obtain the computing power resource value required for the computing task; S103. Calculate the priority coefficient Yi of the computing task according to the following formula: where e and g are preset proportionality coefficients, α is the GPU requirement coefficient preset according to the task type, d is the computing power resource value required for the computing task, and T is the task completion time limit; S104: Obtain the preset computing priority. The computing priority includes a first-level computing level, a second-level computing level and a third-level computing level. Each computing priority is preset with a corresponding priority range. Find the corresponding priority range according to the priority coefficient of the computing task to obtain the computing priority.
3. The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption according to claim 1, wherein The specific process of obtaining the estimated computing task expansion volume is as follows: S201: Obtain multiple groups of task-related information. The task-related information includes the address of the target virtual machine, the identifier of the target GPU device, task resource requirements and task source; S202: Obtain multiple image samples according to the task source. Among them, the image samples include at least one label result of the task source. Merge the image samples with the same task source label result to obtain a training set for each task source. Among them, the name of each training set is the same as the name of the corresponding task source; S203. Input each of the training sets into the shared network to obtain the feature map of the image sample; obtain the feature map of the target to be detected according to the feature map of the image sample; S204. For each of the training sets, input the feature map of the target to be detected into the task prediction model with the same name as the training set, and iteratively train the task prediction model; S205. Obtain the target image according to the task source, input the target image into the multi-task prediction model to obtain the target detection result and the corresponding calculation task trend of the target detection result, and input the calculation task trend into the pre-trained neural network model to obtain the estimated calculation task expansion amount.
4. The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption according to claim 1, characterized in that All heterogeneous GPU resources in the virtual heterogeneous resources form a tree structure, and a root node and several child nodes are correspondingly marked in the tree structure. The child nodes include priority nodes and standby nodes. Among them, the root resource pool is the resource pool corresponding to the root node, the priority resource pool is the heterogeneous resource pool corresponding to the priority node, and the extended resource pool is the heterogeneous resource pool corresponding to the standby node. Among them, the heterogeneous GPU resources in the extended resource pool and the priority resource pool can be scheduled with each other.
5. The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption according to claim 1, wherein The specific process of hierarchically matching the corresponding resource types is as follows: S301. Obtain the computing power resource value and computing priority required for the computing task, and obtain the remaining computing power resource value and performance coefficient of each GPU device, and calculate the priority matching coefficient of each GPU device as follows: The performance coefficient of the GPU device includes the real-time operation score C of the GPU device and the average operation score of the GPU device obtained based on historical data as well as the average computing power resource value for the GPU device to complete historical computing tasks d is the computing power resource value required for the computing task; Calculate the priority matching coefficient Xi of the GPU device according to the following formula: where β and γ are preset proportionality coefficients; S302. Match the priority matching coefficient of each GPU device with the preset priority matching threshold one by one, and mark the GPU device that meets the matching standard as the GPU device to which the computing task is to be allocated; S303. Obtain the type of the resource pool where the GPU device to which the computing task is to be allocated is located, and then obtain the current state of the resource pool, and calculate the weight coefficient of the GPU device to which the computing task is to be allocated occupying the resource pool. Sort the allocated resource pools according to the weight coefficient, and further select the priority resource pool according to the sorting result and the resource allocation situation.
6. The scheduling method for heterogeneous GPU resource pooling and dynamic computing power preemption according to claim 1, wherein The specific process of evaluating the resource expansion trend in combination with the computing power resource value required for the computing task is as follows: S401. Obtain the estimated computing task expansion amount. If the task expansion amount shows a downward trend compared with the original computing task amount, keep the original tree-shaped resource pool structure; S402. If the task expansion amount shows an upward trend compared with the original computing task amount, calculate the trend coefficient according to the task expansion amount and the original computing task amount, and calculate the pre-estimated computing power resource value according to the original computing task amount and the trend coefficient; S403. Match the suitable GPU devices from the extended resource pool according to the estimated computing power resource value and mark them as schedulable GPU devices, and divide the schedulable GPU devices into the priority resource pool.
Citation Information
Cited By
Heterogeneous multitask computing power dynamic scheduling method and system
CN120469784A
Computing resource allocation method and system for heterogeneous GPU (Graphics Processing Unit) hybrid scheduling
CN120723468A
A method and system for allocating computing resources through heterogeneous GPU hybrid scheduling
CN120723468B