Dynamic graphics processor sharing method and device and storage medium

Through dynamic priority evaluation and resource adjustment mechanisms, the problem that the existing GPU scheduling mechanism cannot effectively utilize resources and ensure that tasks are completed on time is solved, and more efficient resource utilization and task completion rate are achieved.

CN120216178APending Publication Date: 2025-06-27HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510280819.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing GPU scheduling mechanism cannot make full use of GPU resources and cannot guarantee that tasks will be completed within the specified time, resulting in waste of resources and delayed tasks.

Method used

By receiving submissions of multiple deep learning tasks, a dynamic priority evaluation model is established, and priority scores are calculated based on the task's historical priority, completion rate and remaining execution time, and a dynamic priority queue is generated. Based on the real-time thread block usage of high-priority tasks, the resource state is passed to the low-priority tasks through the broadcast mechanism, and resource adjustment is triggered according to the preset threshold, and the number of executable thread blocks of the low-priority tasks is dynamically adjusted.

Benefits of technology

It improves the overall efficiency of the system, reduces resource waste, ensures that high-priority tasks are completed on time, and maximizes the use of idle resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216178A_ABST
    Figure CN120216178A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graphics processors, and provides a dynamic graphics processor sharing method, which comprises the following steps of: receiving submission of a plurality of deep learning tasks; generating a dynamic priority queue containing high-priority tasks and low-priority tasks; transmitting a resource state to the low-priority task through a broadcast mechanism, and triggering resource adjustment according to a comparison result of the real-time thread block usage amount of the high-priority task and a preset threshold value; according to the difference value between the real-time thread block usage amount of the high-priority task and a preset threshold value, the number of executable thread blocks of the low-priority task is dynamically adjusted; checking whether the number of thread blocks required by the kernel is smaller than or equal to the number of executable thread blocks of the current low-priority task, if yes, starting execution, and if not, reserving the kernel in a kernel request queue to wait for next cycle scheduling; and periodically updating the dynamic priority queue until all deep learning tasks are completed. According to the technical scheme, the overall efficiency of the system can be improved, and resource waste is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graphics processing units, and particularly relates to a method, device, and storage medium for sharing a dynamic graphics processing unit. Background Art

[0002] The sharing of a Graphics Processing Unit (GPU) plays a key role in computing resource management. By allowing multiple tasks or users to share the same GPU, resource utilization is increased and costs are reduced. This flexibility and dynamic allocation nature enable the system to more effectively meet the needs of different workloads. At the same time, by effectively scheduling and managing the shared GPU resources, system efficiency is improved, resource waste is reduced, and a higher overall performance level is achieved.

[0003] With the wide application of deep learning technology, training neural network models has become increasingly dependent on efficient GPU resources. Deep learning tasks usually require huge computing power. Especially when faced with large-scale datasets and complex models, the computing and memory resources of the GPU become crucial. However, due to the extremely large computing requirements during the training process, GPU resources often face competition when multiple concurrent tasks are executed. Especially in a cloud computing environment or a data center, when multiple tasks are trained simultaneously, how to efficiently allocate and schedule GPU resources has become a key problem to be solved urgently.

[0004] Existing GPU scheduling mechanisms often cannot fully utilize GPU resources and cannot ensure that tasks are completed within the specified time, resulting in resource waste and task delays. Summary of the Invention

[0005] This application provides a method, device, and storage medium for sharing a dynamic graphics processing unit, which can improve the overall efficiency of the system and reduce resource waste.

[0006] On the one hand, this application provides a method for sharing a dynamic graphics processing unit, and the method includes:

[0007] Step S101: Receive the submission of multiple deep learning tasks, and each task of the multiple deep learning tasks is associated with a deadline parameter;

[0008] Step S102: Establish a dynamic priority evaluation model, calculate a priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks;

[0009] Step S103: Based on the real-time thread block usage of high-priority tasks, transmit the resource status to low-priority tasks through the broadcast mechanism, and trigger resource adjustment according to the comparison result between the real-time thread block usage of the high-priority tasks and a preset threshold;

[0010] Step S104: Dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of the high-priority tasks and the preset threshold;

[0011] Step S105: Maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle of scheduling;

[0012] Step S106: Periodically recalculate the priority scores of all deep learning tasks, update the dynamic priority queue, and repeat Steps S103 to S105 based on the updated dynamic priority queue until all deep learning tasks are completed.

[0013] On the other hand, the present application provides a dynamic graphics processor sharing device, and the device includes:

[0014] A receiving module, configured to receive the submission of multiple deep learning tasks, and each of the multiple deep learning tasks is associated with a deadline parameter;

[0015] A generating module, configured to establish a dynamic priority evaluation model, calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks;

[0016] A triggering module, configured to transmit the resource status to low-priority tasks through the broadcast mechanism based on the real-time thread block usage of high-priority tasks, and trigger resource adjustment according to the comparison result between the real-time thread block usage of the high-priority tasks and a preset threshold;

[0017] An adjusting module, configured to dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of the high-priority tasks and the preset threshold;

[0018] An execution module, configured to maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle of scheduling;

[0019] An update module, configured to periodically recalculate the priority scores of all deep learning tasks, update the dynamic priority queue, and repeatedly execute the trigger module, the adjustment module, and the execution module based on the updated dynamic priority queue until all deep learning tasks are completed.

[0020] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the dynamic graphics processor sharing method as described above are implemented.

[0021] In a fourth aspect, the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the technical solution of the dynamic graphics processor sharing method as described above are implemented.

[0022] As can be seen from the technical solutions provided in the present application above, on the one hand, the priority score is calculated based on the historical priority, completion rate, and remaining execution time of each task, and a dynamic priority queue including high-priority tasks and low-priority tasks is generated. That is, the importance of tasks is distinguished through the dynamic priority queue, avoiding long-term exclusive use of resources by low-value tasks. On the other hand, the resource status is transmitted to low-priority tasks through a broadcast mechanism, and resource adjustment is triggered according to the comparison result between the real-time thread block usage of high-priority tasks and a preset threshold. According to the difference between the real-time thread block usage of high-priority tasks and the preset threshold, the number of executable thread blocks of low-priority tasks is dynamically adjusted. While ensuring that critical tasks (such as those approaching the deadline) obtain resources first, low-priority tasks are allowed to gradually resume to B ideal to maximize the utilization of idle resources. In the third aspect, when there is a kernel to be executed in the kernel request queue, it is checked whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, execution is started; otherwise, the kernel is retained in the kernel request queue to wait for the next cycle of scheduling. That is, through kernel queue rate limiting, it is prevented that low-priority tasks over-occupy resources and ensure the continuity of high-priority task execution. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1It is a flowchart of the dynamic graphics processor sharing method provided by an embodiment of the present application;

[0025] Figure 2 It is a structural schematic diagram of the dynamic graphics processor sharing device provided by an embodiment of the present application;

[0026] Figure 3 It is a DynGPU sharing architecture diagram for implementing the dynamic graphics processor sharing method or device provided by an embodiment of the present application;

[0027] Figure 4 It is a comparison schematic diagram of the throughput when the DynGPU sharing architecture provided by an embodiment of the present application executes tasks such as ResNet50, MobileNetV2, ResNetl52, and AlexNet and when executing the above tasks exclusively;

[0028] Figure 5 It is a comparison schematic diagram of the throughput when the DynGPU sharing architecture provided by an embodiment of the present application executes tasks such as ResNet50, MobileNetV2, ResNet152, and AlexNet and when executing the above tasks under the Non architecture;

[0029] Figure 6 It is a comparison schematic diagram of the throughput when the DynGPU sharing architecture provided by an embodiment of the present application executes tasks such as ResNet50, MobileNetV2, ResNet152, and AlexNet and when executing the above tasks under the MPS architecture;

[0030] Figure 7 It is a comparison schematic diagram of the throughput when the DynGPU sharing architecture provided by an embodiment of the present application executes tasks such as ResNet50, MobileNetV2, ResNet152, and AlexNet and when executing the above tasks under the TGS architecture;

[0031] Figure 8 It is a comparison schematic diagram of the completion time of task execution under the DynGPU sharing architecture provided by an embodiment of the present application and the completion time of executing the above tasks under the MPS and TGS architectures;

[0032] Figure 9 It is a schematic diagram of the GPU utilization rate of the DynGPU sharing architecture provided by an embodiment of the present application;

[0033] Figure 10 It is a schematic diagram showing that the DynGPU sharing architecture provided by an embodiment of the present application can ensure the timely completion of high-priority tasks compared with the Non, MPS, and TGS architectures;

[0034] Figure 11 It is a structural schematic of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0036] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.

[0037] In this specification, for ease of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0038] The Graphics Processing Unit (GPU) sharing plays a key role in computing resource management. By allowing multiple tasks or users to share the same GPU, resource utilization is improved and costs are reduced. This flexibility and dynamic allocation nature enable the system to more effectively meet the requirements of different workloads. At the same time, by effectively scheduling and managing the shared GPU resources, system efficiency is improved, resource waste is reduced, and a higher level of overall performance is achieved.

[0039] With the widespread application of deep learning technology, training neural network models has become increasingly dependent on efficient GPU resources. Deep learning tasks usually require huge computing power, especially when facing large-scale data sets and complex models, the computing and memory resources of GPUs become crucial. However, due to the extremely large computing requirements of the training process, GPU resources often face competition when multiple concurrent tasks are executed. Especially in cloud computing environments or data centers, when multiple tasks are trained simultaneously, how to efficiently allocate and schedule GPU resources has become a key issue that needs to be solved urgently. The current GPU sharing mechanism usually adopts a spatiotemporal allocation method to improve resource utilization by binding multiple training tasks to the same GPU. This method divides GPU resources into multiple spatiotemporal segments, allowing multiple tasks to share the computing power and memory resources of the GPU in different time periods, thereby improving the task throughput of the cluster. In some cloud computing environments, this sharing method can significantly improve resource utilization and avoid the low training efficiency of a single task due to resource limitations. However, this solution also has some problems that cannot be ignored. Some sharing methods require modifying the deep learning framework to support sharing GPU resources, which requires users to switch to a specific version of the framework, which may limit the user's choice. In addition, in a multi-version framework environment, users are also required to pay special attention to the compatibility between frameworks. In addition, for the sharing solutions provided by GPU manufacturers, MPS (multi-process service) and MIG (multi-instance GPU) are two common hardware-supported sharing mechanisms. MPS allows multiple processes to execute simultaneously through hardware support, but it requires applications to explicitly set resource limits and does not provide fault isolation, which means that if a process fails, it may affect the stability of other processes. On the other hand, MIG provides certain performance isolation by statically dividing GPU hardware resources to form multiple independent GPU instances to meet the needs of different tasks. However, MIG's resource division is fixed and does not support dynamic adjustment, which limits its flexibility, especially when the task load and demand change greatly, which may lead to uneven resource utilization or waste. Existing GPU scheduling mechanisms often cannot fully utilize GPU resources, nor can they ensure that tasks are completed within the specified time, resulting in resource waste and task delays. This shows that the current GPU resource scheduling scheme needs to be improved urgently, especially in dealing with changes in task priorities and task deadlines, dynamic scheduling is particularly important.

[0040] In summary, implementing GPU sharing at the framework level requires significant modifications to the existing framework, requiring users to switch to the modified version, which may cause considerable interference to the existing system. For users who are not familiar with the internal mechanisms of the deep learning framework, modifying the framework may be complicated. There is a high demand for customization of the framework, and the implementation of different frameworks may be different, making it difficult to provide unified support across frameworks. Using API technology to implement GPU sharing or virtualization through API interception and resource management may bring certain performance overheads, especially when task requests are very frequent or resources are tight, and management overhead may affect overall performance. GPU manufacturers continue to launch advanced software and hardware solutions that rely on specific hardware architectures. For example, a higher version of a GPU from a well-known chip company may not be suitable for older devices or GPUs from different manufacturers. Resource partitioning by technologies such as MIG is usually static and may not adapt well to dynamic task requirements. For example, when the load of a task changes, static resource partitioning may lead to waste or insufficient resources.

[0041] In view of the above problems in the prior art, this application proposes a dynamic graphics processor sharing method, the flow chart of which is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S106, which are described in detail as follows:

[0042] Step S101: receiving submissions of multiple deep learning tasks, wherein each of the multiple deep learning tasks is associated with a deadline parameter.

[0043] In the embodiment of the present application, each deep learning task will have a deadline (DDL) parameter when it is submitted. The deadline parameter is the core basis for dynamic priority evaluation. Without this parameter, the urgency of the task cannot be judged. The system tracks the progress and remaining time of each deep learning task in real time and dynamically adjusts the priority of each deep learning task based on its own deadline parameter to ensure that high-priority deep learning tasks can be completed on time.

[0044] In addition to the associated deadline parameter, each task can also contain a minimum number of thread blocks B min , ideal number of thread blocks B ideal and the maximum number of thread blocks B max , etc., where the minimum number of thread blocks B min represents the minimum guaranteed resource, that is, the minimum number of thread blocks (or video memory / computing units) required for task execution to ensure that the task is not completely interrupted or deadlocked. The ideal number of thread blocks B ideal represents the ideal amount of resources, that is, the number of thread blocks expected to be allocated to the task when there is no resource competition and the system load is sufficient to achieve the optimal training speed or efficiency. The maximum number of thread blocks B maxRepresents the maximum tolerable resources, that is, the upper limit of the number of thread blocks that a task is allowed to occupy. Exceeding this value may lead to resource waste or have a negative impact on other high-priority tasks. The deadline parameter is the core basis for dynamic priority evaluation. Without this parameter, it is impossible to judge the urgency of a task, while the minimum number of thread blocks B min , the ideal number of thread blocks B ideal and the maximum number of thread blocks B max and other parameters define the upper and lower limits of the task's resource requirements, preventing excessive occupation or waste during resource allocation.

[0045] Step S102: Establish a dynamic priority evaluation model, calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue containing high-priority tasks and low-priority tasks.

[0046] In the system of the embodiment of the present application, a user can submit any number of tasks. One task is designated as a high-priority task, while other tasks are regarded as low-priority tasks and can utilize any idle GPU resources. This distinction of task importance ensures that high-priority tasks can access GPU resources first to meet their urgent computing needs, while maximizing the overall utilization rate of system resources. Traditional scheduling methods, such as Round-Robin and Fair Scheduling, although simple to implement, lack the awareness of task urgency and priority and cannot provide sufficient performance guarantees for high-priority tasks. In addition, in an actual production environment, users usually specify a deadline for tasks and expect the tasks to be completed within the specified time. To avoid performance bottlenecks caused by uneven resource allocation or delays in critical tasks, in the embodiment of the present application, while dynamically adjusting the current high-priority tasks, performance guarantees can be provided for high-priority tasks. Therefore, a dynamic priority calculation method can be designed to determine in real time which task should be selected as the high-priority task to ensure its completion before the deadline. This priority calculation mechanism is the core part of the system, helping to dynamically adjust the priorities of tasks and achieve an organic balance between performance and timeliness in a multi-task concurrent scenario.

[0047] As an embodiment of the present application, calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task. Specifically, the priority score Score of each task can be calculated according to the following calculation formula:

[0048] Score = w1 * (1 - x1) + w2 * (1 - x2) + w3 * x3

[0049] Among them, x1 is the normalized value of the historical priority of the task, x2 is the normalized value of the completion rate of the task, x3 is the normalized value of the remaining time for task execution, and w1, w2, and w3 are the weights corresponding to (1 - x1), (1 - x2), and x3 respectively, and w1 + w2 + w3 = 1. In the above embodiments, the reason for normalizing these variables is to ensure the comparability and scale invariance of task characteristics. For each variable such as historical priority, completion rate, and remaining execution time, calculating its normalized value can be to subtract the mean of each variable from it and then divide by the standard deviation. Calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task. The goal is to prioritize tasks with a lower completion rate and insufficient remaining time to complete on time, making them high-priority tasks. At the same time, the historical priority of the task will also be an important reference. Based on the calculated priority score, the tasks can be sorted in descending order. Then, apply the greedy scheduling method to select the task with the highest priority score as the high-priority task. The scheduling mechanism runs in an iterative manner and dynamically updates the priority score of the task according to the real-time state of the task. In this way, it can be ensured that tasks approaching the deadline or lagging behind in progress obtain the necessary resources to complete the goal, while less urgent tasks utilize the remaining resources.

[0050] As can be seen from step S102 of the above embodiments, by distinguishing task priorities through multi-dimensional scoring (historical priority, completion rate, remaining execution time of each task, etc.), the resource supply of high-value tasks (such as tasks approaching the deadline) can be preferentially guaranteed, while allowing low-priority tasks to execute when resources are abundant, reducing the risk of task delay.

[0051] Step S103: Based on the real-time thread block usage of high-priority tasks, transmit the resource status to low-priority tasks through the broadcast mechanism, and trigger resource adjustment according to the comparison result between the real-time thread block usage of high-priority tasks and the preset threshold.

[0052] High-priority tasks play a crucial role in the scheduling process. To ensure their performance, embodiments of this application use a broadcast mechanism to provide real-time status updates of high-priority tasks to low-priority tasks. That is, the broadcast mechanism periodically calculates the number of GPU thread blocks used by high-priority tasks, denoted as BH, and broadcasts this information, i.e., BH, along with the current GPU resource usage status to all low-priority tasks. The purpose of the broadcast is to provide a global view of GPU resource utilization to low-priority tasks, allowing them to dynamically adjust their thread block execution limits to avoid interfering with the performance of high-priority tasks. Specifically, as an embodiment of this application, based on the real-time thread block usage of high-priority tasks, the resource status is transmitted to low-priority tasks through the broadcast mechanism, and resource adjustment is triggered according to the comparison result between the real-time thread block usage of high-priority tasks and a preset threshold, which can be achieved through step S1031 and step S1032, detailed as follows:

[0053] Step S1031: Calculate the thread block usage BH of the current high-priority task in real time, and transmit the value of BH to all low-priority tasks through the broadcast mechanism.

[0054] Initially, BH is used to store the number of thread blocks executed by the high-priority task within the current time interval. Thereafter, the system periodically calculates the number of thread blocks executed within a time interval to accurately track the thread block usage during GPU kernel execution. Once the calculation is completed, i.e., ensuring that the value accurately reflects the GPU resources consumed by the high-priority task, the system broadcasts the thread block usage BH of the current high-priority task to all low-priority tasks. This broadcast ensures that low-priority tasks can promptly learn about the GPU usage status of high-priority tasks, enabling them to adjust their behavior as needed. After the broadcast is completed, the high-priority task enters the sleep state and waits for the specified broadcast time interval TH. During this time interval, the original high-priority task continues to execute and counts the GPU kernel usage or the number of GPU thread blocks occupied by the high-priority task during this period for the next broadcast. This design ensures that the broadcast operation does not frequently interrupt the execution of high-priority tasks while still updating low-priority tasks in a timely manner. Through this broadcast mechanism, high-priority tasks can convey their execution status to all low-priority tasks with minimal scheduling overhead without directly interfering with the execution of other tasks.

[0055] Step S1032: Compare the thread block usage BH of the current high-priority task with the preset threshold MH. If |BH - MH| exceeds the preset difference threshold, trigger resource adjustment for low-priority tasks.

[0056] If the thread block usage BH of the current high-priority task is compared with the preset threshold MH, and the thread block usage BH of the high-priority task is less than the preset threshold MH, that is, the total number of thread blocks of the high-priority task within a period of time, and |BH - MH| exceeds the preset difference threshold, it indicates that other priority tasks may be competing for resources, thus limiting the throughput of the high-priority task. It is necessary to trigger resource adjustment for low-priority tasks.

[0057] As can be seen from the above embodiments, by calculating the thread block usage BH of high-priority tasks in real time, the accuracy of the resource occupancy status is ensured. The broadcast mechanism enables low-priority tasks to quickly respond to resource changes, while threshold comparison (instead of continuous adjustment) reduces scheduling overhead and balances real-time performance and system stability. In other words, the broadcast mechanism enables low-priority tasks to perceive the resource competition status in real time, and threshold triggering avoids frequent adjustments, reduces scheduling overhead, and improves response efficiency. For example, when high-priority tasks have sudden resource requirements, low-priority tasks can quickly contract resources to prevent system overload.

[0058] The above embodiments provide resource scheduling based on priority scores. However, this resource scheduling based on priority scores may lead to jitter phenomena, that is, high-priority tasks switch frequently, bringing additional overhead. To solve the above problems, one solution is to ensure that each high-priority task executes at least one complete time slice during scheduling, thus effectively reducing unnecessary task switching and improving system efficiency. Specifically, after each scheduling of a high-priority task, the scheduling thread is blocked for a specified period of time, such as 30 seconds, to ensure that the new deep learning task can execute for at least the specified time.

[0059] Step S104: Dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of high-priority tasks and the preset threshold.

[0060] After receiving the real-time thread block usage BH of high-priority tasks, low-priority tasks can dynamically adjust resource usage according to the requirements of the scheduling algorithm, so as to ensure that high-priority tasks can maintain their performance, that is, in order to avoid high-priority task blocking and gradually restore resources to prevent oscillation, gradually approach B ideal To maximize the utilization rate of idle GPU resources by low-priority tasks, in the embodiments of the present application, according to the difference between the real-time thread block usage BH of high-priority tasks and the preset threshold MH, the number of executable thread blocks B exce [i] of low-priority tasks is dynamically adjusted. Specifically, as an embodiment of the present application, according to the difference between the real-time thread block usage of high-priority tasks and the preset threshold, the number of executable thread blocks B exce[i]The dynamic adjustment can be achieved through step S1041 and step S1042, and the details are as follows:

[0061] Step S1041: When |BH - MH| exceeds the preset difference threshold, reduce the number of executable thread blocks B of the low-priority task exce [i]to 1 / 2 of the current value.

[0062] If the remaining execution time of the low-priority task with ID i is long and the resource requirement of the high-priority task is large, then the number of executable thread blocks B of the low-priority task exce [i], to avoid interfering with the high-priority task. Specifically, if the real-time thread block usage BH of the high-priority task is less than the preset threshold MH, and |BH - MH| exceeds the set threshold, it means that other priority tasks may be competing for resources, thus limiting the throughput of the high-priority task. Here, MH can be the throughput of the high-priority task when it exclusively occupies GPU resources, usually measured when the high-priority task is initialized. In this case, other priority tasks will reduce the number of executable thread blocks by B exce [i]←B exce [i] / 2, that is, by reducing the number of executable thread blocks B of the low-priority task exce [i]to 1 / 2 of the current value to reduce the number of executable thread blocks, so as to release GPU resources for the high-priority task.

[0063] Step S1042: When |BH - MH| does not exceed the preset difference threshold, gradually increase the number of executable thread blocks B of the low-priority task in an exponential increasing manner exce [i], until B exce [i]reaches the ideal number of thread blocks B of the low-priority task ideal .

[0064] If the resource requirement of the high-priority task is small or there are more idle GPU resources, the number of executable thread blocks B of the low-priority task with ID i can be increased exce [i], so that the low-priority task can make full use of the idle resources. Specifically, if |BH - MH| does not exceed the preset difference threshold, for example, when BH reaches or is close to MH, it means that the resource requirement of the high-priority task has been met. At this time, other priority tasks will increase the number of their executable thread blocks by calling relevant APIs to make the best use of the idle GPU resources as much as possible. This mechanism enables other low-priority tasks to dynamically adapt to the changes in the resource requirements of the high-priority task, while maximizing the utilization of idle GPU resources while ensuring the performance of the high-priority task.

[0065] As an embodiment of the present application, gradually increase the number of executable thread blocks B of the low-priority task in an exponential increasing manner exceThe specific implementation of [i] is as follows: Among them, B' exec [i] is the number of executable thread blocks after the adjustment of low-priority tasks, t is the number of cycles since the last resource adjustment, and after adjustment, it is necessary to satisfy B min ≤B' exec [i]≤B max When B' exec [i] reaches the ideal number of thread blocks B ideal of low-priority tasks, stop incrementing and lock it as the B ideal value. The adjustment of the number of cycles t can be specifically as follows: If resource adjustment has not been triggered for 3 consecutive cycles, reset t = 0, and the number of executable thread blocks B exce [i] of low-priority tasks starts incrementing again from its B min .

[0066] As can be seen from the above embodiments, by halving the resources when |BH - MH| exceeds the preset difference threshold, it is ensured that high-priority tasks can preempt successfully. And when |BH - MH| does not exceed the preset difference threshold, the resources are gradually restored by exponential increment to balance resource utilization and stability. For example, the resource recovery process avoids the oscillation cycle of "allocation - overrun - recovery".

[0067] Step S105: Maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle scheduling.

[0068] To avoid kernel startup conflicts caused by mutations in the number of executable thread blocks B exce [i] of low-priority tasks and ensure that resource allocation is strictly controlled to prevent over-occupation, in the embodiments of the present application, an independent kernel request queue can be maintained for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle scheduling. The above method finely adjusts the GPU kernel startup scheduling of each other-priority task according to the number of executable thread blocks B exce [i] of low-priority tasks. By monitoring the changes in the kernel request queue Q[i] and the changes in the number of executable thread blocks B exce [i], for the low-priority task with ID i, if the kernel request queue Q[i] is not empty (that is, there is a kernel request to be processed), the system enters an internal loop to process each kernel in the queue. If the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks B of the current low-priority taskexce [i], the kernel is considered ready to start, otherwise, it will remain in the queue and wait for the next time period to be scheduled. Through these mechanisms, it is ensured that high-priority tasks can always obtain the required performance guarantee and meet the requirements of critical tasks. At the same time, for other low-priority tasks, the startup rate of their GPU kernels can be dynamically adjusted according to the actual usage of system resources, providing flexible scalability, thereby not only avoiding the impact of resource competition on the performance of high-priority tasks, but also making full use of idle GPU resources, improving overall resource utilization and system efficiency.

[0069] Furthermore, the execution rules of the kernel request queue Q[i] of the low priority task also include:

[0070] If the number of thread blocks required by the kernel exceeds the number of executable thread blocks of the current low-priority task B exce [i], then reinsert the kernel into the tail of the kernel request queue Q[i] and record its waiting times; when the waiting times exceed the preset upper limit, force the kernel to start and allow the current low-priority task to execute thread blocks] B exce [i] Break through the maximum number of thread blocks B of the current low-priority task max In the above embodiment, the method for determining the preset upper limit is: And the upper limit value does not exceed 10 times, among which, T total Indicates the total number of low-priority tasks in the current system. When the kernel is forced to start, the number of executable thread blocks B of low-priority tasks is temporarily allowed exce [i] The breakthrough amplitude does not exceed the maximum number of thread blocks B of the current low-priority task max 30% of.

[0071] Step S106: Periodically recalculate the priority scores of all deep learning tasks and update the dynamic priority queue, and repeat steps S103 to S105 based on the updated dynamic priority queue until all deep learning tasks are completed.

[0072] In order to correct the deviation caused by task progress or environmental changes, avoid the rigidity of static strategies, and adapt to long-term dynamic loads, the priority scores of all deep learning tasks can be periodically recalculated and the dynamic priority queue can be updated, and steps S103 to S105 can be repeatedly executed based on the updated dynamic priority queue until all deep learning tasks are completed. Here, the trigger condition for updating the dynamic priority queue after periodically recalculating the priority scores of all deep learning tasks can be to trigger the update of the dynamic priority queue at a fixed time interval; or to immediately trigger the update of the dynamic priority queue when the remaining execution time of any task is lower than the preset alarm threshold.

[0073] From the above-described dynamic graphics processor sharing method of the example, on the one hand, a priority score is calculated based on the historical priority, completion rate, and remaining execution time of each task, and a dynamic priority queue including high-priority tasks and low-priority tasks is generated. That is, by Figure 1 distinguishing the importance of tasks through the dynamic priority queue, long-term exclusive occupation of resources by low-value tasks is avoided; on the other hand, the resource status is transmitted to low-priority tasks through the broadcast mechanism, and resource adjustment is triggered according to the comparison result between the real-time thread block usage of high-priority tasks and a preset threshold. According to the difference between the real-time thread block usage of high-priority tasks and the preset threshold, the number of executable thread blocks of low-priority tasks is dynamically adjusted. While ensuring that critical tasks (such as those approaching the deadline) obtain resources first, low-priority tasks are allowed to gradually resume to B

[0074] when resources permit, maximizing the utilization of idle resources; thirdly, when there is a kernel to be executed in the kernel request queue, it is checked whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, execution is started; otherwise, the kernel is retained in the kernel request queue to wait for the next cycle of scheduling. That is, through kernel queue rate limiting, overcrowding of resources by low-priority tasks is prevented, and the continuity of high-priority task execution is guaranteed. ideal

[0075] Please refer to the appendix Figure 2 which is a dynamic graphics processor sharing device provided by an embodiment of the present application. The device may include a receiving module 201, a generating module 202, a triggering module 203, an adjusting module 204, an executing module 205, and an updating module 206, which are described in detail as follows:

[0076] The receiving module 201 is configured to receive the submission of multiple deep learning tasks, where each task of the multiple deep learning tasks is associated with a deadline parameter;

[0077] The generating module 202 is configured to establish a dynamic priority evaluation model, calculate a priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks;

[0078] The triggering module 203 is configured to transmit the resource status to low-priority tasks through the broadcast mechanism based on the real-time thread block usage of high-priority tasks, and trigger resource adjustment according to the comparison result between the real-time thread block usage of the high-priority tasks and a preset threshold;

[0079] The adjusting module 204 is configured to dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of the high-priority tasks and the preset threshold;​

[0080] The execution module 205 is used to maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, it checks whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, it starts the execution; otherwise, it retains the kernel in the kernel request queue and waits for the next cycle scheduling.

[0081] The update module 206 is used to periodically recalculate the priority scores of all deep learning tasks, update the dynamic priority queue, and repeat the execution of the trigger module 203, the adjustment module 204, and the execution module 205 based on the updated dynamic priority queue until all deep learning tasks are completed.

[0082] From the above-attached Figure 2 It can be seen from the example dynamic graphics processor sharing device that, on the one hand, the priority score is calculated based on the historical priority, completion rate, and remaining execution time of each task, and a dynamic priority queue including high-priority tasks and low-priority tasks is generated. That is,

[0083] The dynamic priority queue distinguishes the importance of tasks and avoids low-value tasks monopolizing resources for a long time. On the other hand, the resource status is transmitted to low-priority tasks through the broadcast mechanism, and the resource adjustment is triggered according to the comparison result between the real-time thread block usage of high-priority tasks and the preset threshold. According to the difference between the real-time thread block usage of high-priority tasks and the preset threshold, the number of executable thread blocks of low-priority tasks is dynamically adjusted. While ensuring that critical tasks (such as those approaching the deadline) obtain resources first, it allows low-priority tasks to gradually recover to B ideal when the resource allows, maximizing the utilization of idle resources. Thirdly, when there is a kernel to be executed in the kernel request queue, it checks whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, it starts the execution; otherwise, it retains the kernel in the kernel request queue and waits for the next cycle scheduling, that is, through the kernel queue rate limit, preventing low-priority tasks from overcrowding resources and ensuring the continuity of high-priority task execution.

[0084] The above Figure 2 The example dynamic graphics processor (DynGPU) sharing device can be based on Figure 3 The example DynGPU sharing architecture is designed specifically for dynamic, priority-aware GPU resource scheduling in deep learning tasks. Figure 3 The example DynGPU sharing architecture is divided into three layers, namely, the application layer, the management system layer (OS layer), and the hardware layer. Each layer plays a crucial role in efficient resource allocation and task priority management. In Figure 3In the DynGPU shared system corresponding to the exemplary DynGPU shared architecture, users can submit deep learning tasks at any time, and each task will specify a corresponding deadline (DDL). The DynGPU shared system intercepts each GPU kernel call initiated by a deep learning task and stores it in the corresponding task queue. Through the dynamic priority scheduling mechanism, the DynGPU shared system can identify high-priority tasks and provide performance guarantees for these deep learning tasks through the dynamic resource sharing module. On this basis, the DynGPU shared system will also allocate idle GPU resources to other low-priority tasks, thereby improving the overall resource utilization rate. It should be particularly noted that the DynGPU shared system is completely transparent to end users, and users do not need to modify any content in the API or the development process. The implementation described in this application is based on a single GPU, but in a multi-GPU scenario, users need to specify the GPU to be used at runtime.

[0085] Figure 3 The exemplary DynGPU shared architecture mainly includes two core modules: dynamic GPU resource sharing and dynamic priority scheduling. Among them, the core goal of the dynamic GPU resource sharing module is to flexibly allocate the computing and memory resources of the GPU according to the priority and resource requirements of deep learning tasks. By monitoring the resource requirements of each deep learning task in real time, the system can dynamically allocate idle GPU resources to other low-priority tasks while ensuring that high-priority tasks have sufficient resources. To achieve this, the dynamic GPU resource sharing module dynamically adjusts the resource allocation of tasks, avoids resource idleness, and improves the overall efficiency of the system as much as possible. The dynamic priority scheduling module is responsible for adjusting the execution order of deep learning tasks according to the urgency and remaining time of the tasks. Each deep learning task is submitted with a deadline (DDL). The DynGPU shared system dynamically adjusts the priority of each deep learning task by tracking the progress and remaining time of the task in real time, ensuring that high-priority tasks in the deep learning tasks can be completed on time. At the same time, low-priority tasks in the deep learning tasks can continue to execute using idle GPU resources when resources permit, thereby reducing the idle time of the DynGPU shared system and maximizing the utilization rate of resources.

[0086] Figure 3One of the greatest advantages of the exemplary DynGPU sharing architecture or DynGPU sharing system is its transparency to the end user. When using it, the user does not need to modify the existing deep learning framework, nor does the user need to change the API calls. The DynGPU sharing system automatically intercepts the calls to the GPU kernels and optimizes the resource allocation through the built-in scheduling and resource sharing mechanisms. This transparency not only reduces the development burden on the user, but also improves the usability of the DynGPU sharing system, enabling the user to still enjoy efficient GPU resource management even without in-depth understanding of the underlying resource scheduling mechanism. Although the implementation of this application is mainly based on a single GPU, DynGPU can also be applied to multi-GPU environments. In a multi-GPU scenario, the DynGPU sharing system requires the user to specify the GPUs to be used at runtime. Through the resource management and scheduling of different GPUs, the DynGPU sharing system can implement more complex resource allocation strategies, further improving the performance and resource utilization rate of the DynGPU sharing system. In the DynGPU sharing system, dynamic GPU resource sharing and dynamic priority scheduling complement each other. The dynamic GPU resource sharing module is mainly responsible for allocating GPU resources according to the resource requirements of deep learning tasks and the availability of idle resources, while the dynamic priority scheduling module adjusts the execution order and resource allocation of deep learning tasks reasonably according to the deadlines and priorities of deep learning tasks. The coordinated work of these two enables the DynGPU sharing system to execute low-priority tasks by making the best use of idle resources without affecting the performance of high-priority tasks, ensuring the overall efficient operation of the DynGPU sharing system. The design of the DynGPU sharing system realizes efficient resource utilization and task scheduling by handling GPU resource scheduling and priority management in a hierarchical manner and through the synergistic effect of the two core modules of dynamic GPU resource sharing and dynamic priority scheduling. Its transparency and flexibility enable the user to enjoy efficient GPU resource management without additional modification, and its ability to adapt to multi-GPU environments also provides strong support for larger-scale deep learning tasks.

[0087] To verify the dynamic graphics processor sharing method of the above embodiments, this application is implemented using approximately 2,600 lines of C and Python code Figure 3Example DynGPU sharing architecture or DynGPU sharing system. The core design concept of the DynGPU sharing system is to encapsulate each deep learning task in an independent Docker container, thereby achieving isolation between deep learning tasks and independent resource management. This design ensures that the execution environment of each deep learning task is completely independent of other deep learning tasks, effectively avoiding interference between deep learning tasks, and thus improving the reliability, scalability, and fault tolerance of deep learning tasks. In deep learning frameworks, GPU kernels are launched through libraries such as the CUDA runtime API, which provides efficient implementations for common deep neural network (DNN) operations. To precisely control the launch of GPU kernels, this application uses wrapper functions to intercept the kernel launch process by overriding relevant CUDA APIs. These wrapper functions are responsible for passing the necessary scheduling information to the corresponding task queues. The intercepted APIs include cuLaunch, cuLaunchKernel, cudaAlloc, and cudaFree. By this method, the GPU resources of each deep learning task can be precisely managed, and deep learning tasks can be effectively scheduled. The designed wrapper functions are sufficient to support all DNN workloads used in the evaluation. In the current implementation, the management of the dynamic GPU resource scheduling queue and resource reception is achieved through different threads in the same process. This design ensures memory sharing and fast communication between threads, thereby improving the overall performance and responsiveness of deep learning tasks. In the dynamic priority scheduling module, the DynGPU sharing system periodically collects the runtime data of each deep learning task and calculates a priority score based on this data. The calculated priority score is used to determine the ID of the currently highest-priority task, which is stored in a predefined shared memory area. Each deep learning task contains a daemon thread that periodically reads the shared memory to check whether its priority needs to be adjusted. For example, when the deep learning task with ID 2 needs to be promoted to a high-priority task, the daemon thread will trigger a change in the state of the task and re-initialize the scheduling thread associated with the task.

[0088] In the embodiments of this application, the dynamic graphics processor sharing method or Figure 3 Example DynGPU sharing architecture is evaluated by answering the following questions:

[0089] Q1: Can dynamic GPU resource sharing guarantee the performance of high-priority tasks, and what are the advantages of dynamic GPU resource sharing compared to exclusive GPUs?

[0090] Q2: What is the completion time of tasks after using dynamic GPU resource sharing?

[0091] Q3: While ensuring high-priority tasks, does dynamic GPU resource sharing have a high enough GPU utilization rate?

[0092] Q4: What are the benefits of dynamic priority scheduling?

[0093] Q5: How much overhead does the framework have?

[0094] 1) Experimental environment: The experimental hardware environment of the dynamic graphics processor sharing method or Figure 3 example's DynGPU sharing architecture is as follows: Intel(R) Xeon(R) Platinum 8358P CPU, NVIDIA RTX A6000 GPU, and 512GB system memory. The software environment includes Ubuntu 22.04.5 LTS, Python 3.11, CUDA 12.5, etc. When executing the workload, multiple deep learning frameworks were used, including PyTorch and TensorFlow. The workload mainly includes ResNet50, MobileNet, and ShuffleNet. These models are representative, widely used, and are standard benchmarks for evaluating deep learning systems. In the experiment, DynGPU was compared with MPS, Non (directly running multiple tasks in parallel), and TGS. The evaluation metrics include task completion time, task completion rate, and throughput (number of iterations per second). In all experiments, sufficient GPU memory was ensured to be available.

[0095] 2) Dynamic GPU resource sharing: DynGPU utilizes the dynamic GPU resource sharing mechanism to ensure the performance of high-priority tasks while maximizing the performance of other priority tasks. In this experiment, four tasks were executed: ResNet50(1), MobileNetV2(2), ResNet152(3), and AlexNet(4), with a batch size of 64 for each task and 10,000 iterations. Static priorities (numbers in parentheses) were assigned to these tasks, and the task with the smallest number was initialized as a high-priority task at the beginning. When a task is completed, the task with the smallest priority number is selected as the new high-priority task. At the beginning of the experiment, these four tasks were encapsulated in containers and executed simultaneously. The throughput was measured at different time points to answer Q1.

[0096] As Figure 4As shown, from time 0 to 1645, ResNet50 runs as a high-priority task with a throughput of 6.124. When it executes exclusively on the same GPU, the throughput increases to 6.6. Compared with exclusive execution, ResNet50 retains 92.8% of its throughput. Similarly, MobileNetV2 and ResNet152 retain 97.5% and 95.1% of their throughput respectively. Since AlexNet always runs as other-priority tasks until the end, its throughput retention rate compared to exclusive execution is not calculated. It can be seen that in Figure 3 the DynGPU sharing architecture of the example, compared with exclusive execution, the throughput loss is at most 7.2%, thus basically ensuring the performance of high-priority tasks.

[0097] Figure 5 , Figure 6 , Figure 7 respectively show the performance of four tasks under the Non, MPS, and TGS architectures (it should be noted that TGS needs to specify high-priority tasks before execution, so ResNet50 is selected. In addition, since TGS cannot assign priorities when there are more than two tasks, resource contention will occur). It can be seen that at each time point, these methods cannot effectively guarantee the performance of high-priority tasks. Figure 3 Compared with the Non, MPS, and TGS architectures, the throughput of the DynGPU sharing architecture of the example is increased by up to 8.12 times, 8.67 times, and 31.54 times respectively. This improvement is due to Figure 3 the DynGPU sharing architecture of the example can ensure the performance of high-priority tasks in a multi-task concurrent environment. It can be concluded that the dynamic GPU resource sharing mechanism of the DynGPU sharing architecture can guarantee the performance of high-priority tasks and retain most of the throughput compared with exclusive execution. Therefore, question Q1 has been answered.

[0098] 3) Task completion time: The above comprehensively analyzes the task completion time in all experiments, such as Figure 8As shown. The results show that the task completion time based on the DynGPU shared architecture is comparable to that of other architectures such as Baseline, MPS, or TGS, and shows significant advantages when compared with traditional scheduling strategies such as MPS and TGS. Specifically, the task completion time based on the DynGPU shared architecture is 8.4% and 6% higher than that of MPS and TGS respectively. The key to this improvement lies in the fact that the DynGPU shared architecture can not only ensure that the throughput and performance requirements of high-priority tasks are met, but also maximize the resource utilization of other low-priority tasks by intelligently utilizing idle GPU resources. This method enables the DynGPU shared architecture to improve the overall task parallelism and execution efficiency without affecting the performance of high-priority tasks. In addition, it is worth noting that this application has been significantly optimized in task scheduling, especially in the resource allocation of low-priority tasks. Compared with MPS and TGS, the DynGPU shared architecture can dynamically adjust task priorities, flexibly utilize idle GPU resources, and ensure that tasks are completed as quickly as possible, thus effectively reducing the average completion time of the DynGPU shared system. This advantage not only improves the overall efficiency of task scheduling, but also demonstrates the powerful potential of the DynGPU shared architecture in multi-task parallel processing. Therefore, Q2 can be clearly answered: Compared with MPS and TGS, the DynGPU shared architecture has significant advantages in task completion time, which stems from its efficient resource scheduling and optimization capabilities.

[0099] 4) GPU utilization: In the above example of dynamic GPU resource sharing, it is shown that the DynGPU shared architecture can effectively ensure the performance of high-priority tasks. However, in order to further improve resource utilization, it is desired that the DynGPU shared architecture can allow other priority tasks to utilize idle GPU resources as much as possible. To verify this, the following experiment was designed: The ResNet152 task was started at time 0, and two MobileNetV2 tasks were started at time 50, with batch sizes set to 4 and 16 respectively. The experimental results are as Figure 9As shown. During the execution of the ResNet152 task, the GPU utilization rate remained at around 50%-60%. Subsequently, when two MobileNetV2 tasks were launched at time 50, the GPU utilization rate gradually increased within approximately 10 to 20 seconds and finally stabilized at about 95.7%. This experiment shows that the DynGPU sharing architecture can effectively utilize idle GPU resources while executing high-priority tasks, thereby maximizing the resource utilization rate of other priority tasks. Based on the results of this experiment, we can answer Q3: DynGPU can enable other priority tasks to utilize idle GPU resources as much as possible while maintaining a high GPU utilization rate. This feature not only improves the efficiency of GPU resource utilization but also further verifies the flexibility and optimization ability of DynGPU in multi-task scheduling.

[0100] 5) Dynamic priority scheduling: In previous experiments, it was assumed that the priorities of all tasks were static. However, this assumption is unrealistic in actual production environments because the priorities of tasks change as tasks progress and new tasks are introduced. Especially when new tasks have urgent requirements, dynamic priority adjustment must be made to ensure that these urgent tasks receive sufficient performance guarantees. Therefore, in this experiment, dynamic priority scheduling was introduced based on the dynamic GPU resource sharing mechanism. To verify the effectiveness of this mechanism, an experiment involving five models was designed: ResNet1 52, ResNet50, MobileNetV2, AlexNet, and ShuffleNetV2. These models were launched at 0 seconds, 0 seconds, 2200 seconds, 1600 seconds, and 1200 seconds respectively. The batch size and number of iterations for all models were set to 64 and 10,000 respectively, and the deadlines (remaining time after the task starts) were 6700 seconds, 4800 seconds, 900 seconds, 420 seconds, and 780 seconds respectively. The experimental results are shown in the figure. From Figure 10It can be seen that all five tasks in the DynGPU shared architecture can be completed before their respective deadlines. In contrast, when using the Non and MPS architectures, significant delays occurred in the last three models, with the maximum delay rates reaching 247% and 343% respectively. Additionally, under the TGS architecture, the delay of MobileNetV2 was 180%. Although there were also some timeouts in the last two models, the amount of timeouts was still relatively small compared to the deadline requirements. From the experimental results, it can be concluded that the DynGPU shared architecture, through its dynamic priority scheduling mechanism, provides the necessary performance guarantee for urgent tasks, enabling them to complete as close to the deadline as possible. In contrast, traditional static priority scheduling strategies (such as Non, MPS, and TGS) failed to effectively handle the dynamic changes in task priorities and resource requirements. Therefore, Q4 can be answered: Through dynamic priority scheduling, the DynGPU shared architecture effectively ensures the performance of high-priority tasks, ensures that tasks are completed on time, and significantly improves the on-time completion rate of tasks.

[0101] 6) Overhead of the DynGPU shared architecture: As shown in the task completion time section, compared with the Non strategy, the task completion time using DynGPU increased by approximately 2.8%. Therefore, it can be concluded that introducing DynGPU in the "Dynamic GPU Resource Sharing" module increased the overhead by approximately 2.8%. This overhead mainly stems from the dynamic allocation and adjustment of GPU resources. Although relatively small, it still had a certain impact on the overall task completion time. In the dynamic priority scheduling module, when the task priority switches, the initialization of the thread state takes approximately 0.46 milliseconds. For most modern GPU systems, this time overhead can be ignored and has a negligible impact on system performance in practical applications. Therefore, even though dynamic priority scheduling was introduced to enhance the flexibility and performance guarantee of task scheduling, the additional overhead is extremely small for the overall system. In summary, although DynGPU introduced some overhead, these overheads were effectively balanced by the improvement in task scheduling accuracy and performance, resulting in a significant overall performance improvement. Therefore, Q5 can be answered: The overhead of the DynGPU shared architecture is small.

[0102] From the above experimental results, it can be seen that the DynGPU sharing architecture aims to optimize GPU resource utilization while ensuring that high-priority tasks (referred to as high-priority tasks) can meet their deadline requirements. By encapsulating each deep learning task in an independent Docker container, DynGPU ensures resource isolation and independent management, thereby improving the reliability, scalability, and fault tolerance of the DynGPU sharing system. The DynGPU sharing system achieves transparent GPU resource sharing, where high-priority tasks are first allocated the necessary resources, while other priority tasks utilize the remaining GPU capacity. In addition, the DynGPU sharing architecture also introduces a dynamic priority scheduling mechanism that adjusts task priorities based on factors such as completion rate, remaining time, and historical priority to ensure that critical tasks can be completed in a timely manner. Through these mechanisms, the DynGPU sharing architecture maximizes the overall system performance, reduces resource contention, and increases the task completion rate, making it a powerful solution for managing concurrent deep learning tasks in a GPU environment.

[0103] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 11 shown, the electronic device 11 in this embodiment mainly includes: a processor 110, a memory 111, and a computer program 112 stored in the memory 111 and executable on the processor 110, such as a program for the dynamic graphics processor sharing method. When the processor 110 executes the computer program 112, it implements the steps in the above-mentioned embodiment of the dynamic graphics processor sharing method, such as Figure 1 the steps S101 to S106 shown. Alternatively, when the processor 110 executes the computer program 112, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the receiving module 201, the generating module 202, the triggering module 203, the adjusting module 204, the executing module 205, and the updating module 206 shown.

[0104] Exemplarily, the computer program 112 of the dynamic graphics processor sharing method mainly includes: Step S101: Receive the submission of multiple deep learning tasks, where each of the multiple deep learning tasks is associated with a deadline parameter; Step S102: Establish a dynamic priority evaluation model, calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks; Step S103: Based on the real-time thread block usage of high-priority tasks, transmit the resource status to low-priority tasks through the broadcast mechanism, and trigger resource adjustment according to the comparison result between the real-time thread block usage of high-priority tasks and a preset threshold; Step S104: Dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of high-priority tasks and the preset threshold; Step S105: Maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle scheduling; Step S106: Periodically recalculate the priority scores of all deep learning tasks and update the dynamic priority queue, and repeat Steps S103 to S105 based on the updated dynamic priority queue until all deep learning tasks are completed. The computer program 112 can be divided into one or more modules / units. One or more modules / units are stored in the memory 111 and executed by the processor 110 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 112 in the electronic device 11.For example, the computer program 112 can be divided into the functions of a receiving module 201, a generating module 202, a triggering module 203, an adjusting module 204, an executing module 205, and an updating module 206 (modules in the virtual device). The specific functions of each module are as follows: The receiving module 201 is used to receive the submission of multiple deep learning tasks, where each task of the multiple deep learning tasks is associated with a deadline parameter; the generating module 202 is used to establish a dynamic priority evaluation model, calculate a priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks; the triggering module 203 is used to transmit the resource status to low-priority tasks through a broadcast mechanism based on the real-time thread block usage of high-priority tasks, and trigger resource adjustment according to the comparison result between the real-time thread block usage of the high-priority tasks and a preset threshold; the adjusting module 204 is used to dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of the high-priority tasks and the preset threshold; the executing module 205 is used to maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle of scheduling; the updating module 206 is used to periodically recalculate the priority scores of all deep learning tasks, update the dynamic priority queue, and repeat the execution of the triggering module 203, the adjusting module 204, and the executing module 205 based on the updated dynamic priority queue until all deep learning tasks are completed.

[0105] The electronic device 11 may include but is not limited to a processor 110 and a memory 111. Those skilled in the art can understand that Figure 11 merely examples of the electronic device 11, which do not constitute a limitation on the electronic device 11, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0106] The so-called processor 110 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0107] The memory 111 may be an internal storage unit of the electronic device 11, such as the hard disk or memory of the electronic device 11. The memory 111 may also be an external storage device of the electronic device 11, such as a plug-in hard disk equipped on the electronic device 11, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 111 may also include both the internal storage unit of the electronic device 11 and the external storage device. The memory 111 is used to store computer programs and other programs and data required by the electronic device. The memory 111 may also be used to temporarily store data that has been output or will be output.

[0108] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0109] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0110] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0111] In the embodiments provided in this application, it should be understood that the disclosed devices / apparatuses and methods can be implemented in other ways. For example, the device / apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0112] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0113] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0114] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program for the dynamic graphics processor sharing method can be stored in a storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments, that is, step S101: Receive the submission of multiple deep learning tasks, where each of the multiple deep learning tasks is associated with a deadline parameter; step S102: Establish a dynamic priority evaluation model, calculate the priority score based on the historical priority, completion rate, and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks; step S103: Based on the real-time thread block usage of high-priority tasks, transmit the resource status to low-priority tasks through a broadcast mechanism, and trigger resource adjustment according to the comparison result between the real-time thread block usage of high-priority tasks and a preset threshold; step S104: Dynamically adjust the number of executable thread blocks of low-priority tasks according to the difference between the real-time thread block usage of high-priority tasks and the preset threshold; step S105: Maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, check whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If it is satisfied, start the execution; otherwise, retain the kernel in the kernel request queue and wait for the next cycle of scheduling; step S106: Periodically recalculate the priority scores of all deep learning tasks and update the dynamic priority queue, and repeat steps S103 to S105 based on the updated dynamic priority queue until all deep learning tasks are completed. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.

[0115] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. The specific implementation manners described above have further elaborated on the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above is only the specific implementation manner of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included in the protection scope of the present invention.

Claims

1. A dynamic graphics processor sharing method, characterized in that: The method comprises: Step S101: receiving submissions of multiple deep learning tasks, each of the multiple deep learning tasks being associated with a deadline parameter; Step S102: establishing a dynamic priority evaluation model, calculating a priority score based on the historical priority, completion rate and remaining execution time of each task, and generating a dynamic priority queue including high-priority tasks and low-priority tasks; Step S103: Based on the real-time thread block usage of the high-priority task, the resource state is transmitted to the low-priority task through a broadcast mechanism, and resource adjustment is triggered according to the comparison result between the real-time thread block usage of the high-priority task and a preset threshold; Step S104: dynamically adjusting the number of executable thread blocks of the low priority task according to the difference between the real-time thread block usage of the high priority task and the preset threshold; Step S105: maintaining an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, checking whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task is checked. If so, the execution is started; otherwise, the kernel is retained in the kernel request queue and waits for the next cycle scheduling. Step S106: Periodically recalculate the priority scores of all deep learning tasks and update the dynamic priority queue, and repeat steps S103 to S105 based on the updated dynamic priority queue until all deep learning tasks are completed.

2. The dynamic graphics processor sharing method according to claim 1, characterized in that: The calculating of the priority score based on the historical priority, completion rate and remaining execution time of each task includes calculating the priority score Score of each task according to the following calculation formula: Score=w1*(1-x1)+w2*(1-x2)+w3*x3, where x1 is the standardized value of the historical priority of the task, x2 is the standardized value of the completion rate of the task, x3 is the standardized value of the remaining time of the task execution, w1, w2 and w3 are the weights corresponding to (1-x1), (1-x2) and x3 respectively, and w1+w2+w3=1.

3. The dynamic graphics processor sharing method according to claim 1, characterized in that: The method of transmitting resource status to low-priority tasks through a broadcast mechanism based on the real-time thread block usage of the high-priority task, and triggering resource adjustment according to a comparison result between the real-time thread block usage and a preset threshold, includes: Calculate the thread block usage BH of the current high-priority task in real time, and pass the BH value to all low-priority tasks through the broadcast mechanism; The thread block usage BH of the current high priority task is compared with the preset threshold MH, and if |BH-MH| exceeds the preset difference threshold, the low priority task resource adjustment is triggered.

4. The dynamic graphics processor sharing method according to claim 1, characterized in that: The dynamically adjusting the number of executable thread blocks of the low priority task according to the difference between the real-time thread block usage of the high priority task and the preset threshold value includes: When |BH-MH| exceeds the preset difference threshold, the number of executable thread blocks B of the low priority task is increased. exce [i] Reduce to 1 / 2 of the current value; When |BH-MH| does not exceed the preset difference threshold, the number of executable thread blocks B of the low priority task is gradually increased in an exponential manner. exce [i], until the B exce [i] Achieve the ideal number of thread blocks B for the low-priority task ideal .

5. The dynamic graphics processor sharing method according to claim 4, characterized in that: The number of executable thread blocks B of the low priority task is gradually increased in an exponential manner. exce The specific implementation of [i] is: The t is the number of cycles since the last resource adjustment, and after the adjustment, B min ≤B′ exec [i]≤B max , when B′ exec [i] Achieve B ideal When the value is set to B, the increment stops and the value is locked. ideal value.

6. The dynamic graphics processor sharing method according to claim 1, characterized in that: The execution rules of the kernel request queue also include: If the number of thread blocks required by the kernel exceeds the number of executable thread blocks of the current low-priority task B exce [i], reinsert the kernel into the tail of the kernel request queue and record its waiting times; When the number of waiting times exceeds a preset upper limit, the kernel is forced to start and the number of executable thread blocks of the current low priority task is allowed to be increased. exce [i] Exceed the maximum number of thread blocks B of the current low priority task max restrictions.

7. The dynamic graphics processor sharing method according to claim 1, characterized in that: The triggering condition for updating the dynamic priority queue after periodically recalculating the priority scores of all deep learning tasks is: Trigger updates at fixed intervals; or When the remaining execution time of any task falls below the preset alarm threshold, an update is triggered immediately.

8. A dynamic graphics processor sharing device, characterized in that: The device comprises: A receiving module, configured to receive submissions of a plurality of deep learning tasks, each of the plurality of deep learning tasks being associated with a deadline parameter; A generation module, used to establish a dynamic priority evaluation model, calculate the priority score based on the historical priority, completion rate and remaining execution time of each task, and generate a dynamic priority queue including high-priority tasks and low-priority tasks; A trigger module, configured to transmit resource status to low-priority tasks through a broadcast mechanism based on the real-time thread block usage of high-priority tasks, and trigger resource adjustment according to a comparison result between the real-time thread block usage of the high-priority tasks and a preset threshold; An adjustment module, configured to dynamically adjust the number of executable thread blocks of a low priority task according to a difference between the real-time thread block usage of the high priority task and the preset threshold; An execution module is used to maintain an independent kernel request queue for each low-priority task. When there is a kernel to be executed in the kernel request queue, it is checked whether the number of thread blocks required by the kernel is less than or equal to the number of executable thread blocks of the current low-priority task. If so, the execution is started; otherwise, the kernel is retained in the kernel request queue and waits for the next cycle scheduling; An update module is used to periodically recalculate the priority scores of all deep learning tasks and then update the dynamic priority queue, and repeatedly execute the trigger module, adjustment module and execution module based on the updated dynamic priority queue until all deep learning tasks are completed.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Process scheduling method and device, storage medium and computer program product

    CN120508398A

  • A process scheduling method, device, storage medium and computer program product

    CN120508398B

  • Distributed content distribution network operation method and system, and storage medium

    CN120750926A

  • IO flow speed limit control method and device, electronic equipment and storage medium

    CN120891990A