GPU task allocation method based on NVIDIA graphics card

By setting up task client module, preload module, matching module, execution module and monitoring module on NVIDIA graphics card, the problem of singleness and insufficient remedial measures of GPU task allocation scheme in the prior art is solved, and reasonable allocation of tasks and high fault tolerance are achieved.

CN120029750APending Publication Date: 2025-05-23HANGZHOU ARCVIDEO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311570144.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing GPU task allocation scheme based on NVIDIA graphics cards has a single allocation method, inflexible replacement and expansion, and lacks effective remedial measures for the failure of operation after task allocation, and has low fault tolerance.

Method used

A GPU task allocation method based on NVIDIA graphics card is provided, and the task client module, preload module, matching module, execution module and monitoring module are implemented to achieve reasonable assignment and status monitoring. This method supports custom allocation policies and remediates for exception tasks.

Benefits of technology

The reasonable allocation of different GPU tasks is realized on NVIDIA graphics cards, supports custom allocation policies, monitors task status in real time, and performs appropriate remediation of abnormal tasks, improving the system's fault tolerance and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029750A_ABST
    Figure CN120029750A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU (Graphics Processing Unit) task allocation method based on an NVIDIA graphics card, which comprises the following steps of: setting a task client module to define and assemble GPU tasks executable by the NVIDIA graphics card, standardizing task attributes, generating task information, and dispatching and executing the task information by a preloading module, a matching module and an execution module; a pre-loading module is arranged to perform resource occupation pre-analysis on GPU tasks and obtain NVIDIA existing resources; setting a matching module to match a GPU task allocation rule, and pre-configuring an allocation strategy adopted by a current system; setting a task execution module to perform task starting and task closing according to an allocation strategy of the matching module, performing state marking on the GPU task to distinguish a task state, and providing a corresponding interface for a subsequent monitoring module to query; a monitoring module is arranged to monitor the task state of the execution module in real time, and whether the abnormal GPU task needs to be remedied or not is judged according to the state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular relates to a GPU task allocation method based on an NVIDIA graphics card. Background Art

[0002] NVIDIA defines GPUs (Graphics Processing Units), and tasks running on NVIDIA graphics cards are referred to as GPU tasks. When handling different GPU tasks during daily development, it's important to consider how to efficiently allocate them to the target NVIDIA graphics card. Currently, there are few solutions for allocating tasks based on NVIDIA graphics cards. Most existing solutions use simple methods like round-robin or random allocation, resulting in a single allocation method and inflexible replacement and expansion. Furthermore, there are few remedial measures or differentiated handling for task failures after assignment, resulting in low fault tolerance. Summary of the Invention

[0003] In view of the above problems, the present invention provides a GPU task allocation method based on NVIDIA graphics card, which can reasonably allocate different GPU tasks to be executed on the NVIDIA graphics card.

[0004] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0005] A GPU task allocation method based on NVIDIA graphics card includes the following steps:

[0006] Set up the task client module to define and assemble GPU tasks executable by NVIDIA graphics cards, standardize task attributes, generate task information, and then hand it over to the preloading module, matching module, and execution module for scheduling and execution; task information includes task execution implementation information and task content information;

[0007] Set up a preload module to pre-analyze GPU task resource usage and obtain existing NVIDIA resources;

[0008] Setting the matching module to match GPU task allocation rules and pre-configuring the allocation strategy used by the current system, including polling, designated graphics card, minimum usage, and implementation of custom allocation rules;

[0009] Set the task execution module to start and close tasks according to the allocation strategy of the matching module, mark the GPU tasks to distinguish the task status, and provide the corresponding interface for subsequent monitoring module to query;

[0010] The monitoring module is set to monitor the task status of the execution module in real time, and determine whether remediation is needed for abnormal GPU tasks based on the status.

[0011] In one possible implementation, polling is a round-robin method within a specified boundary, and tasks are assigned to specific NVIDIA graphics cards in sequence according to the corresponding card numbers on the NVIDIA graphics cards.

[0012] In a possible implementation, specifying a graphics card is a method of strongly binding a task to a corresponding card number on an NVIDIA graphics card.

[0013] In one possible implementation, the card number is selected by detecting the remaining ratio of NVIDIA video memory and sorting the cards by the minimum usage rate, and the card with the lowest video memory usage rate is selected as the target for executing the task.

[0014] In a possible implementation, the custom allocation rule is implemented as follows: the system provides the task information in the preloaded information, and the implementer provides a corresponding open interface to complete the custom rule.

[0015] In a possible implementation, the task status includes waiting for execution, executing, successful execution, and failed execution.

[0016] In a possible implementation, setting the monitoring module to perform real-time monitoring on the task status of the execution module includes: the status monitoring module of the monitoring module periodically obtains all executed task information and task execution status information from the execution module.

[0017] In one possible implementation, remediation of abnormal GPU tasks includes retrying upon failure, placing the failed tasks into a retry queue and retrying them according to a gradient of times.

[0018] In one possible implementation, remediation of abnormal GPU tasks includes recording an alarm, making a failure record for a failed task, marking a task as having too many failures when the task exceeds a certain number of failure records, and providing an alarm output.

[0019] The adoption of the present invention has the following beneficial effects: different GPU tasks can be reasonably allocated for execution on NVIDIA graphics cards, not only supporting customized GPU task allocation strategies, but also monitoring the status of running tasks in real time, and providing appropriate remedial measures for tasks with abnormal running status. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flowchart of the steps of the GPU task allocation method based on the NVIDIA graphics card according to an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of obtaining video memory from an NVIDIA graphics card in a specific application example. DETAILED DESCRIPTION

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0023] Reference Figure 1 , which is a flowchart of a method for allocating GPU tasks based on an NVIDIA graphics card according to an embodiment of the present invention, includes the following steps:

[0024] S10, setting the task client module to define and assemble GPU tasks executable by the NVIDIA graphics card, standardize task attributes, generate task information, and then hand it over to the preloading module, matching module, and execution module for scheduling and execution; the task information includes task execution implementation information and task content information;

[0025] S20, setting a pre-loading module to pre-analyze the resource usage of GPU tasks and obtain existing NVIDIA resources;

[0026] S30, setting a matching module to match GPU task allocation rules, and pre-configuring an allocation strategy used by the current system, wherein the allocation strategy includes polling, specifying a graphics card, minimum usage, and implementation of a custom allocation rule;

[0027] S40, set the task execution module to start and close tasks according to the allocation strategy of the matching module, and at the same time mark the GPU task status to distinguish the task status, and provide the corresponding interface for subsequent monitoring module to query; the task status includes waiting for execution, executing, successful execution and failed execution.

[0028] S50: Setting a monitoring module to monitor the task status of the execution module in real time, and judging whether it is necessary to remedy the abnormal GPU task according to the status.

[0029] In an embodiment of the present invention, the GPU task allocation method based on NVIDIA graphics cards requires that the task execution implementation in S10 corresponds to the execution module. For example, if the video transcoding task plug-in is named transcode, then the transcode task client needs to fill in the task startup and task shutdown execution implementations. The task content refers to the specific content that needs to be executed. Taking the video transcoding task as an example, the task content needs to include the input source, transcoding parameters, output address, etc. The task content is not limited. Different tasks have different task contents and corresponding execution implementations. After matching the specific strategy, the matching module can submit the task information to the execution module.

[0030] In a GPU task allocation method based on NVIDIA graphics cards according to an embodiment of the present invention, in S20, if the NVIDIA graphics card resources are insufficient to meet the execution of the current GPU task, feedback and subsequent processing are performed in a timely manner. There is no fixed resource for obtaining the GPU task resource occupancy. It mainly combines matching strategies to obtain corresponding resources. Taking the video transcoding task plug-in named transcode as an example, it is sensitive to the NVIDIA video memory occupancy. Therefore, both the formulated matching strategy and the pre-loading module obtain the GPU task video memory occupancy and the current NVIDIA existing resource amount. The following is an example of the NVIDIA graphics card obtaining video memory, and the result is shown as Figure 2 shown below:

[0031] Command: nvidia-smi returns information about card 0: total video memory 8119MiB, used 7681MiB.

[0032] In a GPU task allocation method based on NVIDIA graphics cards according to an embodiment of the present invention, in S30, polling is a way of轮流循环 within a specified boundary. According to the corresponding card numbers on the NVIDIA graphics card, the task allocation is sequentially executed to the specific NVIDIA graphics card in order. For example, the corresponding card numbers on the NVIDIA graphics card are 0-3, and the task allocation will be sequentially executed to the specific card in order. Designating a graphics card is a way of strongly binding a task to the corresponding card number on the NVIDIA graphics card, which is a one-to-one binding method. If a GPU task must be executed on a certain card, then a strong binding relationship can be established through this way of designating the card for execution. The minimum utilization rate selects the card number by sorting the detected remaining ratio of NVIDIA video memory. The one with the lowest video memory utilization rate is elected as the target for executing the task. The implementation of the custom allocation rule is as follows: The system provides the task information in the pre-loading information, and the implementer of the task information provides the corresponding open interface to complete the custom rule. The implementation of the custom allocation rule is a reserved implementation method for complex allocation scenarios. The system will provide the task information in the pre-loading information, and the implementer of the task information needs to implement the corresponding open interface to complete the custom rule. The following is a Java pseudo-code example of the interface:

[0033] public interface BalanceAlgorithm<GPU task information, NVIDIA all graphics card usage information>{

[0034] Selected graphics card information balance(GPU task information, NVIDIA all graphics card usage information);

[0035] }

[0036] In an embodiment of the present invention, a GPU task allocation method based on an NVIDIA graphics card is provided. In S40, the task execution module provides two methods, namely task start run() and task stop stop(). These methods need to be filled in and implemented in the task client module. Different tasks correspond to different filling implementations. The implementation form is not limited to the purpose of correctly starting and stopping GPU tasks. A pseudo code example is as follows:

[0037] public void stop(){

[0038] TaskStop();

[0039] }

[0040] public void run(){

[0041] TaskStart();

[0042] }

[0043] In one embodiment of the present invention, a method for allocating GPU tasks based on an NVIDIA graphics card is provided. In S50, a monitoring module is set to perform real-time monitoring on the task status of the execution module, including: the status monitoring module of the monitoring module periodically obtains all executed task information and task execution status information from the execution module. Remediation of abnormal GPU tasks includes retrying failures. For tasks that fail to execute, the tasks are placed in a retry queue and retried according to a gradient of times. The gradient of times is to reasonably control the retry frequency and reduce resource waste. An example of the gradient is: retries 3 times within 1 minute and still fails, to retrying once within 2 minutes, to retrying once within 10 minutes, to retrying once within 30 minutes. Remediation of abnormal GPU tasks includes recording alarms, making failure records for failed tasks, marking too many failures when the task exceeds a certain number of failure records, and providing an alarm information output.

[0044] Through the above-set GPU task allocation method based on NVIDIA graphics cards, different GPU tasks can be reasonably allocated for execution on NVIDIA graphics cards. It not only supports customized GPU task allocation strategies, but also monitors the status of running tasks in real time and provides appropriate remedial measures for tasks with abnormal running status.

[0045] It should be understood that the exemplary embodiments described herein are illustrative and not restrictive. Although one or more embodiments of the present invention have been described in conjunction with the accompanying drawings, it should be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present invention as defined by the appended claims.

Claims

1. A GPU task allocation method based on NVIDIA graphics card, It is characterized in that The following steps are involved: Set up the task client module to define and assemble GPU tasks executable by NVIDIA graphics cards, standardize task attributes, generate task information, and then hand it over to the preloading module, matching module, and execution module for scheduling and execution; task information includes task execution implementation information and task content information; Set up a preload module to pre-analyze GPU task resource usage and obtain existing NVIDIA resources; Setting a matching module to match GPU task allocation rules, and pre-configuring the allocation strategy adopted by the current system, wherein the allocation strategy includes polling, specifying a graphics card, minimum usage rate, and implementation of a custom allocation rule; Set the task execution module to start and close tasks according to the allocation strategy of the matching module, mark the GPU tasks to distinguish the task status, and provide corresponding interfaces for subsequent monitoring modules to query; The monitoring module is set to monitor the task status of the execution module in real time, and determine whether it is necessary to remedy the abnormal GPU task according to the status.

2. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 1, It is characterized in that Polling is a round-robin method within a specified boundary. Tasks are assigned to specific NVIDIA graphics cards in sequence according to the corresponding card number on the NVIDIA graphics card.

3. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 1, It is characterized in that Specifying a graphics card is a way to strongly bind a task to the corresponding card number on the NVIDIA graphics card.

4. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 1, It is characterized in that The minimum usage rate is selected by detecting the remaining ratio of NVIDIA video memory and sorting it. The card number with the lowest video memory usage rate is selected as the target for executing the task.

5. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 1, It is characterized in that The implementation of the custom allocation rule is as follows: the system provides the task information in the preloaded information, and the implementer provides the corresponding open interface to complete the custom rule.

6. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 1, It is characterized in that The task status includes waiting for execution, executing, successful execution, and failed execution.

7. The GPU task allocation method based on NVIDIA graphics card according to any one of claims 1 to 6, It is characterized in that The setting of the monitoring module to perform real-time monitoring on the task status of the execution module includes: the status monitoring module of the monitoring module periodically obtains all executed task information and task execution status information from the execution module.

8. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 7, It is characterized in that Remediation of abnormal GPU tasks includes failure retry. For tasks that failed to execute, the tasks are placed in a retry queue and retried according to a gradient of times.

9. The GPU task allocation method based on NVIDIA graphics card as claimed in claim 7, It is characterized in that Remediation of abnormal GPU tasks includes recording alarms, making failure records for failed tasks, marking too many failures when a task exceeds a certain number of failure records, and providing alarm information output.