Task processing method, multi-core graphics processing unit, electronic device, storage medium and program product

By actively determining the target subtasks in the multi-core graphics processor, the problems of global scheduling module limitations and task differences between cores are solved, and more efficient GPU task processing is achieved.

WO2025139791A1PCT designated stage expired Publication Date: 2025-07-03MOORE THREADS TECH CO LTD

Patent Information

Application Number
PCT/CN2024/138457
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-11
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the subtask scheduling efficiency of multi-core graphics processors is limited by performance bottlenecks caused by global scheduling modules or task differences between cores, affecting GPU processing performance.

Method used

In a multi-core graphics processor, each core actively determines the target subtasks based on its own processing capabilities and the number of assigned subtasks, realizes autonomous scheduling of subtasks, and reduces load imbalance between cores.

Benefits of technology

It improves the task processing performance of GPU, reduces the accumulation of subtasks in part of cores, and improves the balance of load between cores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138457_03072025_PF_FP_ABST
    Figure CN2024138457_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are a task processing method for a multi-core graphics processing unit (GPU), the multi-core GPU, an electronic device, a computer-readable storage medium and a computer program product. The method comprises: in the process of a GPU processing a first task, determining a target sub-task of a first core from the first task on the basis of the current processing capability of the first core and the number of allocated sub-tasks in the first task that have been allocated to each core, wherein the first core is one of multiple cores; and by means of the first core, acquiring the target sub-task and processing same.
Need to check novelty before this filing date? Find Prior Art

Description

Task processing method, multi-core graphics processor, electronic device, storage medium and program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The embodiments of this application are based on the Chinese patent application with application number 202311868583.5, application date December 29, 2023, and application name “Task processing method, multi-core graphics processor, electronic device and storage medium”, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field

[0003] The present application relates to, but is not limited to, the field of computer technology, and in particular to a task processing method, a multi-core graphics processor, an electronic device, a storage medium, and a program product. Background Art

[0004] Large-scale graphics processing units (GPUs) typically consist of multiple cores. GPU tasks can be divided into multiple subtasks and assigned to each core for execution. However, assigning subtasks to each core requires a globally unique scheduling module, which limits GPU processing efficiency. Alternatively, this process requires pre-setting the set of subtasks that each core needs to execute. Because each core takes different amounts of time to execute its own set of subtasks, the runtimes of each core vary significantly, leading to poor GPU processing performance. Summary of the Invention

[0005] The embodiments of the present application provide a task processing method, a multi-core graphics processor, an electronic device, a storage medium, and a program product, which enable each core to actively acquire subtasks based on its own processing capabilities, implement subtask scheduling, make the subtasks of multiple cores more balanced, and improve GPU processing performance.

[0006] The technical solution of this application is achieved as follows:

[0007] An embodiment of the present application provides a task processing method for a multi-core graphics processor (GPU), comprising: during processing of a first task by the graphics processor, determining a target subtask for the first core from the first task based on the current processing capability of the first core and the number of allocated subtasks in the first task that have been allocated to each core; the first core being one of the multiple cores; and obtaining and processing the target subtask through the first core.

[0008] An embodiment of the present application provides a multi-core graphics processor, comprising: a first core, configured to, during processing of a first task by the multi-core graphics processor, determine a target subtask for the first core from the first task based on a current processing capability of the first core and a number of subtasks already allocated to each core in the first task; the first core being one of the multiple cores; and obtaining and processing the target subtask.

[0009] An embodiment of the present application provides an electronic device including the multi-core graphics processor described above.

[0010] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the above-mentioned task processing method are implemented.

[0011] An embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the above-mentioned task processing method.

[0012] Embodiments of the present application provide a task processing method, a multi-core graphics processor, an electronic device, a computer storage medium, and a program product. Since the cores in the GPU can determine and promptly acquire the target subtask in the first task based on the number of subtasks that can be processed for the subtask and the number of assigned subtasks that have been assigned to the cores; in this way, the faster the core processes the subtasks and the more subtasks it completes, the more subtasks it can acquire, thereby reducing the accumulation of subtasks in some cores, improving the balance of load between cores, and thereby improving the task processing performance of the GPU.

[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0015] FIG1 is a schematic diagram of a global scheduling process provided by an embodiment of the present application;

[0016] FIG2 is a schematic diagram of an autonomous scheduling process in a related technology provided by an embodiment of the present application;

[0017] FIG3 is a process diagram of a task processing method provided in an embodiment of the present application;

[0018] FIG4 is a process diagram of another task processing method provided in an embodiment of the present application;

[0019] FIG5 is a schematic diagram of a multi-core autonomous scheduling process provided by an embodiment of the present application;

[0020] FIG6 is a process diagram of another task processing method provided in an embodiment of the present application;

[0021] FIG7 is a process diagram of another task processing method provided in an embodiment of the present application;

[0022] FIG8 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0024] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0025] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0027] To facilitate understanding of this solution, before describing the embodiments of the present application, the application background of the embodiments of the present application will be described.

[0028] In related art, subtask scheduling among multiple GPU cores can include global scheduling and autonomous scheduling. Global scheduling requires a unified hardware module configured to perform subtask scheduling within the GPU. This module then assigns subtasks to all cores within the GPU. As shown in Figure 1, after the central processing unit (CPU) 2 issues a task to GPU 1, the task is stored in the GPU's memory 13. The task includes multiple subtasks. The global scheduling module 11 can read the subtasks from memory 13 and assign them to multiple cores (exemplarily shown as cores 12-1, 12-2, 12-3, and 12-4). Because the number of cores in a GPU far exceeds the number shown in the figure, the global scheduling module 11 may not be able to keep up with the cores' processing capabilities when assigning tasks to multiple cores. This may result in some cores' subtasks being processed before the global scheduling module 11 has time to assign them new subtasks, resulting in GPU performance loss. Autonomous scheduling, on the other hand, lacks a global scheduling module and requires each core within the GPU to determine and process its own set of subtasks according to a pre-agreed method. FIG2 is a schematic diagram of the autonomous scheduling process in a related technology provided by an embodiment of the present application. As shown in FIG2 , after CPU4 sends a task to GPU3, the task is stored in the memory 33 of the GPU, and the task includes multiple subtasks. Multiple cores within GPU3 (core 32-1, core 32-2, core 32-3, and core 32-4 are shown as examples) can obtain a subtask set from the memory 33 and process the subtasks in the subtask set. It should be noted that there is no dependency between the subtasks, and each core can independently process the subtasks in its own subtask set. Here, each core can obtain a subtask set based on the position of the subtask in the vector space, or it can obtain a subtask set based on the position of the vector space and the subtask identifier; for example, the subtask set of core 32-1 can include subtasks with an even subtask identifier in the upper left position. Since the number of subtasks in the subtask set of each core can be different, and the processing time of each subtask can also be different, the time required for each core to complete the subtasks in the subtask set may also be very different; thus, the total running time of each core is different, and the difference may be large. The core with the longest running time will become the performance bottleneck of the GPU, affecting the GPU's processing performance.

[0029] In order to solve the above problems, an embodiment of the present application provides a task processing method, which can be executed by a multi-core graphics processor of an electronic device. Among them, the electronic device may refer to a server, a laptop, a tablet computer, a desktop computer, a smart TV, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a portable gaming device), etc., which has a display function and requires a graphics processor to process tasks. Figure 3 is a process diagram of a task processing method provided by an embodiment of the present application. As shown in Figure 3, the method includes: S101 to S102.

[0030] S101. During GPU processing of a first task, determine a target subtask of the first core from the first task based on a current processing capability of the first core and a number of allocated subtasks allocated to each core in the first task; the first core is one of multiple cores.

[0031] In an embodiment of the present application, when a graphics processing unit (GPU) processes a first task, it does so by having multiple cores of the GPU process subtasks of the first task. Each core's processing capacity for subtasks can be characterized by its own preset maximum number of tasks, that is, the maximum number of tasks that can be processed simultaneously. Here, the processing capacity of each core can be the same or different; this can be set as needed and is not limited by the embodiment of the present application. As each core processes subtasks, subtasks are continuously completed. At this time, the current processing capacity of each core is the number of subtasks that can currently be processed.

[0032] In this embodiment of the present application, after the GPU receives the first task issued by the CPU, each core in the GPU can obtain and process subtasks from the first task. The first core can determine its own target subtask from the first task based on its current processing capacity and the number of assigned subtasks, obtain the target subtask, and process it. Here, the first core is one of the multiple cores.

[0033] In this embodiment of the present application, all subtasks acquired by the first core, whether currently being processed or pending, are considered unfinished tasks for the first core; the number of unfinished tasks should be less than or equal to the maximum number of tasks for the first core. The number of subtasks that the first core can process can be determined based on the maximum number of tasks and the number of unfinished tasks; here, the sum of the number of processable subtasks and the number of unfinished tasks should be less than or equal to the maximum number of tasks.

[0034] In some embodiments, the difference between the maximum number of tasks and the number of uncompleted tasks is equal to the number of processable subtasks.

[0035] For example, the multiple cores include four cores: core 1 to core 4. The maximum number of tasks for core 1 is 5. If the number of subtasks in core 1 is 0, the number of unfinished tasks in core 1 is also 0. In this case, the number of subtasks that can be processed is 5. After core 1 obtains and begins processing five subtasks, if one subtask is completed, the number of unfinished tasks for core 1 is 4, and the number of subtasks that can be processed is 1.

[0036] In an embodiment of the present application, after the GPU receives the first task sent by the CPU, it can store the first task in the memory of the GPU; the multiple cores in the GPU need to obtain subtasks from the memory. Here, the subtasks that have been obtained by each core are the allocated subtasks in the first task, and the subtasks that have not yet been allocated to each core are the unallocated subtasks. The electronic device needs to maintain the operating performance of each core so that the number of subtasks in each core can continue to maintain the maximum task volume. Therefore, in the process of the GPU processing the first task, each core needs to obtain and execute subtasks from the unallocated subtasks in a timely manner according to the status of its own processing subtasks.

[0037] In an embodiment of the present application, the total number of subtasks of the first task is determined, and the allocation order of all subtasks is determined. The first core can obtain subtasks of the number of processable subtasks from the unassigned subtasks of the first task. If the number of assigned subtasks can be determined, the first core can determine the number of processable subtasks that can be assigned to the core from the unassigned subtasks based on the allocation order.

[0038] It should be noted that each subtask has its own task identifier. The task identifier can represent the order in which the subtasks are assigned. In some embodiments, the task identifier can be represented by letters, and the assignment order can be arranged in alphabetical order. In some embodiments, the task identifier can be represented by numbers, and the assignment order can be arranged in numerical order.

[0039] For example, task A includes M subtasks, and the task identifiers of the subtasks can be: a0 to a M-1 ; In this way, the order of subtask allocation can be allocated in the order of the task identifier from small to large. If 3 subtasks have been allocated, the unallocated subtasks include a3 to a M-1 At this time, if one of the multi-cores requires two subtasks, the core needs to obtain a3 and a4.

[0040] In some embodiments, the task identifier may be a 64-bit integer identifier; in this way, all subtasks may have a globally unique task identifier.

[0041] In an embodiment of the present application, the first task acquired by the GPU may be split, that is, the first task acquired by the GPU may be composed of multiple subtasks, and the first core may read the subtasks one by one as needed. The first task acquired by the GPU may also be unsplit, and the first core may split and read the subtasks as needed; this can be set as needed and is not limited in the embodiment of the present application.

[0042] In some embodiments, multiple subtasks of the first task can be stored in a first-in-first-out (FIFO) queue in the order of allocation. In this way, the number of processable subtasks arranged at the front of the FIFO is the target subtask, and the first core can obtain the subtask from the first-in-first-out queue based on the number of processable subtasks. In some embodiments, the GPU can record the number of allocated subtasks, and the first core can obtain the number of allocated subtasks and determine the target subtask from the first task in combination with the number of processable subtasks. In some embodiments, the first core can determine the identifier of the target subtask from the first task based on the number of processable subtasks and the number of allocated subtasks, and then determine the target subtask.

[0043] S102: Obtain and process the target subtask through the first core.

[0044] In the embodiment of the present application, after determining the target subtask, the first core may obtain the target subtask in the first task and process the target subtask.

[0045] In the embodiment of the present application, as long as not all subtasks in the first task have been assigned to all cores, that is, there are still unassigned subtasks in the GPU, the first core can obtain the target subtask from the unassigned subtasks when the number of subtasks being processed is less than the maximum task capacity. When all subtasks of the first task are assigned and all cores in the GPU have completed processing the subtasks of the first task, the processing of the first task is completed.

[0046] It can be understood that since the first core in the GPU can determine and promptly obtain the target subtask in the first task based on the number of processable subtasks and the number of allocated subtasks allocated to each core; in this way, among the cores in the GPU, the faster the core processes subtasks, the more subtasks it completes and the more subtasks it can obtain, thereby reducing the possibility of subtasks piling up in some cores, improving the load balance between multiple cores, and thus improving the GPU's ability to process tasks.

[0047] In some embodiments of the present application, the multi-core GPU includes a task counting module, and the number of allocated subtasks of each core is determined by a count value of the task counting module.

[0048] In an embodiment of the present application, a multi-core GPU is provided with a task counting module. The number of assigned subtasks of a first task that have been assigned to each core can be determined by the task counting module. Here, as each core acquires subtasks of the first task, the task counting module can conveniently accumulate the number of subtasks acquired by the core, thereby obtaining the number of assigned subtasks, thereby improving the efficiency of determining the number of assigned subtasks.

[0049] In some embodiments, the task counting module can be set in the memory of the GPU. In this way, when each core retrieves the subtasks of the first task from the memory, the task counting module can quickly accumulate the count value, thereby improving the efficiency of determining the number of assigned subtasks. In some embodiments, the task counting module can be set in any core, or in other components of the GPU outside of the memory and cores; this can be set as needed and is not limited by the embodiments of the present application. In some embodiments, the task counting module can be a counter set in the GPU.

[0050] Figure 4 is a process diagram of another task processing method provided in an embodiment of the present application. In some embodiments of the present application, determining the implementation of the target subtask of the first core from the first task in S101 may include: S201 to S203 as shown in Figure 4.

[0051] S201 : Send a subtask acquisition instruction to a task counting module via a first core; the subtask acquisition instruction includes the current processing capability of the first core, where the current processing capability is represented by the number of processable subtasks.

[0052] In an embodiment of the present application, when the first core needs to obtain subtasks of the first task, it can send a subtask acquisition instruction to the task counting module. This subtask acquisition instruction notifies the task counting module of the first core's current processing capacity, thereby facilitating the task counting module to accumulate the number of assigned subtasks. Here, the first core's current processing capacity is the number of subtasks that the first core can process; the first core can carry the number of subtasks that can be processed in the task acquisition instruction and send it to the task counting module.

[0053] In an embodiment of the present application, when the CPU issues a first task to the GPU, each core typically has not yet acquired a subtask. In this case, if the first core needs to acquire a subtask, it can detect the number of subtasks that can be processed by the first core at a preset acquisition time interval. If the number of subtasks that can be processed is greater than 0, it determines that a subtask needs to be acquired. The first core can also determine that a subtask needs to be acquired when it detects that a subtask has been processed within the first core. This can be set as needed and is not limited in the embodiment of the present application.

[0054] S202: Respond to the subtask acquisition instruction through the task counting module and send the count value to the first core.

[0055] In the embodiment of the present application, the task counting module records the number of assigned subtasks, ie, the count value. After receiving the subtask acquisition instruction, the task counting module can respond to the subtask acquisition instruction and send the count value to the first core.

[0056] S203 : Determine, by the first core, a target subtask from unassigned subtasks in the first task according to the count value and the current processing capability.

[0057] In an embodiment of the present application, after receiving the count value from the task counting module, the first core can determine the target subtask from the unassigned subtasks based on the count value and the number of processable subtasks, where the number of target subtasks is the number of processable subtasks.

[0058] For example, FIG5 is a schematic diagram of a multi-core autonomous scheduling process provided by an embodiment of the present application. As shown in FIG5 , after CPU 6 issues a task to GPU 5, the task is stored in GPU memory 53. The task includes multiple subtasks. Multiple cores within GPU 5 (exemplarily shown are cores 52-1, 52-2, 52-3, and 52-4) can send subtask acquisition instructions to task counting module 54, receive the count value returned by task counting module 54, and acquire and process the subtasks from memory 53 based on the count value and the number of processable subtasks.

[0059] It can be understood that the GPU can record the number of assigned subtasks through the task counting module, and the core can obtain the number of assigned subtasks from the task counting module, and determine the target subtasks that can be assigned to the core among the unassigned tasks based on the number of assigned subtasks; in this way, the accuracy of the core in determining the target subtasks can be improved.

[0060] In some embodiments of the present application, after receiving the subtask acquisition instruction and sending the count value to the first core, the task counting module may also increase the count value by the number of processable subtasks to obtain an updated count value.

[0061] In the embodiment of the present application, each time the task counting module receives a task acquisition instruction and obtains the number of processable subtasks from a core, it needs to feed back the count value to the core. Then, the number of processable subtasks is added to the count value to obtain an updated count value. In other words, the first core can obtain the number of processable subtasks, and at the same time, the number of assigned subtasks recorded by the task counting module is increased by the number of processable subtasks. In this way, the count value can be updated in a timely manner, improving the accuracy of the count value.

[0062] In some embodiments of the present application, the first core may send a subtask acquisition instruction to the task counting module upon receiving subtask acquisition instruction information from the central processing unit.

[0063] In an embodiment of the present application, when the CPU sends the first task to the GPU, it will also send subtask acquisition indication information to each core; in this way, when the first core receives the subtask acquisition indication information, it determines that it needs to obtain the subtask of the first task and sends the subtask acquisition instruction to the task counting module.

[0064] In the embodiment of the present application, before the CPU issues the first task, no subtasks of the first task are present in each core of the GPU. When the CPU issues the first task and the first core receives the subtask acquisition instruction information, the first core can set the number of subtasks that can be processed to the preset maximum number of tasks.

[0065] In some embodiments, the CPU sends the first task to the GPU, and the CPU may send subtask acquisition instruction information to each core through broadcasting.

[0066] It is understood that the first core can determine, through the subtask acquisition indication information, that the CPU has sent the first task to the GPU, thereby determining that the number of subtasks that can be processed is the preset maximum number of tasks, and then begin acquiring and processing the target subtasks. In this way, the core can obtain the subtasks of the first task in a timely manner, thereby improving the processing efficiency of the first task.

[0067] In some embodiments of the present application, when a first number of subtasks in the first core are processed, a subtask acquisition instruction can be sent to the task counting module through the first core; the number of processable subtasks is the number of the first number of subtasks that have been processed.

[0068] In an embodiment of the present application, the first core can process multiple subtasks at the same time, but the processing time of each subtask is different and the time to complete the processing is different; as long as there are subtasks that have been processed, the resources of the processed subtasks can be freed up to process newly acquired subtasks.

[0069] In an embodiment of the present application, the first core can detect the processing status of subtasks. When it is determined that a first number of subtasks within the first core have been processed, it is determined that the first core needs to acquire a new subtask and sends a subtask acquisition instruction to the task counting module. Here, the first number is less than or equal to a preset maximum number of tasks, and the embodiment of the present application does not impose any restrictions on the value of the first number.

[0070] In some embodiments, the first core can monitor the subtask processing status in real time. Upon detecting that a subtask has been processed, the first core determines the number of subtasks that have been processed as a first number. The first core then sends a subtask acquisition instruction to the task counting module to initiate acquisition of subtasks. The first number can be one or more, and the first number can be less than or equal to a preset maximum number of tasks.

[0071] It is understandable that the first core can promptly acquire and process subtasks when a subtask has been completed and there are idle resources available to process new subtasks. This enables the first core to promptly acquire the target subtask, reducing the possibility of idle core resources, thereby improving the processing efficiency of the first task.

[0072] In some embodiments of the present application, the first core may send a subtask acquisition instruction to the task counting module if the time interval between the last time the task acquisition instruction was sent is greater than or equal to the preset acquisition time interval and the number of processable subtasks is greater than 0; the number of processable subtasks is the number of subtasks processed and completed after the last time the task acquisition instruction was sent.

[0073] In an embodiment of the present application, after sending a subtask acquisition instruction, the first core may record the time interval between the subtask acquisition instructions and the first core, obtain the number of processable subtasks at a preset acquisition time interval, and send a subtask acquisition instruction to the task counting module based on the number of processable subtasks. The number of processable subtasks is the number of subtasks that have been processed and completed between two subtask acquisition instructions.

[0074] In some embodiments, each time the first core sends a subtask acquisition instruction, it may reset a timer. In this case, the time recorded by the timer is the time interval between the last time the task acquisition instruction was sent. If the time on the timer is greater than or equal to the preset acquisition interval, the first core may obtain the number of processable subtasks and send a subtask acquisition instruction to the task counting module based on the number of processable subtasks.

[0075] It should be noted that when the time interval between the last time a task acquisition instruction was sent is greater than or equal to the preset acquisition time interval, if the number of processable subtasks is 0, the core can send a subtask acquisition instruction to the task counting module to obtain 0 target subtasks; in this way, the task counting module can ignore the subtask acquisition instruction after receiving the subtask acquisition instruction.

[0076] It can be understood that the core can determine the number of processable subtasks according to a preset acquisition time interval, send a subtask acquisition instruction to the task counting module based on the number of processable subtasks, and obtain the target subtasks of the number of processable subtasks; in this way, the core can obtain the target subtasks in a timely manner, reduce the possibility of idle core resources, and thus improve the processing efficiency of the first task.

[0077] In some embodiments of the present application, the first core can determine the number of processable subtasks when the number of processed subtasks is greater than or equal to the preset number of processable subtasks, and send a subtask acquisition instruction to the task counting module based on the number of processable subtasks to start acquiring subtasks. The number of processable subtasks is the number of processed subtasks. The preset number of processable subtasks is less than the preset maximum number of tasks, and the preset number of processable subtasks can be set as needed, which is not limited by the embodiments of the present application. In this way, while achieving intra-core load balancing, the frequency with which the first core sends subtask acquisition instructions to the task counting module is reduced, thereby reducing the resource consumption of the first core.

[0078] In some embodiments of the present application, Figure 6 is a process diagram of another task processing method provided in an embodiment of the present application. In S203, the implementation of the target subtask is determined from the unassigned subtasks in the first task based on the count value and the current processing capacity, as shown in Figure 6, which may include: S301 to S302.

[0079] S301 : When the count value is less than the total number of subtasks of the first task, determine a target task identifier of a target subtask according to the count value and the number of processable subtasks.

[0080] In an embodiment of the present application, the task counting module continuously updates a count value as it receives a subtask acquisition instruction, with the count value continuously increasing. After receiving the count value fed back by the task counting module, the first core can first determine whether the first task has any unassigned subtasks based on the count value. In this embodiment of the present application, the first core can compare the count value with the total number of subtasks in the first task. If the count value is less than the total number of subtasks in the first task, the first core determines that the number of unassigned subtasks is greater than 0, indicating that the first task has any unassigned subtasks. In this case, the first core can determine the target task identifier from the unassigned subtasks based on the count value and the number of processable subtasks. If the count value is greater than or equal to the total number of subtasks in the first task, the number of unassigned subtasks is 0, indicating that the first task has no unassigned subtasks. In this case, the first core can determine that the target subtask does not exist. If the target task does not exist, the first core does not need to determine the target subtask identifier and stops acquiring the target subtask. This allows the first core to acquire the target subtask only if there are any unassigned subtasks in the first task, thereby improving the accuracy of the core's acquisition of the target subtask.

[0081] In the embodiment of the present application, the first core may determine the identifier of the next subtask waiting to be assigned based on the count value, and further determine the number of processable subtasks and the identifier of the target subtask, ie, the target task identifier.

[0082] In some embodiments of the present application, the subtasks of the first task include task identifiers; the order of the task identifiers is used to characterize the allocation order of the subtasks of the first task; the first core can determine the starting identifier in the target task identifier based on the count value, and obtain the number of task identifiers of processable subtasks in the allocation order as the target task identifier.

[0083] In an embodiment of the present application, the subtasks of the first task are provided with task identifiers, and the order of the task identifiers can represent the order in which the subtasks are assigned. In some embodiments, the task identifiers can be numbers, and the order in which the subtasks are assigned is in ascending order of the numbers. In some embodiments, the task identifiers can be even numbers, and the order in which the subtasks are assigned is in ascending order of the even numbers. In some embodiments, the task identifiers can be letters, and the order in which the subtasks are assigned is in alphabetical order, from first to last; this can be set as needed, and the embodiment of the present application does not impose any restrictions.

[0084] In an embodiment of the present application, after obtaining the count value, the first core may determine the starting identifier in the target task identifier, and sequentially obtain the identifiers of the number of processable subtasks in the allocation order as the target task identifier.

[0085] For example, task A includes M subtasks, namely a0 to a M-1, in the case that N subtasks have been assigned, the unassigned subtasks include a N to a M-1 , where N is a positive integer less than M-1. The first core sends a subtask acquisition instruction to the counter, carrying the number of processable subtasks as 2. The counter updates the count value from N to N+2 and feeds N back to the first core. The first core can determine that the identifier of the first subtask of the target subtask is a N , the next subtask is identified as a N+1 , so the target task identification includes: a N and a N+1 .

[0086] It can be understood that since the task identifier can represent the allocation order of subtasks, the first core can determine the starting identifier of the target subtask based on the count value, and then determine the target task identifier based on the number of processable subtasks, which can improve the efficiency of determining the identifier of the target subtask.

[0087] S302: Determine the subtask corresponding to the target task identifier among the unassigned subtasks as the target subtask.

[0088] In the embodiment of the present application, after determining the target task identifier, the first core may obtain a subtask with the same task identifier as the target task identifier from unassigned subtasks, which is the target subtask.

[0089] It is understandable that the first core may first determine the target task identifier and then obtain the target subtask according to the target task identifier. In this way, the accuracy of obtaining the target subtask can be improved.

[0090] In some embodiments of the present application, when the task counting module receives multiple task acquisition instructions from multiple cores, it sorts them according to the number of processable subtasks in the task acquisition instructions to obtain a response order; and responds to the task acquisition instructions in sequence according to the response order.

[0091] In an embodiment of the present application, multiple cores in a GPU may simultaneously send subtask acquisition instructions to the task counting module. At this time, the task counting module can obtain the number of subtasks that can be processed by each core, sort the number of subtasks that can be processed by each core by size, and obtain a response order. Then, according to the response order, the subtask acquisition instructions are responded to one by one. Here, the response order can be the order of the number of subtasks that can be processed from small to large, or the order of the number of subtasks that can be processed from large to small. This can be set as needed and is not limited by the embodiment of the present application.

[0092] It is understandable that the task counting module can only respond to one subtask acquisition instruction at a time. When multiple subtask acquisition instructions are received at the same time, the response order is determined according to the number of processable subtasks, reducing response errors and improving the accuracy of the core acquiring the target subtask, thereby improving the performance of processing tasks.

[0093] In some embodiments, the task counting module may respond to multiple cores sequentially, in ascending order of the number of subtasks they can process. That is, the task counting module may respond first to cores with a smaller number of subtasks they can process, allowing them to obtain a smaller number of target subtasks first. This reduces the probability of a single core obtaining a large number of subtasks while other cores have no subtasks to obtain, thereby improving core load balancing.

[0094] In some embodiments of the present application, the implementation of obtaining and processing the target subtask in S102 may include: feeding back idle information to the central processing unit through the first core when all subtasks of the first core are processed.

[0095] In an embodiment of the present application, while processing subtasks, the first core will promptly obtain unassigned subtasks in the first task for processing until all subtasks of the first task are assigned; the first core will not be able to obtain new subtasks until all subtasks are processed. If the first core has no subtasks to process, it can feedback idle information to the CPU.

[0096] In an embodiment of the present application, all cores in the GPU feed back idle information to the CPU, indicating that the GPU has completed processing the first task and is in an idle state; that is, the first core can inform the CPU of its own processing status by feeding back idle information, and then inform the CPU of the processing status of the GPU; in this way, the CPU can continue to issue the next task when the GPU is idle, thereby improving the efficiency of the CPU in issuing tasks.

[0097] In some embodiments of the present application, in S101, during the process of the GPU processing the first task, based on the current processing capability of the first core and the number of allocated subtasks allocated to each core in the first task, the implementation of the target subtask of the first core is determined from the first task, and it may also include: receiving the first task issued from the central processing unit; the first task includes the total number of subtasks in the first task.

[0098] In an embodiment of the present application, the GPU receives a first task from the CPU, including the total number of subtasks of the first task. Thus, upon receiving the first task, the GPU can determine the total number of subtasks of the first task. The first core can then determine whether there are any unassigned subtasks for the first task based on the total number of subtasks and the count value. Consequently, the first core can promptly stop acquiring target subtasks when all the first tasks have been assigned, thereby reducing core resource consumption.

[0099] In some embodiments of the present application, in S101, during the process of the GPU processing the first task, based on the current processing capability of the first core and the number of allocated subtasks allocated to each core in the first task, determining the implementation of the target subtask of the first core from the first task can include: responding to the clear instruction from the central processing unit through the task counting module to clear the count value.

[0100] In an embodiment of the present application, the GPU may receive a reset instruction from the CPU, the reset instruction being used to instruct the GPU to reset the count value of the task counting module. The GPU responds to the reset instruction and resets the count value of the task counting module, thereby updating the count value to 0.

[0101] In some embodiments, the clear instruction and the first task may be received simultaneously. The GPU receives the clear instruction at the same time as the first task. The GPU then responds to the clear instruction by clearing the count value. In some embodiments, the clear instruction may be received before or after the first task. In some embodiments, the clear instruction may be issued by the first task. This can be configured as needed and is not limited in this embodiment of the present application.

[0102] It is understandable that the GPU needs to clear the count value every time it receives a first task. In this way, the count value of the task counting module is the count value of the number of tasks assigned to the first task, which can improve the accuracy of the count value.

[0103] Based on the above-mentioned task processing method, an embodiment of the present application shows a process of a task processing method; FIG7 is a process diagram of another task processing method provided by an embodiment of the present application. As shown in FIG7 , the method may include:

[0104] S11, the CPU sends a first task and a clear instruction to the GPU, and sends a broadcast subtask acquisition instruction to each core in the GPU; wherein the first task includes the total number of subtasks;

[0105] S12, the GPU responds to the reset instruction through the task counting module and resets the count value to zero;

[0106] S13. The GPU responds to the subtask acquisition instruction information through the first core and sends a subtask acquisition instruction to the task counting module; the number of processable subtasks carried in the subtask acquisition instruction is the preset maximum number of tasks;

[0107] S14. The GPU responds to the subtask acquisition instruction through the task counting module, sends the count value to the first core, and increases the count value by the number of processable subtasks to obtain an updated count value;

[0108] S15. The GPU determines through the first core whether the count value is less than the total number of subtasks; if so, execute S16; otherwise, execute S19;

[0109] S16. The GPU determines the target task identifier through the first core according to the count value and the number of processable subtasks;

[0110] S17, the GPU uses the first core to obtain a subtask with the same task ID as the target task ID from the unassigned subtasks as the target subtask;

[0111] S18. The GPU determines whether a subtask in the core has been processed through the first core; if so, execute S14; otherwise, continue to execute S18;

[0112] S19. The GPU sends idle information to the CPU through the first core after completing processing of all subtasks of the first core.

[0113] In an embodiment of the present application, after receiving the idle information of all cores, the CPU can determine that the first task processing is completed and can send the next task to the GPU.

[0114] An embodiment of the present application provides a multi-core graphics processor (GPU), wherein a first core is configured to, during processing of a first task by the GPU, determine a target subtask for the first core from the first task based on the first core's current processing capability and the number of allocated subtasks in the first task that have been allocated to each core; the first core is one of the multiple cores; and the first core is further configured to obtain and process the target subtask.

[0115] In some embodiments, the GPU further includes a task counting module configured to determine a count value; the count value is used to represent the number of allocated subtasks of each core.

[0116] In some embodiments, the first core is further configured to send a subtask acquisition instruction to the task counting module; the subtask acquisition instruction includes the current processing capability of the first core, and the current processing capability is represented by the number of processable subtasks; the task counting module is further configured to respond to the subtask acquisition instruction and send the count value to the first core; the first core is further configured to determine the target subtask from the unassigned subtasks in the first task based on the count value and the current processing capability.

[0117] In some embodiments, the task counting module is further configured to increase the count value by the number of the processable subtasks to obtain an updated count value.

[0118] In some embodiments, the first core is further configured to send the subtask acquisition instruction to the task counting module upon receiving subtask acquisition indication information from the central processing unit.

[0119] In some embodiments, the first core is further configured to send the subtask acquisition instruction to the task counting module when a first number of subtasks within the first core are processed; the number of processable subtasks is the number of the first number of subtasks that have been processed.

[0120] In some embodiments, the first core is further configured to send the subtask acquisition instruction to the task counting module when the time interval between the last sending of the subtask acquisition instruction is greater than or equal to the preset acquisition time interval and the number of processable subtasks is greater than 0; the number of processable subtasks is the number of subtasks processed and completed after the last sending of the task acquisition instruction.

[0121] In some embodiments, the first core is further configured to, when the count value is less than the total number of subtasks of the first task, determine the target task identifier of the target subtask based on the count value and the number of processable subtasks; and determine the subtask corresponding to the target task identifier in the unassigned subtasks as the target subtask.

[0122] In some embodiments, the subtasks of the first task include task identifiers; the order of the task identifiers is used to characterize the allocation order of the subtasks of the first task; the first core is also configured to determine the starting identifier in the target task identifier based on the count value, and obtain the number of task identifiers of the processable subtasks in sequence according to the allocation order as the target task identifier.

[0123] In some embodiments, the first core is further configured to determine that the number of unassigned subtasks is 0 and the target subtask does not exist if the count value is greater than or equal to the total number of subtasks of the first task.

[0124] In some embodiments, the task counting module is further configured to, upon receiving multiple subtask acquisition instructions from multiple cores, sort them according to the number of processable subtasks in the subtask acquisition instructions to obtain a response order; and respond to the subtask acquisition instructions in sequence according to the response order.

[0125] In some embodiments, the first core is further configured to feed back idle information to the central processing unit when all subtasks of the first core are processed.

[0126] In some embodiments, the task counting module is further configured to clear the count value to zero in response to a clear instruction from the central processing unit.

[0127] In some embodiments, the task counting module is located in the memory of the GPU

[0128] It should be noted that, in the embodiment of the present application, the above-mentioned task processing method can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling the graphics processor of an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0129] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0130] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is run in an electronic device, a graphics processor in the electronic device executes some or all of the steps for implementing the above method.

[0131] An embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the above method.

[0132] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented by hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.

[0133] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0134] The embodiment of the present application also provides a structure of an electronic device. FIG8 is a schematic diagram of the hardware structure of an electronic device provided by the embodiment of the present application. As shown in FIG8 , the electronic device 700 includes a processor 701, a communication interface 702, and a memory 703; wherein the processor 701 includes a central processing unit 7011 and a multi-core graphics processor 7012.

[0135] The central processing unit 7011 is configured to send a first task to the multi-core graphics processor 7012 .

[0136] The multi-core graphics processor 7012 is configured to execute the above task processing method.

[0137] The communication interface 702 enables the electronic device to communicate with other terminals or servers through a network.

[0138] The memory 703 is configured to store instructions and applications executable by the processor 701 and can also cache data to be processed or processed by the processor 701 and various modules in the electronic device 700 (for example, image data, audio data, voice communication data, and video communication data). This can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between the processor 701, the communication interface 702, and the memory 703 via a bus 704.

[0139] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0140] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0142] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0143] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0144] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0145] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0146] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application. Industrial Applicability

[0147] Embodiments of the present application provide a task processing method for a multi-core graphics processor (GPU), a multi-core graphics processor, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: during the GPU processing of a first task, determining a target subtask for the first core from the first task based on the current processing capacity of the first core and the number of subtasks already assigned to each core in the first task; the first core is one of multiple cores; and the target subtask is obtained and processed by the first core. In this way, the faster a core processes subtasks and the more subtasks it completes, the more subtasks it can obtain, thereby reducing the accumulation of subtasks within some cores, improving the load balance between cores, and thereby improving the task processing performance of the GPU.

Claims

1. A task processing method for a multi-core graphics processing unit (GPU), comprising: During the process of the GPU processing a first task, based on the current processing capacity of a first core and the number of assigned subtasks of the first task that have been assigned to each core, determine a target subtask of the first core from the first task; The first core is one of the multi-core; Obtain and process the target subtask through the first core.

2. The method according to claim 1, wherein, The multi-core GPU includes a task counting module, and the number of assigned subtasks of each core is determined by the count value of the task counting module.

3. The method according to claim 2, wherein, The determining the target subtask of the first core from the first task includes: Through the first core, send a subtask acquisition instruction to the task counting module; the subtask acquisition instruction includes the current processing capacity of the first core, and the current processing capacity is characterized by the number of subtasks that can be processed; In response to the subtask acquisition instruction by the task counting module, send the count value to the first core; Through the first core, determine the target subtask from the unassigned subtasks of the first task according to the count value and the current processing capacity.

4. The method according to claim 3, wherein, The method further includes: increasing the count value by the number of subtasks that can be processed through the task counting module to obtain an updated count value.

5. The method according to claim 3, wherein The sending the subtask acquisition instruction to the task counting module includes: When receiving subtask acquisition indication information from a central processing unit, send the subtask acquisition instruction to the task counting module.

6. The method according to claim 3, wherein The sending the subtask acquisition instruction to the task counting module includes: When a first number of subtasks within the first core are processed and completed, send the subtask acquisition instruction to the task counting module; the number of subtasks that can be processed is the number of the first number of subtasks that have been processed and completed.

7. The method according to claim 3, wherein, The sending the subtask acquisition instruction to the task counting module includes: When the time interval between the current sending of the subtask acquisition instruction and the previous sending is greater than or equal to a preset acquisition time interval, and the number of subtasks that can be processed is greater than 0, send the subtask acquisition instruction to the task counting module; the number of subtasks that can be processed is the number of subtasks that have been processed and completed after the previous sending of the task acquisition instruction.

8. The method according to claim 3, wherein The determining the target subtask from the unassigned subtasks of the first task according to the count value and the current processing capacity includes: When the count value is less than the total number of subtasks of the first task, determine the target task identifier of the target subtask according to the count value and the number of subtasks that can be processed; Determine the subtask corresponding to the target task identifier among the unassigned subtasks as the target subtask.

9. The method according to claim 8, wherein The subtasks of the first task include task identifiers; the order of the task identifiers is used to represent the assignment order of the subtasks of the first task; the determining the target task identifier of the target subtask according to the count value and the number of subtasks that can be processed includes: Determine the starting identifier in the target task identifier according to the count value, and sequentially obtain the number of task identifiers of the processable subtasks according to the allocation order as the target task identifier.

10. The method according to claim 3, wherein The determining of the target subtask from the unallocated subtasks in the first task according to the count value and the current processing capacity includes: When the count value is greater than or equal to the total number of subtasks of the first task, determine that the number of unallocated subtasks is 0 and the target subtask does not exist.

11. The method according to claim 3, wherein, The method further includes: When the task counting module receives multiple subtask acquisition instructions from multiple cores, sort them according to the number of processable subtasks in the subtask acquisition instructions to obtain a response order; respond to the subtask acquisition instructions sequentially according to the response order.

12. The method according to any one of claims 1 to 11, wherein The method further includes: Through the first core, when all the subtasks of the first core are processed, feedback idle information to the central processing unit.

13. The method according to any one of claims 1 to 11, wherein, The method further includes: Receive the first task issued by the central processing unit; the first task includes the total number of subtasks of the first task.

14. The method according to claim 13, wherein, The method further includes: Through the task counting module, respond to the clear instruction from the central processing unit to clear the count value.

15. The method according to claim 9, wherein, The task identifier is represented by a 64-bit integer.

16. The method according to claim 2, wherein The task counting module is located in the memory of the GPU.

17. A multi-core graphics processing unit (GPU), comprising: A first core configured to determine a target subtask of the first core from the first task based on the current processing capacity of the first core and the number of allocated subtasks allocated to each core in the first task during the process of the GPU processing the first task; The first core is one of the multi-core. The first core is further configured to obtain and process the target subtask.

18. An electronic device, comprising: The multi-core graphics processing unit (GPU) according to claim 17.

19. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the task processing method according to any one of claims 1 to 16 are implemented.

20. A computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps in the task processing method according to any one of claims 1 to 16 are implemented.

Citation Information

Patent Citations

  • Technique for improving performance in multi-threaded processing units

    CN103729167A

  • Optimized task partitioning through data mining

    CN106802878A

  • Task processing method and device, electronic terminal and readable storage medium

    CN108255607A

  • Task state management method and device

    CN109558237A

  • Task scheduling system, method and equipment and storage medium

    CN114579286A

Cited By

  • Data processing method, graphics processor, storage medium and program product

    CN120807266A

  • Request aggregation and dynamic scheduling method based on lock competition, electronic equipment and program product

    CN122412173A