GPU thread load balancing method, device, chip, and electronic device

By establishing dynamic mapping relationships between threads and tasks in the GPU and building global and local queues, the problem of load imbalance among GPU threads is solved, and the performance of GPU programs is significantly improved.

CN114579299BActive Publication Date: 2025-05-16INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210112283.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-29
Publication Date
2025-05-16
Estimated Expiration
2042-01-29

AI Technical Summary

Technical Problem

When existing GPUs deal with scenarios that require a combination of graphics processing algorithms and non-graphical algorithms such as artificial intelligence and scene management, there is a problem of load imbalance between threads, resulting in limited performance improvement.

Method used

By establishing dynamic mapping relationships between threads and tasks and building local queues and global queues, dynamic allocation according to task load is achieved to ensure load balancing among GPU threads. The specific methods include fixed the number of workgroups and threads opened by the GPU program, building a global task queue, and managing task allocation through global atomic variables.

Benefits of technology

It effectively eliminates the problem of load imbalance between GPU threads and greatly improves the performance of GPU programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114579299B_ABST
    Figure CN114579299B_ABST
Patent Text Reader

Abstract

The present invention proposes a GPU thread load balancing method, device, medium, and electronic device, the method comprising: fixing the number of working groups opened by the GPU program and the number of threads opened by each working group, fixing the total number of threads opened by the GPU program; grouping all computing tasks that need to be processed and putting them all into a command queue, building a global task queue, and allowing each of the working groups to have access to the global task queue; according to the current computing task load of the threads that are always opened by the GPU, obtaining the computing tasks that each of the working groups needs to perform from the global task queue. The method realizes dynamic allocation according to task load by establishing a dynamic mapping relationship between threads and tasks and building local queues and global queues, and finally realizes load balancing between GPU threads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPU technology, and in particular to a GPU thread load balancing method, device, chip, and electronic device. Background Art

[0002] Graphics processing unit (GPU) is an important component of artificial intelligence. As an accelerator, it has been widely used in many fields such as high-performance computing and image processing.

[0003] Existing GPU chips have large-scale fine-grained parallel architecture features, which are suitable for processing large-scale data parallel tasks; at the same time, in traditional GPU programming, the mapping between threads and tasks is static, and the mapping relationship between threads and tasks does not change during the entire program execution process. Existing GPUs are very friendly to applications with large-scale data parallel features and can greatly improve the performance of applications. However, for scenarios in games that require the combination of graphics processing algorithms with non-graphics algorithms such as artificial intelligence and scene management, the current GPU architecture will have load imbalance problems between threads when using the GPU for multi-threaded parallel optimization.

[0004] To address the above problems, the program splitting method and the coarse-grained parallel method are currently commonly used. Among them, the program splitting method splits the application into several parts according to the load conditions, and writes a GPU kernel for each part, trying to ensure that the load of each thread in each GPU kernel is balanced. However, this method does not fundamentally eliminate the phenomenon of GPU thread load imbalance, and introduces multiple kernels, which also increases kernel startup and synchronization overhead, and has limited performance improvement. The coarse-grained parallel method reduces the number of GPU threads enabled and increases the workload of each thread, thereby reducing the performance bottleneck caused by the complex and unbalanced GPU threads, but this method still does not fundamentally eliminate the problem of load imbalance between GPU threads. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes a GPU thread load balancing method, device, chip, and electronic device. By establishing a dynamic mapping relationship between threads and tasks and constructing local queues and global queues, dynamic allocation according to task loads is achieved, and finally load balancing among GPU threads is achieved.

[0006] In order to achieve the above object, the present invention provides a GPU thread load balancing method, comprising:

[0007] By fixing the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, the total number of threads opened by the GPU is fixed;

[0008] Group all computing tasks that need to be processed and put them all into a command queue, build a global task queue, and allow each of the working groups to have access to the global task queue;

[0009] According to the current computing task load of the threads always enabled by the GPU, the computing tasks that each working group needs to execute are obtained from the global task queue.

[0010] Optionally, obtaining from the global task queue the computing tasks that each of the working groups needs to perform further includes:

[0011] Set the initial global atomic variable as the global task queue access switch;

[0012] In the case where the working group obtains access rights to the initial global atomic variable, locking the unique access rights to the initial global atomic variable, and

[0013] According to the variable value of the initial global atomic variable, a group of computing tasks that the working group needs to execute is obtained from the global task queue.

[0014] Optionally, obtaining a group of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable further includes:

[0015] When the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of the computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution;

[0016] Releases access to the initial global atomic variable.

[0017] Optionally, obtaining a group of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable further includes:

[0018] When the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it indicates that all the computing tasks in the global task queue have been executed, and the working group exits execution.

[0019] Optionally, the setting of the initial global atomic variable as a global task queue access switch further includes:

[0020] The variable value of the initial global atomic variable is initialized to 0, and thread No. 0 of each working group accesses the initial global atomic variable.

[0021] Optionally, the method further comprises: dynamically allocating the computing tasks to be processed by each thread of the working group.

[0022] Optionally, the dynamically allocating the computing tasks to be processed by each thread of the workgroup further includes:

[0023] Build a local task queue;

[0024] A set of computing tasks to be executed obtained by each thread opened by the working group is stored in the local task queue;

[0025] All threads opened by the working group adopt a coarse-grained parallel approach to collaboratively process each computing task in the local task queue.

[0026] Another aspect of the present invention further provides a GPU thread load balancing device, comprising:

[0027] Fixed number of open threads module: used to fix the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, and fix the total number of threads opened by the GPU;

[0028] A global task queue construction module: used to group all computing tasks that need to be processed and put them all into the command queue, construct a global task queue, and allow each of the working groups to have access to the global task queue;

[0029] The load balancing module is used to obtain the computing tasks that each working group needs to execute from the global task queue according to the current computing task load of the threads that are always enabled by the GPU.

[0030] Another aspect of the present invention further provides a GPU chip, the chip comprising a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the GPU thread load balancing method.

[0031] Another aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned GPU thread load balancing method when executing the computer program.

[0032] It can be seen from the above scheme that the advantages of the present invention are:

[0033] The GPU thread load balancing method provided by the present invention changes the static mapping relationship between traditional GPU programming threads and tasks into a dynamic mapping relationship by fixing the number of GPU threads enabled and the GPU global command queue based on global atomic variables. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of load imbalance between GPU threads, thereby greatly improving the performance of the GPU program. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flowchart of a GPU thread load balancing method provided in Embodiment 1 of the present invention;

[0035] Figure 2 A flowchart of a GPU thread load balancing method provided in Embodiment 2 of the present invention;

[0036] Figure 3 This is a specific flow chart of step S4 of Example 2;

[0037] Figure 4 It is a framework diagram of the GPU thread load balancing device;

[0038] Figure 5 It is a schematic diagram of the structure of a computer device;

[0039] Figure 6 A schematic diagram of the hardware structure of an electronic device;

[0040] in:

[0041] 400-GPU thread load balancing device;

[0042] 401-Fixed thread number module;

[0043] 402-global task queue construction module;

[0044] 403-load balancing module;

[0045] 404-local task queue building module;

[0046] 500-computer equipment;

[0047] 501- storage medium;

[0048] 502 - processor;

[0049] 600-Electronic equipment;

[0050] 601- radio frequency unit;

[0051] 602-network module;

[0052] 603- audio output unit;

[0053] 604-input unit;

[0054] 641-graphics processor;

[0055] 642-Microphone;

[0056] 605-Sensor;

[0057] 606-display unit;

[0058] 6061-display panel;

[0059] 607-user input unit;

[0060] 6071-touch panel;

[0061] 6072-Other input devices;

[0062] 608-interface unit;

[0063] 609-Memory;

[0064] 610 - Processor. DETAILED DESCRIPTION

[0065] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.

[0066] As mentioned above, in the optimization of GPU programs, the load imbalance between GPU threads will seriously affect the program performance and become the performance bottleneck of GPU programs. The main technical means for CPU multi-threaded programs to solve the load imbalance between threads is dynamic task allocation, that is, dynamically allocating tasks according to the execution status of threads. However, in GPU programs, due to the static mapping relationship between threads and tasks, it is difficult to achieve dynamic task allocation. Neither the program splitting method nor the coarse-grained parallel method fundamentally eliminates the performance bottleneck caused by the load imbalance between GPU threads. The reason is that it does not fundamentally change the GPU programming feature of static mapping between threads and tasks, resulting in unbalanced task load between GPU threads.

[0067] To this end, an embodiment of the present invention provides a GPU thread load balancing method, which realizes dynamic allocation according to task load by establishing a dynamic mapping relationship between threads and tasks and constructing local queues and global queues, and ultimately achieves load balancing among GPU threads.

[0068] Specifically, refer to Figure 1As shown, Figure 1 A schematic diagram of the process flow of a GPU thread load balancing method provided in Embodiment 1 of the present invention is shown.

[0069] A GPU thread load balancing method, comprising:

[0070] S1. By fixing the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, the total number of threads opened by the GPU is fixed.

[0071] In a specific implementation, the GPU includes multiple control units CU, each of which can execute multiple working groups at the same time. In this embodiment, the number of working groups opened by the GPU program and the number of threads opened by each working group are fixed, and the total number of threads opened by the GPU is fixed. For example, assuming that the GPU has M CUs, the number of working groups running on each CU is N, and the number of threads opened by each working group is 32 (NVIDIA GPU) or 64 (AMD GPU), the total number of threads opened is M*N*32 / 64. At this time, the number of opened threads is much smaller than the number of tasks, and each thread will process multiple tasks.

[0072] S2. Group all computing tasks that need to be processed and put them all into a command queue, build a global task queue, and allow each of the working groups to have access to the global task queue.

[0073] In the specific implementation, the global task queue is located in the Global Memory of the GPU and is implemented in a stack. During the execution process, all computing tasks that need to be processed are grouped and put into the command queue. Each group contains m computing tasks, and there are n groups of computing tasks in total. Each work group is allowed to have access to the global task queue, that is, all work groups can access the task queue, and the threads opened by the work group are dynamically mapped to the computing tasks in the global task queue.

[0074] S3. According to the current computing task load of the threads always enabled by the GPU, the computing tasks that need to be executed by each working group are obtained from the global task queue.

[0075] In the specific implementation,

[0076] An initial global atomic variable is set as a global task queue access switch. The variable value of the initial global atomic variable set in this embodiment is initialized to 0, and thread No. 0 of each working group accesses the initial global atomic variable.

[0077] In the case where the working group obtains access to the initial global atomic variable,

[0078] The work group having access rights to the initial global atomic variable locks the sole access rights to the initial global atomic variable to prevent other work groups from accessing the command queue at the same time.

[0079] Then, according to the variable value of the initial global atomic variable, a group of computing tasks that the working group needs to execute is obtained from the global task queue. Specifically,

[0080] In the case where the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution, and the access rights to the initial global atomic variable are released; in the case where the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it means that all the computing tasks in the global task queue have been fully executed, there are no executable computing tasks, and the working group exits execution. When all working groups exit execution, the GPU program is executed. Unlike traditional GPU threads, the GPU thread opened in this embodiment will not exit immediately after processing a computing task, but will continue to process computing tasks until all assigned tasks are processed. At the same time, in this embodiment, by introducing global atomic variables, conflicts in multiple working groups accessing the global command queue at the same time are avoided.

[0081] Compared with the prior art, in this embodiment, by fixing the number of GPU threads enabled and the GPU global command queue based on global atomic variables, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thus realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, and thus greatly improving the performance of the GPU program.

[0082] refer to Figure 2 As shown, Figure 2 A schematic diagram of the flow of a GPU thread load balancing method provided in Embodiment 2 of the present invention is shown.

[0083] A GPU thread load balancing method, comprising:

[0084] S1. By fixing the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, the total number of threads opened by the GPU is fixed.

[0085] In a specific implementation, the GPU includes multiple control units CU, each of which includes multiple running work-groups. In this embodiment, the number of work-groups opened by the GPU program and the number of threads opened by each work-group are fixed, and the total number of threads opened by the GPU is fixed. For example, assuming that the GPU has M CUs, the number of work-groups running on each CU is N, and the number of threads opened by each work-group is 32 (NVIDIA GPU) or 64 (AMD GPU), the total number of threads opened is M*N*32 / 64. At this time, the number of opened threads is much smaller than the number of tasks, and each thread will process multiple tasks.

[0086] S2. Group all computing tasks that need to be processed and put them all into a command queue, build a global task queue, and allow each of the working groups to have access to the global task queue.

[0087] In the specific implementation, the global task queue is located in the Global Memory of the GPU and is implemented in a stack. During the execution process, all computing tasks that need to be processed are grouped and put into the command queue. Each group contains m computing tasks, and there are n groups of computing tasks in total. Each work group is allowed to have access to the global task queue, that is, all work groups can access the task queue, and the threads opened by the work group are dynamically mapped to the computing tasks in the global task queue.

[0088] S3. According to the current computing task load of the threads always enabled by the GPU, the computing tasks that need to be executed by each working group are obtained from the global task queue.

[0089] In the specific implementation,

[0090] An initial global atomic variable is set as a global task queue access switch. The variable value of the initial global atomic variable set in this embodiment is initialized to 0, and thread No. 0 of each working group accesses the initial global atomic variable.

[0091] In the case where the working group obtains access to the initial global atomic variable,

[0092] The work group having access rights to the initial global atomic variable locks the sole access rights to the initial global atomic variable to prevent other work groups from accessing the command queue at the same time.

[0093] Then, according to the variable value of the initial global atomic variable, a group of computing tasks that the working group needs to execute is obtained from the global task queue. Specifically,

[0094] In the case where the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution, and the access rights to the initial global atomic variable are released; in the case where the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it means that all the computing tasks in the global task queue have been fully executed, there are no executable computing tasks, and the working group exits execution. When all working groups exit execution, the GPU program is executed. Unlike traditional GPU threads, the GPU thread opened in this embodiment will not exit immediately after processing a computing task, but will continue to process computing tasks until all assigned tasks are processed. At the same time, in this embodiment, by introducing global atomic variables, conflicts in accessing the global command queue by different working groups at the same time are avoided.

[0095] S4, dynamically allocate the computing tasks to be processed by each thread of the working group, such as Figure 3 As shown, Figure 3 A specific flow chart of step S4 is shown;

[0096] Specifically include:

[0097] S41, build a local task queue;

[0098] In the specific implementation, the local task queue is located in the local memory on the chip and is also implemented in a stack manner. The local queue has a very high memory access efficiency.

[0099] S42, storing a group of computing tasks to be executed obtained by each thread opened by the working group in the local task queue;

[0100] S43. All threads opened by the working group adopt a coarse-grained parallel approach to collaboratively process each computing task in the local task queue.

[0101] Based on the first embodiment, this embodiment aims to solve the problem of unbalanced task load among threads within a single workgroup by constructing a local task queue. The computing tasks required to be executed obtained by each thread opened by the workgroup are stored in the local task queue. A coarse-grained parallel method is adopted for all opened threads to collaboratively process each computing task in the local task queue, thereby realizing dynamic distribution of task load among threads within a single workgroup.

[0102] In summary, the present invention changes the static mapping relationship between traditional GPU programming threads and tasks into a dynamic mapping relationship by fixing the number of GPU threads enabled, building a GPU global command queue based on global atomic variables, and a local task queue. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, thereby greatly improving the performance of the GPU program.

[0103] Reference Figure 4 , Figure 4 A GPU thread load balancing device 400 is shown, which can be applied to a GPU thread load balancing method and can be applied to a personal terminal and a host terminal device, which can be implemented as shown in FIG1- Figure 3 The GPU thread load balancing method shown in the figure, the setting device provided in the embodiment of the present application can implement each process of the above-mentioned GPU thread load balancing method, at least including a fixed start thread number module 401, a global task queue construction module 402, a load balancing module 403, and a local task queue construction module 404, that is, specifically:

[0104] A GPU thread load balancing device 400 includes:

[0105] Fixed open thread number module 401: used to fix the total number of GPU open threads by fixing the number of work groups opened by the GPU program and the number of threads opened by each work group;

[0106] The global task queue construction module 402 is used to group all the computing tasks to be processed and put them all into the command queue, construct a global task queue, and allow each of the working groups to have access to the global task queue;

[0107] The load balancing module 403 is used to obtain the computing tasks that each working group needs to execute from the global task queue according to the current computing task load of the threads that are always enabled by the GPU.

[0108] Optionally, obtaining from the global task queue the computing tasks that each of the working groups needs to perform further includes:

[0109] Set the initial global atomic variable as the global task queue access switch;

[0110] In the case where the working group obtains access rights to the initial global atomic variable, locking the unique access rights to the initial global atomic variable, and

[0111] According to the variable value of the initial global atomic variable, a group of computing tasks that the working group needs to execute is obtained from the global task queue.

[0112] Optionally, obtaining a group of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable further includes:

[0113] When the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of the computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution;

[0114] Releases access to the initial global atomic variable.

[0115] Optionally, obtaining a group of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable further includes:

[0116] When the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it indicates that all the computing tasks in the global task queue have been executed, and the working group exits execution.

[0117] Optionally, the setting of the initial global atomic variable as a global task queue access switch further includes:

[0118] The variable value of the initial global atomic variable is initialized to 0, and thread No. 0 of each working group accesses the initial global atomic variable.

[0119] Optionally, the device local task queue construction module 404 is used to dynamically allocate the computing tasks to be processed by each thread of the work group.

[0120] Optionally, the dynamically allocating the computing tasks to be processed by each thread of the workgroup further includes:

[0121] Build a local task queue;

[0122] A set of computing tasks to be executed obtained by each thread opened by the working group is stored in the local task queue;

[0123] All threads opened by the working group adopt a coarse-grained parallel approach to collaboratively process each computing task in the local task queue.

[0124] Therefore, according to the GPU thread load balancing device 400 of the embodiment of the present application, by fixing the number of GPU threads enabled, building a GPU global command queue based on global atomic variables and a local task queue, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of load imbalance between GPU threads, thereby greatly improving the performance of the GPU program.

[0125] It should be understood that the descriptions of the GPU thread load balancing method are also applicable to the GPU thread load balancing device 400 according to the embodiment of the present application, and will not be described in detail to avoid repetition.

[0126] In addition, it should be understood that in the GPU thread load balancing device 400 according to the embodiment of the present application, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the GPU thread load balancing device can be divided into functional modules different from the modules illustrated above to complete all or part of the functions described above.

[0127] Figure 5 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application.

[0128] like Figure 5 As shown in, an embodiment of the present application also provides a computer device 500, including a storage medium 501, a processor 502, and a program or instruction stored on the storage medium 501 and executable on the processor 502. When the program or instruction is executed by the processor 502, the steps of the above-mentioned GPU thread load balancing method are implemented and the same technical effect can be achieved.

[0129] Therefore, according to the computer device 500 of the embodiment of the present application, by fixing the number of GPU threads enabled, constructing a GPU global command queue based on global atomic variables and a local task queue, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, thereby greatly improving the performance of the GPU program.

[0130] To avoid repetition, other technical effects of the computer device 500 according to the embodiment of the present application are not described in detail here.

[0131] It should be noted that the electronic devices in the embodiments of the present application may include mobile electronic devices and non-mobile electronic devices.

[0132] Figure 6 It is a schematic diagram of the specific hardware structure of the electronic device provided in the embodiment of the present application.

[0133] Reference Figure 6 The electronic device 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610 and other components.

[0134] It should be understood that in the embodiment of the present application, the radio frequency unit 601 can be used for receiving and sending signals during information transmission or calls. Specifically, after receiving downlink data from the base station, it is sent to the processor 610 for processing; in addition, uplink data is sent to the base station. Generally, the radio frequency unit 601 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the radio frequency unit 601 can also communicate with the network and other devices through a wireless communication system.

[0135] The electronic device 600 provides the user with wireless broadband Internet access through the network module 602, such as helping the user to send and receive emails, browse web pages, and access streaming media.

[0136] The audio output unit 603 can convert the audio data received by the RF unit 601 or the network module 602 or stored in the memory 609 into an audio signal and output it as sound. Moreover, the audio output unit 603 can also provide audio output related to a specific function performed by the electronic device 600 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 603 includes a speaker, a buzzer, a receiver, etc.

[0137] The input unit 604 is used to receive audio or video signals. It should be understood that in the embodiment of the present application, the input unit 604 may include a graphics processing unit (GPU) 641 and a microphone 642, and the graphics processor 641 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode.

[0138] The electronic device 600 also includes at least one sensor 605, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 6061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 6061 and / or the backlight when the electronic device 600 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 605 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.

[0139] The display unit 606 is used to display information input by the user or information provided to the user. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0140] The user input unit 607 can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the electronic device. Specifically, the user input unit 607 includes a touch panel 6071 and other input devices 6072. The touch panel 6071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by users using fingers, styluses, or any other suitable objects or accessories on or near the touch panel 6071). The touch panel 6071 may include two parts: a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, and joysticks, which will not be repeated here. The interface unit 608 is an interface for connecting an external device to the electronic device 600. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device having an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 608 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the electronic device 600 or may be used to transmit data between the electronic device 600 and an external device.

[0141] The memory 609 can be used to store software programs and various data. The memory 609 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 609 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0142] The processor 610 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 609, and calling data stored in the memory 609, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 610 may include one or more processing units; preferably, the processor 610 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 610. It can be understood by those skilled in the art that the electronic device 600 may also include a power supply (such as a battery) for powering various components, and the power supply may be logically connected to the processor 610 through a power management system, so that the power management system can realize functions such as management of charging, discharging, and power consumption management. Figure 6 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here. In the embodiments of the present application, the electronic device includes but is not limited to mobile phones, tablet computers, laptop computers, PDAs, vehicle-mounted terminals, wearable devices (such as bracelets, glasses), and pedometers.

[0143] Specifically, the processor 610 is used to: fix the total number of threads opened by the GPU by fixing the number of workgroups opened by the GPU program and the number of threads opened by each workgroup;

[0144] Group all computing tasks that need to be processed and put them all into a command queue, build a global task queue, and allow each of the working groups to have access to the global task queue;

[0145] According to the current computing task load of the threads always enabled by the GPU, the computing tasks that each working group needs to execute are obtained from the global task queue.

[0146] Therefore, according to the electronic device 600 of the embodiment of the present application, by fixing the number of GPU threads enabled, constructing a GPU global command queue based on global atomic variables and a local task queue, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, thereby greatly improving the performance of the GPU program.

[0147] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the above-mentioned GPU thread load balancing method are implemented and the same technical effect can be achieved.

[0148] Therefore, according to the readable storage medium of the embodiment of the present application, by fixing the number of GPU threads enabled, building a GPU global command queue based on global atomic variables and a local task queue, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, thereby greatly improving the performance of the GPU program.

[0149] To avoid repetition, other technical effects of the readable storage medium according to the embodiments of the present application are not described here.

[0150] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0151] An embodiment of the present application also provides a GPU chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the above-mentioned GPU thread load balancing method and achieve the same technical effect.

[0152] Therefore, according to the GPU chip of the embodiment of the present application, by fixing the number of GPU threads enabled, building a GPU global command queue based on global atomic variables and a local task queue, the static mapping relationship between traditional GPU programming threads and tasks is changed to a dynamic mapping relationship. The GPU thread can obtain the computing tasks that each work group needs to perform from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, thereby realizing the dynamic mapping of threads and tasks and the on-demand allocation of tasks, thereby effectively eliminating the problem of unbalanced load between GPU threads, thereby greatly improving the performance of the GPU program.

[0153] To avoid repetition, other technical effects of the chip according to the embodiments of the present application are not described here.

[0154] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0155] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be applied, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0156] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0157] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A GPU thread load balancing method, characterized in that: include: By fixing the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, the total number of threads opened by the GPU is fixed; Group all computing tasks that need to be processed and put them all into a command queue, build a global task queue, and allow each of the working groups to have access to the global task queue; According to the current computing task load of the threads always enabled by the GPU, the computing tasks to be performed by each working group are obtained from the global task queue, including: Set the initial global atomic variable as the global task queue access switch; In the case where the working group obtains access to the initial global atomic variable, Locking the sole access right to the initial global atomic variable, and acquiring a set of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable; When the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of the computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution; Release the access right to the initial global atomic variable; When the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it indicates that all the computing tasks in the global task queue have been executed, and the working group exits execution.

2. The method according to claim 1, characterized in that The step of setting the initial global atomic variable as a global task queue access switch further includes: The variable value of the initial global atomic variable is initialized to 0, and thread No. 0 of each working group accesses the initial global atomic variable.

3. The method according to claim 1, characterized in that Also includes: The computing tasks to be processed by each thread of the working group are dynamically allocated.

4. The method according to any one of claims 1 to 3, characterized in that: The dynamic allocation of the computing tasks to be processed by each thread of the working group further includes: Build a local task queue; A set of computing tasks to be executed obtained by each thread opened by the working group is stored in the local task queue; All threads opened by the working group adopt a coarse-grained parallel approach to collaboratively process each computing task in the local task queue.

5. A GPU thread load balancing device, characterized in that: include: Fixed number of open threads module: used to fix the number of workgroups opened by the GPU program and the number of threads opened in each workgroup, and fix the total number of threads opened by the GPU; A global task queue construction module: used to group all computing tasks that need to be processed and put them all into the command queue, construct a global task queue, and allow each of the working groups to have access to the global task queue; The load balancing module is used to obtain the computing tasks that each work group needs to execute from the global task queue according to the current computing task load of the threads that are always enabled by the GPU, including: Set the initial global atomic variable as the global task queue access switch; In the case where the working group obtains access to the initial global atomic variable, Locking the sole access right to the initial global atomic variable, and acquiring a set of computing tasks that the working group needs to execute from the global task queue according to the variable value of the initial global atomic variable; When the variable value of the initial global atomic variable is less than the number of groups of the computing tasks, a group of the computing tasks corresponding to the variable value of the initial global atomic variable is obtained from the global task queue for execution; Release the access right to the initial global atomic variable; When the variable value of the initial global atomic variable is greater than or equal to the number of groups of the computing tasks, it indicates that all the computing tasks in the global task queue have been executed, and the working group exits execution.

6. A GPU chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, characterized in that: The processor is used to run a program or instruction to implement the GPU thread load balancing method according to any one of claims 1 to 4.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the GPU thread load balancing method described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • KNN-GPU acceleration method based on OpenCL

    CN104020983A

  • Parallel data processing method and parallel processor for effectively eliminating data access delay

    CN112732416A